Multi-modal medical AI cooperative system and hierarchical training method thereof

Through the integration of three-level nested federated learning framework and multimodal data, the shortcomings of medical AI models in geographical differences and multi-group needs are solved, personalized diagnosis and multi-scenario applications are realized, and the adaptability and synergistic efficiency of the model are improved.

CN120388752APending Publication Date: 2025-07-29马恩冕
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510474355.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The existing medical AI models have not effectively solved the problems of insufficient data regional differences and generalization capabilities, incomplete coverage of multi-group demands, and low data coordination and model integration efficiency, especially in different regions, disease differences and multi-group demands.

Method used

A three-level nested federated learning framework is adopted, combined with multimodal data integration, and a multimodal medical AI collaboration system is built through dynamic geographical weights, cross-domain knowledge distillation and multi-particle size model fusion to realize regional adaptive training and data collaboration.

Benefits of technology

The model's learning ability of disease characteristics in different regions has been improved, personalized diagnostic suggestions are provided, and the learning needs of multiple groups has been supported, the model's generalization ability and data collaboration efficiency have been improved, and the problem of traditional education and resource inequality has been solved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388752A_ABST
    Figure CN120388752A_ABST
Patent Text Reader

Abstract

The invention relates to the field of artificial intelligence of computer medicine, in particular to a multi-modal medical AI collaboration system and a hierarchical training method thereof, which can solve the problems of data region difference, insufficient generalization ability, incomplete multi-group demand coverage and low data collaboration and model integration efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Specification Technical Field

[0001] The present invention relates to the field of computer medical artificial intelligence, and specifically to a multi-modal medical AI collaborative system and its hierarchical training method. Background Art

[0002] With the gradual in-depth application of artificial intelligence in the medical field, many deficiencies have emerged in existing medical AI models. On the one hand, the amount of data is relatively limited, and most are trained based on the case data of only a small number of hospitals or departments, making it difficult to cover the differences of various diseases in different regions and the complexity and diversity of diseases, resulting in poor generalization ability of the model. Due to differences in culture, habits, etc. in different regions, there are significant differences in the occurrence, treatment methods, and prognosis of diseases, and existing models pay insufficient attention to regional characteristics and cannot accurately meet the regionalized medical needs.

[0003] On the other hand, the existing applications of AI models are mainly oriented towards patient diagnosis, and less attention is paid to the needs of medical staff, teachers, and students. In the field of medical education, the traditional teaching mode relies on hospital internships and limited teaching materials, and it is difficult for students to access the latest clinical knowledge and advanced experiences of different medical institutions, resulting in a large lag in theory. At the same time, medical staff in grass-roots areas lack learning resources suitable for their own levels, and existing learning materials are too research-oriented, with high costs and insufficient practicality.

[0004] In addition, in terms of data integration and collaborative training, existing technologies mostly adopt a single federated learning framework, do not fully consider the impact of geographical factors on disease characteristics, and lack mechanisms for cross-domain knowledge distillation and multi-granularity model fusion, making it difficult to form an efficient medical AI collaborative system covering the whole country.

[0005] Therefore, developing a multi-modal medical AI collaborative system and its hierarchical training method can solve the problems of data regional differences and insufficient generalization ability, incomplete coverage of multi-group needs, and low efficiency of data collaboration and model integration. Summary of the Invention

[0006] The present invention provides a multi-modal medical AI collaborative system and its hierarchical training method, which realizes regionalized adaptive training through a three-level nested federated learning framework, combines multi-modal data integration and multi-scenario applications, and constructs a medical AI collaborative ecosystem covering the whole country.

[0007] The present invention provides a system architecture of a multi-modal medical AI collaborative system and its hierarchical training method, including the following content:

[0008] For the first-level sub-consortium (city level), an adaptive federated learning algorithm based on dynamic geographic weights is constructed. A geographic attenuation factor is set for each medical institution within the city-level sub-consortium, calculated as W = 1 / (1 + α × d), where d is the straight-line distance between medical institutions and α is the regional disease correlation coefficient. This weighting factor strengthens data interaction between geographically adjacent medical institutions and enhances the ability to learn city-level regional disease characteristics. The dynamic calculation method for the regional disease correlation coefficient α is automatically calibrated through a linear regression model based on data such as referral rates and disease co-occurrence frequencies between city-level medical institutions over the past three years, avoiding the subjectivity of manual settings.

[0009] For the second-level sub-joint (provincial level), a cross-domain knowledge distillation mechanism is developed to extract the feature mapping relationship between different first-level sub-joints through a bidirectional attention mechanism. Specifically, it includes: ① self-attention encoding of the output features of the city-level model to identify local key disease features; ② mutual attention alignment of cross-city features to calculate the feature similarity matrix between different cities. Based on the above mechanism, a provincial knowledge graph is established, containing more than 100,000 disease-region-treatment plan triplets, realizing cross-domain fusion of city-level model features, and capturing the commonalities and differences of diseases within the provincial scope (such as the correlation between diabetes incidence and climate factors in a certain province).

[0010] For the overall model (national level), a multi-granularity model aggregator was created, and a dynamic capsule routing algorithm was used to intelligently aggregate the parameters of the three-level models. This algorithm is implemented through the following steps: ① Encoding the characteristics of the city-level sub-models into basic capsules, and encoding the provincial knowledge graph into advanced capsules; ② Iteratively calculating the coupling coefficients between capsules through dynamic routing, prioritizing the activation of capsule nodes that match the regional characteristics of the input case (for example, when processing cases in Tibet, capsules related to altitude sickness are prioritized); ③ Capsule vector weighted aggregation is used to generate national-level diagnostic decisions, balancing regional specificity with national applicability.

[0011] Furthermore, a multimodal data integration module application is provided to integrate unstructured data such as electronic medical record text data, medical images (such as CT, MRI), and test reports, and use natural language processing (NLP), image recognition and other technologies to achieve unified representation and feature extraction of multimodal data, providing a rich input source for model training.

[0012] Furthermore, a listed APP or non-listed system application is provided to provide patients with functions such as intelligent medical consultation, health management, and medical recommendations; a personalized learning platform is provided for medical staff and students, including a case library, diagnosis and treatment guidelines, and skill training modules, supporting hierarchical learning based on sub-models and overall models.

[0013] Furthermore, an offline training AI model application is provided, which is independently controlled by hospitals or schools, supports importing local case data and medical literature databases for offline training, ensures data security and privacy protection, and meets the personalized needs of grass-roots institutions.

[0014] Furthermore, an online model application is provided. By using federated learning technology, it integrates the data of each offline model to form regional and national training pools, realizing cross-institutional data collaboration and model iteration.

[0015] Furthermore, a data transfer station application is provided. A data update review mechanism is established to manage the version iteration and data synchronization between the online model and the APP / system and the hospital internal system, ensuring the security and efficiency of data flow.

[0016] The present invention provides a hierarchical training method for a multi-modal medical AI collaborative system. The specific method is as follows:

[0017] The first-level sub-consortium training includes a data input layer, which contains multi-dimensional data such as patient basic information (age, gender, BMI), medical history (previous surgical history, allergy history), diagnosis results (ICD-10 coding), and geographical features (latitude and longitude, altitude, average annual temperature and humidity). Among them, the geographical features automatically capture meteorological and demographic data of the region where the medical institution is located through the API.

[0018] Furthermore, it includes a federated learning layer, which adopts an asynchronous update mechanism. After each medical institution completes the training of 50 cases, it encrypts and uploads the model gradient (gradient compression rate ≥ 70%). The municipal coordination node aggregates the gradients according to the dynamic geographical weight matrix W, and the weight ratio of neighboring institutions (distance < 20 km) is ≥ 60%, ensuring the priority learning of regional disease characteristics.

[0019] The second-level sub-consortium training collects the training parameters and feature vectors of each first-level sub-consortium, analyzes the feature differences and correlations of different municipal models through a bidirectional attention mechanism, and extracts common features (such as the regional high-incidence disease pattern) and individual features (such as the diagnosis and treatment differences in different cities) at the provincial level. A provincial knowledge graph is constructed, integrating disease classification, treatment guidelines, and epidemiological data to provide cross-domain knowledge support for the total model. The formula is: where f i , f j is the output feature of the municipal model, and d k is the feature dimension. Provincial common features (such as the proportion of patients with hypertension in a certain province with an average BMI ≥ 28 reaching 75%) and differential features (such as City A preferring to use ACEI drugs and City B more commonly using ARB drugs) are extracted based on the similarity matrix.

[0020] Further, it includes knowledge graph construction. Using the Neo4j graph database, a three-layer graph is constructed, including disease nodes, regional nodes (34 provincial-level administrative regions), and diagnosis and treatment plan nodes. The nodes are connected by relationships such as "high incidence in" and "applicable to" to support complex queries.

[0021] Total model training: Using the dynamic capsule routing algorithm, the knowledge graph of the secondary sub-units and the model parameters of each province are fused at multiple granularities. Through the dynamic routing of capsule nodes, feature-level aggregation from the city level to the provincial level and then to the national level is achieved, generating a total model adapted to different regions across the country. The total model supports diagnosis decision-making for patients (ToC) and collaborative training for medical institutions (ToB), forming a virtuous cycle of "sub-model - sub-unit - total model".

[0022] Further, it includes defining a three-layer structure of basic capsules (city-level model features), advanced capsules (provincial-level knowledge graph entities), and output capsules (disease diagnosis results). By iterating the routing algorithm 3 times, the coupling coefficient C between capsules is calculated ij , and the formula is: where b ik is the initial routing score. The final output capsule vector u j = ∑ i c ij ·u j∣i realizes feature-level aggregation from regional specificity to national generality.

[0023] The technical effects and advantages of the present invention:

[0024] Through dynamic geographical weights and hierarchical training, the learning of regional disease characteristics is strengthened, enabling the model to provide more accurate diagnostic suggestions for the disease spectrum, treatment habits, and prognosis effects in different regions, and distinguishing the prevalent diseases and personalized treatment plans in different regions.

[0025] Provide medical treatment recommendations and prognosis evaluations based on regions and hospital levels for patients to avoid blind referrals; provide advanced diagnostic models across institutions for medical staff to learn (for example, the neurosurgery model of Beijing Tiantan Hospital can be learned by doctors in provincial hospitals), supporting downward-compatible skill improvement; provide real-time updated clinical knowledge bases and virtual teaching platforms for medical students and teachers to solve the problems of theoretical lag and uneven resources in traditional education.

[0026] The three - level federated learning framework realizes hierarchical aggregation of data from the municipal level to the national level. Under the premise of protecting data privacy (such as local training of offline models and data transfer station audit mechanisms), it maximally utilizes the regional diversity of medical data, reduces the cost of repeated training, and improves the generalization ability and iteration speed of the overall model. By incorporating the specialized models of hospitals (such as the complex surgery model of Xiehe Hospital) into the teaching system, students can learn the diagnosis and treatment logic of different levels of hospitals in a virtual environment, quickly adapt to the work requirements of the target institution after graduation, shorten the training cycle, and alleviate the shortage of grass - roots medical talents. Brief Description of the Drawings

[0027] Figure 1 , A schematic diagram of the three - level nested federated learning framework, showing the hierarchical relationship and data flow at the municipal level (the first - level sub - consortium layer), provincial level (the second - level sub - consortium layer), and national level (the total model layer);

[0028] Figure 2 , An architecture diagram of the multi - modal data integration module, explaining the processing flow of inspection data;

[0029] Figure 3 , The first - level sub - consortium architecture, the second - level sub - consortium architecture, and the total model architecture, explaining the data processing flow. Detailed Implementation Manner

[0030] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present application.

[0031] Embodiment 1

[0032] The present invention discloses a multi - modal medical AI collaborative system and its hierarchical training method, including the first - level sub - consortium training process, second - level sub - consortium knowledge distillation, and total model dynamic routing aggregation.

[0033] As Figure 1, the first-level sub-consortium training process shown in 103 includes dynamic geographic weight calculation. The straight-line distance d is calculated based on the coordinates of medical institutions, and combined with the regional disease correlation coefficient α (set by the local health commission or medical association according to historical data. For example, α takes 0.8 in high-incidence infectious disease areas and 0.5 in ordinary areas), a weight matrix W is generated. It also includes federated learning iteration. Medical institutions within the municipal sub-consortium regularly upload model gradients or parameters, and perform weighted aggregation based on the weight W, giving priority to fusing the updates of neighboring institutions to form a municipal sub-model. For example, when three hospitals in a certain city train a diabetes sub-model, the weights of the relatively close A hospital and B hospital are 0.6 and 0.5, and the weight of the relatively far C hospital is 0.3, improving the consistency of diabetes typing diagnosis within the region.

[0034] Such as Figure 1 , the second-level sub-consortium knowledge distillation shown in 102 includes feature mapping extraction. A bidirectional attention mechanism is used to analyze the output features of different municipal sub-models (such as disease probability distribution, treatment plan vector), and identify common features across cities (such as common complications of hypertension within a certain province) and differential features (such as drug preferences for hypertension in different cities); It also includes the construction of a provincial knowledge graph, which fuses the extracted features with medical knowledge bases (such as UpToDate, clinical guidelines) to construct a knowledge graph containing disease-symptom-treatment-region associations, providing cross-domain knowledge support for the overall model.

[0035] Such as Figure 1 , the overall model dynamic routing aggregation shown in 101 includes capsule node initialization. The knowledge graph of the second-level sub-consortium is transformed into capsule nodes, and each node represents provincial disease characteristics (such as a high-incidence node of stroke in the Northeast region, a dengue fever warning node in the South China region); It also includes a dynamic routing algorithm. By iteratively calculating the coupling coefficient between capsules, the connection weights between provincial nodes and national-level nodes are dynamically adjusted to achieve multi-granularity feature aggregation. For example, when processing cases from Guangdong Province, the overall model preferentially activates the dengue fever node in the South China region, and combines the common features of the national stroke node to generate personalized diagnostic suggestions.

[0036] Example 2

[0037] Based on Example 1, in this example, the training method further includes the application of derivative entities and data security and ethics.

[0038] Such as Figure 2As shown, the derivative entity applications include medical education scenarios. Students at a medical college can select models through an APP for learning. The system provides the diagnostic logic, typical case libraries, and virtual diagnosis and treatment training of the models. Students can simulate learning in an offline environment. When they graduate and enter a provincial hospital, they only need to supplement the differential training of the local sub-model (such as the treatment of endemic diseases common in the province) to quickly start working. It also includes clinical diagnosis scenarios. A doctor in a primary hospital receives a patient with chest pain and initially judges it to be angina by calling the local sub-model. Synchronously referring to the high incidence characteristics of coronary heart disease in the provincial sub-consortium and the national diagnosis and treatment guidelines of the total model, a treatment plan with regional characteristics is finally formulated.

[0039] Data security and ethics include privacy protection. The differential privacy technology of federated learning is adopted to add noise during data uploading and aggregation to ensure that the data of medical institutions is not leaked. The offline model is stored locally by the hospital, and sensitive data is not transmitted externally. It also includes ethical review. An ethical committee for the model is established to regularly review the diagnostic suggestions of the total model and the training data of the sub-models to ensure that the analysis of regional differences does not involve discriminatory features and guarantee medical fairness.

[0040] Example 3

[0041] As Figure 3 shown, the training of the first-level sub-consortium includes data preparation. Five municipal hospitals are included, and a total of 200,000 diabetes case data are uploaded, including fields such as blood glucose indicators, complication records, and medication plans. The regional characteristic data includes an average altitude of 50 meters, an average annual humidity of 75%, and a diabetes incidence rate of 8.5% (higher than the national average).

[0042] Furthermore, it includes the federated learning process. Each hospital locally trains a diabetes sub-model, adopting a ResNet+LSTM hybrid architecture. The input includes data such as patient age, BMI, and glycated hemoglobin. After each round of training, the model gradients are uploaded. The municipal coordination node aggregates them according to W = 1 / (1 + 0.6×d). The weights of the closest A hospital and B hospital (5 kilometers apart) are 0.4 and 0.35 respectively, and the weight of the farther C hospital (30 kilometers apart) is 0.15. After 10 rounds of iteration, a municipal diabetes sub-model is formed, and the diagnostic accuracy rate for local gestational diabetes reaches 92%, which is 18% higher than that of the single-hospital model.

[0043] Furthermore, the communication rounds of federated learning are set to 50 - 100 rounds and are dynamically adjusted according to the data volume. The gradient compression algorithm adopts Top-K sparsification to retain the top 20% of the important gradients.

[0044] The secondary sub-concatenation knowledge distillation includes feature extraction, collecting hypertension feature vectors of 12 municipal sub-models in the province, and identifying cross-city difference features such as "high diuretic usage rate in high-salt diet areas" and "ACEI drug resistance rate in cold areas reaches 15%" through a two-way attention mechanism; constructing a provincial hypertension knowledge graph, which includes 10 common complications, 8 types of first-line drugs, and 5 regional-related risk factors (such as a 20% increase in the incidence of hypertension in areas with an altitude of >1000 meters).

[0045] Furthermore, the training batch size of the bidirectional attention mechanism is 64, and the learning rate uses the Adam optimizer with an initial value set to 1e-4.

[0046] The overall model application includes inputting patient data: a 65-year-old male living in Guangdong Province with sudden numbness of the right limbs, and CT showing low-density lesions in the basal ganglia; dynamic capsule routing activates dengue fever-related capsules in the South China region (to rule out neurological symptoms caused by dengue fever) and the national stroke core capsule (to extract basal ganglia lesion characteristics); outputs a diagnosis of acute ischemic stroke (95% confidence level), and recommends a treatment plan: based on the rt-PA thrombolysis recommended by the national guidelines and combined with the Guangdong Province stroke quality control standards, a Cantonese version of the rehabilitation guidance video is added.

[0047] The multimodal medical AI collaborative system and its hierarchical training method provided by this invention can be directly deployed in areas with relatively well-developed medical information infrastructure, and can be quickly connected to hospital HIS systems, medical imaging platforms, and education management systems through API interfaces. The system hardware requirements are compatible with mainstream server clusters (such as Alibaba Cloud ECS clusters), and the software level supports Docker containerized deployment. It has high scalability and compatibility, can effectively improve the regional adaptability and multi-scenario application capabilities of medical AI, and has significant industrial application value.

[0048] In summary, this invention uses a three-level nested federated learning framework and multimodal data integration to build a medical AI collaborative system that adapts to regional differences and supports multi-group applications, improves data collaboration and model integration, and provides innovative solutions for improving medical diagnosis accuracy, promoting the development of medical education, and optimizing medical resource allocation.

[0049] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and substitutions can be made without departing from the technical principles of the present invention. These improvements and substitutions should also be regarded as the scope of protection of the present invention.

Claims

1. A hierarchical training method for a multi-modal medical AI collaborative system, characterized in that, It includes the following steps: Construct a three-level nested federated learning framework, including a first-level sub-alliance, a second-level sub-alliance and a total model, to achieve regionalized adaptive training and multi-granularity model fusion; First-level sub-alliance training: within the city-level sub-alliance, adopt an adaptive federated learning algorithm based on dynamic geographical weights, and strengthen the data interaction of geographically adjacent medical institutions by setting a geographical location attenuation factor. The calculation formula of the geographical location attenuation factor is W = 1 / (1 + α×d), where d is the straight-line distance between medical institutions, and α is the regional disease correlation coefficient. The regional disease correlation coefficient is automatically calibrated through a linear regression model based on data such as the referral rate and disease co-occurrence frequency among city-level medical institutions in the past 3 years; Each medical institution adopts an asynchronous update mechanism, and encrypts and uploads the model gradient after every 50 cases of training are completed. The city-level coordination node aggregates the gradients according to the dynamic geographical weight matrix W, and the weight ratio of adjacent institutions is ≥60%. Second-level sub-alliance training: Collect the training parameters and feature vectors of each first-level sub-alliance, analyze the feature differences and correlations of different city-level models through a bidirectional attention mechanism, and extract common features and individual features at the provincial level; Construct a provincial knowledge graph, integrating disease classification, diagnosis and treatment guidelines, and epidemiological data. The knowledge graph uses a Neo4j graph database, including disease nodes, geographical nodes, and diagnosis and treatment plan nodes, and the nodes are connected by "high incidence in" and "applicable to" relationships. Total model training: Adopt a dynamic capsule routing algorithm to perform multi-granularity fusion of the knowledge graph of the second-level sub-alliance and the parameters of each provincial model. The dynamic capsule routing algorithm includes: encoding the features of the city-level sub-model as basic capsules and encoding the provincial knowledge graph as advanced capsules; iteratively calculating the coupling coefficient between capsules through dynamic routing, and preferentially activating the capsule nodes that match the geographical features of the input cases; generating a national-level diagnostic decision through weighted aggregation of capsule vectors, taking into account regional specificity and national generality.

2. The method according to claim 1, wherein The data input layer of the first-level sub-alliance training includes multi-dimensional data such as patient basic information, medical history, diagnosis results, and geographical features. Among them, the geographical features automatically capture meteorological and demographic data of the area where the medical institution is located through an API.

3. The method according to claim 1, wherein The bidirectional attention mechanism includes self-attention encoding of the output features of the city-level model and mutual attention alignment of cross-city features, and extracts the feature mapping relationship by calculating the feature similarity matrix between different cities to achieve i , where f j , f k are the output features of the city-level model, and d k is the feature dimension.

4. The method according to claim 1, wherein The dynamic capsule routing algorithm defines a three-layer structure of basic capsules, advanced capsules, and output capsules, and calculates the coupling coefficient between capsules through three iterations of the routing algorithm. Among them, b ik is the initial routing score, and the final output capsule vector u j = ∑ i c ij ·u j∣i realizes feature-level aggregation.

5. The method according to claim 1, characterized in that, It also includes a multi-modal data integration module, which integrates unstructured data such as electronic medical record text data, medical images, and test reports, and uses technologies such as natural language processing and image recognition to achieve unified representation and feature extraction of multi-modal data.

6. The method according to claim 5, wherein The medical image processing includes DICOM parsing, image normalization, lesion localization and feature extraction, and uses algorithms such as U-Net, ResNet or ViT for lesion segmentation and deep learning feature extraction.

7. The method according to claim 1, characterized in that, It also includes functions such as intelligent consultation, health management, and medical treatment recommendation for patients, as well as a personalized learning platform for medical staff and students. The learning platform includes a case library, diagnosis and treatment guidelines, and a skill training module, and supports hierarchical learning based on sub-models and the total model.

8. The method according to claim 1, characterized in that, It also includes the application of an offline training AI model, which is independently controlled by hospitals or schools, supports importing local case data and medical literature databases for offline training, and ensures data security and privacy protection.

9. The method according to claim 1, characterized in that, It also includes an online model application that integrates the data of each offline model through federated learning technology to form regional and national training pools, realizing cross-institutional data collaboration and model iteration, as well as a data transfer station application that establishes a data update review mechanism to manage the version iteration and data synchronization between the online model and the APP / system and the hospital internal system.

10. The method according to claim 1, characterized in that, The gradient compression technology is adopted in the federated learning process, and the gradient compression rate is ≥70%, and noise is added in combination with differential privacy technology to ensure the privacy and security of medical institution data.