Ophthalmology remote intelligent consultation system and method based on multi-modal data fusion

Through the multimodal data fusion and knowledge graph update remote intelligent ophthalmology consultation system, the problem of incomplete information and inaccurate expert matching in traditional consultations is solved, accurate diagnosis and personalized treatment are achieved, and the intelligence and precision level of ophthalmology medical care is improved.

CN120260980AActive Publication Date: 2025-07-04NORTHWEST WOMEN & CHILDREN HOSPITAL

Patent Information

Application Number
CN202510653842.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-21
Publication Date
2025-07-04
Estimated Expiration
2045-05-21

AI Technical Summary

Technical Problem

In traditional ophthalmic consultations, information transmission is incomplete or inaccurate, expert resources are unevenly distributed, and the existing intelligent diagnostic system cannot effectively integrate multimodal data, resulting in insufficient comprehensive and accurate diagnosis and inaccurate expert matching, which affects the effectiveness of the consultation.

Method used

The remote intelligent consultation system of ophthalmology based on multimodal data fusion generates a multimodal feature set by collecting images, text and detection data, using the improved cross-modal attention model and graph neural network for deep fusion analysis, combining local knowledge graphs for symptom-disease association reasoning, quantifying expert information and calculating matching degrees, generating electronic consultation sheets, and updating the knowledge graph through federated learning.

Benefits of technology

It has achieved comprehensive exploration and accurate diagnosis of the patient's condition, ensured that the patient received the most suitable expert consultation, improved diagnostic accuracy and personalized treatment suggestions, and promoted knowledge sharing and improved medical level.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120260980A_ABST
    Figure CN120260980A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of ophthalmology medical intelligence, and provides an ophthalmology remote intelligent consultation system and method based on multi-modal data fusion, and the method comprises the steps: collecting and integrating ophthalmology multi-modal data, generating a multi-modal feature set, and inputting an improved cross-modal attention model, generating a final fusion feature through dynamic weight fusion and a two-way interaction mechanism, outputting the ophthalmic disease probability distribution of the patient in combination with a local ophthalmic knowledge graph, quantifying detailed information of ophthalmic experts as expert feature vectors, calculating the matching degree between the ophthalmic disease probability distribution of the patient and the expert feature vectors, and determining the ophthalmic disease probability distribution of the patient. The method comprises the following steps of: screening high-matching-degree ophthalmology experts for consultation, calculating recommended values of treatment schemes provided by the ophthalmology experts and generating an electronic consultation sheet, supplementing local case data to train a local knowledge graph, and generating a global knowledge graph and performing feedback updating on the local knowledge graph by a central server through weighting and aggregating the local knowledge graphs of multiple mechanisms.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of intelligent ophthalmic medical technology, and specifically relates to an ophthalmic remote intelligent consultation system and method based on multi-modal data fusion. Background Art

[0002] In the traditional ophthalmic medical consultation process, patients usually need to go to specific medical institutions, where local ophthalmologists conduct preliminary examinations and diagnoses. In case of difficult diseases, local doctors will transfer the patient's condition information to superior or more professional ophthalmic experts for consultation through relatively traditional methods such as written reports and image film transmission. This mode has many significant drawbacks. Written reports may not comprehensively and accurately describe key information such as the patient's symptoms and medical history. Image films are also prone to wear and loss during transmission, resulting in incomplete or inaccurate information received by experts, thus affecting the accuracy of diagnosis. Patients need to travel long distances to the consultation location, consuming a large amount of time and energy. Especially for patients with severe conditions and limited mobility, this undoubtedly increases the difficulty and pain of seeking medical treatment. At the same time, the distribution of expert resources is uneven, and it is difficult for patients in some areas to obtain the consultation opinions of authoritative experts in a short time.

[0003] With the rapid development of information technology, some ophthalmic intelligent diagnosis systems have gradually emerged. However, these systems still have obvious technical defects in dealing with complex ophthalmic disease consultations. In the prior art, the diagnosis of ophthalmic diseases often requires integrating information of multiple modalities such as image data, text data, and detection data. However, most existing systems can only process single or a few modalities of data and cannot effectively integrate and deeply analyze multi-modal data, resulting in incomplete and inaccurate disease judgments.

[0004] On the other hand, intelligent diagnosis systems in the prior art mainly rely on preset rules or simple machine learning models for disease diagnosis, lacking in-depth understanding and reasoning ability of ophthalmic medical knowledge. Therefore, they cannot accurately judge the disease type and development trend, nor can they provide personalized treatment plans for patients. In the case of requiring expert consultation, the prior art often simply assigns patients to a certain expert without fully considering the matching degree between the patient's condition and the expert's professional field, experience, etc. This may lead to patients not being able to get the most suitable expert for consultation, affecting the consultation effect and treatment effect.

[0005] In view of the above problems, the present invention proposes an ophthalmic remote intelligent consultation system and method based on multi-modal data fusion. Summary of the Invention

[0006] In order to make up for the deficiencies of the prior art and solve at least one of the technical problems proposed in the background art.

[0007] The technical solution adopted by the present invention to solve the technical problem is: an ophthalmic remote intelligent consultation system based on multi-modal data fusion, including: Data acquisition module: collect and integrate ophthalmic multi-modal data to generate a multi-modal feature set; Fusion diagnosis module: input the multi-modal feature set into an improved cross-modal attention model, generate final fusion features through a dynamic weight fusion and two-way interaction mechanism, perform graph neural network reasoning in combination with a local ophthalmic knowledge graph, and output the probability distribution of the patient's ophthalmic diseases; Consultation matching module: quantify the detailed information of ophthalmic experts into expert feature vectors, calculate the matching degree between the probability distribution of the patient's ophthalmic diseases and the expert feature vectors, screen ophthalmic experts with high matching degrees for consultation, calculate the recommendation value of the treatment plans provided by the ophthalmic experts, and generate an electronic consultation form containing a diagnosis conclusion and treatment suggestions based on the recommendation value; Iterative update module: desensitize the electronic consultation form and supplement local case data to train the local knowledge graph. The central server generates a global knowledge graph by weighted aggregation of multi-institutional local knowledge graphs, and performs feedback update on the local knowledge graph according to the generated global knowledge graph.

[0008] Further preferably, the generation method of the multi-modal feature set is: Collect and preprocess the ophthalmic multi-modal data of the patient to be treated. The ophthalmic multi-modal data includes image data, text data and detection data. Perform lesion segmentation and geometric feature extraction on the preprocessed image data to obtain an image feature vector, perform semantic feature extraction on the text data to obtain a text feature vector, and perform structured feature mapping on the detection data to obtain a detection feature vector; Integrate the three modalities of the image feature vector, text feature vector and detection feature vector to generate a multi-modal feature set in a unified format.

[0009] Further preferably, the acquisition method of the probability distribution of the patient's ophthalmic diseases is: Based on clinical guidelines and local case data, construct a directed local ophthalmic knowledge graph containing three types of core entities. Among them, the three types of core entities include disease entities, symptom entities and detection entities. Obtain the final fusion feature as the initial feature vector of the disease entity node feature vector. Design a multi-modal attention aggregation function to iteratively update the node feature vector starting from the initial feature vector, and apply a fully connected layer and a Softmax function to the iteratively updated node feature vector to obtain the probability distribution of the patient's ophthalmic diseases.

[0010] Further preferably, the acquisition method of the final fusion feature is: Obtain the weight matrix. According to the weight matrix, design a bidirectional fusion mechanism. In the multi-modal feature set, each modal feature serves as both a query to extract other modal features and a key to provide its own features. Combine the weight matrix to calculate the intra-modal fusion and inter-modal fusion. The intra-modal fusion retains the core features of each modality itself, and the inter-modal fusion captures cross-modal associations; Combine the intra-modal fusion and inter-modal fusion, and through layer normalization, stably train to calculate the final fused features.

[0011] Further preferably, the way to obtain the weight matrix is as follows: Through a fully connected layer, map the image feature vector, text feature vector, and detection feature vector in the multi-modal feature set to a shared latent space, convert the multi-modal feature set into a shared latent space representation of the same dimension, define a tri-modal interaction matrix to represent the similarity between modalities, calculate the attention weights between any two modal features, and integrate to obtain the weight matrix.

[0012] Further preferably, the way to obtain the electronic consultation form is as follows: In a remote consultation, the consulting ophthalmologists who respond to the request respectively provide treatment plans. Use the Jaccard similarity to calculate the similarity between any two treatment plans, classify the two treatment plans with a similarity greater than or equal to the similarity threshold into the same similar plan set, normalize the number of treatment plans in each similar plan set, and mark it as the consensus eigenvalue of the treatment plans in the similar plan set; Based on any treatment plan, obtain the matching degree of the ophthalmologist who provides the treatment plan as the weight of the treatment plan. Combine the consensus eigenvalue of the treatment plan, and calculate the recommended value of the treatment plan by weighted calculation. Sort the treatment plans in descending order according to the recommended value of the plan, and fill the sorted treatment plans and the corresponding diagnosis results into the electronic consultation form to generate the electronic consultation form.

[0013] Further preferably, the matching method of the consulting ophthalmologists who respond to the request is as follows: Obtain the matching degrees of all ophthalmologists in the platform, sort the ophthalmologists in descending order according to the matching degrees, and mark the selection serial numbers for the ophthalmologists according to the order to obtain the sorted ophthalmologist list. Select the ophthalmologists whose selection serial numbers in the ophthalmologist list are less than or equal to the selection threshold to form a set of consulting ophthalmologists; Send a consultation request to the ophthalmologists in the set of consulting ophthalmologists to match the consulting ophthalmologists who respond to the request.

[0014] Further preferably, the way to obtain the matching degree is as follows: Collect the detailed information of all ophthalmology experts in the platform, extract and quantify their features to obtain expert feature vectors, obtain the probability distribution of the ophthalmic diseases of the patients to be treated, construct the probability distribution of the ophthalmic diseases of the patients as the disease feature vectors of the patients to be treated, and perform data processing on the disease feature vectors of the patients to be treated and the expert feature vectors of each ophthalmology expert to obtain the matching degree between the probability distribution of the ophthalmic diseases of the patients and the expert feature vectors.

[0015] Further preferably, the way to perform feedback update on the local knowledge graph is as follows: Obtain the ophthalmic multimodal data of the patients to be treated and the generated electronic consultation form and perform desensitization processing to generate desensitized ophthalmic treatment data, continuously record the desensitized ophthalmic treatment data, generate a desensitized data set and add it to the local case data. Each medical institution performs incremental training on the local knowledge graph based on the local case data to optimize the symptom-disease association strength. After the training is completed, extract the graph structure parameters and node feature parameters to update the local knowledge graph; The central server uses the FedAvg algorithm to aggregate the parameters of the local knowledge graphs of multiple institutions, calculates the aggregation weights according to the local case data volume of each medical institution and performs parameter fusion to obtain the global edge weight matrix and global node embedding vector of the global knowledge graph, and distributes them to each medical institution to replace the local knowledge graph parameters for iterative feedback update of the local knowledge graph.

[0016] An ophthalmic remote intelligent consultation method based on multimodal data fusion includes the following steps: Collect and integrate ophthalmic multimodal data to generate a multimodal feature set; Input the multimodal feature set into an improved cross-modal attention model, generate the final fusion feature through dynamic weight fusion and bidirectional interaction mechanism, and perform graph neural network reasoning in combination with the local ophthalmic knowledge graph to output the probability distribution of the ophthalmic diseases of the patients; Quantify the detailed information of ophthalmology experts as expert feature vectors, calculate the matching degree between the probability distribution of the ophthalmic diseases of the patients and the expert feature vectors, screen out ophthalmology experts with high matching degrees for consultation, calculate the recommendation value of the treatment plans provided by the ophthalmology experts, and generate an electronic consultation form containing diagnostic conclusions and treatment suggestions based on the recommendation value; Perform desensitization processing on the electronic consultation form and supplement the local case data to train the local knowledge graph. The central server aggregates the local knowledge graphs of multiple institutions through weighting to generate a global knowledge graph, and performs feedback update on the local knowledge graph according to the generated global knowledge graph.

[0017] The beneficial effects of the present invention are as follows: 1. Through collecting and integrating multi-modal ophthalmology data and conducting in-depth fusion analysis, the present invention can comprehensively explore the patient's condition information, perform symptom-disease association reasoning in combination with the local ophthalmology knowledge graph, accurately generate the disease probability distribution, greatly improving the accuracy of ophthalmology disease diagnosis. At the same time, it quantifies expert information and calculates the matching degree with the patient's condition, accurately screening experts for consultation with high matching degrees to ensure that patients receive the most suitable diagnosis and treatment suggestions, effectively avoiding the problem of inaccurate expert matching in traditional consultations and providing patients with higher-quality and personalized medical services.

[0018] 2. By desensitizing the electronic consultation form and supplementing local case data, the present invention trains and updates the local knowledge graph based on incremental data, enabling the knowledge graph to promptly reflect the latest clinical experience and medical achievements, continuously enhancing the diagnostic and treatment capabilities. The central server uses the FedAvg algorithm to aggregate the local knowledge graphs of multiple institutions to generate a global knowledge graph, breaking down the data barriers between institutions, realizing knowledge sharing and collaboration, promoting the common improvement of the medical levels of each institution, and driving the overall progress of the ophthalmology medical field towards intelligence and precision. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The present invention will be further described below in conjunction with the drawings.

[0020] Figure 1 is the system module architecture diagram of the ophthalmology remote intelligent consultation system based on multi-modal data fusion according to the embodiment of the present invention; Figure 2 is the step flow chart of the ophthalmology remote intelligent consultation method based on multi-modal data fusion according to the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0021] In order to make the technical means, creative features, achieved purposes, and effects of the present invention easy to understand, the present invention will be further described below in conjunction with specific embodiments.

[0022] Embodiment 1 Please refer to Figure 1 As shown, the ophthalmology remote intelligent consultation system based on multi-modal data fusion according to the embodiment of the present invention includes the following modules: Data acquisition module: Collect and integrate multi-modal ophthalmology data to generate a multi-modal feature set; Collect the multi-modal ophthalmology data of the patient to be treated. The multi-modal ophthalmology data includes three modalities, namely image data, text data, and detection data. Among them, the image data includes fundus OCT images and fundus color photos, the text data includes medical record texts, and the detection data includes visual field detection data and intraocular pressure detection data; Preprocess the collected multi-modal ophthalmology data to obtain the preprocessed multi-modal ophthalmology data; Specifically, for spatial alignment of image data, for fundus OCT images and fundus color photos, a rigid registration algorithm based on mutual information is used. First, the SIFT algorithm is used to extract corresponding feature points in the fundus OCT image and the fundus color photo and match them into matching point pairs. The RANSAC algorithm is used to remove mismatched point pairs, and the rotation matrix and translation vector are calculated by the least squares method to achieve geometric alignment of the fundus OCT image and the fundus color photo in the global retinal coordinate system. For images with inconsistent resolutions, bicubic interpolation is used to uniformly scale them to 512×512 pixels, preserving edge details. When the resolution is insufficient, zero-padding and Gaussian smoothing are used to avoid artifacts; After spatial alignment of the image data, time alignment is performed on the ophthalmic multimodal data of the patient to be treated. Based on the visit time of the patient to be treated, ophthalmic multimodal data with a detection time interval less than or equal to the time difference threshold from the visit time is screened to ensure that the ophthalmic multimodal data reflects the same disease stage. Format standardization, noise filtering, and outlier processing are respectively performed on the image data, text data, and detection data included in the ophthalmic multimodal data to obtain preprocessed ophthalmic multimodal data; Feature extraction is performed on the preprocessed ophthalmic multimodal data to construct a multimodal feature set; Specifically, for lesion segmentation and geometric feature extraction of image data, the U-Net++ network is trained using a public ophthalmic dataset combined with in-hospital annotated data. The preprocessed image data is input into the U-Net++ network for pixel-level segmentation. The network structure includes 4 layers of encoders, a bottleneck layer, and 4 layers of decoders, and outputs a lesion probability segmentation mask of the same size as the input. Based on the lesion probability segmentation mask, core lesion indicators are calculated. The core lesion indicators include basic features, morphological features, gray-level features, interlayer features of the fundus OCT image, and color photo features; Among them, the basic features include lesion area, lesion perimeter, and lesion volume. The morphological features include lesion circularity and lesion eccentricity. The gray-level features include the average gray value of the lesion, the gray standard deviation of the lesion, and the contrast between the lesion and the surrounding normal tissue. The interlayer features of the fundus OCT image include the thickness of the nerve fiber layer and the integrity score of the outer nuclear layer. The color photo features include lesion vascular density and the number of lesion exudates. The core lesion indicators are integrated to generate an image feature vector V; Semantic feature extraction is performed on the text data. The medical record text is input into the clinical BERT model pre-trained on the MIMIC-III medical corpus, and word-level embeddings are generated through multiple layers of Transformer encoders. The text feature vector T is obtained through average pooling; Perform structured feature mapping on the detection data, normalize the visual field detection data and convert it into a feature vector. Take the average of the intraocular pressure values of both eyes in the intraocular pressure detection data and perform interval mapping, mapping the normal range to [0, 0.7], and linearly extending the abnormal values to [0, 1]. Integrate to obtain the detection feature vector D; Integrate the three modalities of the image feature vector V, the text feature vector T, and the detection feature vector D to generate a unified format multi-modal feature set {V, T, D}; It should be noted that the function of this module is to achieve the geometric alignment of OCT and color photos through the SIFT-RANSAC algorithm, solve the problem of image space misalignment, ensure the cross-modal consistency of the lesion location, design the core indicators of the lesion for ophthalmic data, and combine the U-Net++ and BERT models to achieve cross-modal feature extraction from pixel-level segmentation to semantic-level understanding, and retain the key clinical information; Fusion diagnosis module: Input the multi-modal feature set into the improved cross-modal attention model, generate the final fusion feature through the dynamic weight fusion and two-way interaction mechanism, and perform graph neural network reasoning in combination with the local ophthalmic knowledge graph to output the probability distribution of the patient's ophthalmic diseases; Specifically, map the image feature vector V, the text feature vector T, and the detection feature vector D in the multi-modal feature set to the shared hidden space through the fully connected layer, and convert the multi-modal feature set into a shared hidden space representation of the same dimension ; Among them, 、 、 respectively represent the shared hidden space representations of the same dimension obtained after the image feature vector V, the text feature vector T, and the detection feature vector D are mapped to the shared hidden space through the fully connected layer, and respectively represent the feature vectors of the image feature vector V, the text feature vector T, and the detection feature vector in the shared hidden space; Define the three-modal interaction matrix S, calculate the attention weights between any two modal features, and reflect the feature dependence relationship;

[0023] Among them, is the element in the i-th row and j-th column of the three-modal interaction matrix S, representing the similarity of the i-th modal feature to the j-th modal feature, i, j ∈ {1, 2, 3}, where the first modal feature corresponds to the image feature vector, the second modal feature corresponds to the text feature vector, and the third modal feature corresponds to the detection feature vector, 、 、 respectively represent the similarities of the image feature vector to the three modal features, 、 、 respectively represent the similarity of the text feature vector to the three modal features, and and respectively represent the similarity of the detection feature vector to the three modal features; Perform normalization calculation on the similarity to obtain the weight matrix where represents the attention weight of the i-th modal feature to the j-th modal feature, and the formula is:

[0024] where a, b ∈ {1, 2, 3}, represents the element in the a-th row and b-th column of the three-modal interaction matrix S, representing the similarity of the a-th modal feature to the b-th modal feature, and d represents the dimension of the shared latent space, represents the matrix element in the i-th row and j-th column of the weight matrix ; The weight matrix satisfies:

[0025] Exemplarily, is the attention weight of the image feature to the text feature, representing the degree of attention of the image feature to the text feature, is the attention weight of the detection feature to the image feature, representing the degree of attention of the detection feature to the image feature; According to the weight matrix, design a bidirectional fusion mechanism, where each modal feature serves as both a query to extract other modal features and a key to provide its own features: Calculate the intra-modal fusion The formula is:

[0026] where represents the matrix element in the first row and first column, represents the matrix element in the second row and second column, is the matrix element in the third row and third column; Calculate the inter-modal fusion The formula is:

[0027] where the intra-modal fusion retains the core features of each modal feature, and the inter-modal fusion captures the cross-modal feature correlations; Combine the intra-modal fusion and the inter-modal fusion to calculate the final fusion feature The formula is:

[0028] Among them, LayerNorm represents layer normalization. By using layer normalization, the training is stabilized, and the final fusion features containing intra-modal characteristics and cross-modal interaction information are output. ; Construct a local ophthalmology knowledge graph. Based on clinical guidelines and local case data, construct a directed local ophthalmology knowledge graph containing three types of core entities. Among them, the three types of core entities include disease entities, symptom entities, and detection entities. Among them, the symptom entities include imaging features, text features, and detection features. Define the associations between the core entities, define the association between the symptom entity and the disease entity, and set the weight indicating the degree of support of the symptom entity for the disease entity, define the association between the detection entity and the disease entity, and set the weight indicating the correlation between the detection entity and the disease entity, define the association between the disease entity and the disease entity, and set the weight indicating the comorbidity probability; Store the local ophthalmology knowledge graph using triples {head entity, association, tail entity, weight}. The total number of graph nodes is about 500, and the number of edges is about 3,000. Store it through the Neo4j graph database to support efficient graph query and reasoning; Take the final fusion features as the initial features of the disease entity. The features of the symptom entity and the detection entity are predefined by the clinical knowledge base. For each node , the initial feature vector contains disease entities, symptom entities, and detection entities, where the disease entity corresponds to the final fusion features , the symptom entity corresponds to one-hot encoding, and the detection entity corresponds to the normalized detection index value. Among them, k represents the node number; Design a multi-modal attention aggregation function:

[0029] Iteratively update the node feature vectors to capture the association between symptoms and diseases. Among them, represents the feature vector of node m at the L + 1 layer, representing the information representation of node m after aggregation and update, represents the feature vector of neighbor node n at the L layer, represents the activation function, which is used to introduce non-linearity so that the model can learn complex patterns, represents all neighbor nodes n of node m, represents the weight reflecting the association strength between node m and neighbor node n, represents the modal weight coefficient, i ∈ {1, 2, 3}, corresponding to imaging features, text features, and detection features respectively. W represents the weight matrix for linearly transforming the neighbor node features ; Apply a fully connected layer and the Softmax function to the iteratively updated node feature vectors, and output the probability distribution P of the ophthalmic diseases of the patient to be treated;

[0030] Among them, represents the probability that the patient to be treated has the ophthalmic disease corresponding to the disease entity with serial number c, c = {1, 2,..., q}, and q represents the total number of disease entities in the local ophthalmic knowledge graph; It should be noted that the function of this module is to map multi-modal features to a shared latent space through a fully connected layer, calculate cross-modal dependence weights using a three-modal attention mechanism, generate final fusion features containing intra-modal characteristics and cross-modal interactions in combination with a bidirectional fusion mechanism, construct a local ophthalmic knowledge graph, iteratively update node features through a graph neural network, output the probability distribution P of ophthalmic diseases, and realize the correlation reasoning from symptoms to diseases; Consultation matching module: Quantify the detailed information of ophthalmic experts into expert feature vectors, calculate the matching degree between the probability distribution of the patient's ophthalmic diseases and the expert feature vectors, screen out ophthalmic experts with high matching degrees for consultation, calculate the recommendation value of the treatment plans provided by the ophthalmic experts, and generate an electronic consultation form containing a diagnosis conclusion and treatment suggestions based on the recommendation value; Obtain the probability distribution P of the ophthalmic diseases of the patient to be treated, and accurately match ophthalmic experts for the patient for remote consultation; Specifically, collect the detailed information of all ophthalmic experts on the platform, including professional fields, years of practice, types of ophthalmic diseases they are good at, and the number of historical consultation cases. Extract and quantify the detailed information of ophthalmic experts, and convert the detailed information of ophthalmic experts into computable expert feature vectors. For each ophthalmic expert, the corresponding expert feature vector consists of a professional field feature, a years-of-practice feature, and a historical consultation case feature: Among them, the professional field feature uses one-hot encoding to represent the types of diseases that the ophthalmic expert is good at. For the ophthalmic diseases corresponding to the q disease entities covered in the local ophthalmic knowledge graph, if the ophthalmic expert is good at the ophthalmic disease corresponding to the c-th disease entity, the c-th element of the feature vector is 1, and the rest are 0. The years-of-practice feature is obtained by normalizing the years of practice of the ophthalmic expert. The historical consultation case feature is obtained by calculating the proportion of the number of consultation cases of the ophthalmic expert in each ophthalmic disease to the total number of cases and the consultation success rate; Perform weighted fusion processing on the professional field feature, years-of-practice feature, and historical consultation case feature in the expert feature vector to obtain the expert feature vector of the ophthalmic expert ;

[0031] where f represents the serial number of the ophthalmic expert in the system, It represents the fusion feature value of an ophthalmologist on the ophthalmic diseases corresponding to the disease entity with serial number c, where c = {1, 2, ……, q}, and q represents the total number of disease entities in the local ophthalmic knowledge graph; Obtain the probability distribution P of the ophthalmic diseases of the patient to be treated, and construct the probability distribution P of the ophthalmic diseases as the disease feature vector of the patient to be treated , and the disease feature vector directly uses the disease probability distribution P. ; Use the cosine similarity to calculate the matching degree between the disease feature vector of the patient to be treated and the expert feature vector of each ophthalmologist. The matching degree between the patient to be treated and the ophthalmologist with serial number f The calculation formula is:

[0032] Among them, represents the dot product of the disease feature vector and the expert feature vector, and respectively represent the norms of the disease feature vector and the expert feature vector; Calculate the matching degrees of all ophthalmologists, sort the ophthalmologists in descending order according to the matching degrees, and mark the selection serial numbers for the ophthalmologists according to the order to obtain the sorted ophthalmologist list. Select the ophthalmologists whose selection serial numbers are less than or equal to the selection threshold in the ophthalmologist list to form a consultation ophthalmologist set; The system automatically sends a consultation request to the ophthalmologists in the consultation ophthalmologist set, including detailed information such as the basic information of the patient, the disease probability distribution, and the multi-modal features, and arranges a remote consultation for the ophthalmologists who respond to the request; In the remote consultation, the ophthalmologists who respond to the request respectively provide treatment plans. Based on any two treatment plans, use the Jaccard similarity to calculate the similarity between the two treatment plans. Specifically, obtain the ratio of the number of the same measures in the two treatment plans to the total number of measures in the two treatment plans to obtain the similarity between the two treatment plans; Classify the two treatment plans with a similarity greater than or equal to the similarity threshold into the same similar plan set, summarize and integrate all treatment plans to obtain several similar plan sets, and perform normalization processing on the number of treatment plans in each similar plan set, and mark it as the consensus feature value of the treatment plans in the similar plan set; Based on any treatment plan, use the matching degree of the ophthalmologist who provides the treatment plan as the weight of the treatment plan, and combine it with the consensus feature value of the plan to calculate the plan recommendation value of the treatment plan by weighted calculation; Sort the treatment plans in descending order according to the plan recommendation value, and fill the sorted treatment plans and the corresponding diagnosis results into the electronic consultation form to generate an electronic consultation form for the patient to be treated; It should be noted that the function of this module is to quantify expert information into a feature vector containing professional fields, years of practice, and case experience, calculate the matching degree between the disease characteristics of patients and expert characteristics through cosine similarity, and screen experts with high matching degrees to form a consultation set. Based on the Jaccard similarity, cluster treatment plans, calculate the recommendation value by combining expert weights and plan consensus, and generate an electronic consultation form containing diagnostic conclusions and treatment suggestions; Iterative update module: desensitize the electronic consultation form and supplement local case data to train the local knowledge graph. The central server generates a global knowledge graph by weighted aggregation of multi-institutional local knowledge graphs, and feeds back and updates the local knowledge graph according to the generated global knowledge graph; According to the generated electronic consultation form, treat the patients to be treated in combination with the actual situation, obtain the ophthalmic multimodal data of the patients to be treated and the generated electronic consultation form, remove the patient information, use homomorphic encryption technology to blur the coordinates of the lesion area to ensure pixel-level privacy protection, identify and replace sensitive information such as names and medical record numbers through natural language processing (NLP), retain symptom descriptions and diagnostic terms, generate desensitized ophthalmic treatment data, meet the requirements of privacy protection while retaining clinical characteristics, and provide secure input for federated learning; Continuously record the desensitized ophthalmic treatment data, generate a desensitized dataset and add it to the local case data. Each medical institution performs incremental training on the local knowledge graph based on the local case data, optimizes the symptom-disease association strength, and after the training is completed, extracts the graph structure parameters and node feature parameters to update the local knowledge graph; The central server uses the FedAvg algorithm to aggregate the parameters of multi-institutional local knowledge graphs, calculates the aggregation weights according to the local case data volume of each medical institution to ensure that institutions with rich data contribute more to the global model, performs parameter fusion in combination with the aggregation weights to obtain the global edge weight matrix and global node embedding vectors of the global knowledge graph, and distributes the global edge weight matrix and global node embedding vectors to the local knowledge graphs of each medical institution to replace the local knowledge graph parameters and iteratively update the local knowledge graph; It should be noted that the function of this step is to desensitize the electronic consultation form data, generate a secure dataset available for federated learning, each institution performs incremental training on the knowledge graph based on local case data, extracts graph structure and node feature parameters, and the central server aggregates multi-institutional parameters through the FedAvg algorithm, generates a global knowledge graph and distributes and updates it to achieve cross-institutional collaborative optimization; The technical solution of the embodiment of the present invention is as follows: collect and integrate ophthalmic multimodal data, generate a multimodal feature set through standardized processing and feature extraction, which includes image feature vectors, text feature vectors, and detection feature vectors. Deeply fuse the multimodal feature set, input it into the constructed local ophthalmic knowledge graph, perform association reasoning on symptoms and diseases for the multimodal feature set, generate the probability distribution of ophthalmic diseases of the patient to be treated, quantify the detailed information of ophthalmic experts into expert feature vectors, calculate the matching degree between the probability distribution of the patient's ophthalmic diseases and the expert feature vectors, screen ophthalmic experts with high matching degrees for consultation, calculate the recommendation value of the treatment plan provided by the ophthalmic experts, generate an electronic consultation form including a diagnosis conclusion and treatment suggestions based on the recommendation value, perform desensitization processing on the electronic consultation form and supplement the local case data to incrementally train the local knowledge graph. The central server weights and aggregates the local knowledge graphs of multiple institutions through the FedAvg algorithm to generate a global knowledge graph and updates the local knowledge graph.

[0033] Embodiment 2 As Figure 2 shown, the ophthalmic remote intelligent consultation method based on multimodal data fusion according to the embodiment of the present invention includes the following steps: Step 1: Collect and integrate ophthalmic multimodal data to generate a multimodal feature set; Step 2: Input the multimodal feature set into an improved cross-modal attention model, generate the final fusion feature through dynamic weight fusion and bidirectional interaction mechanism, and perform graph neural network reasoning in combination with the local ophthalmic knowledge graph to output the probability distribution of the patient's ophthalmic diseases; Step 3: Quantify the detailed information of ophthalmic experts into expert feature vectors, calculate the matching degree between the probability distribution of the patient's ophthalmic diseases and the expert feature vectors, screen ophthalmic experts with high matching degrees for consultation, calculate the recommendation value of the treatment plan provided by the ophthalmic experts, and generate an electronic consultation form including a diagnosis conclusion and treatment suggestions based on the recommendation value; Step 4: Perform desensitization processing on the electronic consultation form and supplement the local case data to train the local knowledge graph. The central server weights and aggregates the local knowledge graphs of multiple institutions to generate a global knowledge graph and feedback-update the local knowledge graph according to the generated global knowledge graph.

[0034] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification only illustrates the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.

Claims

1. An ophthalmic remote intelligent consultation system based on multimodal data fusion, characterized in that: Including: Data acquisition module: Collect and integrate multi-modal ophthalmology data to generate a multi-modal feature set; Fusion diagnosis module: Input the multi-modal feature set into an improved cross-modal attention model, generate final fusion features through a dynamic weight fusion and bidirectional interaction mechanism, perform graph neural network inference in combination with the local ophthalmology knowledge graph, and output the probability distribution of the patient's ophthalmic diseases; Consultation matching module: Quantify the detailed information of ophthalmology experts into expert feature vectors, calculate the matching degree between the probability distribution of the patient's ophthalmic diseases and the expert feature vectors, screen out ophthalmology experts with high matching degrees for consultation, calculate the recommendation value of the treatment plans provided by the ophthalmology experts, and generate an electronic consultation form containing diagnostic conclusions and treatment suggestions based on the recommendation value; Iterative update module: Desensitize the electronic consultation form and supplement local case data to train the local knowledge graph. The central server generates a global knowledge graph by weighted aggregation of multi-institutional local knowledge graphs, and performs feedback update on the local knowledge graph according to the generated global knowledge graph.

2. The ophthalmic remote intelligent consultation system based on multi-modal data fusion according to claim 1, wherein: The generation method of the multi-modal feature set is as follows: Collect and preprocess the multi-modal ophthalmology data of the patient to be treated. The multi-modal ophthalmology data includes image data, text data, and detection data. Perform lesion segmentation and geometric feature extraction on the preprocessed image data to obtain an image feature vector, perform semantic feature extraction on the text data to obtain a text feature vector, and perform structured feature mapping on the detection data to obtain a detection feature vector; Integrate the three modalities of the image feature vector, text feature vector, and detection feature vector to generate a multi-modal feature set in a unified format.

3. The ophthalmology remote intelligent consultation system based on multimodal data fusion according to claim 1, characterized in that: The acquisition method of the probability distribution of the patient's ophthalmic diseases is as follows: Based on clinical guidelines and local case data, construct a directed local ophthalmology knowledge graph containing three types of core entities. Among them, the three types of core entities include disease entities, symptom entities, and detection entities. Obtain the final fusion feature as the initial feature vector of the disease entity node feature vector. Design a multi-modal attention aggregation function to iteratively update the node feature vector starting from the initial feature vector, and apply a fully connected layer and a Softmax function to the iteratively updated node feature vector to obtain the probability distribution of the patient's ophthalmic diseases.

4. The ophthalmology remote intelligent consultation system based on multi-modal data fusion according to claim 3, characterized in that: The acquisition method of the final fusion feature is as follows: Obtain a weight matrix. According to the weight matrix, design a bidirectional fusion mechanism. Each modality feature in the multi-modal feature set serves as both a query to extract other modality features and a key to provide its own features. Calculate intra-modal fusion and inter-modal fusion in combination with the weight matrix. Intra-modal fusion retains the core features of each modality itself, and inter-modal fusion captures cross-modal associations; Combine intra-modal fusion and inter-modal fusion and calculate through layer normalization to obtain the final fusion feature stably.

5. The ophthalmic remote intelligent consultation system based on multi-modal data fusion according to claim 4, characterized in that: The acquisition method of the weight matrix is as follows: Map the image feature vector, text feature vector, and detection feature vector in the multi-modal feature set to a shared latent space through a fully connected layer, convert the multi-modal feature set into a shared latent space representation of the same dimension, define a three-modal interaction matrix to represent the similarity between modalities, calculate the attention weights between any two modality features, and integrate to obtain the weight matrix.

6. The ophthalmology remote intelligent consultation system based on multimodal data fusion according to claim 1, characterized in that: The acquisition method of the electronic consultation form is as follows: In remote consultation, the consulting ophthalmologists who respond to the request respectively provide treatment plans. The Jaccard similarity is used to calculate the similarity between any two treatment plans. Two treatment plans with a similarity greater than or equal to the similarity threshold are grouped into the same similar plan set. The number of treatment plans in each similar plan set is normalized and marked as the consensus eigenvalue of the treatment plans in the similar plan set. Based on any treatment plan, the matching degree of the ophthalmologist who provides the treatment plan is obtained as the weight of the treatment plan. Combining with the consensus eigenvalue of the treatment plan, the recommended value of the treatment plan is calculated by weighted calculation. The treatment plans are sorted in descending order according to the recommended value of the plan, and the sorted treatment plans and the corresponding diagnosis results are filled into the electronic consultation form to generate the electronic consultation form.

7. The ophthalmology remote intelligent consultation system based on multi-modal data fusion according to claim 6, characterized in that: The matching method of the consulting ophthalmologists who respond to the request is as follows: Obtain the matching degrees of all ophthalmologists in the platform, sort the ophthalmologists in descending order according to the matching degrees, and mark the selection serial numbers for the ophthalmologists according to the order to obtain the sorted list of ophthalmologists. Select the ophthalmologists whose selection serial numbers in the list of ophthalmologists are less than or equal to the selection threshold to form a set of consulting ophthalmologists. Send a consultation request to the ophthalmologists in the set of consulting ophthalmologists to match the consulting ophthalmologists who respond to the request.

8. The ophthalmic remote intelligent consultation system based on multi-modal data fusion according to claim 7, wherein: The method for obtaining the matching degree is as follows: Collect the detailed information of all ophthalmologists in the platform, perform feature extraction and quantization to obtain the expert feature vector. Obtain the probability distribution of the ophthalmic diseases of the patient to be treated, construct the probability distribution of the ophthalmic diseases of the patient to be treated as the disease feature vector of the patient to be treated. Perform data processing on the disease feature vector of the patient to be treated and the expert feature vector of each ophthalmologist to obtain the matching degree between the probability distribution of the ophthalmic diseases of the patient and the expert feature vector.

9. The ophthalmology remote intelligent consultation system based on multimodal data fusion according to claim 1, wherein: The method for feedback updating the local knowledge graph is as follows: Obtain the ophthalmic multimodal data of the patient to be treated and the generated electronic consultation form, and perform desensitization processing to generate desensitized ophthalmic treatment data. Continuously record the desensitized ophthalmic treatment data, generate a desensitized data set and add it to the local case data. Each medical institution performs incremental training on the local knowledge graph based on the local case data to optimize the symptom-disease association strength. After the training is completed, extract the graph structure parameters and node feature parameters to update the local knowledge graph. The central server uses the FedAvg algorithm to aggregate the parameters of the local knowledge graphs of multiple institutions, calculates the aggregation weights according to the local case data volumes of each medical institution and performs parameter fusion to obtain the global edge weight matrix and global node embedding vector of the global knowledge graph, and distributes them to each medical institution to replace the local knowledge graph parameters for iterative feedback updating of the local knowledge graph.

10. An ophthalmic remote intelligent consultation method based on multi-modal data fusion, applied to the ophthalmic remote intelligent consultation system based on multi-modal data fusion according to any one of claims 1-9, characterized in that: It includes the following steps: Collect and integrate ophthalmic multimodal data to generate a multimodal feature set. Input the multimodal feature set into an improved cross-modal attention model, generate the final fusion feature through dynamic weight fusion and bidirectional interaction mechanism, and perform graph neural network reasoning in combination with the local ophthalmic knowledge graph to output the probability distribution of the patient's ophthalmic diseases. Quantify the details of ophthalmology experts into expert feature vectors, calculate the matching degree between the probability distribution of the patient's ophthalmic diseases and the expert feature vectors, screen out ophthalmology experts with high matching degrees for consultation, calculate the recommended values of the treatment plans provided by the ophthalmology experts, and generate an electronic consultation form containing diagnostic conclusions and treatment suggestions based on the recommended values; Perform desensitization processing on the electronic consultation form and supplement local case data to train the local knowledge graph. The central server generates a global knowledge graph by weighted aggregation of multi-institutional local knowledge graphs, and performs feedback update on the local knowledge graph according to the generated global knowledge graph.

Citation Information

Patent Citations

  • Self-adaptive remote medical expert recommendation method

    CN115238168A

  • Cross-mechanism medical knowledge graph representation learning method and system

    CN116821375A

  • Gout disease staging prediction method and system based on federal learning and knowledge graph, and storage medium

    CN118866216A

  • Medical diagnosis intelligent decision-making system based on multi-modal data fusion

    CN119495423A

  • Construction method and construction system of telemedicine expert recommendation model, expert recommendation method and electronic equipment

    CN119560169A

Cited By

  • Intelligent terminal multi-mode sentiment analysis method, device and server

    CN120951171A

  • Intelligent terminal multi-modal sentiment analysis method, device and server

    CN120951171B

  • Multi-modal training data desensitization and traceability management method

    CN121278775A

  • Expert matching result generation system based on multi-modal feature fusion

    CN121682724A

  • Expert matching result generation system based on multi-modal feature fusion

    CN121682724B