Large model inquiry system based on multi-modal feature embedding and key point feature alignment
Through a large-modal feature embedding and key point features alignment, the problem of feature alignment and fusion of medical images and text data is solved, and more efficient and accurate medical diagnosis is achieved, and the robustness and diagnostic performance of the system are enhanced.
Patent Information
- Application Number
- CN202510003787.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-02
- Publication Date
- 2025-06-06
AI Technical Summary
The existing multimodal learning is difficult to directly apply in medical diagnosis, especially when processing medical image data and text data, and it is difficult to effectively align and fuse the characteristics of different modalities, resulting in limited diagnostic accuracy and efficiency.
A large-modal feature embedding is used to align with key point features, and feature conversion is performed through convolutional neural networks and recurrent neural networks, feature fusion is performed by combining attention enhancement modules and multi-layer perception machines, and medical knowledge graphs are used to assist in consultation to achieve deep alignment and fusion of image and text features.
It improves the robustness and diagnostic accuracy of the large-model consultation system, can better utilize the complementarity between different modal features, provide high-quality consultation instructions, and enhances the performance and user-friendliness of the system.
Smart Images

Figure CN120108686A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of large models in the medical industry, and specifically relates to a large model consultation system based on multimodal feature embedding and key point feature alignment. Background Art
[0002] Modality refers to some ways of expressing or perceiving things. Each source or form of information can be called a modality. When a research problem contains multiple modalities, it is described as multimodal. In the era of big data, there are a large number of different types of data in the medical field. Common medical data include: disease data, electronic medical record data, imaging data, patient report data, etc. Obviously, the source of data in a real diagnosis and treatment environment is multimodal. Finding the relationship and correspondence between subcomponents of instances from two or more modalities is defined as multimodality. For example, given an image and a title, it is hoped to find the image area that corresponds to the words or phrases in the title.
[0003] Medical image data often contains rich disease semantic information. At the same time, there is often redundancy between data of different medical modalities, which brings challenges to the alignment of multimodal features in medical data.
[0004] That is to say, for the relatively professional scenario of medical diagnosis, current multimodal learning is difficult to be directly applied to the medical diagnosis process. Summary of the invention
[0005] In order to solve the above technical problems, the present invention discloses a large model consultation system based on multimodal feature embedding and key point feature alignment, including:
[0006] A feature conversion module, which is used to: realize the conversion of image features of an original image I from the image modality to text features and the conversion of text features of an original text T from the text modality to image features according to the image modality data and text modality data of the multimodal aspects of the same medical research object, and obtain deep image features and deep text features and a personalized information feature library based on multimodal feature embedding;
[0007] The feature fusion module is used to: align and fuse the deep image features, deep text features and personalized information feature library based on key point features to generate high-quality medical consultation instruction information.
[0008] Preferably,
[0009] The feature conversion module also includes:
[0010] Convolutional Neural Networks and Recurrent Neural Networks.
[0011] Preferably,
[0012] The feature conversion module also includes:
[0013] Text encoder and attention enhancement module.
[0014] Preferably,
[0015] The feature fusion module also includes:
[0016] A multi-layer perceptron is used to map deep image features and deep text features into a joint semantic space.
[0017] Preferably,
[0018] The feature fusion module is also used to combine medical knowledge graphs to assist in diagnosis.
[0019] In addition, the present invention also discloses a computer storage medium, wherein the storage medium includes computer instructions, and when the computer is run on the computer, the computer executes the following method:
[0020] Step S100: based on the image modality data and text modality data of the same medical research object in multimodal aspects, the image features of the original image I from the image modality are transformed into text features, and the text features of the original text T from the text modality are transformed into image features, and deep image features, deep text features and a personalized information feature library are obtained based on multimodal feature embedding;
[0021] Step S200: For the deep image features, deep text features and personalized information feature library, align and fuse features based on key point features to generate high-quality medical consultation instruction information.
[0022] In addition, the present invention also discloses an electronic device, wherein the electronic device comprises:
[0023] A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein:
[0024] When the processor executes the program, the following method is implemented:
[0025] Step S100: based on the image modality data and text modality data of the same medical research object in multimodal aspects, the image features of the original image I from the image modality are transformed into text features, and the text features of the original text T from the text modality are transformed into image features, and deep image features, deep text features and a personalized information feature library are obtained based on multimodal feature embedding;
[0026] Step S200: For the deep image features, deep text features and personalized information feature library, align and fuse features based on key point features to generate high-quality medical consultation instruction information.
[0027] The present invention has the following characteristics:
[0028] The present invention is based on multimodal feature embedding and key point feature alignment. By optimizing transmission, the features of different modalities are aligned and merged, which can better utilize the "complementarity" between different modal features and realize the combination of semantic information under different modalities from multiple angles. It not only enhances the robustness of the large model consultation system, but also improves the performance of the large model consultation system by using rich data. In addition, combining with medical knowledge graphs can help diagnose diseases and ultimately obtain high-quality related indicators that characterize the personalized characteristics of patients. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] By reading the detailed description of the preferred specific embodiments below, various other advantages and benefits of the present invention will become clear to those of ordinary skill in the art. The description figures are only used for the purpose of illustrating the preferred embodiments and are not considered to be limitations of the present invention. Obviously, the figures described below are only some embodiments of the present invention, and for those of ordinary skill in the art, other figures can also be obtained based on these figures without paying creative work. Moreover, the same figure marks are used throughout the figures to represent the same parts.
[0030] Figure 1 It is a structural schematic diagram of a large model medical consultation system based on multimodal feature embedding and key point feature alignment in one embodiment of the present invention;
[0031] Figure 2 It is a schematic diagram of the workflow of a large model consultation system based on multimodal feature embedding and key point feature alignment in another embodiment of the present invention. DETAILED DESCRIPTION
[0032] The following will refer to Figure 1 to Figure 2 Specific embodiments of the present invention are described in detail. Although specific embodiments of the present invention are shown in the figures, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to enable a more thorough understanding of the present invention and to fully convey the scope of the present invention to those skilled in the art.
[0033] It should be noted that certain words are used in the specification and claims to refer to specific components. Those skilled in the art should understand that technicians may use different nouns to refer to the same component. This specification and claims do not use the difference in nouns as a way to distinguish components, but use the difference in the functions of the components as the criterion for distinction. As mentioned throughout the specification and claims, "including" or "comprising" is an open term, so it should be interpreted as "including but not limited to". The subsequent description of the specification is a preferred embodiment of the present invention, but the description is based on the general principles of the specification and is not intended to limit the scope of the present invention. The scope of protection of the present invention shall be determined by the attached claims.
[0034] To facilitate understanding of the embodiments of the present invention, further explanation will be given below using specific embodiments as examples in conjunction with the drawings, and each drawing does not constitute a limitation on the embodiments of the present invention.
[0035] See also Figure 1 In one embodiment, the present invention discloses a large model consultation system based on multimodal feature embedding and key point feature alignment, comprising:
[0036] A feature conversion module, which is used to: realize the conversion of image features of an original image I from the image modality to text features and the conversion of text features of an original text T from the text modality to image features according to the image modality data and text modality data of the multimodal aspects of the same medical research object, and obtain deep image features and deep text features and a personalized information feature library based on multimodal feature embedding;
[0037] The feature fusion module is used to: align and fuse the deep image features, deep text features and personalized information feature library based on key point features to generate high-quality medical consultation instruction information.
[0038] This embodiment is based on multimodal feature embedding and key point feature alignment, and aligns and fuses features of different modalities through optimized transmission. It can better utilize the "complementarity" between features of different modalities and realize the combination of semantic information in different modalities from multiple angles. It not only enhances the robustness of the large model consultation system, but also utilizes rich data to improve the performance of the large model consultation system.
[0039] In another embodiment,
[0040] The feature conversion module also includes:
[0041] Convolutional Neural Networks and Recurrent Neural Networks.
[0042] In another embodiment,
[0043] The feature conversion module also includes:
[0044] Text encoder and attention enhancement module.
[0045] In another embodiment,
[0046] The feature fusion module also includes:
[0047] A multi-layer perceptron is used to map deep image features and deep text features into a joint semantic space.
[0048] In another embodiment,
[0049] The feature fusion module is also used to combine medical knowledge graphs to assist in diagnosis.
[0050] Obviously, the above embodiments can help diagnose diseases by combining medical knowledge graphs, and ultimately obtain high-quality relevant indicators that characterize the patient's personalized characteristics.
[0051] See also Figure 2 , in another embodiment, regarding the feature conversion module:
[0052] The feature conversion module disclosed in the present invention is used for:
[0053] For multimodal data of the same medical research object (including image modality data and text modality data), the image features of the original image I\ are transformed into text features, and the text features of the original text T from the text modality are transformed into image features, and the transformed text feature weighted vector and channel weighted vector are output to obtain deep image features and deep text features. The specific steps are as follows:
[0054] 1. Conversion of image features to text features
[0055] Image feature extraction: Based on the original image I, a convolutional neural network (CNN) is used to extract image features f(I)∈R H ×W×C .
[0056] Image feature processing: The image features are processed by the fully connected layer (FC) and the recurrent neural network (RNN) to generate text features f′(T)∈R D .
[0057] Text feature normalization: Normalize the generated text features, that is:
[0058]
[0059] Text feature weighting: According to the normalized text features, the text feature weighting vector is obtained and multiplied with the feature vector of the original text T to generate the enhanced text feature f(T* ):
[0060] f(T * )=f′(t i )*f(t i )={t 1 *t′ 1 ,t 2 *t′ 2 ,…,t i *t′ i}∈R D
[0061] Deep text feature generation: Use BERT to process the enhanced text features to obtain deep text features F(T * ).
[0062] 2. Conversion of text features to image features
[0063] Text encoding: The input text is passed through a text encoder (including a word embedding layer and a bidirectional recurrent neural network (BiRNN)) to generate sentence features and word features.
[0064] Sentence features to image features: Sentence features undergo a series of convolution and upsampling (or deconvolution) operations to generate low-resolution image features f low ∈R H×W×C :
[0065] f low =UConv(Conv(S))
[0066] From word features to image features: word features are mapped to obtain image space features. The attention weight of each word is determined through the attention mechanism. Each word feature is weighted and summed to form a comprehensive word feature vector, which is combined with the low-resolution image features to generate optimized high-resolution image features.
[0067] The word features are mapped to obtain image space features, which are used together with the low-resolution image features to generate optimized high-resolution image features, and the high-resolution image features are further processed into image features finally obtained based on the text; wherein,
[0068] Exemplarily, the word features are mapped to obtain image space features, specifically including: first, for each word, after extracting its features separately, calculating the correlation between these word features and the low-resolution image features, and determining the attention weight of each word (i.e., which words are the key to generating image features) through an attention mechanism (such as softmax normalized dot product attention); then, based on the attention weight of each word, performing weighted summation on each word feature to form a comprehensive word feature vector, wherein the comprehensive word feature vector reflects which details in the original text should be enhanced or reflected in the image.
[0069] In order to further explore how the present invention uses Bayesian neural networks (BNNs) to handle uncertainty in data and enhance the robustness of the model, we first need to understand the basic principles of BNNs and their specific application in the present invention. Bayesian neural networks are a special type of neural network that not only learns the mapping relationship between input and output, but also quantifies the uncertainty of model predictions. This uncertainty mainly comes from the uncertainty of the model parameters themselves and the noise in the observed data.
[0070] Basics of Bayesian Neural Networks
[0071] In this invention, in order to achieve effective conversion and fusion between image and text features and deal with uncertainty in data, we use Bayesian neural networks (BNNs). The prior distribution p(θ) of the model parameter θ is assumed to be a Gaussian distribution:
[0072]
[0073] Here, μ 0 and Σ 0 Represent the mean vector and covariance matrix of the prior distribution, respectively, which reflect our belief or knowledge about the initial state of the model parameters. In the absence of specific information, a broad prior is usually chosen to represent a wide range of uncertainty about the parameter values.
[0074] Data observation and posterior distribution
[0075] When we observe the data set D = {(x 1 ,y 1 ),(x 2 ,y 2 ),…,(x N ,y N )}, according to Bayes’ theorem, we can update our knowledge about the model parameters, that is, calculate the posterior distribution p(θ│D):
[0076] p(θ│D)∝p(D│θ)p(θ)
[0077] Here, p(D│θ) is the likelihood function, which measures the probability of observing data D given the parameters θ. If we assume that the error of each observation y_i is independent and identically distributed (iid) Gaussian noise, the likelihood function can be written as:
[0078]
[0079] Among them, f(x i ; θ) represents the neural network model for input x i The predicted output, and σ 2 is the variance of the observation noise. This equation shows that for each observation y i are all around their predicted value f(x i ;θ) distribution, and the distribution is composed of the noise variance σ 2 Decide.
[0080] Quantification of certainty
[0081] Through the above process, BNNs can not only provide point estimates of the prediction results, but also give a measure of uncertainty in the prediction results. This is because the posterior distribution p(θ│D) describes the probability distribution of the possible values of the model parameters after considering all available data. Therefore, for a new input x * , we can get a predictive distribution:
[0082] p(y * │x * ,D)=∫p(y * │x * ,θ)p(θ│D)d
[0083] This integral expresses the condition that the data D is known, for the new input x * The predicted value y * In practice, since the above integrals are often difficult to solve analytically, approximate methods such as variational inference or Monte Carlo sampling are usually used to estimate them.
[0084] The present invention realizes the effective conversion and fusion between image and text features through the above-mentioned Bayesian neural network method, and at the same time handles the uncertainty in the data through the Bayesian neural network, thereby improving the robustness and reliability of the model. In particular, in the medical consultation system, this technology can better deal with the vague or incomplete information provided by the patient and provide more accurate and reliable diagnostic suggestions. In addition, by quantifying the uncertainty of the model, the present invention can also help identify those cases that the model is confident of, thereby guiding doctors to take further inspection measures to ensure that the final diagnostic decision is more accurate and safe.
[0085] Word feature extraction and correlation calculation:
[0086] For each word feature (where d w is the word feature dimension, and the word feature set is n w is the number of words), first extract them separately and map them to the potential image feature space to obtain the word image feature vector set where d i is the image feature dimension.
[0087] Application of attention mechanism: Calculate the image feature IW of each word j The correlation between the low-resolution image feature f is calculated using the dot product attention form, and then normalized by the softmax function to obtain the attention weight of the word
[0088]
[0089] Among them, f flat This means flattening f into a one-dimensional vector for dot product operations.
[0090] Comprehensive word feature vector generation: weighted sum of word image features according to attention weight a to form a comprehensive word feature vector
[0091]
[0092] Exemplarily, the high-resolution image features are processed through a post-processing or refining stage, which may include fine-tuning of details or further alignment with text features to ensure that the image features finally obtained based on the text are highly consistent with the description of the original text. That is, the image features finally obtained based on the text can be better aligned with the features of the original text, which facilitates the feature fusion of subsequent multimodal features. Typically, first, the comprehensive word feature vector is combined with the low-resolution image features as the image space feature (for example, by simple splicing, weighted summation or designing a more complex fusion module to achieve its combination), so that the subsequent high-resolution image features retain the original structure while absorbing the specific details guidance from the text; then, the combined comprehensive word features and low-resolution image features are further improved in resolution through upsampling operations (such as deconvolution), and these features are further refined through a multi-layer convolutional network to finally generate high-resolution image features. It can be found that each step is aimed at increasing the detail richness of the image according to the refined text guidance.
[0093] Regarding the complementary fusion that the present invention can further adopt:
[0094] Fusion of low-resolution image features and comprehensive word features: Fusion of comprehensive word features CW and low-resolution image features f. A fusion operation F is used usion , which can be concatenation, weighted summation or other complex fusion strategies. Taking weighted summation as an example, let W f is the fusion weight, then the fused features FusedFeatures are:
[0095] FusedFeatures=W f ·CW+(1―Wf)·Flatten(f)
[0096] Among them, Flatten(f) flattens f into a one-dimensional vector to match the dimension of CW, W f is a scalar between 0 and 1 that controls the contribution ratio of the two sets of features.
[0097] Feature resolution improvement and refinement: The fused features are improved in resolution through upsampling (deconvolution). The deconvolution operation is set to UpSample, and the parameter is the upsampling multiple s, to obtain the preliminary high-resolution feature Prelim HR :
[0098] Prelim HR =UpSample(FusedFeatures,s)
[0099] Then, these features are further refined through a multi-layer convolutional network (MultiLayerConvNet). Assume that the network contains L layers of convolution, and the number of convolution kernels in each layer is c 1 ,c 2 ,…,c L , the convolution operation can be expressed as Conv l , and finally obtain the high-resolution image feature f′(I):
[0100]
[0101] f′(I)=Conv_L(BatchNorm(ReLU(HR_{temp}^{(L―1)})))
[0102] Among them, BatchNorm means batch normalization and ReLU is the activation function.
[0103] Feature alignment: In order to ensure that the final image features are highly aligned with the original text, an additional loss term Loss can be introduced align, for example, using cosine similarity loss to maximize the similarity between CW and feature vectors extracted from HR, or using other suitable alignment strategies:
[0104] Loss align =1―CosineSimilarity(CW,ExtractFeature(f′(I)))
[0105] Among them, ExtractFeature means extracting a vector matching the text feature dimension from the high-resolution image feature HR for comparison and alignment.
[0106] Feature resizing: The high-resolution image feature f'(I) is resized by upsampling or downsampling operations so that the feature map can match the dimensions of the original image I for subsequent operations. Exemplarily, a convolution operation is used to match the sizes of the two to obtain the resized image feature f"(I):
[0107] f″(I)=Conv(f′(I)),f″(I)∈R H×W×C
[0108] Attention enhancement module: The resized image feature f″(I) is first passed through the attention enhancement module to obtain the attention enhanced feature f″′(I); then the attention enhanced feature f″′(I) is processed by the activation function to obtain the channel weight vector f(I^*), and then the channel weight vector is embedded into the image feature f(I)∈R of the original image. H×W×C middle:
[0109] f″′(I)=Atten(f″(I)),f″′(I)∈R C
[0110] f(I^*)=σ(f″′(I)·f(I))
[0111] The attention enhancement module includes a pooling layer, two fully connected layers and a sigmoid activation function in sequence. The pooling layer is used for dimensionality reduction to reduce the computational complexity, and the fully connected layer is used to keep the feature dimension in C dimension. Exemplarily, the corresponding formulas involved in the attention enhancement module are as follows:
[0112] f″′(I)=Atten(f″(I)),f″′(I)∈R C
[0113] f(I * )=σ(f″′(I)·f(I))
[0114] Deep image feature extraction: The channel weight vector f(I *) and the image features f(I)∈R of the original image I H×W×C Vector multiplication is performed, and then further processed by ResNet to obtain deep image features F(I * ).
[0115] Under the diagnosis and treatment model, this paper organically combines multimodal feature fusion and uncertainty processing through a systematic modeling method. Specifically, a multi-level uncertainty resolution framework is designed, which can gradually reduce uncertainty at different levels. The following is a detailed model introduction:
[0116] Feature extraction layer
[0117] At the feature extraction layer, the uncertainty at the feature level is reduced by finely aligning multi-resolution image features and multi-level text semantic features.
[0118] Image feature extraction: Use convolutional neural network (CNN) to extract multi-resolution features from medical images. Let f i is the feature map of the i-th layer, and features of different resolutions can be extracted through multi-scale convolution operations:
[0119] f i =Conv i (I),i=1,2,…,N
[0120] Where I is the input image, Conv i is the convolution operation of the i-th layer, and N is the number of layers.
[0121] Text feature extraction: Use natural language processing (NLP) technology to extract multi-granular semantic features from the patient's symptom description. j is the text feature of the jth layer, and features of different granularities can be extracted through multi-layer encoders (such as Transformer):
[0122] t j =Enc j (T),j=1,2,…,M
[0123] Among them, T is the input text, Enc j is the encoder of the jth layer, and M is the number of layers.
[0124] Feature alignment: Through key point detection and alignment technology, we can ensure the accurate correspondence between image features and text features in the multimodal space. Let A be the alignment operation, and the aligned feature f a and t a for:
[0125] f a =A(f i ),t a=A(t j )
[0126] Model prediction layer
[0127] In the model prediction layer, Bayesian neural networks (BNNs) are used to quantify and process the uncertainty of model parameters to improve the robustness of the model.
[0128] Uncertainty quantification of image features: Assuming that the model parameters θ I The prior distribution p(θ I ) is a Gaussian distribution:
[0129]
[0130] Observe that the dataset D I After that, update the posterior distribution p(θ I │D I ):
[0131] p(θ I │D I )∝p(D I │θ I )p(θ I )
[0132] Among them, p(D I │θ I ) is the likelihood function, assuming that each observation y i The errors are independent and identically distributed (iid) Gaussian noise:
[0133]
[0134] Uncertainty quantification of text features: The uncertainty of text features is also quantified by Bayesian neural networks (BNNs). Assume that the model parameter θ T The prior distribution p(θ T ) is a Gaussian distribution:
[0135]
[0136] Observe that the dataset D T After that, update the posterior distribution p(θ T │D T ):
[0137] p(θ T │D T )∝p(D T │θ T )p(θ T )
[0138] Among them, p(D T │θ T) is the likelihood function, assuming that each observation y i The errors are independent and identically distributed (iid) Gaussian noise:
[0139]
[0140] Uncertainty in the predictive distribution: For a new input x * , the predicted distribution can be calculated through the posterior distribution p(θ│D):
[0141] p(y * │x * ,D)=∫p(y * │x * ,θ)p(θ│D)d
[0142] This integral expresses the condition that the data D is known, for the new input x * The predicted value y * In practice, approximate methods such as variational inference or Monte Carlo sampling are usually used to estimate .
[0143] Decision-making level
[0144] At the decision-making level, the uncertainty in the final decision is further eliminated through the complementary fusion of multimodal uncertainties, ensuring the accuracy and reliability of diagnostic recommendations.
[0145] Uncertainty quantification: Calculate the uncertainty of image features and text features separately. Let u I and u T They are the uncertainties of image features and text features respectively:
[0146] u I =Uncertainty(f a ),u T =Uncertainty(t a )
[0147] Complementary fusion: In the fusion process, the uncertainty of image features is used to supplement the uncertainty of text features, and vice versa. Specifically, when the text information is unclear, the image features can provide additional contextual information; when the image information is unclear, the text features can provide additional semantic information. In this way, the present invention can more comprehensively handle the uncertainty in the data and improve the robustness and reliability of the model. Let F usion For the fusion operation, the fused features FusedFeatures are:
[0148] FusedFeatures=F usiohn (f a ,t a ,uI ,u T )
[0149] Decision: Final diagnostic recommendation pred Generated by fused features FusedFeatures:
[0150] y pred =Decision(FusedFeatures)
[0151] Application Effect
[0152] Through the above multi-level uncertainty elimination framework, the present invention has achieved remarkable results in the following aspects:
[0153] Accurate and reliable: Through multimodal feature fusion and uncertainty processing, the present invention can provide more accurate and reliable diagnostic suggestions, especially when dealing with ambiguous or incomplete information.
[0154] Convenient and efficient: The design of the present invention makes the medical consultation system more user-friendly, capable of quickly responding to user input and providing instant diagnostic suggestions.
[0155] High-quality services: Through systematic modeling and multi-level uncertainty resolution, the present invention can provide high-quality medical services, help doctors make more accurate diagnostic decisions, and improve the quality of medical services.
[0156] See also Figure 2 In another embodiment, regarding the feature fusion module:
[0157] The inventors have noticed that the association between image data and related text data in the medical field can be regarded as a primary-secondary relationship. Images and texts, as primary and auxiliary information sources, work together to support the extraction of richer features. The introduction of this implicit relationship helps to more deeply explore the complex information in the medical field and improve the performance and application potential of deep learning models in this field. Therefore, in order to enhance the learning ability of the technical solution disclosed by the present invention, the present invention adopts two key technologies, namely feature alignment based on optimal transmission of key points and redundancy removal based on threshold suppression, to make full use of the complementarity between multimodal data.
[0158] In the joint semantic space, there are complex features composed of multiple modalities of images and texts. Generally speaking, the data of text modalities contain features from different angles and multiple levels. Among them, the confidence of misdiagnosis, the degree of trust of patients in the diagnosis of large models, the degree of patient friendliness, the degree of diagnosability based on current information, the degree of emotional state of patients, the degree of understanding of large models by patients, the tolerance of the diagnostic inquiry process and the credibility of patient information are often difficult to obtain from the data of patients' image modalities. At the same time, there are direct or indirect connections between the above information, which brings challenges to the data fusion between different modalities. Therefore, a personalized information feature library is constructed based on text features. In the process of fusion of different modal features, the personalized information features corresponding to different features are clarified through the response between features and databases, which provides good interpretability for the process of the diagnosis system.
[0159] F(I*)=Resnet(f(I*))
[0160] F(T*)=Bert(f(T*))
[0161] For the deep image feature F(I * ) and deep text features F(T * ),in,
[0162] Deep text features F(T * ), and construct a personalized information feature library Y through a bidirectional long short-term memory network BiLSTM, that is,
[0163] Y = BiLSTM(F(I*))
[0164] Then, since the feature distribution of different modal data is often related to its own modality, it is impossible to directly perform feature fusion. Therefore, it is necessary to map the deep image features F(I*) and deep text features F(T*) to the joint semantic space Z through the multi-layer perceptron MLP, that is,
[0165] I z =MLP(F(I*))
[0166] T Z =MLP(F(T*))
[0167] Among them, IZ and TZ are the projections of deep features F(I*) and F(T*) in the joint space Z.
[0168] Finally, in order to improve the fusion effect of different modal features, the features of different modalities in the joint projection space need to be aligned and redundantly removed. The process is as follows:
[0169] Perform feature decomposition in each mode to obtain each main feature gi1 and many secondary features gij, where i = 1, 2, ..., m, j = 2, 3, ..., n, m is the number of modes, and n is the number of features. Match the features of each mode to obtain the matching matrix P:
[0170] P ij kl =G(g ki ,g lj )
[0171] Among them, G is the matching function, and k and l represent different modes.
[0172] In order to reduce the redundancy between different modal data, threshold suppression is used to remove redundancy in the matching matrix P. Specifically, the similarity S between the secondary features of different modalities is calculated based on the matching matrix:
[0173] S ij kl =sim(P ih k ,P hj l )
[0174] Where h=1,2,…,n. If S ij kl is greater than the threshold ε, indicating that g k i and g l j The similarity between the two features is high, and there is information redundancy. Select any one of them to remove to reduce the complexity of the feature parameters.
[0175] Furthermore, the mask matrix M is obtained by using the features of each modality and the personalized information feature library Y:
[0176]
[0177] Among them, Y d Represents different features in the personalized information feature library, including confidence, patient trust, patient friendliness, diagnosability, and patient emotional state.
[0178] The mask matrix M reflects the key point pairing situation. Taking it as the core, the optimal transmission method is further used to measure the cost LOT of aligning different modal features, namely:
[0179]
[0180] Among them, 1m is an m-dimensional all-1 matrix, and c is the transmission cost matrix.
[0181] The mask matrix M reflects the key point pairing situation. Taking it as the core, the optimal transmission method is further used to measure the cost L of aligning different modal features. OT ,Right now:
[0182]
[0183] Among them, 1 m is an m-dimensional all-1 vector, and c is the transmission cost matrix. The cost matrix C is used in this context to measure the cost of mapping the key points of one modality to the corresponding key points of another modality. Each element c ij represents the cost of mapping the i-th key point in the source modality to the j-th key point in the target modality. This cost can be defined based on different metrics, such as the Euclidean distance, the Mahalanobis distance, or the distance calculated based on feature similarity. In terms of specific definition, if we represent the key point features of the two modalities as sets and Then the elements of the cost matrix C can be calculated as follows:
[0184]
[0185] Here, d(·,·) is a distance metric function, for example:
[0186] If the Euclidean distance is used, then
[0187] If based on feature similarity, then where similarity(·,·) can be cosine similarity, correlation coefficient, or other similarity measures.
[0188] Introducing uncertainty considerations
[0189] In practical applications, the position of key points often has a certain degree of uncertainty, which may be caused by image noise, modal conversion error or other factors. In order to make the model more robust, we can add uncertainty considerations based on the original linear optimization transfer (LOT) framework. This uncertainty can be expressed through a probabilistic model, for example, assuming that the position of each key point follows a certain probability distribution.
[0190] Mathematical representation of uncertainty
[0191] Assumptions and are the positions of the key points in the source mode and the target mode, respectively, and their position uncertainties can be expressed using the probability density function and At this time, the transmission cost c ijThis uncertainty should be reflected, which can be achieved by calculating the expected transmission cost:
[0192]
[0193] This means that the new transmission cost c′ ij Not only the direct distance or similarity between two key points is considered, but also the uncertainty of their positions is comprehensively considered.
[0194] Adjustment of the optimization problem
[0195] After taking uncertainty into account, the original optimization problem is adjusted to:
[0196]
[0197] in,
[0198]
[0199] Here c′ ij This formula describes a new linear programming problem, which aims to find an optimal transportation plan P that considers the mask matrix M and the transmission cost c′ after uncertainty correction. ij Under the condition of , the total transmission cost from the source modality feature to the target modality feature is minimized. At the same time, P needs to satisfy the constraint of probability distribution, that is, the total weight from the source modality to the target modality must remain unchanged, which ensures the quality conservation principle in the matching process.
[0200] The personalized information feature library can help construct the mask matrix M, so that data of different modalities with similar personalized information features can be better integrated. The medical knowledge graph can help diagnose diseases and ultimately obtain relevant indicators that characterize the patient's personalized characteristics.
[0201] Although real medical data is complex and diverse, existing image-based disease diagnosis and treatment models do not consider non-image information including patients' clinical data, biomarker data, and pathological data, and their single modality features lead to limited representation capabilities. However, the present invention establishes a large-model consultation system based on multimodal feature embedding and key point feature alignment, which provides high-quality indication information through a large language model, thereby ensuring excellent consultation results.
[0202] In another embodiment,
[0203] The system processes real medical data in multiple modalities, including text and images, from patients, and follows the following process:
[0204] Firstly, a multimodal feature embedding method based on feature transformation is adopted to make full use of the connection between image and text, and generate text feature weighted vectors and channel weighted vectors for image modality data and text modality data respectively, so as to obtain deep image features and deep text features.
[0205] Next, the deep features of different modalities are projected into the joint semantic space to generate the cost matrix and the matching matrix; and the deep text features are used to build a personalized information feature library and establish a mask matrix in the semantic space;
[0206] Then, the deep features are aligned to obtain fused features by using feature alignment based on optimal transmission of key points and redundancy removal based on threshold suppression. Finally, high-quality prompt words such as diseases, emotions, and languages are generated to ensure that the large language model can make correct feedback.
[0207] In addition, in another embodiment, the present invention also discloses a computer storage medium, wherein the storage medium includes computer instructions, and when the computer is executed on the computer, the computer executes the following method:
[0208] Step S100: based on the image modality data and text modality data of the same medical research object in multimodal aspects, the image features of the original image I from the image modality are transformed into text features, and the text features of the original text T from the text modality are transformed into image features, and deep image features, deep text features and a personalized information feature library are obtained based on multimodal feature embedding;
[0209] Step S200: For the deep image features, deep text features and personalized information feature library, align and fuse features based on key point features to generate high-quality medical consultation instruction information.
[0210] In addition, in another embodiment, the present invention further discloses an electronic device, wherein the electronic device includes:
[0211] A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein:
[0212] When the processor executes the program, the following method is implemented:
[0213] Step S100: based on the image modality data and text modality data of the same medical research object in multimodal aspects, the image features of the original image I from the image modality are transformed into text features, and the text features of the original text T from the text modality are transformed into image features, and deep image features, deep text features and a personalized information feature library are obtained based on multimodal feature embedding;
[0214] Step S200: For the deep image features, deep text features and personalized information feature library, align and fuse features based on key point features to generate high-quality medical consultation instruction information.
[0215] In summary, the present invention has the following characteristics:
[0216] In the data-driven large-model medical consultation system, firstly, the features of each modality are extracted in a noise-resistant, accurate and robust manner for massive cross-modal real medical data, and a unified embedded expression is efficiently iterated and derived; secondly, based on the optimal transmission of weakly connected key points and the suppression of strongly disentangled representation chains, the inherent correlation and complementarity of real multi-modal data are utilized to carry out semantic-oriented deep alignment and fusion between multi-view and cross-modal features; in addition, based on the real medical data with unified embedding expression and semantic alignment, an efficient medical consultation prompt generation model supported by a cascade training framework is constructed, and the optimal large-model medical consultation prompt words that can carry out high-performance diagnosis and treatment decisions are derived in a data-driven manner;
[0217] By aligning the features of different modalities through optimized transmission, we can better utilize the "complementarity" between the features of different modalities and realize the combination of semantic information in different modalities from multiple angles, which not only enhances the robustness of the model but also improves the performance of the model by utilizing rich data.
[0218] Although the embodiments of the present invention are described above in conjunction with the figures, the present invention is not limited to the above specific embodiments and application fields, and the above specific embodiments are only illustrative and instructive, rather than restrictive. A person of ordinary skill in the art can make many forms under the guidance of this specification and without departing from the scope of protection of the claims of the present invention, all of which belong to the protection of the present invention.
Claims
1. A large-model medical consultation system based on multimodal feature embedding and key point feature alignment, characterized by: include: A feature conversion module, which is used to: realize the conversion of image features of an original image I from the image modality to text features and the conversion of text features of an original text T from the text modality to image features according to the image modality data and text modality data of the multimodal aspects of the same medical research object, and obtain deep image features and deep text features and a personalized information feature library based on multimodal feature embedding; The feature fusion module is used to: align and fuse the deep image features, deep text features and personalized information feature library based on key point features to generate high-quality medical consultation instruction information.
2. The system according to claim 1, characterized in that Preferably, the feature conversion module further includes: Convolutional Neural Networks and Recurrent Neural Networks.
3. The system according to claim 1, characterized in that The feature conversion module also includes: Text encoder and attention enhancement module.
4. The system according to claim 1, characterized in that The feature fusion module also includes: A multi-layer perceptron is used to map deep image features and deep text features into a joint semantic space.
5. The system according to claim 1, characterized in that The feature fusion module is also used to combine medical knowledge graphs to assist in diagnosis.
6. A computer storage medium, wherein: The storage medium includes computer instructions, which, when executed on a computer, enable the computer to perform the following method: Step S100: based on the image modality data and text modality data of the same medical research object in multimodal aspects, the image features of the original image I from the image modality are transformed into text features, and the text features of the original text T from the text modality are transformed into image features, and deep image features, deep text features and a personalized information feature library are obtained based on multimodal feature embedding; Step S200: For the deep image features, deep text features and personalized information feature library, align and fuse features based on key point features to generate high-quality medical consultation instruction information.
7. An electronic device, wherein: The electronic device comprises: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the following method is implemented: Step S100: based on the image modality data and text modality data of the same medical research object in multimodal aspects, the image features of the original image I from the image modality are transformed into text features, and the text features of the original text T from the text modality are transformed into image features, and deep image features, deep text features and a personalized information feature library are obtained based on multimodal feature embedding; Step S200: For the deep image features, deep text features and personalized information feature library, align and fuse features based on key point features to generate high-quality medical consultation instruction information.
Citation Information
Cited By
Medical image feature conduction identification method and system
CN121983221A
A medical image feature conduction recognition method and system
CN121983221B