Multi-mode-based allergic rhinitis diagnosis method and system and electronic equipment
By employing a multimodal fusion method and utilizing a multi-path attention model to process diagnostic image and text features, combined with tensor shrinkage, we can achieve standardized and precise diagnosis of allergic rhinitis, overcoming the limitations of traditional diagnostic methods and improving the reliability and efficiency of diagnosis.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-28
- Publication Date
- 2026-03-31
AI Technical Summary
In existing technologies, the diagnosis of allergic rhinitis mainly relies on a single Western or Traditional Chinese Medicine diagnosis, which has problems such as obvious drug side effects, long treatment courses, or high dependence on the subjective experience of physicians. Moreover, there is little discussion on the combination of Traditional Chinese Medicine theory and artificial intelligence, making it difficult to promote on a large scale.
A multimodal fusion method is adopted to acquire sample images and text of the diagnostic subjects, and to perform self-attention and cross-attention processing using a multi-path attention model, combined with tensor contraction, to achieve standardized and accurate diagnosis of allergic rhinitis.
Standardized and precise diagnosis of allergic rhinitis can be achieved without human intervention, improving the reliability and efficiency of diagnosis and overcoming the limitations of traditional methods.
Smart Images

Figure CN121768628A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, and in particular to a multimodal diagnostic method, system, and electronic device for allergic rhinitis. Background Technology
[0002] Allergic rhinitis is a common allergic disease. Although not fatal, it can severely impact patients' quality of life. Current technologies typically utilize deep learning models to study allergic rhinitis through methods such as disease symptom feature classification and pathological text processing. However, these methods largely focus on solving single Western medical problems, with limited exploration of how to organically combine traditional Chinese medicine theory with artificial intelligence. This restricts the development of intelligent traditional Chinese medicine and ultimately hinders its large-scale promotion. Summary of the Invention
[0003] The following is an overview of the subject matter described in detail in this disclosure. This overview is not intended to limit the scope of the claims.
[0004] This disclosure provides a multimodal diagnostic method for allergic rhinitis, which can achieve standardization and accuracy in the diagnosis of allergic rhinitis without human intervention, effectively improving the reliability and efficiency of the diagnosis of allergic rhinitis, and is conducive to the promotion of this diagnostic method.
[0005] On one hand, embodiments of this disclosure provide a multimodal diagnostic method for allergic rhinitis, including: Obtain the sample diagnostic images and sample diagnostic text of the diagnostic object; Feature extraction is performed on the sample diagnostic image to obtain diagnostic image features, and feature extraction is performed on the sample diagnostic text to obtain diagnostic text features; A multi-path attention model is invoked to perform self-attention processing on the diagnostic image features and the diagnostic text features respectively to obtain image enhancement features and text enhancement features. Cross-attention processing is then performed on the image enhancement features and the text enhancement features to obtain image association features and text association features. During the training of the multi-path attention model, the contrast loss is determined based on the similarity between the image enhancement features and the text enhancement features. Tensor shrinkage is performed based on the image association features and the text association features to obtain target fusion features, and rhinitis classification results are output based on the target fusion features.
[0006] On the other hand, embodiments of this disclosure also provide a multimodal allergic rhinitis diagnostic system, including: The sample acquisition module is used to acquire sample diagnostic images and sample diagnostic text of the diagnostic object; The feature extraction module is used to extract features from the sample diagnostic image to obtain diagnostic image features, and to extract features from the sample diagnostic text to obtain diagnostic text features; The multimodal fusion module is used to call a multi-path attention model to perform self-attention processing on the diagnostic image features and the diagnostic text features respectively to obtain image enhancement features and text enhancement features. The image enhancement features and the text enhancement features are then subjected to cross-attention processing to obtain image association features and text association features. During the training of the multi-path attention model, the contrast loss is determined based on the similarity between the image enhancement features and the text enhancement features. The classification module is used to perform tensor shrinkage based on the image association features and the text association features to obtain target fusion features, and output rhinitis classification results based on the target fusion features.
[0007] The embodiments of this disclosure include at least the following beneficial effects: acquiring sample diagnostic images and sample diagnostic text of the diagnostic object; extracting features from the sample diagnostic images to obtain fixed-dimensional diagnostic image features; extracting features from the sample diagnostic text to obtain fixed-dimensional diagnostic text features; and outputting fixed-dimensional diagnostic image features and diagnostic text features to provide a prerequisite for subsequent multimodal feature fusion. Based on this, self-attention processing is applied to the diagnostic image features and diagnostic text features respectively, further enhancing their expressive power to obtain image enhancement features and text enhancement features. Simultaneously, cross-attention processing is applied to the image enhancement features and text enhancement features to achieve modal interaction between the fusion of image enhancement features and text enhancement features, resulting in image association features and text association features. During the training of the multi-path attention model, a contrast loss is determined based on the similarity between image enhancement features and text enhancement features. This contrast loss strengthens the consistency of matching image and text features and weakens the correlation between mismatched image and text features, achieving multimodal spatial alignment. Finally, tensor shrinking is performed on the image association features and text association features to obtain target fusion features. Based on the target fusion features, the rhinitis classification results are output. The standardization and accuracy of allergic rhinitis diagnosis can be achieved without human intervention, which effectively improves the reliability and efficiency of allergic rhinitis diagnosis.
[0008] Other features and advantages of this disclosure will be set forth in the following description and will be apparent in part from the description or may be learned by practicing this disclosure. Attached Figure Description
[0009] The accompanying drawings are provided to further understand the technical solutions of this disclosure and constitute a part of the specification. They are used together with the embodiments of this disclosure to explain the technical solutions of this disclosure and do not constitute a limitation on the technical solutions of this disclosure.
[0010] Figure 1 A schematic diagram illustrating an optional implementation environment provided for an embodiment of this disclosure; Figure 2 An optional flowchart for a multimodal-based diagnostic method for allergic rhinitis provided in this disclosure embodiment; Figure 3 An optional schematic diagram of sample data provided in an embodiment of this disclosure; Figure 4 This is a schematic diagram of an optional structure of the multi-path attention model provided in an embodiment of the present disclosure; Figure 5 A schematic diagram of an optional overall framework for a multimodal-based diagnostic method for allergic rhinitis provided in this disclosure embodiment; Figure 6 This is a schematic diagram of the structure of a multimodal allergic rhinitis diagnostic system provided in an embodiment of this disclosure. Detailed Implementation
[0011] To make the objectives, technical solutions, and advantages of this disclosure clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and are not intended to limit the scope of this disclosure.
[0012] It should be noted that in the various specific embodiments of this disclosure, when processing is required based on data related to the characteristics of the target object, such as target object attribute information or a set of attribute information, the permission or consent of the target object will be obtained first. Furthermore, the collection, use, and processing of this data will comply with relevant laws, regulations, and standards. The target object can be a user. In addition, when embodiments of this disclosure require obtaining target object attribute information, separate permission or consent from the target object will be obtained through pop-ups or redirection to a confirmation page. Only after obtaining the target object's separate permission or consent will the necessary target object-related data for the normal operation of the embodiments of this disclosure be obtained.
[0013] In this disclosure, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.
[0014] To facilitate understanding of the technical solutions provided in the embodiments of this disclosure, some key terms used in the embodiments of this disclosure will be explained below: Allergic rhinitis, also known as hay fever, is a chronic non-infectious inflammatory disease of the nasal mucosa caused by contact with allergens such as pollen, dust mites, and pet dander. Its pathogenesis is closely related to the abnormal response of the body's immune system. The typical clinical manifestations are sneezing, runny nose with clear discharge, nasal congestion, and nasal itching. Some patients may also experience local discomfort such as itchy eyes and throat. The condition is easily affected by environmental factors and recurs frequently. It not only affects the patient's daily life and sleep quality, but may also induce complications such as asthma in severe cases.
[0015] Tensor shrinking is a core algebraic operation for high-order tensors. Essentially, it involves summing and reducing the dimensionality of a tensor across its corresponding dimensions to extract inter-dimensional relationships. Typically, it involves specifying one or more pairs of indices (corresponding dimensions) within the tensor, multiplying the elements of these dimensions at their corresponding positions, and then summing the results. This achieves dimensional reduction and aggregation of key information. This operation preserves the inherent relationships between modalities and feature dimensions within a tensor and is widely used in fields such as multimodal feature fusion and deep learning model optimization.
[0016] Attention mechanisms are a core technology in deep learning that simulates the characteristics of human visual attention allocation. Essentially, they assign differentiated levels of attention to different parts of the input information through a dynamic weight allocation mechanism. This highlights key information, suppresses redundant interference, and accurately captures the correlation features between information. The mechanism first calculates the attention weights of the input elements, and then aggregates the input information based on these weights to obtain the output. It is suitable for key feature extraction from single-modal data and can also efficiently handle cross-domain correlation modeling of multimodal data, making it widely used in natural language processing, computer vision, and multimodal fusion.
[0017] Allergic rhinitis is a common allergic disease that, while not fatal, significantly impacts patients' quality of life. Current diagnostic and intervention methods for allergic rhinitis are mainly divided into two major systems: Traditional Chinese Medicine (TCM) and Western medicine. Western medicine uses allergen testing combined with symptom scales for diagnosis, and medication can quickly relieve acute symptoms. TCM follows the principle of "syndrome differentiation and treatment," diagnosing based on the patient's symptoms, tongue and pulse signs, and constitution, and treating through herbal remedies, acupuncture, and acupressure, offering advantages such as low recurrence rates and high safety. However, Western medicine approaches have inherent problems such as significant drug side effects and lengthy treatment courses, while TCM diagnosis relies heavily on the physician's subjective experience, lacking unified, objective, and quantifiable standards for surface biological characteristics, resulting in limitations in both diagnostic methods.
[0018] In recent years, with the development of artificial intelligence, machine learning models or deep learning models have emerged to study allergic rhinitis through methods such as blood testing, exploring air pollutants, studying cell genes, classifying disease symptom characteristics, and processing pathological texts. However, most of these methods focus on solving single Western medical problems and rarely explore how to organically combine traditional Chinese medicine theory with artificial intelligence, which limits the development of intelligent traditional Chinese medicine and ultimately makes it difficult to promote it on a large scale.
[0019] Based on this, the present disclosure provides a multimodal diagnostic method for allergic rhinitis, which can achieve standardization and accuracy in the diagnosis of allergic rhinitis without human intervention, effectively improving the reliability and efficiency of the diagnosis of allergic rhinitis.
[0020] Reference Figure 1 , Figure 1 This is a schematic diagram of an optional implementation environment provided by an embodiment of the present disclosure. The implementation environment includes a terminal 101 and a server 102. Specifically, the server 102 acquires sample diagnostic images and sample diagnostic text of the diagnostic object collected from the terminal 101. It performs feature extraction on the sample diagnostic images to obtain diagnostic image features, and performs feature extraction on the sample diagnostic text to obtain diagnostic text features. It performs self-attention processing on the diagnostic image features and diagnostic text features respectively to obtain image enhancement features and text enhancement features. It determines the contrast loss based on the similarity between the image enhancement features and text enhancement features, performs cross-attention processing on the image enhancement features and text enhancement features to obtain image association features and text association features, performs tensor shrinkage based on the image association features and text association features to obtain target fusion features, outputs rhinitis classification results based on the target fusion features, and sends the rhinitis classification results to the terminal 101 for physicians to diagnose and predict rhinitis based on the rhinitis classification results.
[0021] Reference Figure 2 , Figure 2 This is an optional flowchart of a multimodal-based allergic rhinitis diagnostic method provided in an embodiment of the present disclosure. The multimodal-based allergic rhinitis diagnostic method includes, but is not limited to, the following steps S201 to S204.
[0022] Step S201: Obtain the sample diagnostic image and sample diagnostic text of the diagnostic object.
[0023] The diagnostic object refers to the object that provides sample diagnostic images and sample diagnostic text. The sample diagnostic images are images of the face and tongue, and the sample diagnostic text is text describing the sample diagnostic images, such as... Figure 3As shown, the left image is a sample diagnostic image, showing the facial and tongue condition of the patient being diagnosed. The right image is a sample diagnostic text, describing the basic condition of the sample diagnostic image in the left image, specifically "pale red tongue with teeth marks, thin white tongue coating, deficiency of both Qi and Yin, Qi deficiency constitution".
[0024] Specifically, clinical samples of the diagnostic subjects are collected using tongue surface pulse and acupoint information collection and management equipment. After screening and labeling, sample diagnostic images are obtained. At the same time, biochemical data, case information or consultation information corresponding to the diagnostic subjects are collected. This information is then screened and organized to obtain sample diagnostic text. A sample dataset is constructed based on the sample diagnostic images and sample diagnostic text.
[0025] It should be noted that the sample diagnostic images and sample diagnostic text obtained from the sample dataset in step S201 can be matched images and text, i.e., clinical samples or medical record information of the same diagnostic subject, or they can be mismatched text, i.e., clinical samples or case information of different diagnostic subjects. By randomly obtaining sample diagnostic text and sample diagnostic images, the multi-path attention model can better learn the association between image-text pairs.
[0026] Step S202: Extract features from the sample diagnostic image to obtain diagnostic image features, and extract features from the sample diagnostic text to obtain diagnostic text features.
[0027] The diagnostic image features are obtained by extracting features from the sample diagnostic images using an image encoder. Specifically, the image encoder is the image encoder in the CLIP (Contrastive Language–Image Pre-training) model. The image encoder includes multiple processing layers, specifically a preprocessing layer, an embedding layer, and a feature extraction layer. The processing procedure of the image encoder can be described by the following formula. , For fixed-dimensional diagnostic image features, The diagnostic image for the sample is a tensor of shape (C,H,W), where C is the number of channels, H is the image height, and W is the image width. This is the image encoder in the CLIP model, used to convert image data into feature representations. The model parameters of the CLIP model are learned during its pre-training phase through comparative learning on a large amount of image-text data. Therefore, when training the target model, the model parameters of the image encoder are frozen and the model parameters (such as weights) are not updated during training. This training method can effectively reduce the number of parameters for training the target model, thereby reducing training costs. In addition, this training method can also avoid overfitting and ensure the stability of the target model in visual modality performance.
[0028] The diagnostic text features are obtained by extracting features from the sample diagnostic text using a text encoder, specifically a Large Language Model (LLM). The processing of the text encoder can be described by the following formula: , For fixed-dimensional diagnostic text features, The sample diagnostic text is typically a sequence where each element represents a word or sub-token in the sample diagnostic text. The LLM (Large Language Model) calls a function with the sample diagnostic text as output. Large language models are pre-trained on large amounts of text data to learn rich language representations, enabling them to understand complex language structures and semantics. Therefore, when training the target model, the text encoder's model parameters are frozen; these parameters (such as weights) are not updated during training. This training method effectively reduces the number of parameters required for training the target model, thus lowering training costs. Furthermore, this training method effectively avoids the risk of overfitting that may occur during end-to-end training of the target model, addressing the issue of small text datasets in rhinitis classification tasks. By freezing the text encoder's model parameters and fully utilizing its generalization and semantic representation capabilities to adapt to the rhinitis classification scenario, the risk of overfitting during small-sample training can be reduced.
[0029] In one possible implementation, during the process of extracting features from the sample diagnostic image to obtain diagnostic image features, the sample diagnostic image can be segmented into multiple sample image blocks, feature mapping can be performed on the multiple sample image blocks to obtain the image block embedding vectors corresponding to the image blocks, self-attention processing can be performed on the multiple image block embedding vectors to obtain the image attention scores corresponding to the multiple image block embedding vectors, and the multiple image block embedding vectors can be weighted and summed based on the image attention scores to obtain the diagnostic image features.
[0030] Specifically, the sample diagnostic image is input into the image encoder. In the preprocessing layer, the sample diagnostic image is preprocessed by normalization, scaling, and cropping to convert it into a data format compatible with the image encoder. Next, the preprocessed sample diagnostic image is segmented into multiple non-overlapping sample image patches of the same size. In the embedding layer, these patches are convolved, and a linear transformation maps each patch to a high-dimensional feature space, yielding a high-dimensional image patch embedding vector for each patch. These embedding vectors are arranged according to the segmentation order, and their sequence can be represented as [Patch1, Patch2, ..., PatchN]. This sequence is then converted into an image patch embedding matrix to be compatible with the image encoder's data format. The image patch embedding matrix can be represented as [Batchsize, Patches, Dim], where Batchsize is the batch size (representing the amount of data the image encoder can process at one time), Patches is the number of image patches, and Dim is the dimension of the image patches. The image patch embedding matrix is input into multiple feature extraction layers for feature extraction, yielding diagnostic image features. These feature extraction layers can specifically be Transformer layers, including a self-attention mechanism and a feedforward network. The self-attention mechanism guides the image encoder to focus on different regions of the sample diagnostic image; specifically, it is a multi-head self-attention mechanism. The feedforward network is used for further processing. The calculation process of diagnostic image features can be represented by the following formula.
[0031] Among them, the PatchEmbed function is a segmentation function used to segment the sample diagnostic image into multiple image patches and embed the image patches into a high-dimensional feature space; MultiHead is a multi-head self-attention mechanism that takes the query matrix, key matrix and value matrix as input and outputs a weighted summation feature representation.
[0032] In the feature extraction layer, self-attention processing is implemented by a multi-head attention mechanism. Based on the image patch embedding matrix, a linear transformation is performed to obtain the corresponding query matrix, key matrix, and value matrix. According to the preset number of attention heads N, the query matrix, key matrix, and value matrix are divided into N query sub-matrices, key sub-matrices, and value sub-matrices of the same shape. For any attention head, the image attention score among all image patch embedding vectors in the image patch embedding matrix is calculated based on the query sub-matrices and key matrices. After normalizing the image attention score, it is weighted and summed with the value sub-matrices to obtain the image patch attention features corresponding to each image patch. The image patch attention features corresponding to all image patches are integrated to obtain the image attention features corresponding to the current attention head.
[0033] Next, the image attention features output from multiple attention heads are summed and then input into the feedforward network for further feature extraction. This yields the first intermediate feature output by the feature extraction layer. The first intermediate feature is then input into the next feature extraction layer for image feature extraction, continuing until the last feature extraction layer extracts image features and outputs fixed-dimensional diagnostic image features. These diagnostic image features can capture high-level semantic information of the sample diagnostic images and achieve cross-modal semantic alignment through contrastive learning of the CLIP model, forming image features that are highly adapted to text features. This provides strong support for subsequent interaction and fusion of image and text modalities.
[0034] In one possible implementation, during the feature extraction process of the sample diagnostic text to obtain diagnostic text features, the sample diagnostic text can be divided into multiple word texts. These word texts are then converted into word text embedding vectors, and positional encoding is performed on the word text embedding vectors to preserve their positional information within the sample diagnostic text. Self-attention processing is then applied to the positionally encoded word text embedding vectors to obtain text attention scores. Based on these text attention scores, a weighted sum of multiple word text embedding vectors is performed to obtain the diagnostic text features. Here, word texts are either words or tokens. For example, if the sample diagnostic text is "pale red tongue, with teeth marks, thin white tongue coating, qi and yin deficiency, qi deficiency constitution," the resulting text sequence after segmentation could be ["pale red tongue", "teeth marks", "thin white tongue coating", "qi and yin deficiency", "qi deficiency constitution"].
[0035] Specifically, the sample diagnostic text is segmented into multiple word texts. These word texts are then input into the embedding layer of the text encoder for mapping, resulting in multiple fixed-dimensional word text embedding vectors. These vectors are then integrated into a sequence of word text embedding vectors. It should be noted that the embedding layer has already implemented semantic association modeling during the pre-training phase, making semantically similar word texts closer together in the vector space. Next, the word text embedding vectors are positionally encoded, and the positional encoding results are superimposed on the corresponding word text embedding vectors to preserve the positional information of each word text in the sample diagnostic text. Self-attention processing is then applied to the multiple word text embedding vectors, calculating the text attention score of each word text embedding vector with respect to the other word text embedding vectors. Based on these text attention scores, a weighted sum is performed on the multiple word text embedding vectors to obtain the word text attention features corresponding to each word text. The self-attention processing of the word text embedding vectors can be represented by the following formula.
[0036] in, Let K be the word text query vector corresponding to the i-th word text, and K and V be the key matrix and value matrix corresponding to all word text embedding vectors, respectively. Let be the dimension of each key vector in the key matrix K, used to scale the dot product and avoid the gradient vanishing problem when the dimension is large.
[0037] Next, multiple word text attention features are input into a feedforward network for processing to obtain intermediate features containing rich semantic information. This sample feedforward network consists of two linear layers and an activation function. The processing procedure of the feedforward network can be specifically represented by the following formula. , in, For word text attention features, and This is the weight matrix. and For bias terms, The activation function is then applied. Layer normalization is then performed on the intermediate features to reduce covariate shifts within the features, resulting in fixed-dimensional diagnostic text features output by the text encoder. These diagnostic text features integrate global information from the sample diagnostic text.
[0038] To further capture long-range dependencies in the sample diagnostic text, the diagnostic text features are input into the Transformer encoder. A self-attention mechanism is used to accurately associate scattered key information within the sample diagnostic text, thereby capturing long-range dependencies. Simultaneously, the Transformer encoder can perform global processing on the text sequence, enabling a more accurate understanding of the contextual logic of the sample diagnostic text. This results in diagnostic text features that express long-range dependencies and contain contextual information, providing richer semantic support for subsequent multimodal fusion. The formula is expressed as follows: These are the diagnostic text features obtained after processing by the Transformer encoder.
[0039]
[0040] Step S203: Call the attention model to perform self-attention processing on the diagnostic image features and diagnostic text features respectively to obtain image enhancement features and text enhancement features. Perform cross-attention processing on the image enhancement features and text enhancement features to obtain image association features and text association features.
[0041] Specifically, step S203 is performed in the multi-channel attention model, such as... Figure 4As shown, diagnostic image features and diagnostic text features first pass through a normalization layer and a multi-head self-attention layer. The multi-head self-attention layer performs self-attention processing on both the diagnostic image features and diagnostic text features separately, and then feeds them into expert networks for feature processing and analysis. The multi-path attention model can select the corresponding expert network for multi-head self-attention processing based on features of different modalities. The multi-path attention model includes three paths: the first path is the image modality expert network (V-FFN) for processing diagnostic image features; the second path is the text modality expert network (L-FFN) for processing diagnostic text features; and the third path is the image-text modality expert network (VL-FFN) for processing image-text fusion features.
[0042] In one possible implementation, the process of performing self-attention processing on diagnostic image features to obtain image enhancement features can specifically involve performing a linear transformation on the diagnostic image features to obtain an image query matrix, an image key matrix, and an image value matrix. These matrices are then split into multiple query head matrices, multiple key head matrices, and multiple value head matrices. For any given self-attention head, a similarity calculation is performed based on the query head matrix and the key head matrix. A self-attention score is obtained based on the similarity calculation result. The attention head features of the self-attention head are determined by the product of the self-attention score and the value head matrix. These multiple attention head features are then concatenated to obtain the image enhancement features.
[0043] Specifically, in this implementation, self-attention processing is implemented using a multi-head attention mechanism. The image query matrix, image key matrix, and image value matrix are obtained by linearly transforming the image enhancement features using their respective linear transformation matrices. Taking the image query matrix as an example, based on a preset number of self-attention heads, the image query matrix... Split into multiple query header matrices L is the sequence length of the query image matrix, and D is the feature dimension of the vectors in the query image matrix. To query the feature dimensions of the vectors in the header matrix, h represents the number of self-attention heads.
[0044] For any self-attention head, attention processing is performed independently based on the corresponding query head matrix, key head matrix, and value head matrix. Similarity is calculated based on the query head matrix and key head matrix. The ratio of the similarity calculation result to the feature dimension of the key vector is normalized and then multiplied by the value head matrix to obtain the attention head feature of that self-attention head. The processing of self-attention heads can be represented by the following formula. , Where Q is the query submatrix, K is the key submatrix, and V is the value submatrix. Let be the dimension of each key vector in the key matrix K, used to scale the dot product and avoid the gradient vanishing problem when the dimension is large. Softmax The function normalizes the image attention scores, allowing the image attention features output by each attention head to be summed. Finally, the attention head features from multiple attention heads are concatenated to obtain concatenated features. A linear transformation matrix is then applied to the concatenated features to fuse the attention head features from multiple attention heads into a more compact feature representation, resulting in the image enhancement features.
[0045] It should also be noted that the process of performing self-attention processing on diagnostic text features to obtain text enhancement features is similar to the process of obtaining image enhancement features, and will not be elaborated here.
[0046] In one possible implementation, during the training of the multi-path attention model, image enhancement features and text enhancement features are processed in tensor form. In determining the contrast loss based on the similarity between the image enhancement features and the text enhancement features, specifically, the image enhancement features can be linearly transformed to obtain the image attention tensor corresponding to the image enhancement features, and the text enhancement features can be linearly transformed to obtain the text attention tensor corresponding to the text enhancement features. The contrast loss is then determined based on the similarity between the image attention tensor and the text attention tensor.
[0047] Specifically, the cosine similarity between the image attention tensor and the text attention tensor is calculated to obtain the image-text similarity matrix. A contrastive loss is then calculated based on this matrix, maximizing the similarity between matching images and text and minimizing the similarity between mismatched images and text, thereby achieving spatial alignment of multimodal features. Furthermore, the multi-path attention model introduces a tensor-based multimodal attention mechanism to simultaneously explore the multimodal relationships between different modalities. This allows for the natural fusion of multimodal information to calculate the multi-path attention tensor, thus comprehensively capturing the full range of many-to-many interaction paths between modalities, effectively improving the feature representation capability and processing efficiency of the multi-path attention model.
[0048] It is worth noting that the multi-path attention model, based on the tensor multi-path structure, can be flexibly extended to adaptable scenarios with any number of modalities, effectively breaking through the bottleneck of traditional models in terms of the number of modalities, and effectively improving the flexibility and scalability of the multi-path attention model.
[0049] In one possible implementation, the process of performing a linear transformation on the image enhancement features to obtain the corresponding image attention tensor can specifically involve: obtaining the query linear transformation matrix and the key linear transformation matrix; performing tensor decomposition on the query linear transformation matrix and the key linear transformation matrix along their respective dimensions to obtain two low-dimensional query factor tensors and two low-dimensional key factor tensors; and then performing linear transformations on the image enhancement features based on the query factor tensors and the key factor tensors to obtain the corresponding image attention tensor. The two query factor tensors and the two key factor tensors form a circular structure.
[0050] Specifically, the query linear transformation matrix and the key linear transformation matrix are linear transformation matrices used to perform linear transformations on image enhancement features to generate query vectors and key vectors, respectively. It can be understood that a matrix is mathematically defined as a two-dimensional structure; therefore, performing tensor decomposition along the matrix's dimensions yields two corresponding low-order factor tensors, which form a circular structure. For example, the matrix... Perform tensor decomposition on matrix A to obtain factor tensors. factor tensor , and Let be the rank of the tensor ring, and let the ring constraint satisfy... .
[0051] It should also be noted that the process of obtaining the text attention tensor by performing a linear transformation on the text enhancement features is similar to the process of obtaining the image attention tensor by performing a linear transformation on the image enhancement features, and will not be repeated here.
[0052] By performing tensor decomposition on the linear transformation matrix and constructing a ring-connected factor tensor structure, we can more accurately capture the intra-modal feature dependencies corresponding to each factor tensor, thus enhancing the intra-modal interaction effect. Simultaneously, tensor decomposition can break down a high-order matrix into a product of low-order factor tensors. Compared to directly calculating the high-order matrix, low-order factor tensors effectively reduce feature storage density and computational complexity, thereby improving the training efficiency of subsequent contrastive learning and achieving efficient multimodal contrastive learning.
[0053] In one possible implementation, the process of performing linear transformations on the image enhancement features based on the query factor tensor and key factor tensor to obtain the image attention tensor corresponding to the image enhancement features can specifically involve: performing linear transformations on the image enhancement features based on the two query factor tensors to obtain two corresponding query core tensors; performing a Kronecker product operation on the two query core tensors to obtain a query matrix; and reshaping the query matrix to obtain the query tensor, which is the query tensor corresponding to the image enhancement features. It can be calculated using the following formula, where, To enhance image features, and The query factor tensor corresponding to the image enhancement features. This indicates that the Khatri-Rao product is performed on the first dimension of the query core matrix, which is used to merge the first dimensions of the two query core matrices.
[0054]
[0055] Next, linear transformations are performed on the image enhancement features based on the two key factor tensors to obtain two corresponding key core tensors. A Kronecker product is then performed on the two key core tensors to obtain the key matrix. The key matrix is then reshaped to obtain the key tensors, which are the key tensors corresponding to the image enhancement features. It can be calculated using the following formula. and The key factor tensor corresponding to the image enhancement features.
[0056]
[0057] Finally, for the query tensor and key tensors Performing the Hadamard product yields the image attention tensor corresponding to the image enhancement features. It can be calculated using the following formula, where # represents the Hadamard product.
[0058]
[0059] Understandably, the process of obtaining the text attention tensor corresponding to the text enhancement features is similar to the process described above. Specifically, the query tensor corresponding to the text enhancement features... It can be calculated using the following formula. Enhance text features, and The query factor tensor corresponding to the text enhancement features.
[0060]
[0061] Key tensors corresponding to text augmentation features It can be calculated using the following formula. and The key factor tensor corresponding to the text enhancement features.
[0062]
[0063] Finally, for the query tensor Bond tensor Performing the Hadamard product yields the text attention tensor corresponding to the text enhancement features. It can be calculated using the following formula.
[0064]
[0065] In one possible implementation, the process of performing cross-attention processing on image enhancement features and text enhancement features to obtain image association features and text association features can specifically involve: performing a linear transformation on the image enhancement features to obtain the image query vector of the image enhancement features; performing a linear transformation on the text enhancement features to obtain the text key vector and text value vector of the text enhancement features; determining the similarity score between the image query vector and the text key vector; determining the cross-attention score based on the sum of the similarity score and the contrast loss; and weighting the text value vector based on the cross-attention score to obtain the image association features and text association features.
[0066] Specifically, self-attention processing is first applied to the image enhancement features and text enhancement features to obtain their respective image self-attention enhancement features and text self-attention enhancement features. Next, cross-attention processing is applied to these features. A linear transformation is performed on the image self-attention enhancement features to obtain the image query vector, and a linear transformation is performed on the text self-attention enhancement features to obtain the text key vector and text value vector. The similarity score between the image query vector and the text key vector is calculated. The cross-attention score is determined based on the sum of the similarity score and the contrast loss. The cross-attention score, AttentionScore, can be calculated using the following formula.
[0067] Where Q is the image query vector, K is the text key vector, and Z is the contrast loss. Finally, the text value vectors are weighted based on the cross-attention scores to obtain the text association features. The formula is as follows, where V is the text value vector.
[0068]
[0069] Similarly, a linear transformation is performed on the text self-attention enhancement features to obtain the text query vector, and a linear transformation is performed on the image self-attention enhancement features to obtain the image key vector and the image value vector. Based on the text query vector, the image key vector, and the image value vector, the image association features are obtained.
[0070] Step S204: Perform tensor shrinkage based on image association features and text association features to obtain target fusion features, and output rhinitis classification results based on target fusion features.
[0071] Specifically, image association features and text association features are concatenated to obtain multimodal fusion features. Image attention tensors and text attention tensors are concatenated to obtain contrastive fusion tensors. Tensor shrinking is performed on the multimodal fusion features and contrastive fusion features to integrate different modal attention features, resulting in high-dimensional target fusion features. By performing tensor shrinking on the multimodal fusion features and contrastive fusion features, cross-modal complementary information can be integrated through high-order tensor operations, and redundant features can be filtered out by using contrastive fusion features. At the same time, the feature discrimination of different modalities is enhanced to improve the semantic expression ability of the multi-path attention model, effectively improving the robustness of the multi-path attention model, thereby improving the accuracy of rhinitis classification results, and ultimately improving the reliability and efficiency of allergic rhinitis diagnosis.
[0072] Reference Figure 5 , Figure 5 This is a schematic diagram of an optional overall framework for a multimodal allergic rhinitis diagnostic method provided in this disclosure. The multimodal allergic rhinitis diagnostic method provided in this disclosure can be applied to allergic rhinitis diagnostic scenarios. Furthermore, this diagnostic method can be extended to other intelligent traditional Chinese medicine diagnostic methods, such as tongue diagnosis, pulse diagnosis, and facial diagnosis. The principle of the multimodal allergic rhinitis diagnostic method in this disclosure is described in its entirety from the perspective of the training process: The multimodal allergic rhinitis diagnosis method provided in this disclosure can be executed by a target model, which includes an image encoder, a text encoder, a Transformer encoder, and a multi-path attention model.
[0073] First, sample diagnostic images and sample diagnostic texts are obtained from the rhinitis dataset, such as... Figure 5 As shown, the sample diagnostic images are images of the face and tongue, and the sample diagnostic text is a description of the sample diagnostic images: "Pale red tongue with teeth marks, thin white tongue coating, deficiency of both Qi and Yin, Qi deficiency constitution." The model parameters of the image encoder and text encoder are frozen, and the sample diagnostic images are input into the image encoder for feature extraction to obtain the diagnostic image features. The sample diagnostic text is input into a text encoder for feature extraction to obtain diagnostic text features. These features are then input into a Transformer encoder to capture long-range dependencies, resulting in diagnostic text features that express long-range dependencies and contain contextual information. .
[0074] Next, the diagnostic image features and diagnostic text features are input into the multi-path attention model, and the diagnostic image features... After passing through a normalization layer and a multi-head self-attention layer, the diagnostic image features are compared. By fusing, intermediate features are obtained. Diagnostic image features and intermediate features After fusion, the intermediate features are obtained through a normalization layer. Based on the modality, select the corresponding visual expert network for intermediate features. Focus on processing and analyzing visual information, combining the features output by the visual expert network with intermediate features. Intermediate features are obtained after fusion. The intermediate feature F3 is then fed into the first iterative layer (self-attention layer and feedforward network) for 8 iterations to obtain the image enhancement features. Similarly, the diagnostic text features are sequentially processed through a normalization layer and a multi-head self-attention layer, then through a language expert network for focused processing and analysis of linguistic information before being input into the first iteration layer for eight iterations to obtain the text enhancement features. .
[0075] Next, image-text loss comparison learning is performed on the multi-channel attention model, specifically targeting image enhancement features. Or text enhancement features The query linear transformation matrix and the key linear transformation matrix are obtained. Tensor decomposition is then performed on both matrices along their respective dimensions to obtain two low-dimensional query factor tensors and two low-dimensional key factor tensors. Image enhancement features are then applied based on these query factor tensors and key factor tensors. Text enhancement features Perform a linear transformation to obtain the image attention tensor. and text attention tensor Based on image attention tensor and text attention tensor The contrast loss is determined by the similarity between the two, and the multi-path attention model is trained based on the contrast loss.
[0076] Next, image enhancement features are applied. and text enhancement features The input is processed through eight iterations in the second iteration layer (self-attention layer, cross-attention layer, and feedforward network). In the second iteration layer, image enhancement features are applied respectively. and text enhancement features Perform intra-modal attention and then enhance image features. and text enhancement features After intermodal attention, the image association features are obtained through a feedforward network. Text-related features .
[0077] Finally, image association features Textual association features Image attention tensor and text attention tensor Tensor contraction is performed to obtain target fusion features, and rhinitis classification results are obtained based on these features. The rhinitis classification results are then analyzed to diagnose and predict allergic rhinitis.
[0078] In summary, the multimodal-based allergic rhinitis diagnostic method provided in this disclosure ensures data accuracy and consistency by establishing a scientific sample database and screening and labeling clinical samples, thus providing high-quality input for subsequent target model training. During the target model training phase, the parameters of the image encoder and text encoder are frozen. For a small sample database, this approach leverages the generalization capabilities of the image and text encoders for feature extraction while reducing the number of training parameters and avoiding overfitting issues associated with small sample training. Furthermore, the intramodal and intermodal interactions of the multimodal data are learned to improve the target model's ability to model multimodal data, effectively enhancing the reliability and efficiency of allergic rhinitis diagnosis.
[0079] Reference Figure 6 , Figure 6 This is a schematic diagram of the structure of a multimodal allergic rhinitis diagnostic system 600 provided in an embodiment of the present disclosure. The multimodal allergic rhinitis diagnostic system 600 includes: The sample acquisition module 601 is used to acquire sample diagnostic images and sample diagnostic text of the diagnostic object; The feature extraction module 602 is used to extract features from the sample diagnostic image to obtain diagnostic image features, and to extract features from the sample diagnostic text to obtain diagnostic text features. The multimodal fusion module 603 is used to call the multi-path attention model to perform self-attention processing on the diagnostic image features and diagnostic text features respectively to obtain image enhancement features and text enhancement features. The image enhancement features and text enhancement features are then subjected to cross-attention processing to obtain image association features and text association features. During the training of the multi-path attention model, the contrast loss is determined based on the similarity between the image enhancement features and the text enhancement features. The classification module 604 is used to perform tensor shrinkage based on image association features and text association features to obtain target fusion features, and output rhinitis classification results based on target fusion features.
[0080] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in this disclosure and the foregoing drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or sequence. It should be understood that such data can be interchanged where appropriate to describe embodiments of this disclosure, for example, those that can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatuses.
[0081] It should be understood that in this disclosure, "at least one item" means one or more, and "more than one" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0082] It should be understood that in the description of the embodiments of this disclosure, "multiple" means two or more, "greater than", "less than", "exceeding" etc. are understood to exclude the number itself, and "above", "below", "within" etc. are understood to include the number itself.
[0083] In the embodiments provided in this disclosure, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0084] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0085] Furthermore, the functional units in the various embodiments of this disclosure can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0086] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this disclosure. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0087] It should also be understood that the various implementation methods provided in this disclosure can be combined arbitrarily to achieve different technical effects.
[0088] The above is a detailed description of the preferred embodiments of this disclosure. However, this disclosure is not limited to the above embodiments. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of this disclosure. All such equivalent modifications or substitutions are included within the scope defined by the claims of this disclosure.
Claims
1. A multimodal diagnostic method for allergic rhinitis, characterized in that, include: Obtain the sample diagnostic images and sample diagnostic text of the diagnostic object; Feature extraction is performed on the sample diagnostic image to obtain diagnostic image features, and feature extraction is performed on the sample diagnostic text to obtain diagnostic text features; A multi-path attention model is invoked to perform self-attention processing on the diagnostic image features and the diagnostic text features respectively to obtain image enhancement features and text enhancement features. Cross-attention processing is then performed on the image enhancement features and the text enhancement features to obtain image association features and text association features. During the training of the multi-path attention model, the contrast loss is determined based on the similarity between the image enhancement features and the text enhancement features. Tensor shrinkage is performed based on the image association features and the text association features to obtain target fusion features, and rhinitis classification results are output based on the target fusion features.
2. The multimodal-based diagnostic method for allergic rhinitis according to claim 1, characterized in that, During the training of the multi-path attention model, the image enhancement features and the text enhancement features are processed in tensor form. The step of determining the contrast loss based on the similarity between the image enhancement features and the text enhancement features includes: A linear transformation is performed on the image enhancement features to obtain the image attention tensor corresponding to the image enhancement features; a linear transformation is performed on the text enhancement features to obtain the text attention tensor corresponding to the text enhancement features. The contrast loss is determined based on the similarity between the image attention tensor and the text attention tensor.
3. The multimodal-based diagnostic method for allergic rhinitis according to claim 2, characterized in that, The step of performing a linear transformation on the image enhancement features to obtain the image attention tensor corresponding to the image enhancement features includes: Obtain the query linear transformation matrix and the key linear transformation matrix, and perform tensor decomposition on the query linear transformation matrix and the key linear transformation matrix along the dimensions respectively to obtain two low-dimensional query factor tensors and two low-dimensional key factor tensors, wherein the two query factor tensors form a ring structure and the two key factor tensors form a ring structure. Based on the query factor tensor and the key factor tensor, linear transformations are performed on the image enhancement features to obtain the image attention tensor corresponding to the image enhancement features.
4. The multimodal-based diagnostic method for allergic rhinitis according to claim 3, characterized in that, The step of performing linear transformations on the image enhancement features based on the query factor tensor and the key factor tensor respectively to obtain the image attention tensor corresponding to the image enhancement features includes: Linear transformations are performed on the image enhancement features based on the two query factor tensors to obtain two corresponding query core tensors. Kronecker product is performed on the two query core tensors to obtain a query matrix. The query matrix is then dimensionally reshaped to obtain a query tensor. Linear transformations are performed on the image enhancement features based on the two key factor tensors to obtain two corresponding key core tensors. Kronecker product is performed on the two key core tensors to obtain a key matrix. The key matrix is then dimensionally reshaped to obtain a key tensor. Perform a Hadamard product on the query tensor and the key tensor to obtain the image attention tensor corresponding to the image enhancement feature.
5. The multimodal-based diagnostic method for allergic rhinitis according to claim 1, characterized in that, Cross-attention processing is performed on the image enhancement features and the text enhancement features to obtain image association features, including: A linear transformation is performed on the image enhancement features to obtain the image query vector of the image enhancement features. A linear transformation is then performed on the text enhancement features to obtain the text key vector and text value vector of the text enhancement features. The similarity score between the image query vector and the text key vector is determined. A cross-attention score is determined based on the sum of the similarity score and the contrast loss. The text value vector is then weighted based on the cross-attention score to obtain the image association features.
6. The multimodal-based diagnostic method for allergic rhinitis according to claim 1, characterized in that, The step of extracting features from the sample diagnostic image to obtain diagnostic image features includes: The sample diagnostic image is segmented into multiple sample image blocks, and feature mapping is performed on the multiple sample image blocks to obtain the image block embedding vector corresponding to the image block; Self-attention processing is performed on multiple image patch embedding vectors to obtain image attention scores corresponding to the image patch embedding vectors. Based on the image attention scores, multiple image patch embedding vectors are weighted and summed to obtain diagnostic image features.
7. The multimodal-based diagnostic method for allergic rhinitis according to claim 1, characterized in that, The step of extracting features from the sample diagnostic text to obtain diagnostic text features includes: The sample diagnostic text is divided into multiple word texts, the word texts are converted into word text embedding vectors, and the word text embedding vectors are positionally encoded so that the word text embedding vectors retain their position information in the sample diagnostic text. Self-attention processing is performed on the word text embedding vectors after position encoding to obtain the text attention score corresponding to the word text embedding vectors. Based on the text attention score, multiple word text embedding vectors are weighted and summed to obtain diagnostic text features.
8. The multimodal-based diagnostic method for allergic rhinitis according to claim 1, characterized in that, Self-attention processing is applied to the diagnostic image features to obtain image enhancement features, including: A linear transformation is performed on the diagnostic image features to obtain an image query matrix, an image key matrix, and an image value matrix. The image query matrix, the image key matrix, and the image value matrix are then split into multiple query header matrices, multiple key header matrices, and multiple value header matrices. For any self-attention head, a similarity calculation is performed based on the query head matrix and the key head matrix. A self-attention score is obtained based on the similarity calculation result. The attention head feature of the self-attention head is determined based on the product of the self-attention score and the value head matrix. Multiple attention head features are concatenated to obtain image enhancement features.
9. A multimodal diagnostic system for allergic rhinitis, characterized in that, include: The sample acquisition module is used to acquire sample diagnostic images and sample diagnostic text of the diagnostic object; The feature extraction module is used to extract features from the sample diagnostic image to obtain diagnostic image features, and to extract features from the sample diagnostic text to obtain diagnostic text features; The multimodal fusion module is used to call a multi-path attention model to perform self-attention processing on the diagnostic image features and the diagnostic text features respectively to obtain image enhancement features and text enhancement features. The image enhancement features and the text enhancement features are then subjected to cross-attention processing to obtain image association features and text association features. During the training of the multi-path attention model, the contrast loss is determined based on the similarity between the image enhancement features and the text enhancement features. The classification module is used to perform tensor shrinkage based on the image association features and the text association features to obtain target fusion features, and output rhinitis classification results based on the target fusion features.
10. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the multimodal-based allergic rhinitis diagnostic method according to any one of claims 1 to 8.