System for positioning and answering surgical visual questions

By constructing an expert surgical hybrid model, utilizing the cross-modal feature extraction and fusion of the large language vision model and the small language vision model, and combining the MoE fusion module and the LoRA-MoE module, the problem that the existing system cannot provide location information is solved, and efficient positioning and answering of surgical vision questions are achieved, thereby improving the automation and accuracy of surgical guidance.

CN120671836APending Publication Date: 2025-09-19ZHEJIANG ACAD OF TRADITIONAL CHINESE MEDICINE

Patent Information

Application Number
CN202510782015.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing surgical visual question answering systems cannot provide positional information, and visual-language embedding has rarely been studied in surgery, resulting in a lack of effective automated guidance for medical students and junior surgeons during the learning process.

Method used

An expert surgical hybrid model consisting of a large language vision model and a small language vision model is constructed. Through cross-modal feature extraction and fusion, combined with the MoE fusion module and the LoRA-MoE module, the positioning and answering of visual questions are achieved.

Benefits of technology

It significantly improves the accuracy and reasoning ability of surgical visual question positioning and answering, and enhances the automation and personalized guidance capabilities of surgical visual question answering tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120671836A_ABST
    Figure CN120671836A_ABST
Patent Text Reader

Abstract

The invention discloses a system for positioning and answering surgical visual questions. The system is characterized in that an expert surgical hybrid model composed of a large language visual model and a small language visual model is constructed; in a surgical hybrid model, respectively extracting image features and text features by using the large language visual model, then performing cross-modal feature extraction and fusion on the image features and the text features to output features, and then converting the output features into text embedding; image features extracted by the large-language visual model and text embedding of the large-language visual model are obtained through the small-language visual model, then cross-modal feature extraction and fusion are conducted on the image features and the text embedding to output fusion features, and finally positioning and answering of surgical visual problems are conducted through the fusion features. According to the method, the reasoning ability, the accuracy and the visual consistency in the operation visual question-answering task can be enhanced, and the excellent question positioning and answering ability is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of robotic surgery, and in particular to a system for locating and answering surgical visual questions. Background Art

[0002] Recent advances in surgical visual question localization and answering (Surgical VQLA) have demonstrated significant potential for application in medical surgical robotics, addressing the urgent need for automated methods for personalized surgical guidance. However, this issue still poses significant challenges. For example, medical students and junior surgeons often rely on senior surgeons and experts to answer their questions during surgical learning, but experts are often busy with clinical and academic work and are unlikely to provide guidance. Existing deep learning-based surgical visual question answering (VQA) systems can only provide simple answers without providing the location of the answers. Furthermore, vision-language (ViL) embeddings have rarely been studied for this type of task. Therefore, a system that can provide visual question localization and answering (VQLA) would be very helpful for medical students and junior surgeons to learn and understand surgical videos. Summary of the Invention

[0003] The purpose of the present invention is to provide a system for locating and answering surgical visual questions. The present invention can enhance the reasoning ability, accuracy and visual consistency in surgical visual question-answering tasks, and has excellent question-locating and answering capabilities.

[0004] The technical solution of the present invention is as follows: a system for locating and answering surgical visual questions, which constructs an expert surgical hybrid model composed of a large language visual model and a small language visual model; in the surgical hybrid model, the large language visual model is used to extract image features and text features respectively, and then the image features and text features are cross-modal feature extracted and fused to output features, and then the output features are converted into text embeddings; the small language visual model is used to obtain the image features extracted by the large language visual model and the text embedding output by the large language visual model, and then the image features and text embedding are cross-modal feature extracted and fused to output fused features, and finally the fused features are used to locate and answer surgical visual questions.

[0005] In the above-mentioned system for surgical visual question localization and answering, the large language vision model includes a ViT network, a word segmenter, a hybrid cross attention module, a MoE fusion module, and an MoE model; the ViT network is used to extract basic image features; the word segmenter tokenizes the input text and then extracts basic text features; the hybrid cross attention module and the MoE fusion module sequentially perform cross-modal feature extraction and fusion to obtain output features; the MoE model is used to convert the output features into text embeddings;

[0006] The small language vision model includes a hybrid cross-attention module, a MoE fusion module, a decoder, a MoE position head and a MoE answer head; the hybrid cross-attention module and the MoE fusion module perform cross-modal feature extraction and fusion in turn to obtain fusion features, the decoder is used to receive the fusion features, and use the MoE position head and the MoE answer head respectively to locate and answer surgical vision questions.

[0007] The aforementioned system for surgical visual question localization and answering, the hybrid cross attention module consists of three layers of MoA, two of which are used to extract image and text features, and one layer is used for feature interaction. Each layer includes a routing network G and N attention experts; wherein, for the query vector q t , the routing network G selects a subset of attention experts and assigns weights, and the MoA output is expressed as:

[0008]

[0009] Where: i is the length of the routing network G, w i,t is the normalized weight, E i is the function for calculating the output of the attention expert, K ​​is the key feature input of the previous layer, V is the value feature input of the previous layer, and t is the routing time;

[0010] Among them, the routing network G calculates the normalized weight w i,t : First use the linear weight W g And softmax calculates the routing probability p at routing time t i :

[0011] p i,t =Softmax(q t ·Wg);

[0012] Where: W g represents linear weight;

[0013] Select the top k attention experts and renormalize their probabilities:

[0014]

[0015] Where: p j represents the renormalized routing probability;

[0016] Each attention expert uses the projection matrix W k 、W v 、 The attention weights are calculated as follows:

[0017]

[0018] Where: is the attention weight; d h is the dimension of the projection matrix, normalized by the dimension, T represents the mathematical operation, rank conversion;

[0019] The weighted sum of the values ​​is:

[0020]

[0021] Where: i,t is the intermediate output;

[0022] The final attention expert output is:

[0023]

[0024] The aforementioned system for surgical visual question localization and answering, the MoE fusion module inputs text features and image features, processes them through the attention network to learn the weights a of the attention expert i , each attention expert is obtained through a linear network and a tanh activation function; among them, for text features and image features, feature fusion is performed using the following formula:

[0025]

[0026] Where: y is the final fused feature, bilinear is the bilinear operator; Expert(I) i is the image feature, Expert(T) i is a text feature.

[0027] In the aforementioned system for surgical visual question localization and answering, the MoE model is provided with a LoRA-MoE module, which combines the hybrid expert MoE with the low-rank adapter LoRA to increase parameters without increasing the amount of computation. The LoRA-MoE module selects the network G(x) to generate weights for the input x:

[0028] G(x)=[G(x)1,...,G(x) n ];

[0029] The mixture of experts MoE output is a weighted sum of the attention experts’ outputs:

[0030]

[0031] Where: E i (x) is the output of attention experts;

[0032] Since only k<n experts are active during the computation, sparsity is ensured through Top-k routing;

[0033] The low-rank adapter LoRA reduces the dimensionality by A(x) and restores the dimensionality using B(x):

[0034] A(x)=xA, B(x)=xB, LoRA(x)=xAB

[0035] Where: A(x) and B(x) are calculation functions, d is the dimension of input A, r is the dimension of input A;

[0036] These modules are integrated by treating the low-rank adapter LoRA as an expert, whose output is:

[0037]

[0038] Where: LoRA i (x) is the low-rank adapter LoRA function;

[0039] Finally, convolution, activation, normalization, and projection operations are performed to convert the output features into text embeddings:

[0040] x=Linear(Conv1D(HardSwish(CoordDWConv(x0))))+x0;

[0041] Where: Linear is the linear layer, Conv 1D is the one-dimensional convolution layer, HardSwish is the Hard Swish activation function, CoordDWConv is the coordinate grouped convolution, and X0 is the input.

[0042] In the aforementioned system for surgical visual question localization and answering, the MoE location head uses a MoE network, a linear projection layer, and a sigmoid activation to model the coordinates of the bounding box, and uses a softmax activation to achieve classification.

[0043] In the aforementioned system for surgical vision question localization and answering, the MoE answer head implements answering through the MoE network and the linear projection layer.

[0044] The aforementioned system for surgical visual question positioning and answering, the MoE position head and MoE answer head use the GIoU loss function L GIoU Perform bounding box regression; among them, the position head MoE uses cross entropy loss L CE For positioning classification, the MoE answer head uses the L1 norm loss function for question answering.

[0045] Compared to existing technologies, our hybrid surgical model is a novel architecture that integrates large and small linguistic and visual models through a mixture of experts (MoE) strategy. This model significantly improves the performance of surgical visual question localization and answering. By introducing a LoRA-MoE module for fine-tuning, a hybrid cross-attention module (MCAM) for learning cross-modal features, an MoE fusion module for cross-feature interaction, and an MoE location head and MoE answer head for expert system decision-making, our model achieves superior results compared to existing state-of-the-art models. Experimental results demonstrate that our model outperforms other models on the corresponding datasets, achieving significant improvements in accuracy, F-score, and mean Intersection Over Union (MIOU), validating its great potential in the field of medical surgical robotics. Furthermore, related research confirms the important contributions of MoE-related modules, particularly the LoRA-MoE module and the MoE fusion module, which significantly impact model performance. These results demonstrate the effectiveness of our model in enhancing reasoning ability, accuracy, and visual consistency in surgical visual question answering tasks, paving the way for more advanced and automated solutions in personalized surgical guidance. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 is a schematic diagram of the principle of the expert surgical hybrid model of the present invention;

[0047] Figure 2 Schematic diagram of the principle of hybrid cross attention module;

[0048] Figure 3 Schematic diagram of the principle of MoA;

[0049] Figure 4 This is a schematic diagram of the principle of the MoE fusion module;

[0050] Figure 5 Figure 2 is a schematic diagram of the MoE location header and the MoE answer header. DETAILED DESCRIPTION

[0051] The present invention will be further described below with reference to the accompanying drawings and examples, but they are not intended to limit the present invention.

[0052] Example: A system for surgical visual question location and answering, such as Figure 1As shown, the steps are to construct a surgical mixed of experts (Surgical MoE) composed of a large language and vision model (LLVM) and a small language and vision model (SLVM); in the surgical mixed model, the large language and vision model is used to extract image features and text features respectively, and then the image features and text features are cross-modal extracted and fused to output features, and then the output features are converted into text embeddings; the small language and vision model is used to obtain the image features extracted by the large language and vision model and the text embedding of the large language and vision model, and then the image features and text embedding are cross-modal extracted and fused to output fused features, and finally the fused features are used to locate and answer surgical vision questions.

[0053] In this embodiment, the large language vision model includes a ViT network, a word segmenter, a hybrid cross attention module, a MoE fusion module and a MoE model; the ViT network is used to extract basic image features, and the ViT network (Vision Transformer) is a model based on the Transformer architecture, which is mainly used for image recognition and classification tasks. The ViT network divides the image into multiple small patches (patches), then linearly embeds these small patches into vectors of fixed dimension, and encodes them through a series of Transformer layers, and finally classifies them through a multi-layer perceptron (MLP) head; the word segmenter tokenizes the input text and then extracts basic text features using a Tokenizer. Tokenizer is a key tool in natural language processing (NLP), and its main function is to divide the input text data into smaller, easier-to-process units, usually words, subwords or characters. In the NLP model, Tokenizer is an important step in the text preprocessing process, which directly affects the input format and performance of the model; the Mixture of Cross-Attention Module (MCAM) and the MoE Fusion Module (MoE Fusion Module) perform cross-modal feature extraction and fusion in turn to obtain output features; the MoE model is used to convert the output features into text embeddings. The MoE model is Qwen1.5-MoE-A2.7B, which is the open source MoE model of the Tongyi Qianwen team.

[0054] The small language vision model includes a hybrid cross-attention module, a MoE fusion module, a decoder, a MoE position head and a MoE answer head; the hybrid cross-attention module and the MoE fusion module perform cross-modal feature extraction and fusion in turn to obtain fusion features, and the decoder (the decoder is DeiT) is used to receive the fusion features and use the MoE position head and MoE answer head respectively to locate and answer surgical vision questions.

[0055] In this embodiment, Figure 2 As shown in the figure, the hybrid cross attention module consists of three layers of MoA (Mixture of Attention), two of which are used to extract image and text features, and one layer is used for feature interaction. Each layer includes a routing network G and N attention experts; for the query vector q t , the routing network G selects a subset of attention experts and assigns weights to generate the output of MoA:

[0056]

[0057] Where: i is the length of the routing network G, w i,t is the normalized weight, E i is the function for calculating the output of the attention expert, K ​​is the key feature input of the previous layer, V is the value feature input of the previous layer, and t is the routing time;

[0058] like Figure 3 As shown, the routing network G calculates the normalized weight w i,t : First use the linear weight W g And softmax calculates the routing probability p at routing time t i :

[0059] p i,t =Softmax(q t ·Wg);

[0060] Where: W g represents linear weight;

[0061] Then select the top k attention experts and renormalize their probabilities:

[0062]

[0063] Where: p j represents the renormalized routing probability;

[0064] Each attention expert uses the projection matrix W k 、W v 、 The attention weights are calculated as follows:

[0065]

[0066] Where: is the attention weight; d h is the dimension of the projection matrix, normalized by the dimension, T represents the mathematical operation, rank conversion;

[0067] The weighted sum of the values ​​is:

[0068]

[0069] Where: i,t is the intermediate output; the final attention expert output is:

[0070]

[0071] Share W k and W v Allows pre-calculation of KW k and VW v , thus reducing computation and memory overhead.

[0072] In this embodiment, Figure 4 As shown, the MoE fusion module inputs text features and image features, which are processed by the attention network to learn the weights a of the attention expert. i , each attention expert is obtained through a linear network and a tanh activation function; among them, for text features and image features, feature fusion is performed using the following formula:

[0073]

[0074] Where: y is the final fused feature, bilinear is the bilinear operator; Expert(I) i is the image feature, Expert(T) i is a text feature.

[0075] In this embodiment, the MoE model is provided with a LoRA-MoE module, which combines the hybrid expert MoE with the low-rank adapter LoRA to increase parameters without increasing the amount of computation. For the input x, the LoRA-MoE module selects the network G(x) to generate the weight:

[0076] G(x)=[G(x)1,...,G(x)n];

[0077] The mixture of experts MoE output is a weighted sum of the attention experts’ outputs:

[0078]

[0079] Where: E i (x) is the output of attention experts;

[0080] Since only k<n experts are active during the computation, sparsity is ensured through Top-k routing;

[0081] The low-rank adapter LoRA reduces the dimensionality by A(x) and restores the dimensionality using B(x):

[0082] A(x)=xA, B(x)=xB, LoRA(x)=xAB

[0083] Where: A(x) and B(x) are calculation functions, d is the dimension of input A, r is the dimension of input A;

[0084] These modules are integrated by treating the low-rank adapter LoRA as an expert, whose output is:

[0085]

[0086] Where: LoRA i (x) is the low-rank adapter LoRA function;

[0087] Finally, convolution, activation, normalization, and projection operations are performed to convert the output features into text embeddings:

[0088] x=Linear(Conv1D(HardSwish(CoordDWConv(x0))))+x0;

[0089] Where: Linear is the linear layer, Conv 1D is the one-dimensional convolution layer, HardSwish is the Hard Swish activation function, CoordDWConv is the coordinate grouped convolution, and x0 is the input.

[0090] In this embodiment, Figure 5 As shown, the MoE location head uses the MoE network, linear projection layer and sigmoid activation to model the coordinates of the bounding box and uses softmax activation to achieve classification. The MoE answer head uses the MoE network and linear projection layer to achieve answering. Among them, the MoE location head and MoE answer head use the GIoU loss function L GIoU Perform bounding box regression; among them, the position head MoE uses cross entropy loss L CE For positioning classification, the MoE answer head uses the L1 norm loss function for question answering.

[0091] Based on the above model, this example conducts experiments on surgical visual question localization and answering. The EndoVis-18-VQLA dataset contains 14 robotic surgery videos from the MICCAI 2018 challenge. VQLA annotations provide instrument locations, bounding boxes, and question-answer pairs for each surgical step.

[0092] The EndoVis-17-VQLA dataset includes 97 frames of surgical videos from the MICCAI 2017 challenge, annotated with surgical actions and instrument positions. It serves as an external validation set to evaluate model generalization in various surgical scenarios.

[0093] The EndoVis Conversational Dataset combines the EndoVis-18-VQLA and EndoVis-17-VQLA datasets, with 19,020 training and 2,151 test image question-answering (QA) pairs. Through diverse QA tasks, it helps improve model understanding of surgical scenarios. In experiments, the MoE-LoRA-based Qwen1.5-MoE-A2.7B was fine-tuned using the EndoVis Conversational Dataset. The proposed surgical MoE was then trained using the training set of the EndoVis-18-VQLA dataset. The model was evaluated on the EndoVis-18-VQLA dataset and the test set of the entire EndoVis-17-VQLA dataset.

[0094] The performance comparison of different models on the ENDOVIS-18-VQLA and ENDOVIS-17-VQLA datasets is shown in Table 1:

[0095]

[0096] Table 1

[0097] The experiments were implemented using the Python-PyTorch framework and performed on an Nvidia10a100-GPU server. The epoch, batch size, and learning rate were set to 200, 32, and 1×10, respectively. -5. The evaluation metrics used in the experiments include accuracy, F-score, and mean intersection over union (mIoU). As shown in Table 1, the experimental results show that the surgical MOE of the present invention performs very well on the EndoVis-18-VQLA and EndoVis-17-VQLA datasets. On the EndoVis-18-VQLA dataset, the surgical MOE of the present invention outperforms all other models in terms of accuracy 0.7086, F-score (0.3790), and mIoU (0.8552), surpassing models such as GVLE LViT, surgical VQLA++, and surgical LVLM. In addition, on the EndoVis-17-VQLA dataset, the surgical MOE of the present invention also achieved impressive results with mIoU (0.7960), outperforming other competing models. These experimental results demonstrate that the surgical MOE significantly improves the reasoning ability, accuracy, and visual consistency of surgical visual question answering tasks, has obvious advantages over existing advanced methods, and demonstrates its strong potential.

[0098] In addition, to verify the effectiveness of MoE-related modules, related experiments were conducted by removing the MoE head, MoE fusion, MoA, and MoE LoRA from the network. Removing the MoE head means replacing the MoE structure in the MoE head with a simple fully connected network. Removing MoE fusion means replacing the entire MoE fusion module with a simple addition operation. Removing MoA means directly replacing the MoA in MCAM with self-attention. Finally, removing MoE LoRA means not fine-tuning the Qwen1.5-MoE-A2.7B model. The results are shown in Table 2:

[0099]

[0100] Table 2

[0101] From Table 2, we can observe the effectiveness of all MoE-related modules, especially MoE LoRA and MoE Fusion, which have the most significant impact on the performance of surgical MoE.

[0102] In summary, our hybrid surgical model is a novel architecture that integrates large and small linguistic and visual models through a mixture of experts (MoE) strategy. This model significantly improves the performance of surgical visual question localization and answering. By introducing a LoRA-MoE module for fine-tuning, a hybrid cross-attention module (MCAM) for learning cross-modal features, an MoE fusion module for cross-feature interaction, and an MoE location head and MoE answer head for expert system decision-making, our model achieves superior results compared to existing state-of-the-art models. Experimental results demonstrate that our model outperforms other models on the corresponding datasets, achieving significant improvements in accuracy, F-score, and mean Intersection Over Union (MIOU), validating its great potential in the field of medical surgical robotics. Furthermore, related research confirms the important contributions of MoE-related modules, particularly the LoRA-MoE module and the MoE fusion module, which significantly impact model performance. These results demonstrate the effectiveness of our model in enhancing reasoning ability, accuracy, and visual consistency in surgical visual question answering tasks, paving the way for more advanced and automated solutions in personalized surgical guidance.

Claims

1. A system for locating and answering surgical visual questions, characterized by: Includes an expert surgical hybrid model consisting of a large language-vision model and a small language-vision model; In the surgical hybrid model, the large language vision model is used to extract image features and text features respectively, and then the image features and text features are cross-modally extracted and fused to output features, and then the output features are converted into text embeddings; the small language vision model is used to obtain the image features extracted by the large language vision model and the text embedding output by the large language vision model, and then the image features and text embeddings are cross-modally extracted and fused to output fused features, and finally the fused features are used to locate and answer surgical vision questions.

2. The system for surgical visual question location and answering according to claim 1, characterized in that: The large language vision model includes a ViT network, a word segmenter, a hybrid cross attention module, a MoE fusion module and an MoE model; the ViT network is used to extract basic image features; the word segmenter tokenizes the input text and then extracts basic text features; the hybrid cross attention module and the MoE fusion module sequentially perform cross-modal feature extraction and fusion to obtain output features; the MoE model is used to convert the output features into text embeddings; The small language vision model includes a hybrid cross-attention module, a MoE fusion module, a decoder, a MoE position head and a MoE answer head; the hybrid cross-attention module and the MoE fusion module perform cross-modal feature extraction and fusion in turn to obtain fusion features, the decoder is used to receive the fusion features, and use the MoE position head and the MoE answer head respectively to locate and answer surgical vision questions.

3. The system for surgical visual question location and answering according to claim 2, characterized in that: The hybrid cross attention module consists of three layers of MoA, two of which are used to extract image and text features, and one layer is used for feature interaction. Each layer includes a routing network G and N attention experts; For the query vector q t , the routing network G selects a subset of attention experts and assigns weights, and the MoA output is expressed as: Where: i is the length of the routing network G, w i,t is the normalized weight, E i is the function for calculating the output of the attention expert, K ​​is the key feature input of the previous layer, V is the value feature input of the previous layer, and t is the routing time; Among them, the routing network G calculates the normalized weight w i,t : First use the linear weight W g And softmax calculates the routing probability p at routing time t i : p i,t =Softmax(q t ·W g ); Where: W g represents linear weight; Select the top k attention experts and renormalize their probabilities: Where: p j represents the renormalized routing probability; Each attention expert uses the projection matrix W k 、W v 、 The attention weights are calculated as follows: Where: is the attention weight; d h is the dimension of the projection matrix, normalized by the dimension, Indicates mathematical operations, rank conversion; The weighted sum of the values ​​is: Where: i,t is the intermediate output; The final attention expert output is:

4. The system for surgical visual question location and answering according to claim 2, characterized in that: The MoE fusion module inputs text features and image features, which are processed by the attention network to learn the weights a of the attention experts. i , each attention expert is obtained through a linear network and a tanh activation function; among them, for text features and image features, feature fusion is performed using the following formula: Where: y is the final fused feature, bilinear is the bilinear operator; Expert(I) i is the image feature, Expert(T) i is a text feature.

5. The system for surgical visual question location and answering according to claim 2, characterized in that: The MoE model is provided with a LoRA-MoE module, which combines the hybrid expert MoE with the low-rank adapter LoRA, and increases parameters without increasing the amount of computation. For the input x, the LoRA-MoE module selects the network G(x) to generate the weight: G(x)=[G(x)1,...,G(x) n ]; The mixture of experts MoE output is a weighted sum of the attention experts’ outputs: Where: E i (x) is the output of attention experts; Since only k<<n experts are active during the computation, sparsity is ensured through Top-k routing; The low-rank adapter LoRA reduces the dimensionality by A(x) and restores the dimensionality using B(x): A(x)=xA, B(x)=xB, LoRA(x)=xAB; Where: A(x) and B(x) are calculation functions, d is the dimension of input A, r is the dimension of input A; These modules are integrated by treating the low-rank adapter LoRA as an expert, whose output is: Where: LoRA i (x) is the low-rank adapter LoRA function; Finally, convolution, activation, normalization, and projection operations are performed to convert the output features into text embeddings: x=Linear(Conv1D(HardSwish(CoordDWConv(x0))))+x0; Where: Linear is the linear layer, Conv1D is the one-dimensional convolution layer, HardSwish is the Hard Swish activation function, CoordDWConv is the coordinate grouped convolution, and x0 is the input.

6. The system for surgical visual question location and answering according to claim 2, characterized in that: The MoE location head uses a MoE network, a linear projection layer, and a sigmoid activation to model the coordinates of the bounding box, and uses a softmax activation to implement classification.

7. The system for surgical visual question location and answering according to claim 6, characterized in that: The MoE answer head answers questions through the MoE network and linear projection layer.

8. The system for surgical visual question location and answering according to claim 7, characterized in that: The MoE position head and MoE answer head use the GIoU loss function LGIoU for bounding box regression; among them, the position head MoE uses the cross entropy loss LCE for positioning classification, and the MoE answer head uses the L1 norm loss function for question answering.

Citation Information

Patent Citations

  • Adapter-based large language model multi-modal lightweight fusion method and system

    CN118364066A

  • Multimodal large language model counterfeit information detection method introducing expert knowledge

    CN118606892A

Cited By

  • Medical question and answer method fusing cross-modal hybrid experts

    CN121189510A

  • A medical question and answer method fusing cross-modal mixed experts

    CN121189510B