Traditional Chinese medicine large model diagnosis and treatment method based on missing modal perception
Patent Information
- Application Number
- CN202611212717.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-11
- Publication Date
- 2026-10-09
AI Technical Summary
[0005]有鉴于此,为了解决现有中医辅助诊疗方法中大多假设模态数据均完整,或者多依赖静态替换或固定权重补偿,进而导致面对模态缺失场景的适用性不高的技术问题,本发明提出一种基于缺失模态感知的中医大模型辩证论治方法,该方法包括以下步骤:
[0010]本发明还提出了一种基于缺失模态感知的中医大模型辩证论治系统,该系统包括:编码模块、对齐模块、补全模块、融合模块和诊疗输出模块。
Smart Images

Figure CN122889342A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of big data reasoning, and in particular to a method and system for dialectical treatment based on a large-scale TCM model with missing modality perception. Background Technology
[0002] Traditional Chinese medicine (TCM) diagnosis and treatment, centered on the four diagnostic methods of "inspection, auscultation and olfaction, inquiry, and palpation," is essentially a comprehensive decision-making process based on multimodal information, encompassing heterogeneous data such as visual images of the tongue / face, temporal waveforms of the pulse, and textual consultations. With the rapid development of deep learning and multimodal large-scale model technology, research aimed at using artificial intelligence to assist TCM diagnosis and prescription generation has gradually become a hot topic.
[0003] However, existing studies generally model each diagnostic method as an independent task, lacking a unified multimodal representation and collaborative fusion mechanism. The complementary information of each modality cannot be fully explored, and the overall diagnostic accuracy is limited. At the same time, TCM diagnosis is highly dependent on personal experience, has a low degree of standardization, and high-quality medical resources are highly concentrated, making it difficult to guarantee the quality of diagnosis and prescription at the grassroots level.
[0004] Furthermore, TCM multimodal data suffers from severe scarcity, imbalance, and missing pairing issues. While text-based medical record data is relatively abundant, objective data such as tongue and pulse examinations are limited in scale, and strictly paired samples collected simultaneously from the four diagnostic methods are extremely scarce. In real clinical settings, especially at the grassroots level, modality missing is commonplace—many patients only have medical records and lack standard visual images or pulse diagnosis data. Existing multimodal methods all assume modality completeness, and their performance deteriorates sharply when faced with data imbalance and modality missing, severely hindering the practical implementation of such systems. Summary of the Invention
[0005] In view of this, in order to address the technical problem that most existing TCM auxiliary diagnostic methods assume that modal data is complete, or rely heavily on static replacement or fixed weight compensation, thus resulting in low applicability to scenarios with missing modalities, this invention proposes a TCM large-scale model dialectical treatment method based on missing modality perception. This method includes the following steps: In the data processing stage, adaptive feature encoding is first performed on the collected multi-source heterogeneous modal data to extract initial feature information. Then, using TCM syndrome types as a semantic reference, cross-modal alignment is performed on the extracted features to unify information from different modalities within the syndrome type semantic space. To address potential missing modal data, the system uses static cue signals to identify missing conditions and invokes a pre-trained model for semantic repair and supplementation to obtain complete feature data. Subsequently, the supplemented features undergo multimodal deep fusion to generate a comprehensive fused representation. Finally, driven by this fused representation, and utilizing a pre-trained generation strategy, a clearly structured TCM syndrome differentiation and treatment report that conforms to clinical standards is automatically output.
[0006] Furthermore, in the adaptive coding, corresponding domain-adaptive encoders are constructed for the three heterogeneous modalities of visual images, pulse signals, and consultation texts. The masked autoencoder (MAE) paradigm is adopted to uniformly map different forms of diagnostic information to a shared latent space of the same dimension. At the same time, the syndrome element perception representation is introduced at the encoding end so that visual, pulse and text features can retain local semantic information related to TCM syndrome differentiation.
[0007] Furthermore, using TCM syndrome types as cross-modal semantic anchors, this is expanded into a three-tiered structure of "syndrome elements - pathogenesis attributes - syndrome type," constructing a hierarchical dialectical anchor alignment mechanism, and utilizing this mechanism for cross-modal alignment. Hierarchical contrastive loss and differentiated semantic distance metrics drive the convergence of each modal feature to the corresponding syndrome type semantic space, achieving multimodal semantic unification under the condition of only weak pairing of "modal data + syndrome type annotation," providing a semantic foundation for subsequent fusion mechanisms.
[0008] Furthermore, a learnable modality missing cue is introduced to semantically fill in the missing modalities. Combined with a two-level dynamic cross-modal fusion mechanism that depends on the certificate type, fine-grained interaction is completed at the local granularity level, and the contribution of each modality is dynamically adjusted at the certificate type semantic level to achieve semantic-level feature completion under missing conditions.
[0009] Furthermore, by combining knowledge from the field of traditional Chinese medicine, a dual-reward reinforcement learning mechanism is constructed, which includes rewards for dialectical consistency and rewards for prescription rationality. The prescription generation strategy is trained by optimizing reinforcement learning preferences, and a structured dialectical treatment plan containing dialectical results, treatment principles, and prescription composition is output.
[0010] This invention also proposes a large-scale TCM diagnostic and treatment system based on missing modality perception, which includes: an encoding module, an alignment module, a completion module, a fusion module, and a diagnosis and treatment output module.
[0011] Based on the above scheme, this invention provides a method and system for dialectical treatment based on a large-scale TCM model with missing modality awareness. It changes the existing paradigm of modeling the four diagnostic methods (inspection, auscultation, inquiry, and palpation) in isolation. By constructing a unified multimodal representation and collaborative fusion mechanism, it fully explores the complementary information among heterogeneous data such as tongue appearance, pulse, and medical records, enabling the integration of multimodal information in TCM diagnosis and treatment to be realistically reproduced in the artificial intelligence system. Furthermore, through the semantic completion and missing information awareness mechanism of the pre-trained model, it can effectively integrate multi-source heterogeneous information even with limited data, reducing the dependence of traditional methods on strictly paired samples collected simultaneously from the four diagnostic methods. This characteristic allows the system to fully utilize limited objective data such as tongue and pulse, based on relatively abundant text-based medical records, expanding the boundaries of training data usage and exhibiting stronger data adaptability. Attached Figure Description
[0012] Figure 1 This is a flowchart of the steps of a TCM large-scale model dialectical treatment method based on missing modality perception according to the present invention; Figure 2 This is a comparison chart of the identification performance under different input conditions; Figure 3 This is a comparison chart of the recognition performance of different frequency evidence elements; Figure 4 This is a comparison chart of dialectical performance under different proportions of missing modes; Figure 5 This is a structural block diagram of a large-scale TCM diagnostic and treatment system based on missing modality perception, according to the present invention. Detailed Implementation
[0013] In addition to the issues mentioned in the background, TCM prescription generation is a structured generation task with strong domain-specific knowledge constraints. Existing methods mostly adopt the supervised fine-tuning (SFT) paradigm, treating prescription generation as a continuation of ordinary text. They fail to explicitly integrate compatibility standards, syndrome-prescription consistency, and drug contraindications into the generation process, which easily leads to various clinical irrationalities and makes it difficult to meet the safety requirements of actual auxiliary diagnosis and treatment.
[0014] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0015] It should be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings. Unless otherwise specified, the embodiments and features described in this application can be combined with each other.
[0016] It should be understood that the terms "system," "apparatus," "unit," and / or "module" used in this application are a method of distinguishing different components, elements, parts, sections, or assemblies at different levels. However, if other terms can achieve the same purpose, they may be replaced by other expressions.
[0017] Furthermore, flowcharts are used in this application to illustrate the operations performed by the system according to embodiments of this application. It should be understood that the preceding or following operations are not necessarily performed precisely in sequence. Instead, the steps can be processed in reverse order or simultaneously. Additionally, other operations can be added to these processes, or one or more steps can be removed from them.
[0018] Current TCM-assisted diagnostic and treatment technologies face three core bottlenecks: First, there is a lack of TCM semantic alignment and dynamic fusion mechanisms for multi-source heterogeneous modalities. Existing multimodal fusion methods often employ a uniform approach, making it difficult to dynamically adjust the contributions of different modalities based on the TCM dialectical context. Second, in scenarios with missing modalities, existing methods rely heavily on static replacement or fixed-weight compensation, failing to integrate prior knowledge and available modalities from areas such as syndrome differentiation for semantic-level completion. Third, the prescription generation process lacks consistency constraints between dialectical logic and prescription decision-making, easily leading to problems such as a disconnect between syndrome differentiation and prescription, and unreasonable combinations. Therefore, there is an urgent need for a multimodal dialectical treatment method that can operate stably under conditions of missing modalities and deeply integrate TCM dialectical knowledge into the prescription generation process, in order to systematically improve the accuracy of syndrome differentiation, the rationality of prescriptions, and clinical usability.
[0019] Reference Figure 1 This is a flowchart illustrating an optional example of the TCM large-scale model dialectical treatment method based on missing modality perception proposed in this invention. The method can be applied to computer devices, and the method proposed in this embodiment may include, but is not limited to, the following steps: Step S1: Acquire multimodal data and perform adaptive encoding to obtain feature information; Step S2: Using TCM syndrome types as semantic anchors, perform cross-modal alignment of feature information; Step S3: Based on static prompts, perform missing modality awareness on the aligned features and combine it with a pre-trained model for semantic completion; Step S4: Perform cross-modal fusion based on the completed feature data; Step S5: Using the fusion representation as input, output a structured dialectical governance report using a pre-trained generation strategy.
[0020] In some feasible embodiments, step S1, multimodal adaptive coding, specifically includes: The three heterogeneous modalities of visual images, pulse signals, and consultation text are mapped to the same dimension of shared latent space through a domain-adaptive encoder, and local semantic units at the syndrome element level are further extracted, laying the foundation for subsequent fine-grained alignment and dynamic fusion.
[0021] For the visual modality, the Vision Transformer (ViT) is used as the backbone network, and domain adaptation pre-training is performed on a public visual dataset, taking the input image as the core. Encoded as d-dimensional visual feature vectors The focus is on capturing fine-grained visual information that is highly relevant to dialectical thinking, such as the shape and color of the tongue.
[0022] For temporal signal modalities, a one-dimensional convolutional network (1D-CNN) is used to extract local morphological features of the pulse waveform, combined with a Transformer module to capture long-range temporal dependencies between signal periods, thus transforming the sequence... Encode as feature vector Effectively characterizes and provides diagnostic elements related to the mode of time-series signals.
[0023] Regarding text modality, the Chinese language model will be further pre-trained using TCM corpora (including structured knowledge bases such as clinical medical records, TCM books, and domain knowledge graphs), incorporating consultation records, chief complaints, and medical records. Encode into semantic feature vectors .
[0024] Self-supervised pre-training is performed independently using the Masked Autoencoder (MAE) paradigm, with the training loss defined as: Where M is a random mask, For each modality, a corresponding decoder is provided. The training process does not require cross-modal pairing annotations, but only relies on the domain corpus of each modality. That should complete the task.
[0025] In this part, the output end further sets up a syndrome element perception mapping layer, which maps the global features of each modality into local semantic units related to TCM syndrome differentiation: The aforementioned local semantic units serve as input for subsequent fine-grained local alignment and dynamic fusion of evidence conditions.
[0026] In some feasible embodiments, step S2, hierarchical dialectical anchor point cross-modal alignment, specifically includes: Traditional Chinese medicine syndrome types are constructed as a hierarchical set of shared anchor points. ,in, As anchor points for syndrome elements, they are used to characterize local diagnostic information such as tongue quality, tongue coating, pulse morphology, and symptoms. As anchor points for pathogenesis attributes, they are used to characterize the dialectical attributes such as cold and heat, deficiency and excess, exterior and interior, qi, blood and body fluids; As anchor points for syndrome types, they are used to represent the final candidate syndrome types. Each modal feature is first aligned to the syndrome elements and pathogenesis attribute layers, and then converged to the syndrome type semantic space, achieving cross-modal dialectical semantic unity under weak pairing conditions.
[0027] In the anchor point alignment loss at the certificate type layer, the InfoNCE contrastive learning objective is used to drive the convergence of each modality feature to its corresponding certificate type anchor point: in, Let i be the feature representation of sample i in mode v. Let be the anchor vector corresponding to the sample type, and τ be the temperature coefficient. This loss enables the semantic space of the type to have good discriminative power.
[0028] Furthermore, differentiated semantic distance metrics are set for different modalities: regional syndrome matching is used for visual modalities, temporal morphological matching is used for pulse modalities, and symptom semantic coverage is used for text modalities, so as to avoid insufficient dialectical semantic expression caused by using a uniform similarity for all modalities.
[0029] Total loss: In some feasible embodiments, step S3 receives the cross-modal alignment features and modal availability markers output by step S2, explicitly identifies and dialectically completes the missing modalities, and outputs completed feature data with fixed modal slots for step S4 to perform dynamic cross-modal fusion. Thus, step S3 is positioned between the semantic alignment of step S2 and the feature fusion of step S4, to avoid missing modalities directly causing incomplete fusion input or semantic shifts. The missing modality awareness and semantic completion in step S3 specifically include: To address the common occurrence of modality loss in real-world clinical settings, a learnable modality-missing prompt token is introduced. This token allows for explicit modeling of modality loss during the training and inference phases, enabling subsequent processing to perceive and respond to different modality combinations.
[0030] Define a binary availability tag for each modality of each patient sample. 1 indicates that sample i is available in mode v, and 0 indicates that the mode is missing. A set of learnable missing cue vectors is maintained independently for each mode. The initial replacement is represented as: Building upon the static cue padding described above, a semantic-level completion mechanism guided by syndrome prior is further introduced. First, based on the modality availability label, aligned features of the currently available modalities are aggregated to form an available modality context vector. Second, the similarity between this context vector and the symptom element anchors, pathogenesis attribute anchors, and syndrome anchors constructed in step S2 is calculated. After Softmax normalization, a syndrome probability distribution is obtained, and the top K syndromes with the highest probabilities are selected as candidate syndromes. Third, the corresponding anchors are weighted and summed using the candidate syndrome probabilities to obtain a syndrome prior vector. Finally, the available modality context vector, syndrome prior vector, and missing cue vector are input into the semantic completion network to update the missing modality representation. This transforms the missing modality from a fixed placeholder vector into a complete representation related to the current patient's dialectical context, providing a complete modality slot input with dialectical semantics for step S4.
[0031] In some feasible embodiments, the training process of the pre-trained model used in step S3 specifically includes: A three-stage learning strategy is adopted, which gradually expands from a single-modal text base to multi-modal joint training, and introduces modality random dropout enhancement in each stage to actively simulate the modality loss distribution in real clinical practice.
[0032] Using a structured text knowledge base as training corpus, the parameters of the visual and temporal signal encoders are frozen, and LoRA is used to fine-tune only the text-side encoder and language model, so that the model can establish a dialectical basis in the text space, while reducing the early dependence on multimodal paired data.
[0033] Weakly paired data with visual and textual annotations are introduced. The visual encoder is activated to participate in training based on textual instruction fine-tuning. Hierarchical dialectical anchor alignment loss and local granular alignment loss are introduced simultaneously to drive visual and textual modal alignment.
[0034] Fine-tuning is performed using a small amount of strictly paired complete trimodal data, while modality dropout augmentation is enabled. During training, arbitrary modalities are randomly masked with a certain probability. This enables the model to maintain stable semantic completion, dynamic fusion, and reasoning capabilities under both full-modal and modality-deficient inputs.
[0035] The objective of the three-stage joint pre-training is defined as a weighted combination of the loss terms, namely: in, This indicates the loss due to missing modal consistency. The weights for each loss can be dynamically adjusted during the training process according to the course stages.
[0036] A three-stage pre-training strategy based on a curriculum is adopted, which gradually expands from a single-modal text base to full-modal joint training, and is combined with modal random dropout enhancement to improve the robustness of the model under arbitrary modal subset inputs.
[0037] In some feasible embodiments, the cross-modal fusion in step S4 specifically includes: Step S4 receives the completed feature data output from step S3 and uses a two-level dynamic cross-modal fusion mechanism with syndrome type condition dependence to perform local interaction and modality contribution allocation, so that the model generates a unified dimension of patient fusion representation under different available modal combinations.
[0038] Specifically, the first layer is a local fine-grained interaction layer, which performs cross-modal attention using text words, image regions, and pulse time sequence segments as basic units to capture the local correspondence between symptom descriptions, tongue regions, and pulse waveforms, forming local interactive representations.
[0039] The second layer is the dynamic fusion layer of syndrome conditions. It obtains the candidate syndrome distribution calculated by the available modal context and hierarchical dialectical anchor points in step S3, and performs weighted aggregation on the corresponding candidate syndrome anchor points according to the candidate syndrome distribution to obtain the syndrome condition vector. Based on the syndrome condition vector, it dynamically calculates the contribution weights of visual, pulse and text modalities to the current dialectical judgment, and uses the contribution weights to generate the patient fusion representation.
[0040] The cross-modal attention fusion process is represented as: in, For the modal dynamic fusion weights generated under the proof condition, This serves as the anchor guide. When a complete modality is input, the model integrates information from multiple sources to complete the fusion; when a modality is missing, the model reduces the impact of missing or low-confidence modalities based on semantic completion representations and dynamic weights.
[0041] In some feasible embodiments, the training process of the generation strategy in step S5 specifically includes: Using the SFT model as the initial policy network and the patient's multimodal fusion representation and candidate syndrome semantics as state inputs, the generation policy is trained through reinforcement learning preference optimization with dialectical-prescription consistency constraints, so that the output is synergistically improved in multiple dimensions such as dialectical logic, prescription rationality and medication safety.
[0042] The first level of reward is the dialectical consistency reward, which evaluates the consistency between the syndrome type, pathogenesis, treatment principles and methods in the generated results and the multimodal fusion representation of patients, based on the hierarchical dialectical anchor space: ; The second level of reward is the prescription rationality reward, which assesses the consistency between the prescription composition and the diagnostic conclusions, treatment methods, and principles based on the drug pair / prescription diagram, the principal-assistant-adjuvant structure, and the compatibility rules: ; Security constraints: ; Composite reward output: ; Where s is the current state, i.e., the multimodal fusion representation and candidate syndrome semantics; a is the strategy action, i.e. the generated dialectical treatment plan; and G is the drug pair relationship graph. The target proof type anchor vector; This is a violation indicator function.
[0043] Strategy optimization goal: ; Furthermore, the reasoning process is similar to the training process, so it will not be described.
[0044] To verify the effectiveness of semantic space and cross-modal fusion, such as Figure 2 As shown, in a dual-modal validation embodiment, 300 samples with tongue image and text description were selected. The text features and tongue image visual features were mapped to a unified evidence element space, and the results were evaluated on a subset of 27 observable tongue image evidence elements. The evaluation metrics used were multi-label micro-F1 and macro-F1, where micro-F1 reflects the overall sample-label prediction performance, and macro-F1 is calculated separately for each evidence element and then averaged to reflect the recognition ability of long-tail evidence elements.
[0045] In the verification embodiment, the micro-F1 score for plain text input was 0.3801 and the macro-F1 score was 0.1964; the micro-F1 score for pure visual input was 0.7331 and the macro-F1 score was 0.6765; after cross-modal fusion using the evidence type semantic condition, the micro-F1 score was 0.7348 and the macro-F1 score was 0.6911. The fusion result improved the macro-F1 score by 0.4947 compared to the plain text input and by 0.0146 compared to the pure visual input, indicating that even with an imbalance in the quality of text and visual information, the method can adjust the modal contribution based on the evidence type semantic condition, ensuring that the fused representation at least maintains and slightly outperforms the evidence element recognition ability of a stronger single modality. The recognition ability was compared after grouping evidence elements according to their frequency of occurrence in the data, such as... Figure 3 As shown, the recognition ability of plain text input on long-tailed syndrome elements is almost ineffective, while the average F1 of long-tailed syndrome elements is significantly improved after cross-modal fusion, indicating that the method can alleviate the recognition difficulties caused by the long-tailed distribution of TCM syndrome elements.
[0046] Furthermore, missing modality simulation can be performed through modal random masking: visual modality, pulse modality, or text modality are masked according to preset ratios, and the results are compared between three settings: no completion, using only fixed missing prompts, and using the prior knowledge of the evidence type of this invention to guide semantic completion. The F1 score for evidence element recognition, the accuracy of evidence type recognition, and the consistency between the complete modality output and the missing modality output are statistically analyzed. This simulation is used to independently evaluate the completion effect of step S3 under different missing rates and different modality combinations.
[0047] go through Figure 4 Experimental analysis shows that the performance of all schemes decreases synchronously as the modality missing ratio increases, while the decrease in the proposed method is significantly smaller; when the missing ratio is around 80%, its certificate type recognition accuracy is improved by 0.25 compared to the comparative scheme 1. This demonstrates that relying on the distribution of candidate certificate types and the available modal context to achieve semantic completion is superior to static replacement and fixed weight compensation, and can significantly improve the robustness of modality missing scenarios.
[0048] like Figure 5 As shown, a large-scale TCM diagnostic and treatment system based on missing modality perception includes: The encoding module is used to execute step S1; Alignment module, used to perform step S2; The completion module is used to execute step S3; The fusion module is used to execute step S4; The diagnosis and treatment output module is used to execute step S5.
[0049] The content of the above method embodiments is applicable to this system embodiment. The specific functions implemented in this system embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.
[0050] A large-scale TCM diagnostic and treatment device based on missing modality perception: At least one processor; At least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements the TCM large-scale model dialectical treatment method based on missing modality perception as described above.
[0051] The content of the above method embodiments is applicable to the device embodiments. The specific functions implemented by the device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0052] A storage medium storing processor-executable instructions, which, when executed by a processor, are used to implement the diagnostic and treatment method of a large-scale TCM model based on missing modality perception as described above.
[0053] The content of the above method embodiments is applicable to this storage medium embodiment. The specific functions implemented in this storage medium embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.
[0054] The above is a detailed description of the preferred embodiments of the present invention. However, the present invention is not limited to the embodiments described. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of the present invention. All such equivalent modifications or substitutions are included within the scope defined by the claims of this application.
Claims
1. A TCM large-scale model-based dialectical treatment method based on missing modality perception, characterized in that, Includes the following steps: Multimodal data is acquired and adaptively encoded to obtain feature information; Using TCM syndrome types as semantic anchors, the feature information is cross-modal aligned to obtain aligned features; Based on static prompts, missing modal awareness is performed on the aligned features, and semantic completion is performed in combination with a pre-trained model to obtain the completed feature data. Cross-modal fusion is performed based on the completed feature data to obtain a fused representation; Using the fused representation as input, a structured dialectical governance report is output using a pre-trained generation strategy.
2. The method for syndrome differentiation and treatment based on a large-scale TCM model using missing modality perception as described in claim 1, characterized in that, The step of acquiring multimodal data and performing adaptive encoding specifically includes: Acquire visual image data, pulse diagnosis signal data, and medical history text data; Taking into account the shape and color of the tongue, the visual image data is encoded into visual features; Considering the local morphological features of the pulse waveform, the pulse diagnosis signal is encoded into sequence features; The consultation text data is encoded into semantic features based on a large language model; The visual features, sequence features, and semantic features are mapped to local semantic units related to TCM syndrome differentiation to obtain feature information.
3. The method for syndrome differentiation and treatment based on a large-scale TCM model using missing modality perception as described in claim 2, characterized in that, The step of performing cross-modal alignment of the feature information using TCM syndrome types as semantic anchors specifically includes: The TCM syndrome types are constructed as hierarchical shared anchor points, which include symptom element anchor points, pathogenesis attribute anchor points, and syndrome type anchor points; Set a differential semantic distance metric; Based on the symptom element anchors, the pathogenesis attribute anchors, the syndrome type anchors, and the differential semantic distance metric, the feature information is aligned across modalities.
4. The method for syndrome differentiation and treatment based on a large-scale TCM model using missing modality perception as described in claim 3, characterized in that, The alignment loss at the type layer is learned using InfoNCE contrastive learning.
5. The method for syndrome differentiation and treatment based on a large-scale TCM model using missing modality perception as described in claim 3, characterized in that, The differential semantic distance metric includes: For visual features, regional syndrome matching is used; for sequence features, temporal morphological matching is used; and for semantic features, symptom semantic coverage is used.
6. The method for syndrome differentiation and treatment based on a large-scale TCM model using missing modality perception as described in claim 1, characterized in that, The process of combining a pre-trained model for semantic completion specifically includes: The aligned features corresponding to the current available modal are masked and aggregated according to the modal availability label to obtain the available modal context vector; The semantic similarity of the available modal context vectors with the symptom element anchors, pathogenesis attribute anchors, and syndrome type anchors is calculated respectively, and the semantic similarity is normalized to obtain the probability value of each syndrome type. Based on the probability values, select the certificate type and the corresponding certificate type anchor point to form a candidate certificate type distribution; Based on the distribution of candidate evidence types, the corresponding hierarchical dialectical anchor points are weighted and aggregated to obtain the evidence type prior vector. The available modality context vector, the proof type prior vector, and the missing cue vector corresponding to the missing modality are input into a pre-trained semantic completion network to update the missing modality representation and generate a complete representation.
7. The method for syndrome differentiation and treatment based on a large-scale TCM model using missing modality perception as described in claim 6, characterized in that, The step of performing cross-modal fusion based on the completed feature data specifically includes: Based on the completed feature data, cross-modal attention is performed using text words, image regions, and pulse time sequence segments as basic units to capture the local correspondence between symptom descriptions, tongue regions, and pulse waveforms, forming local interactive representations. The candidate certificate type distribution is weighted and aggregated with the candidate certificate type anchor points to obtain the certificate type condition vector; Using the aforementioned evidence condition vector as semantic conditions, the contribution weights of visual features, temporal features, and semantic features to the current dialectical judgment are dynamically calculated. A fusion representation is generated based on the local interaction features and the contribution weights.
8. The method for syndrome differentiation and treatment based on a large-scale TCM model using missing modality perception as described in claim 1, characterized in that, The pre-training process of the generation strategy includes: A composite reward system is established, which includes a reward for dialectical consistency and a reward for rational prescription. Using the multimodal fusion representation and candidate certificate semantics in the dataset as state input, and combining the composite reward, security constraints and corresponding labels, a generation strategy is trained.
9. A large-scale TCM diagnostic and treatment system based on missing modality perception, characterized in that, include: The encoding module is used to acquire multimodal data and perform adaptive encoding to obtain feature information; The alignment module is used to perform cross-modal alignment of the feature information with TCM syndrome types as semantic anchors to obtain aligned features. The completion module performs missing modality detection on the aligned features based on static prompts, and performs semantic completion in combination with a pre-trained model to obtain the completed feature data. The fusion module performs cross-modal fusion based on the completed feature data to obtain a fused representation; The diagnosis and treatment output module is used to output a structured dialectical treatment report by taking the fused representation as input and using a pre-trained generation strategy.
10. A large-scale TCM diagnostic and treatment device based on missing modality perception, characterized in that, include: At least one processor; At least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements the TCM large-scale model dialectical treatment method based on missing modality perception as described in any one of claims 1-8.