An auxiliary decision system for traditional Chinese medicine based on semantic topological feature fusion and hidden variable constraint
Patent Information
- Application Number
- CN202611281231.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-24
- Publication Date
- 2026-09-22
AI Technical Summary
[0002]当前领域内广泛采用的深度学习模型在中医智能决策中存在技术缺陷:第一,单一数据驱动模型存在逻辑黑盒与辨证逻辑断层问题
1、解决了深度学习模型在中医决策中的逻辑黑盒问题。
Smart Images

Figure CN122800205A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the fields of natural language processing, knowledge graph and traditional Chinese medicine informatics, and specifically refers to a traditional Chinese medicine auxiliary decision-making system that integrates semantic topological features and latent variable constraints. Background Technology
[0002] The deep learning models widely used in the field currently suffer from technical deficiencies in TCM intelligent decision-making: First, single data-driven models suffer from logical black boxes and gaps in diagnostic logic. Traditional end-to-end architectures forcibly reduce the complex process of diagnosis and treatment to a flattened multi-label classification task, lacking explicit modeling of the TCM diagnostic logic of "cause, mechanism, symptom, and treatment." This opaque reasoning process not only disrupts the interconnected reasoning chain in TCM but also easily generates logical illusions that contradict common medical sense. Second, the knowledge supervision signals for intermediate logical links are sparse. In ancient TCM texts and real medical cases, there are often only descriptions of symptoms, prescriptions, and part of the pathogenesis, while the "treatment method," which is the key link connecting the pathogenesis and prescription, is often omitted or implicit. This sparsity of labels makes it difficult for existing models to capture the complete logical flow during training, seriously affecting the accuracy of feature aggregation and the coherence of the reasoning chain. Third, the ability to mine nonlinear implicit associations in complex clinical cases is insufficient. Discrete standardized entities often lose the syntactic structure and contextual dependencies of the original text, while traditional graph neural networks have difficulty perceiving personalized symptom input in real time during initialization, resulting in limited generalization ability of the system when faced with incomplete information or atypical complex cases. Summary of the Invention
[0003] To address the technical problems existing in the prior art, this invention provides a TCM auxiliary decision-making system based on semantic topological feature fusion and latent variable constraints. The system includes a knowledge-enhanced dual-tower deep reasoning architecture, which includes two parallel feature extraction channels forming a horizontal parallel structure. One side is the BERT semantic perception tower for processing clinical case texts, which uses a pre-trained language model to encode unstructured expressions into high-dimensional semantic vectors to capture implicit context. On the other side is a GNN structure inference tower for processing graph topology. Multi-source knowledge graphs are introduced as prior constraints to construct a heterogeneous graph network including four types of entities: symptoms, pathogenesis, treatment, and prescriptions. A hybrid feature weighting strategy is adopted. Highly personalized symptom nodes are dynamically initialized and the text vectors extracted in real time by the semantic perception tower are directly reused. For pathogenesis, treatment, and prescription nodes that represent objective medical laws, a learnable knowledge graph embedding matrix is initialized based on static topological priors to achieve the bottom-level fusion of heterogeneous spaces. A bridge mechanism based on latent variables is designed in the topology construction to retain treatment nodes as implicit semantic hubs. A two-layer deep heterogeneous graph convolution fusion graph attention mechanism is adopted. By dynamically allocating logical weights, the prescription nodes at the end are forced to aggregate the symptom, pathogenesis, and treatment features of the upstream across levels. Even in the absence of explicit textual mentions, the entire logical loop of TCM reasoning can still be completed in the latent space. In terms of vertical reasoning logic, a three-layer architecture of perception layer, reasoning layer and decision layer is constructed. The perception layer is responsible for capturing and representing multi-source heterogeneous data; the reasoning layer is responsible for mining and reconstructing latent pathogenesis features; and the decision layer is responsible for integrating and outputting multimodal information. Based on the fusion features, a cascaded prediction head for pathogenesis, treatment method and prescription is designed. The treatment method prediction head does not output prediction results and exists as a logical middleware. The pathogenesis and prescription prediction heads output the predicted pathogenesis and prescription respectively. In constructing the total loss function, a knowledge graph consistency loss is introduced. The static adjacency matrix of the knowledge graph is used to detect the logical connectivity of the prediction distribution. Orthogonal penalties are applied to outputs that are interrupted by the path, or where the predicted pathogenesis and prescription are not related in the graph. This forces the model to automatically search for a solution space that is connected in accordance with medical theory during gradient descent, thereby reducing the logical illusion of the neural network internally.
[0004] The beneficial effects of the technical solution provided by this invention include at least the following: 1. It solves the problem of the logical black box in decision-making of traditional Chinese medicine using deep learning models.
[0005] By designing a semantic-topological dual-tower architecture, text semantic perception and graph structure reasoning are organically combined. The progressive logic of "symptom-mechanism-treatment-prescription" in traditional Chinese medicine is explicitly modeled inside the neural network, making the model's reasoning process traceable and logically consistent, and effectively suppressing logical illusions that violate medical common sense.
[0006] 2. It has broken through the technical bottleneck of sparse treatment method labels.
[0007] By using a topological bridge mechanism based on latent variables, the treatment node is designed as an implicit semantic hub rather than an independent prediction target, so that the gradient can update the embedding vector of the treatment node through backpropagation of the prescription node. Even without explicit treatment annotation, the model can still maintain a complete inference chain.
[0008] 3. Significantly improved generalization ability in complex clinical scenarios.
[0009] The hybrid node initialization strategy that combines dynamic and static elements enables the model to perceive personalized symptom input in real time while relying on the objective medical laws of static knowledge graphs, thus exhibiting stronger robustness when facing incomplete information or atypical complex cases.
[0010] 4. It achieves an organic unity between soft logic constraints and data-driven approaches.
[0011] The introduction of consistency loss in knowledge graphs imposes medical logic regularization constraints at the loss function level without relying on external rule engines. This forces the model to spontaneously converge to a connected solution space that conforms to traditional Chinese medicine theory during training, achieving a dual-drive approach of "data-driven + knowledge-constrained". Attached Figure Description
[0012] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0013] Figure 1 This is a block diagram of a TCM auxiliary decision-making system based on semantic topological feature fusion and latent variable constraints provided by an embodiment of the present invention; Figure 2 This is a schematic diagram of the processing procedure (semantic text feature extraction) of the BERT semantic perception tower provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the processing procedure (deeply stacked full-chain feature aggregation) of the GNN structure inference tower provided in this embodiment of the invention; Figure 4 This is a schematic diagram of the ternary feature-level deep fusion and cascaded prediction head provided in an embodiment of the present invention. Detailed Implementation
[0014] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.
[0015] The specific implementation of this invention mainly relies on a computer system to execute the above-mentioned technical solution. The following is an example of coronary heart disease.
[0016] like Figure 1As shown, this embodiment of the invention provides a TCM auxiliary decision-making system based on semantic topological feature fusion and latent variable constraints. The system includes a knowledge-enhanced dual-tower deep reasoning architecture, which includes two parallel feature extraction channels to form a horizontal parallel structure. One side is the BERT semantic perception tower for processing clinical case texts, which uses a pre-trained language model to encode unstructured expressions into high-dimensional semantic vectors to capture implicit context. On the other side is a GNN structure inference tower for processing graph topology. Multi-source knowledge graphs are introduced as prior constraints to construct a heterogeneous graph network including four types of entities: symptoms, pathogenesis, treatment, and prescriptions. A hybrid feature weighting strategy is adopted. Highly personalized symptom nodes are dynamically initialized and the text vectors extracted in real time by the semantic perception tower are directly reused. For pathogenesis, treatment, and prescription nodes that represent objective medical laws, a learnable knowledge graph embedding matrix is initialized based on static topological priors to achieve the bottom-level fusion of heterogeneous spaces. A bridge mechanism based on latent variables is designed in the topology construction to retain treatment nodes as implicit semantic hubs. A two-layer deep heterogeneous graph convolution fusion graph attention mechanism is adopted. By dynamically allocating logical weights, the prescription nodes at the end are forced to aggregate the symptom, pathogenesis, and treatment features of the upstream across levels. Even in the absence of explicit textual mentions, the entire logical loop of TCM reasoning can still be completed in the latent space. In terms of vertical reasoning logic, a three-layer architecture of perception layer, reasoning layer and decision layer is constructed. The perception layer is responsible for capturing and representing multi-source heterogeneous data; the reasoning layer is responsible for mining and reconstructing latent pathogenesis features; and the decision layer is responsible for integrating and outputting multimodal information. Based on the fusion features, a cascaded prediction head for pathogenesis, treatment method and prescription is designed. The treatment method prediction head does not output prediction results and exists as a logical middleware. The pathogenesis and prescription prediction heads output the predicted pathogenesis and prescription respectively. In constructing the total loss function, a knowledge graph consistency loss is introduced. The static adjacency matrix of the knowledge graph is used to detect the logical connectivity of the prediction distribution. Orthogonal penalties are applied to outputs that are interrupted by the path, or where the predicted pathogenesis and prescription are not related in the graph. This forces the model to automatically search for a solution space that is connected in accordance with medical theory during gradient descent, thereby reducing the logical illusion of the neural network internally.
[0017] Optionally, the BERT semantic perception tower is used to perceive features and is responsible for processing the four diagnostic methods information of the text structure in clinical cases. It adopts a BERT-Base-Chinese pre-trained model to encode the complex and ambiguous unstructured expressions of traditional Chinese medicine into continuous high-dimensional semantic vectors, capturing the implicit contextual dependencies between entities. The high-dimensional semantic vectors serve as the representation of the text modality and are projected into the initial features of different nodes in the subsequent graph neural network, realizing semantic injection from the text space to the graph space.
[0018] Optionally, in the pre-training stage, the BERT semantic perception tower, BERT-Base-Chinese, is trained unsupervised on a large-scale Chinese Wikipedia corpus. It employs a character-level Word Piece vocabulary for word segmentation, treating each Chinese character in the input sequence as an independent basic unit token. Joint pre-training is performed through two main tasks: Masked Language Model (MLM) and Next Sentence Prediction (NSP). This effectively captures the deep bidirectional semantic features of Chinese characters in different contexts. The model's training data is a pre-processed set of standardized structured entities. While this discrete, standardized entity model eliminates lexical ambiguity, it loses the syntactic structure and modification relationships (e.g., entity combinations) of the original text. and Although they share the core entity of "chest pain," they point to two completely different pathogenesis directions: "cold congealing the heart vessels" and "blood stasis obstruction." Therefore, the role of this module is to use a deep attention mechanism to reconstruct the context of standardized symptom entity sequences and capture the underlying pathogenesis indicators behind symptom combinations.
[0019] Optionally, such as Figure 2 As shown, the processing procedure of the BERT semantic awareness tower specifically includes: (1) Entity serialization construction based on symptom nodes Due to the highly personalized nature of symptoms, a text serialization strategy based on symptom nodes is designed to focus on constructing high-density symptom representations. This strategy utilizes functionally specific control tags to reassemble discrete symptom entities into pseudo-sentences rich in pathological features. Let there be a training sample Including standardized symptom sets The serialization construction process is formally described as follows: in: As a global classification label, it is used to aggregate the symptom cluster features of the entire case, serving as a core vector for inferring the underlying pathogenesis; correspond Labels, serving as functional cues, guide the model to focus on clinical manifestation features during multi-head attention computation; Use comma separators to separate to The connection enables the model to learn the nonlinear ability to directly map from "symptom combination patterns" to "core pathogenesis"; This is the end marker; For example, a typical sample of ancient book passages is serialized as follows: The pre-standardized entity extraction of training data and the BERT context reconstruction in this step are not redundant work, but rather form a complementary relationship of "semantic dimensionality reduction - logical dimensionality enhancement": the previous standardization work effectively eliminates the expressive noise in natural language, compressing high-dimensional sparse free text into a set of core pathological features with high signal-to-noise ratio, avoiding the model from wasting computing power on processing variations of traditional Chinese medicine language. However, the combination of entities contains infinite changes in pathogenesis. BERT uses positional encoding to distinguish the primary and secondary relationships in the symptom list, representing the distribution characteristics of symptom clusters; and uses the attention mechanism to enhance the learning of the co-occurrence intensity between symptom entities. (2) Multi-head attention mechanism and context reconstruction The Transformer Encoder's multi-head self-attention mechanism is used to reconstruct the context of the serialized pseudo-sentences. In the high-dimensional semantic space, there are strong nonlinear interactions between entities. The model calculates the context of the entities. With entity By assigning attention weights between these elements and capturing implicit associations, this mechanism achieves two semantic functions: feature enhancement (e.g., when "dark purple tongue" (an indicator of blood stasis) appears, the model uses the attention mechanism to shift the semantic vector of "chest pain" in the same sequence toward "stabbing pain") and noise suppression (for some non-specific symptoms, such as "poor appetite," if there is a lack of associated spleen and stomach disease indicators in the context, the attention mechanism will automatically reduce its weight in the global representation, thereby focusing on the core symptoms of coronary heart disease). Through a deep stacking of 12 encoder layers, discrete entity symbols are transformed into a dense vector sequence that includes combined contextual information. (3) Global semantic aggregation and output representation After BERT encoding, the model obtains the last hidden state ( ), perform pooling operations to obtain a fixed-length vector for downstream classification and graph neural network initialization: Select the first one of the sequences The hidden vector corresponding to the label serves as the global semantic representation of the entire case. : in This represents the output of the last layer of the BERT model, corresponding to After the hidden state vector of this special marker has been fully interacted with by the multi-head self-attention mechanism of 12 layers of Transformer, the vector at this position has absorbed and aggregated the global contextual information of all symptom entities and their modification relationships in the current case. and These represent the learnable weight matrix and bias term of the pooling layer, respectively. This step performs a linear feature space transformation, the purpose of which is to map the raw features extracted by BERT to a specific dimensional space suitable for downstream HeteroGNN processing; The hyperbolic tangent activation function performs a nonlinear mapping on the result of the linear transformation, smoothly compressing (normalizing) the values of each dimension of the output vector to the interval [-1, 1]. This enhances the model's nonlinear fitting ability and ensures the stability of the output feature distribution, which is beneficial for the stable training and gradient propagation of the subsequent network. Based on this, the vector... The semantic summary, which condenses the relationships between all symptoms, is directly projected into dynamic feature vectors of symptom nodes carrying specific context in the subsequent HeteroGNN module. It should be noted that this dynamic vector generated in real time based on BERT is only for highly personalized symptom nodes. For the pathogenesis, treatment and prescription nodes that represent objective dialectical treatment rules in the TCM knowledge graph, this invention adopts a completely different static prior coding strategy in the feature extraction layer. The two together constitute the hybrid initialization input layer of the subsequent heterogeneous graph neural network.
[0020] (4) Static coding based on pathogenesis, treatment method, and prescription nodes Different from personalized symptoms, pathogenesis, treatment methods and prescriptions are objective experience summaries accumulated by traditional Chinese medicine (TCM) physicians through past dynasties, and their semantic connotations do not shift with changes of a single case. For example, the core connotation of the pathogenesis "qi stagnation and blood stasis" or the efficacy of the prescription "Xuefu Zhuyu Decoction" will not undergo semantic shift due to changes of a single case. Therefore, if real-time calculation is also performed on such entities, it will not only cause redundant computing power, but also undermine the objectivity and a priori constraint of classical TCM theories. Thus, in the feature extraction stage, a static encoding strategy is adopted for the three types of nodes: pathogenesis, treatment methods and prescriptions. First, knowledge graph triples including these three types of nodes are extracted based on the knowledge graph. Then, a knowledge graph embedding algorithm is used to jointly train the nodes in a high-dimensional topological space, forcing the model to learn implicit medical logic. The finally output static vectors are solidified and stored in the feature library. These vectors have two functions: in the subsequent GNN message passing, they serve as spatial anchors that do not fluctuate with individual cases, guiding dynamic symptom features to converge in the direction of correct medical logic; for "treatment method" nodes that are often omitted in ancient TCM medical records, static pre-trained vectors provide initialization features for them, enabling GNN to still activate the node along topological relations in the latent space even when there is no explicit text mention. So far, the semantic encoding of all nodes has been completed.
[0021] Embodiments of the present invention further include a fine-tuning strategy based on domain adaptation: Although Bert-Base-Chinese has powerful general language understanding capability, the term distribution in the TCM field is significantly different from that in general corpora. For example, "shu (数)" means a numeral in a general context, while refers to a rapid pulse in the TCM context. To address this semantic gap, the present invention adopts a full-parameter supervised fine-tuning strategy. During training, the parameters of the word embedding layer and all Transformer layers of BERT participate in gradient update, the AdamW optimizer is used, and the learning rate is set to , and it is used in conjunction with a linear warm-up learning rate scheduling strategy.
[0022] Through fine-tuning on a large amount of well-annotated standardized data, the model not only learns the vector representation of TCM entities, but more importantly, learns the implicit logic behind entity combinations.
[0023] Optionally, the GNN structure inference tower serves as a logical constraint, introducing a multi-source knowledge graph as a priori topological structure. Based on a heterogeneous graph neural network (HeteroGNN), a heterogeneous graph is constructed including four types of entities: "symptom-pathogenesis-treatment-prescription". For the various edge relationships defined in the graph, independent graph attention networks (GATs) are applied (traditional heterogeneous graph models often use shared attention heads / shared transformation matrices, making it difficult to distinguish the semantic differences between different knowledge associations in traditional Chinese medicine; this invention configures independent GATs for each relationship, such as symptom-pathogenesis, pathogenesis-treatment, and treatment-prescription). The sub-network fully isolates the feature transformation rules of different TCM logical pathways and adapts to the hierarchical knowledge features of TCM syndrome differentiation and treatment. Through layered message transmission, the downstream "prescription" node can explicitly aggregate the feature information of the upstream "pathogenesis" and "symptom" nodes. This process actually simulates the logical deduction path of TCM "observing the pulse and symptoms, knowing what is wrong, and treating accordingly" in a high-dimensional vector space, thereby generating a structural feature vector rich in global topological information.
[0024] Optionally, the GNN structure inference tower, based on the PyG (PyTorch Geometric) framework, constructs a graph inference network including implicit logical constraints, comprising: (1) Construction of heterogeneous topology networks based on knowledge graphs Before using HeteroGNN for deep feature aggregation, the static, discrete TCM structured knowledge graph is first mapped into a high-dimensional heterogeneous topological network space that can be computed by deep learning, including: 1.1 Heterogeneous Topology and the Implicit Hub Mechanism of "Governance" The knowledge graph includes multiple types of nodes and relational edges, such as etiology, pathogenesis, symptoms, treatment methods, and prescriptions. Among them, the treatment method node plays a pivotal role, connecting the preceding and following parts. However, the treatment method labels in the training data are extremely sparse (in ancient texts, there are often only descriptions of symptoms, prescriptions, and parts that can be broken down into pathogenesis, while the intermediate treatment method link is often omitted or implicit, resulting in extremely weak supervision signals and unsatisfactory training effects). A topological bridge mechanism based on latent variables was designed to retain the treatment method as an implicit semantic hub connecting the prediction of pathogenesis and the output of prescriptions, rather than as a node that needs to be predicted: A heterogeneous topological network including multiple types of entities was constructed based on knowledge graphs. Transforming the entity set in the knowledge graph into the node set of the graph neural network. The triple relationship is transformed into a message passing path in a graph neural network. ; To perform entity type identification (in order to handle the heterogeneity of different medical entities), define a node type mapping function: Its function is to combine any node instance in the knowledge graph. Mapped to the specific object type it belongs to , ; The design requires that all pathogenesis must first be mapped to the treatment method space before being passed to the prescription (the feature information of all pathogenesis entities in the graph must first be mapped to the treatment method representation space to complete the feature transformation before being passed to the prescription entity, blocking the direct information transmission path between pathogenesis and prescription). This design forces the model to learn a potential treatment method representation in the intermediate layer. Even without explicit labels, the gradient can update the embedding vector of the treatment method node through backpropagation of the prescription node. Although the treatment method labels are sparse, the model applies a low-weight soft constraint to the node as auxiliary supervision during training, which not only avoids the model overfitting to the "symptom-prescription" surface layer, but also ensures the integrity of the inference chain. 1.2 Definition of Edge Relationships and Construction of Inference Chains Based on the above node design, the edge set Three sets of relationship types following traditional Chinese medicine theory were defined. This forms a hierarchical reasoning chain, ensuring that message passing in the graph neural network is highly isomorphic to the diagnostic thinking of traditional Chinese medicine: Aggregate high-dimensional dynamic symptom features into the hidden state of pathogenesis nodes; The pathogenesis characteristics here are transformed by the matrix Mapped to the governance space; The formula nodes aggregate treatment characteristics, ultimately forming a decision vector that includes information from the entire chain; Meanwhile, in order to enhance the efficiency of gradient backpropagation, the graph explicitly constructs the reverse edges of the above relationships, so that the prescription supervision signal at the end can backtrack to update the upstream pathogenesis and treatment embeddings.
[0025] Optionally, the GNN structure inference tower further includes: (2) Initialization of hybrid nodes combining dynamic and static elements: An initialization strategy combining "static prior" and "dynamic perception" is adopted. "Static prior" refers to the objective medical knowledge in the knowledge graph, which is statically initialized for each node using graph embedding. "Dynamic perception" refers to the model's real-time capture of specific symptom text from the current input case, extracting it as a text vector rich in immediate features. This vector dynamically assigns weights to symptom nodes in the graph neural network, representing personalized symptom input from the case, thus achieving deep fusion of textual and graph modalities. Dynamic semantic injection of symptom nodes, for symptom nodes in the graph. Its initial characteristics Directly reuse the global semantic vector output by the BERT encoder : Because the vector space output by BERT is inconsistent with the vector space required by graph neural networks in terms of dimension and feature distribution, By using linear matrix multiplication, text features are smoothly projected and transformed into dimensions that graph networks can understand; This means that the entry point of the graph is dynamic. When BERT perceives the context of "encountering cold", the features of the symptom nodes will shift, thereby activating "cold coagulation" type pathogenesis nodes more strongly in subsequent message transmission, rather than "heat toxicity" type nodes. For the nodes of pathogenesis, treatment method, and prescription, a learnable embedding matrix is initialized based on static topological priors: These embedding vectors are continuously optimized during the training process of the full dataset. In particular, the treatment node, by aggregating a large amount of forward information of pathogenesis and backward gradient of prescription, its embedding vector gradually learns the global topological semantics of the treatment in the knowledge graph, thus playing a stable anchor role in reasoning. (3) Attention-based heterogeneous message passing: A heterogeneous graph convolutional fusion graph attention mechanism is employed for message passing between nodes to handle multiple medical semantic relationships in heterogeneous networks and simulate the complete clinical diagnosis and treatment path. During message passing, the state update of the target node depends not only on the inherent characteristics of its surrounding neighbor nodes but also on the strength of the logical relationship between them. By dynamically allocating computational weights through the attention mechanism, when faced with complex symptom information in cases, higher feature aggregation weights are automatically assigned to the core main symptom, while attenuating the interference of concurrent symptoms or irrelevant noise. This endows the model with the ability to accurately identify the core pathogenesis in complex high-dimensional features. The HeteroConv operator is used, and in each meta-relation... The above applies a graph attention mechanism to nodes. State updates depend not only on neighboring nodes The characteristics of [the relationship] depend more on the logical strength of the relationship between the two. No. The formula for updating the node features of a layer is: in, Represents the target node After passing through one network layer, at the... The latest feature vector obtained by the layer; Neighboring nodes Features in the current layer; It is for a specific relationship The learnable weight matrix is used to perform linear spatial transformation on neighbor features; These are logical weights calculated using a self-attention mechanism: The overall structure of the formula is a standard Softmax function, which exponentially normalizes and normalizes the scores between the target node and all its neighbors, thus calculating... It is a probability value between 0 and 1: where The target node and neighboring nodes The concatenation of feature vectors combines the features of both. It is for a specific relationship The learnable attention parameter vector is calculated by performing an inner product with the concatenated features, compressing the high-dimensional vector into a scalar score representing the original correlation between the two nodes. Apply to the calculated raw score The activation function preserves a weak negative signal to prevent neuron death; This mechanism ensures that even in the absence of explicit data, neural networks can spontaneously generate logical reasoning paths. (4) Full-chain feature aggregation based on deep stacking: If only a single-layer graph convolution is performed, the prescription node will not be able to perceive the initial symptom node. In order to achieve cross-level reasoning from symptoms to prescriptions, a deep graph convolution stacked structure is designed to complete the progressive feature aggregation of "symptom-mechanism-treatment-prescription" in the vector space. Configure a two-layer heterogeneous graph convolutional network to perform hierarchical transfer of node information: First convolutional layer (Layer=1): Symptom characteristics of upstream aggregation of pathogenesis nodes ; Treatment nodes aggregate upstream pathogenesis characteristics ; Second convolutional layer (Layer=2): Characteristics of Formula Node Aggregation Treatment At this point, the aggregated treatment node has already absorbed the information of "pathogenesis" and "symptoms" in the first convolution stage. Therefore, when the prescription node aggregates treatment features, it also indirectly aggregates symptom features and pathogenesis features. Although all nodes participate in computation in each convolutional layer, the flow of logical information is temporal: the first layer completes the semantic abstraction from symptoms to treatment, at which point the prescription node only aggregates static neighbor features; only in the second layer can the prescription node aggregate dynamic treatment features enhanced by reasoning, thus completing the entire logical loop. Through this cascaded message passing, the final generated prescription node feature vector is obtained. Instead of isolated entity embeddings, it aggregates high-level vectors that feature upstream symptoms, pathogenesis, and treatment methods, and integrates a full-chain structured representation of the complete reasoning path, ultimately extracting global features from the prescription nodes. ,like Figure 3 As shown: The global feature vector As the final topological feature vector of the knowledge graph, it carries structural information that has been logically verified by the knowledge graph, providing logical support for the next stage of integration with BERT semantic feature encoding.
[0026] Optionally, such as Figure 4 As shown, the decision layer includes a prior biased multimodal fusion layer and a cascaded prediction head; The multimodal fusion layer performs ternary feature fusion based on path prior bias, including: Before performing the final classification prediction, the heterogeneous features extracted by the dual-tower network are fused, and a path prior bias mechanism is introduced to handle the significant differences in information completeness and diagnostic complexity among different input cases: For each input case, a discrete prior logical path identifier (Path ID) is assigned based on its clinical characteristics and preliminary rule assessment. This Path ID serves as a global guiding signal and is mapped into a dense prior bias vector through an embedding layer. ; Subsequently, in the multimodal fusion layer, the model performs ternary feature concatenation, combining the text semantic vectors extracted by BERT. Graph topological feature vectors extracted by HeteroGNN and prior bias vector High-dimensional stitching and nonlinear mapping are performed to generate the final fused features. ; The The introduction of this feature provides a conditional constraint from a macro-level medical cognitive perspective to the neural network. It can dynamically adjust the weight distribution within the fusion layer, enabling the model to adaptively generate attention preferences for dynamic textual features or static graph features when facing different types of clinical scenarios, thereby generating… More in line with the current medical context; The cascaded prediction head comprises three parallel prediction heads. Although computationally parallel, they form a hierarchical dependency in terms of logical supervision, clearly defining the auxiliary role of the treatment method, including: The pathogenesis prediction head uses multi-label classification because a single case may have multiple pathogenesis mechanisms simultaneously. The method prediction head, given its design as an implicit hub, does not pursue extremely high classification accuracy, but exists as a logical middleware; The existence of this prediction head requires It must include information that can distinguish different treatment strategies. Even in the case of unsupervised signals, the gradient can be backpropagated through consistency loss to update the parameters here and ensure that the inference chain does not break. The prescription prediction head uses multi-class or Top-K recommendation to output the final decision probability (the prescription prediction result output in this embodiment is only for reference for doctors' clinical diagnosis and is not the final treatment plan).
[0027] Optionally, the knowledge graph consistency loss utilizes the knowledge graph for soft logic constraints, applying regularization constraints at the loss function level (to mitigate the logical illusions generated by deep neural networks): set up This represents the adjacency matrix of "pathogenesis-treatment" in the knowledge graph. For the adjacency matrix of "treatment method-prescription", consistency loss Defined as the orthogonality penalty between the predicted distribution and the true spectral structure, it is used to detect the connectivity of the prediction results: If the model predicts the set of pathogenesis With the predicted set of prescriptions If there is no connection through any "method" in the knowledge graph, i.e., the path is broken, then... The value of approaches 0, causing the loss function to... The tendency toward infinity forces the model to automatically search for solutions that are logically connected in medical terms during gradient descent, thus achieving logical self-consistency within the neural network without relying on an external rule engine.
[0028] Embodiments of the present invention also include a multi-objective joint optimization and training strategy: In order to coordinate the prediction tasks of three different levels—pathogenesis, treatment, and prescription—within a unified computational framework, this invention constructs a hierarchically weighted multi-objective joint optimization system, which not only defines the composition of the total loss function, but also establishes the task priority and parameter update strategy in the gradient descent process.
[0029] Optionally, the total loss function is as follows: in, The formula loss is defined as a multi-class classification problem, and the standard multi-class cross-entropy loss is used. This loss function directly penalizes the probability distribution of the wrong category, driving the model to focus the probability density on the correct prescription label; and The losses are respectively related to pathogenesis and treatment, and a multi-label binary cross-entropy loss with Logits is used. ,in, The total number of label categories. (This is the Sigmoid activation function), which effectively handles the label imbalance problem and ensures that the model remains sensitive to sparse pathogenesis data; For knowledge graph consistency loss; The hyperparameter settings are as follows: As the benchmark primary task for prescription prediction, its weight is implicitly set to 1.0; As a strong supervision weight, pathogenesis identification is the logical premise of prescription selection. If the pathogenesis identification is wrong, the subsequent prescription recommendation is a probabilistic problem even if it is correct. Therefore, the pathogenesis loss is given the highest weight, and deep supervision is implemented to force the model to prioritize the optimization of the underlying pathogenesis extraction capability. To assist in supervising the weights, the treatment method, as the hub connecting the pathogenesis and the prescription, must maintain a certain gradient contribution to maintain the connectivity of the inference chain, even though the labels are sparse. Setting a medium weight will not interfere with the main task, but will also play a regularization role. The soft constraint weights and the knowledge graph consistency loss, as a penalty term, are designed to fine-tune the space rather than dominate the optimization direction. The low weights are sufficient to provide huge gradient feedback when the model produces serious logical violations, while the influence automatically decays as the model gradually converges, avoiding the model from getting stuck in rigid rule matching.
[0030] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A Traditional Chinese Medicine (TCM) auxiliary decision-making system based on semantic topological feature fusion and latent variable constraints, characterized in that, The system includes a knowledge-enhanced dual-tower deep reasoning architecture, which comprises two parallel feature extraction channels forming a horizontally parallel structure. One side is the BERT semantic perception tower for processing clinical case texts, which uses a pre-trained language model to encode unstructured expressions into high-dimensional semantic vectors to capture implicit context. On the other side is a GNN structure inference tower for processing graph topology. Multi-source knowledge graphs are introduced as prior constraints to construct a heterogeneous graph network including four types of entities: symptoms, pathogenesis, treatment, and prescriptions. A hybrid feature weighting strategy is adopted, and highly personalized symptom nodes are dynamically initialized. The text vectors extracted in real time by the semantic perception tower are directly reused. For the nodes representing the pathogenesis, treatment methods, and prescriptions that represent objective medical laws, a learnable knowledge graph embedding matrix is initialized based on static topological priors to achieve the bottom-level fusion of heterogeneous spaces. Furthermore, a bridge mechanism based on latent variables is designed in the topology construction, retaining the treatment method nodes as implicit semantic hubs. A two-layer deep heterogeneous graph convolution fusion graph attention mechanism is adopted, and by dynamically allocating logical weights, the prescription nodes at the end are forced to aggregate the symptom, pathogenesis, and treatment features of the upstream across levels. Even in the absence of explicit textual mentions, the entire logical loop of TCM reasoning can still be completed in the latent space. In terms of vertical reasoning logic, a three-layer architecture of perception layer, reasoning layer and decision layer is constructed. The perception layer is responsible for capturing and representing multi-source heterogeneous data; the reasoning layer is responsible for mining and reconstructing latent pathogenesis features; and the decision layer is responsible for integrating and outputting multimodal information. Based on the fusion features, a cascaded prediction head for pathogenesis, treatment method and prescription is designed. The treatment method prediction head does not output prediction results and exists as a logical middleware. The pathogenesis and prescription prediction heads output the predicted pathogenesis and prescription respectively. In constructing the total loss function, a knowledge graph consistency loss is introduced. The static adjacency matrix of the knowledge graph is used to detect the logical connectivity of the prediction distribution. Orthogonal penalties are applied to outputs that are interrupted by the path, or where the predicted pathogenesis and prescription are not related in the graph. This forces the model to automatically search for a solution space that is connected in accordance with medical theory during gradient descent, thereby reducing the logical illusion of the neural network internally.
2. The system according to claim 1, characterized in that, The BERT semantic perception tower is responsible for perceiving features and processing the four diagnostic methods information of the text structure in clinical cases. It adopts a BERT-Base-Chinese pre-trained model to encode the complex and ambiguous unstructured expressions of traditional Chinese medicine into continuous high-dimensional semantic vectors, capturing the implicit contextual dependencies between entities. The high-dimensional semantic vectors serve as the representation of the text modality and are projected into the initial features of different nodes in the subsequent graph neural network, realizing the semantic injection from the text space to the graph space.
3. The system according to claim 1, characterized in that, The BERT semantic perception tower, in the pre-training stage, uses BERT-Base-Chinese for unsupervised training based on a large-scale Chinese Wikipedia corpus. It employs a word-level Word Piece vocabulary for word segmentation, treating each Chinese character in the input sequence as an independent basic unit token. It is jointly pre-trained through two major tasks: Masked Language Modeling (MLM) and Next Sentence Prediction (NSP). This effectively captures the deep bidirectional semantic features of Chinese characters in different contexts. The model's training data is a set of standardized structured entities after preprocessing. While this discrete standardized entity eliminates lexical ambiguity, it loses the syntactic structure and modification relationships in the original text. Therefore, the role of this module is to reconstruct the context of standardized symptom entity sequences through a deep attention mechanism, capturing the underlying pathological indicators behind symptom combinations.
4. The system according to claim 1, characterized in that, The processing procedure of the BERT semantic awareness tower specifically includes: (1) Entity serialization construction based on symptom nodes Due to the highly personalized nature of symptoms, a text serialization strategy based on symptom nodes is designed to focus on constructing high-density symptom representations. This strategy utilizes functionally specific control tags to reassemble discrete symptom entities into pseudo-sentences rich in pathological features. Let there be a training sample Including standardized symptom sets The serialization construction process is formally described as follows: in: As a global classification label, it is used to aggregate the symptom cluster features of the entire case, serving as the core vector for inferring the underlying pathogenesis; correspond Labels, serving as functional cues, guide the model to focus on clinical manifestation features during multi-head attention computation; Use comma separators to separate to The connection enables the model to learn the nonlinear ability to directly map from "symptom combination patterns" to "core pathogenesis"; This is the end marker; The pre-standardized entity extraction of training data and the BERT context reconstruction in this step are not redundant work, but rather form a complementary relationship of "semantic dimensionality reduction - logical dimensionality enhancement": the previous standardization work effectively eliminates the expressive noise in natural language, compressing high-dimensional sparse free text into a set of core pathological features with high signal-to-noise ratio, avoiding the model from wasting computing power on processing variations of traditional Chinese medicine language. However, the combination of entities contains infinite changes in pathogenesis. BERT uses positional encoding to distinguish the primary and secondary relationships in the symptom list, representing the distribution characteristics of symptom clusters; and uses the attention mechanism to enhance the learning of the co-occurrence intensity between symptom entities. (2) Multi-head attention mechanism and context reconstruction The Transformer Encoder's multi-head self-attention mechanism is used to reconstruct the context of the serialized pseudo-sentences. In the high-dimensional semantic space, there are strong nonlinear interactions between entities. The model calculates the context of the entities. With entity By assigning attention weights between features and capturing implicit associations, this mechanism achieves two semantic functions: feature enhancement and noise suppression. Through a deep stacking of 12 encoder layers, discrete entity symbols are transformed into a dense vector sequence that includes combined contextual information. (3) Global semantic aggregation and output representation After BERT encoding, the model obtains the last hidden state and performs pooling to obtain a fixed-length vector for downstream classification and graph neural network initialization. Select the first one of the sequences The hidden vector corresponding to the label serves as the global semantic representation of the entire case. : in This represents the output of the last layer of the BERT model, corresponding to After the hidden state vector of this special marker has been fully interacted with by the multi-head self-attention mechanism of 12 layers of Transformer, the vector at this position has absorbed and aggregated the global contextual information of all symptom entities and their modification relationships in the current case. and These represent the learnable weight matrix and bias term of the pooling layer, respectively. This step performs a linear feature space transformation, the purpose of which is to map the raw features extracted by BERT to a specific dimensional space suitable for downstream HeteroGNN processing; The hyperbolic tangent activation function performs a nonlinear mapping on the result of the linear transformation, smoothly compressing (normalizing) the values of each dimension of the output vector to the interval [-1, 1]. This enhances the model's nonlinear fitting ability and ensures the stability of the output feature distribution, which is beneficial for the stable training and gradient propagation of the subsequent network. Based on this, the vector... The semantic summary, which condenses the relationships between all symptoms, is directly projected into dynamic feature vectors of symptom nodes carrying specific context in the subsequent HeteroGNN module. (4) Static coding based on pathogenesis, treatment method, and prescription nodes Unlike personalized symptoms, pathogenesis, treatment methods, and prescriptions are objective summaries of experience accumulated by generations of traditional Chinese medicine practitioners. Their semantic shift does not occur with changes in a single case. Therefore, a static encoding strategy is adopted for these three types of nodes during the feature extraction stage: First, knowledge graph triples containing these three types of nodes are extracted. Then, a knowledge graph embedding algorithm is used to jointly train the nodes in a high-dimensional topological space, forcing the model to learn implicit medical logic. The final output static vectors are solidified and stored in the feature library. These vectors serve two purposes: in subsequent GNN message passing, they act as spatial anchors that do not fluctuate with individual cases, guiding dynamic symptom features towards the correct medical logic; for "treatment method" nodes, which are often omitted in ancient Chinese medicine case studies, the static pre-trained vectors provide initialization features, enabling the GNN to activate the node in the latent space according to topological relationships even without explicit textual mention. This completes the semantic encoding of all nodes.
5. The system according to claim 1, characterized in that, The GNN structure inference tower serves as a logical constraint, introducing a multi-source knowledge graph as a priori topological structure. Based on the heterogeneous graph neural network HeteroGNN, a heterogeneous graph is constructed that includes four types of entities: "symptoms, pathogenesis, treatment, and prescriptions." For the various edge relationships defined in the graph, independent graph attention networks (GAT) are applied. Through hierarchical message passing, the downstream "prescription" node can explicitly aggregate the feature information of the upstream "pathogenesis" and "symptom" nodes. This process actually simulates the logical deduction path of traditional Chinese medicine in a high-dimensional vector space: "observe the pulse and symptoms, know what is wrong, and treat accordingly." This generates a structural feature vector rich in global topological information.
6. The system according to claim 1, characterized in that, The GNN structure inference tower, based on the PyG framework, constructs a graph inference network including implicit logical constraints, comprising: (1) Construction of heterogeneous topology networks based on knowledge graphs Before using HeteroGNN for deep feature aggregation, the static, discrete TCM structured knowledge graph is first mapped into a high-dimensional heterogeneous topological network space that can be computed by deep learning, including: 1.1 Heterogeneous Topology and the Implicit Hub Mechanism of "Governance" The knowledge graph includes various types of nodes and relational edges, such as etiology, pathogenesis, symptoms, treatment methods, and prescriptions. Among them, the treatment method node plays a pivotal role, connecting the preceding and following parts. However, the treatment method labels in the training data are extremely sparse. Therefore, a topological bridge mechanism based on latent variables was designed to retain the treatment method as an implicit semantic hub connecting the prediction of pathogenesis and the output of prescriptions, rather than as a node that needs to be predicted. A heterogeneous topological network including multiple types of entities was constructed based on knowledge graphs. Transforming the entity set in the knowledge graph into the node set of the graph neural network. The triple relationship is transformed into a message passing path in a graph neural network. ; To identify entity types, define a node type mapping function: Its function is to combine any node instance in the knowledge graph. Mapped to the specific object type it belongs to , ; The design requires all pathogenesis to be mapped to the treatment space before being passed to the prescription. This design forces the model to learn a potential treatment representation in the intermediate layer. Even without explicit labels, the gradient can update the embedding vector of the treatment node through backpropagation of the prescription node. Although the treatment labels are sparse, the model applies a low-weight soft constraint to the node as auxiliary supervision during training. This not only avoids the model from overfitting to the "symptom-prescription" surface, but also ensures the integrity of the inference chain. 1.2 Definition of Edge Relationships and Construction of Inference Chains Based on the above node design, the edge set Three sets of relationship types following traditional Chinese medicine theory were defined. This forms a hierarchical reasoning chain, ensuring that message passing in the graph neural network is highly isomorphic to the diagnostic thinking of traditional Chinese medicine: Aggregate high-dimensional dynamic symptom features into the hidden state of pathogenesis nodes; The pathogenesis characteristics here are transformed by the matrix Mapped to the governance space; The formula nodes aggregate treatment characteristics, ultimately forming a decision vector that includes information from the entire chain; Meanwhile, in order to enhance the efficiency of gradient backpropagation, the graph explicitly constructs the reverse edges of the above relationships, so that the prescription supervision signal at the end can backtrack to update the upstream pathogenesis and treatment embeddings.
7. The system according to claim 1, characterized in that, The GNN structure inference tower also includes: (2) Initialization of hybrid nodes combining dynamic and static elements: An initialization strategy combining "static prior" and "dynamic perception" is adopted. "Static prior" refers to the objective medical knowledge in the knowledge graph, which is statically initialized for each node using graph embedding. "Dynamic perception" refers to the model's real-time capture of specific symptom text from the current input case, extracting it as text vectors rich in immediate features. These vectors are then dynamically weighted onto symptom nodes in the graph neural network, representing personalized symptom input from the case, thus achieving deep fusion of textual and graph modalities. Dynamic semantic injection of symptom nodes, for symptom nodes in the graph. Its initial characteristics Directly reuse the global semantic vector output by the BERT encoder : Because the vector space output by BERT is inconsistent with the vector space required by graph neural networks in terms of dimension and feature distribution, By using linear matrix multiplication, text features are smoothly projected and transformed into dimensions that graph networks can understand; This means that the entry point of the graph is dynamic. When BERT perceives the context of "encountering cold", the characteristics of the symptom nodes will shift, thereby activating "cold coagulation" type pathogenesis nodes more strongly in subsequent message transmission, rather than "heat toxicity" type nodes. For the nodes of pathogenesis, treatment method, and prescription, a learnable embedding matrix is initialized based on static topological priors: These embedding vectors are continuously optimized during the training process of the full dataset. In particular, the treatment node, by aggregating a large amount of forward information of pathogenesis and backward gradient of prescription, its embedding vector gradually learns the global topological semantics of the treatment in the knowledge graph, thus playing a stable anchor role in reasoning. (3) Attention-based heterogeneous message passing: A heterogeneous graph convolutional fusion graph attention mechanism is employed for message passing between nodes to handle multiple medical semantic relationships in heterogeneous networks and simulate the complete clinical diagnosis and treatment path. During message passing, the state update of the target node depends not only on the inherent characteristics of its surrounding neighbor nodes but also on the strength of the logical relationship between them. By dynamically allocating computational weights through the attention mechanism, when faced with complex symptom information in cases, higher feature aggregation weights are automatically assigned to the core main symptom, while attenuating the interference of concurrent symptoms or irrelevant noise. This endows the model with the ability to accurately identify the core pathogenesis in complex high-dimensional features. The HeteroConv operator is used, and in each meta-relation... The above applies a graph attention mechanism to nodes. State updates depend not only on neighboring nodes The characteristics of [the relationship] depend more on the logical strength of the relationship between the two. No. The formula for updating the node features of a layer is: in, Represents the target node After passing through one network layer, at the... The latest feature vector obtained by the layer; Neighboring nodes Features in the current layer; It is for a specific relationship The learnable weight matrix is used to perform linear spatial transformation on neighbor features; These are logical weights calculated using a self-attention mechanism: The overall structure of the formula is a standard Softmax function, which exponentially normalizes and normalizes the scores between the target node and all its neighbors, thus calculating... It is a probability value between 0 and 1: where The target node and neighboring nodes The concatenation of feature vectors combines the features of both. It is for a specific relationship The learnable attention parameter vector is calculated by performing an inner product with the concatenated features, compressing the high-dimensional vector into a scalar score representing the original correlation between the two nodes. Apply to the calculated raw score The activation function preserves a weak negative signal to prevent neuron death; This mechanism ensures that even in the absence of explicit data, neural networks can spontaneously generate logical reasoning paths. (4) Full-chain feature aggregation based on deep stacking: The design employs a depthwise graph convolution stacked structure to achieve progressive feature aggregation of the TCM concept of "symptom-mechanism-treatment-prescription" in vector space. Configure a two-layer heterogeneous graph convolutional network to perform hierarchical transfer of node information: First convolutional layer: Symptom characteristics of upstream nodes in the pathogenesis process ; Treatment nodes aggregate upstream pathogenesis characteristics ; Second convolutional layer: Characteristics of Formula Node Aggregation Treatment At this point, the aggregated treatment node has already absorbed the information of "pathogenesis" and "symptoms" in the first convolution stage. Therefore, when the prescription node aggregates treatment features, it also indirectly aggregates symptom features and pathogenesis features. Although all nodes participate in computation in each convolutional layer, the flow of logical information is temporal: the first layer completes the semantic abstraction from symptoms to treatment, at which point the prescription node only aggregates static neighbor features; only in the second layer can the prescription node aggregate dynamic treatment features enhanced by reasoning, thus completing the entire logical loop. Through this cascaded message passing, the final generated prescription node feature vector is obtained. Instead of isolated entity embeddings, it aggregates high-level vectors that include features from upstream symptom, pathogenesis, and treatment nodes, and integrates a full-chain structured representation of the complete reasoning path, ultimately extracting global features from the prescription nodes. : The global feature vector As the final topological feature vector of the knowledge graph, it carries structural information that has been logically verified by the knowledge graph, providing logical support for the next stage of integration with BERT semantic feature encoding.
8. The system according to claim 1, characterized in that, The decision layer includes a prior biased multimodal fusion layer and a cascaded prediction head; The multimodal fusion layer performs ternary feature fusion based on path prior bias, including: Before performing the final classification prediction, the heterogeneous features extracted by the dual-tower network are fused, and a path prior bias mechanism is introduced to handle the significant differences in information completeness and diagnostic complexity among different input cases: For each input case, a discrete prior logical path identifier (Path ID) is assigned based on its clinical characteristics and preliminary rule assessment. This Path ID serves as a global guiding signal and is mapped into a dense prior bias vector through an embedding layer. ; Subsequently, in the multimodal fusion layer, the model performs ternary feature concatenation, combining the text semantic vectors extracted by BERT. Graph topological feature vectors extracted by HeteroGNN and prior bias vector High-dimensional stitching and nonlinear mapping are performed to generate the final fused features. ; The The introduction of this feature provides a conditional constraint from a macro-level medical cognitive perspective to the neural network. It can dynamically adjust the weight distribution within the fusion layer, enabling the model to adaptively generate attention preferences for dynamic textual features or static graph features when facing different types of clinical scenarios, thereby generating… More in line with the current medical context; The cascaded prediction head comprises three parallel prediction heads. Although computationally parallel, they form a hierarchical dependency in terms of logical supervision, clearly defining the auxiliary role of the treatment method, including: The pathogenesis prediction head uses multi-label classification because a single case may have multiple pathogenesis mechanisms simultaneously. The method prediction head, given its design as an implicit hub, does not pursue extremely high classification accuracy, but exists as a logical middleware; The existence of this prediction head requires It must include information that can distinguish different treatment strategies. Even in the case of unsupervised signals, the gradient can be backpropagated through consistency loss to update the parameters here and ensure that the inference chain does not break. The prescription prediction head uses multi-class or Top-K recommendation to output the final decision probability.
9. The system according to claim 1, characterized in that, The knowledge graph consistency loss utilizes the knowledge graph for soft logic constraints, applying regularization constraints at the loss function level: set up This is the adjacency matrix of "pathogenesis-treatment" in the knowledge graph. For the adjacency matrix of "treatment method-prescription", consistency loss Defined as the orthogonality penalty between the predicted distribution and the true spectral structure, it is used to detect the connectivity of the prediction results: If the model predicts the set of pathogenesis With the predicted set of prescriptions If there is no connection in the knowledge graph through any "method," i.e., the path is broken, then... The value of approaches 0, causing the loss function to... The tendency toward infinity forces the model to automatically search for solutions that are logically connected in medical terms during gradient descent, thus achieving logical self-consistency within the neural network without relying on an external rule engine.
10. The system according to claim 9, characterized in that, The total loss function is as follows: in, The formula loss is defined as a multi-class classification problem and adopts the standard multi-class cross-entropy loss. This loss function directly penalizes the probability distribution of the wrong class, driving the model to focus the probability density on the correct formula label. and The loss functions are pathogenesis and treatment, respectively. A multi-label binary cross-entropy loss with Logits is used. This function can effectively handle the label imbalance problem and ensure that the model remains sensitive to sparse pathogenesis data. For knowledge graph consistency loss; The hyperparameter settings are as follows: As the benchmark primary task for prescription prediction, its weight is implicitly set to 1.0; As a strong supervision weight, pathogenesis identification is the logical premise of prescription selection. If the pathogenesis identification is wrong, the subsequent prescription recommendation is a probabilistic problem even if it is correct. Therefore, the pathogenesis loss is given the highest weight, and deep supervision is implemented to force the model to prioritize the optimization of the underlying pathogenesis extraction capability. To assist in supervising the weights, the treatment method, as the hub connecting the pathogenesis and the prescription, must maintain a certain gradient contribution to maintain the connectivity of the inference chain, even though the labels are sparse. Setting a medium weight will not interfere with the main task, but will also play a regularization role. The soft constraint weights and the knowledge graph consistency loss, as a penalty term, are designed to fine-tune the space rather than dominate the optimization direction. The low weights are sufficient to provide huge gradient feedback when the model produces serious logical violations, while the influence automatically decays as the model gradually converges, avoiding the model from getting stuck in rigid rule matching.