A medical consumable abnormal data two-type fuzzy screening method based on cross-domain cooperation and heterogeneous atlas
By employing cross-domain collaboration and heterogeneous graph methods, combined with NLP, cross-domain transfer learning, and graph neural networks, the problem of inconsistent features in local medical consumables data was solved, enabling efficient screening and repair of abnormal data, and improving processing speed and accuracy.
Patent Information
- Application Number
- CN202511504814.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-21
- Publication Date
- 2026-01-27
- Estimated Expiration
- 2045-10-21
AI Technical Summary
Local medical consumables data differs from national data in characteristics. Existing technologies are slow to process data and have low accuracy, making it difficult to effectively screen for abnormal data with inconsistent characteristics and logical conflicts.
We employ a cross-domain collaboration and heterogeneous graph approach, combining NLP (Natural Language Processing), cross-domain transfer learning, graph neural networks, and type II fuzzy systems. We use a three-level governance framework for data screening, including data preprocessing, anomaly reasoning, feature repair, and deep discrimination, and dynamically adjust the rule base to improve screening accuracy.
It achieves high precision in anomaly data screening, breaks through the bottleneck of cross-domain heterogeneous data screening, improves the judgment accuracy of cross-domain transfer learning and the accuracy of anomaly data repair, and enhances the decision-maker's adaptability to high-dimensional dynamic data.
Smart Images

Figure CN120974390B_ABST
Abstract
Description
Technical Field
[0001] This invention proposes a type II fuzzy screening method for abnormal medical consumable data based on cross-domain collaboration and heterogeneous graphs, which relates to the field of medical information processing technology. Background Technology
[0002] Currently, local medical consumables data suffers from inconsistencies with national data due to issues such as a disconnect between product iteration and data updates, a lack of governance for heterogeneous local data, and insufficient efficiency in manual governance. The current approach primarily relies on manual methods, which suffer from slow processing speed and low accuracy, hindering the management of local medical consumables data. Furthermore, since inconsistencies can arise from multiple causes, including product feature updates, product withdrawals, and product relisting, identifying the root causes and implementing targeted solutions for each product exhibiting inconsistencies presents a significant challenge to data governance. Summary of the Invention
[0003] To address the technical problems in existing technologies, such as the inability of NLP (Natural Language Processing) to effectively extract information about medical consumables, the low accuracy of traditional cross-border transfer learning, the inability of graph neural networks to handle contextual information effectively, and the immutability of decision rules in type II fuzzy systems, this invention proposes a three-level governance framework of "perception-reasoning-decision." This framework combines NLP, cross-domain transfer learning, knowledge graph and graph neural network propagation modeling, and uncertainty decision-making technology for type II fuzzy systems. Through various deep learning techniques, it filters out abnormal data in local medical consumable databases, including those with inconsistent features, logical conflicts, and risks of illegal use.
[0004] A type II fuzzy screening method for abnormal medical consumable data based on cross-domain collaboration and heterogeneous graphs includes the following steps:
[0005] Step 1: Obtain the required medical consumables dataset and preprocess the data, construct positive sample data, negative sample data and target concept set, process the sample data through the pre-model and provide basic prediction results or features, calculate the weights of the pre-model through the attention mechanism, and finally make the final prediction through the gradient decision tree to generate standardized entity-concept features.
[0006] Step 2: Use a deep domain adversarial anomaly collaborative reasoning model based on a class of transfer learning to perform anomaly reasoning checks on the medical consumables data after feature extraction, and conduct preliminary detection to determine whether the data is anomalous; specifically, use a feature-sharing encoder to extract cross-domain invariant features, use a hyperpolyhedral decision boundary to constrain the normal data distribution, combine the maximum mean difference loss of the feature space to align inter-domain differences, and use adversarial training to make the discriminator unable to distinguish the domain source, thereby achieving anomaly scoring; and conduct preliminary screening through the trained deep domain adversarial collaborative reasoning model.
[0007] Step 3: Based on the heterogeneous semantic knowledge graph information fusion model of knowledge graph and graph neural network, feature repair and completion are performed on the medical consumables features detected as abnormal data; a heterogeneous semantic knowledge graph is constructed and embedded into the graph neural network. The 1- and 2-hop neighbor subgraphs are extracted with anchor triples as the center. The triple-level and graph-level features are fused through the hierarchical graph transformation network to predict the probability distribution of missing entities. Data repair is performed according to the probability distribution.
[0008] Step 4: Based on the type II fuzzy recognition and analysis model, perform in-depth discrimination on the corrected abnormal products: Build a dynamic rule engine, dynamically update the rule base through rule dynamic growth, pruning and fusion, improve interpretability through group-level, rule-level and feature entropy three-order sparse constraints, output category confidence intervals, and make judgments based on the confidence interval data.
[0009] The beneficial effects of this invention are as follows:
[0010] A dedicated NLP language processing model for handling medical consumables text was constructed, enabling the model to effectively extract key entities from medical consumables statements. By combining these entities with relevant concepts, the model parses the specific text. The bottleneck of cross-domain heterogeneous data screening was overcome, improving the accuracy of cross-domain transfer learning and achieving high precision in the initial screening of anomalies. Combining knowledge graphs with graph neural networks, and integrating global and contextual information, improved the accuracy and correctness of anomaly repair. Dynamic adjustments to the rule base enhanced the decision-maker's adaptability to high-dimensional dynamic data, enabling it to better identify data anomalies. Attached Figure Description
[0011] Figure 1 This is a diagram illustrating the overall structure of the method of the present invention.
[0012] Figure 2 This is a flowchart of a system for cleaning and extracting attribute features of medical consumables based on a language processing model.
[0013] Figure 3 This is a system architecture diagram of a deep domain adversarial anomaly collaborative reasoning model based on a class of transfer learning.
[0014] Figure 4 This is a schematic diagram of the heterogeneous knowledge graph information fusion framework based on knowledge graphs and graph neural networks.
[0015] Figure 5 This is a flowchart of the system processing based on a type-two fuzzy recognition and analysis model.
[0016] Figure 6 This is a two-dimensional mapping of the detection results of the deep-domain adversarial anomaly collaborative reasoning model.
[0017] Figure 7 The PR curve for the test set of the deep-domain adversarial anomaly collaborative reasoning model;
[0018] Figure 8 Example image showing the feature repair results of a prediction model based on knowledge graphs and graph neural networks;
[0019] Figure 9 This is a partial screening result graph using this method. Detailed Implementation
[0020] To better understand the purpose, structure, and function of this invention, the following description, in conjunction with the accompanying drawings, provides a more detailed account of a type II fuzzy screening method for abnormal medical consumable data based on cross-domain collaboration and heterogeneous graphs.
[0021] Example 1
[0022] This invention proposes a type II fuzzy screening method for abnormal data of medical consumables based on cross-domain collaboration and heterogeneous graphs. The overall structure is as follows: Figure 1 As shown, the system includes a medical consumables attribute feature cleaning and extraction system based on a language processing model, a heterogeneous semantic knowledge graph information fusion framework based on knowledge graphs and graph neural networks, an anomaly collaborative reasoning model based on a type of transfer learning and deep domain adversarial approach, and a recognition and analysis model based on type II fuzzy logic. The specific steps are as follows:
[0023] Step 1: As Figure 2 As shown, the required medical consumables dataset is obtained and the data is preprocessed to construct positive sample data, negative sample data and target concept set. After the sample data is processed by the pre-model and basic prediction results or features are provided, the weights of the pre-model are calculated based on the attention mechanism. Finally, the final prediction is made through gradient decision tree to generate standardized entity-concept features.
[0024] Step 1.1: Obtain the required training dataset and preprocess the data. The data includes medical consumables data to be processed and standard medical consumables product data. In this embodiment, medical consumables data from the local pharmaceutical and medical device bidding and procurement service center is collected as the data to be processed; medical consumables product data from the current medical insurance information business coding standard database is used as standard consumables data. Preprocessing includes filtering special characters and punctuation, removing tags, and standardizing text format.
[0025] Step 1.2: This mainly includes constructing positive sample data, constructing negative sample data, and generating the target concept set. Basic positive samples are obtained, and using standard labeled data, entities are paired with standard concepts. Then, a three-layer expansion strategy—data transfer expansion, symmetric expansion, and external knowledge expansion—is used to construct the positive sample dataset. Interfering erroneous concepts are filtered out through semantic similarity retrieval to build the negative sample dataset. Based on information retrieval technology, standard concepts related to the input entities are selected to generate the target concept set. In this embodiment, the specific construction process is as follows:
[0026] Step 1.2.1: Construct a positive sample dataset using a three-layer expansion strategy:
[0027] To obtain the basic positive sample D0, we directly used the labeled data provided by the China Health Information Processing Conference (CHIP2025) to match entities with standard concepts (ICDs). Entities mainly include the registration certificate number, registration certificate name, specification information, model information, material information, manufacturer information, and applicant company information from the consumable sample data. Concepts mainly consist of a set of registration certificate number regulations, a set of standard product name concepts, a set of standard specification information concepts, a set of standard model information concepts, a set of standard material information concepts, and a set of standard company information concepts.
[0028] Then, a first-level data transfer extension is performed: new samples are generated based on synonymous entities from the training data. If the standard concept in the two entity → standard concept pairs is the same standard concept, then the two entities are reconstructed.
[0029] Then, a second-order symmetric extension is performed: by reversing the existing entity relationships, the model's understanding of the concept → entity reverse mapping is enhanced, and the transformation from entity → standard concept pair to standard concept → entity pair is performed.
[0030] Finally, a three-level external knowledge expansion is carried out: by integrating the International Medical Terminology System (SNOMED CT), all synonyms under the same concept in SNOMED CT are paired up, and through terminology alignment, SNOMED synonyms are standardized and coded, mapped to the standard concepts of the International Classification of Diseases (ICD-9-CM-3);
[0031] Step 1.2.2: Using a difficult negative sample generation strategy, select distracting erroneous concepts through semantic similarity retrieval to construct a negative sample dataset. The main steps are as follows:
[0032] For the i-th entity in the training set i The ranking algorithm is used to calculate the score of the entity word in the document. Then, it retrieves the top 20 relevant candidate concepts by score from the entire database. The retrieval formula is as follows:
[0033]
[0034] Where IDF(q) i ) represents inverse document frequency, and n represents entity word. Quantity, The word q i Frequency of occurrence in document D k represents the total number of words in document D, and k1 represents a parameter that controls word frequency saturation. This is a parameter that controls the strength of document length normalization; avgdl represents the average length of all documents in the corpus. Correctly labeled concepts are removed, and the remaining candidate concepts are used as negative samples to form... The negative sample set, This represents the concept of negative samples numbered 2-20.
[0035] Step 1.2.3: Efficiently filter out standard concepts that may be related to the input entity to generate the target concept set. The method is based on information retrieval technology and combines the following algorithm with knowledge base enhancement. The specific implementation process is as follows:
[0036] Using Chinese characters as the smallest unit, construct an inverted index for the standard concept ICD. For each input entity... i Calculate its score against all standard concepts, and output the Top n target concepts based on the scores, where n=20. Ternary index fusion; construct the base index using the original text of standard concepts ( ), constructing an extended index for the training set using synonyms in the training set. Then, an extended index is built by linking synonymous entities with their corresponding standard concepts. By using a ternary index to expand the retrieval formula, a fused index is obtained. The expanded formula is:
[0037]
[0038] Step 1.3: Building the fusion architecture model, including the construction of the dual-stream transform attention model, the position-aware enhancement model, the dual-tower concept mapping model, and the implementation of the hybrid architecture;
[0039] This paper develops a permutation language model that overcomes the mask independence assumption of traditional bidirectional compiler processing models (BERT). It also employs a two-stream self-attention mechanism, resulting in improved long text processing capabilities. The permutation language model samples a permutation of all input word lists in the dataset and uses only position z when predicting the target word. t The previous content, z t This represents the position of the target word in the permutation z; the two-stream transform attention model uses both the content stream and the query stream to obtain the position of the content stream and query stream at position z. t The hidden state; in this embodiment, the specific construction steps are as follows:
[0040] The content stream is:
[0041] The query flow is:
[0042] in Indicates the content stream at position The hidden state of the m-th layer, Indicates the location of the query stream. The hidden state of the m-th layer, Let represent the set of hidden states at all positions in the (m-1)th layer; Attention is the self-attention mechanism computation function, where Q, K, and V are the query, key, and value vector parameters, respectively. These are learnable parameters for the self-attention mechanism.
[0043] The location-aware enhancement model addresses the location-content coupling problem and the relative location modeling problem by building a decoupled attention and enhanced mask decoder. In this embodiment, the specific steps are as follows:
[0044] Decoupling attention will be used to predict target words Attention is calculated using the position vector matrix P and the content vector matrix H, respectively, as follows:
[0045]
[0046] in This represents the relative position encoding of position i relative to j in the position vector matrix P. This represents the relative position encoding of position j relative to i in the position vector matrix P. For the i-th and j-th elements in the content vector matrix H, The calculated decoupled attention results are combined into an attention matrix A.
[0047] The enhanced mask decoder incorporates absolute position information before the softmax function, as shown in the following formula:
[0048]
[0049] Where A is the attention matrix, H is the content vector matrix, P is the position vector matrix, c is the learning parameters, and the superscript T denotes transpose. The output mixing vector matrix of the location-aware enhancement model. It is a normalized exponential function.
[0050] The two-tower concept mapping model employs a two-tower architecture, using two encoders with shared weights to independently encode entities and concepts, then comparing their embeddings to generate semantic matching scores for entity-concept pairs. In this embodiment, the specific steps are as follows:
[0051] First, two encoders with shared weights are built to process the input entities and concepts separately.
[0052] Two of the encoders employ a shared weight strategy, where the input entities and concepts do not interact directly through cross-calculation, thus avoiding coupling.
[0053] Then, pooling is performed on the input encoding. A pooling layer is used to pool the encoder output to obtain fixed-dimensional entity embeddings. and concept embedding Their structures are respectively and ,in, This represents the nth hidden state in entity embedding. This represents the nth hidden state in the concept embedding.
[0054] Absolute difference feature fusion: By concatenating the original embeddings and element-wise absolute differences, differential features are explicitly introduced. By calculating the absolute difference in each dimension of the embedding space, local differences are amplified, enhancing the model's sensitivity to subtle differences. The specific implementation formula is as follows:
[0055]
[0056] in, Indicates fusion features, The symbol represents vector concatenation, and |·| represents element-wise absolute value operation. It is an absolute difference calculation performed on each dimension of the embedding space. This indicates that the dimension of the input fusion feature is 3n.
[0057] The fused features are converted into interpretable matching probabilities, providing a quantitative basis for ranking target concepts. The overall steps are as follows:
[0058]
[0059] in It is a trainable weight matrix, and Score is a probability normalization function. Softmax is used to calculate the classification probability, obtaining the probability of the model matching each entity-concept. In this case, K represents the binary classification output dimension (label=1 or label=0). , Let v be the score value when label=1 and 0 respectively. The matching score v is calculated through linear transformation, as shown in the following formula:
[0060]
[0061] Step 1.4: The weights of the three preceding models are calculated and weighted averaged using an attention mechanism to obtain a mixed feature vector. Finally, a gradient decision tree is used for the final prediction. In this embodiment, the specific steps are as follows:
[0062] Step 1.4.1: Assign a weight to the output of each preceding model, which depends on the current sample and the output of all base models.
[0063] By designing a multilayer perceptron attention network, the output of each preceding model is mapped to a score, and then the weights are obtained through the softmax function.
[0064] Specifically, the calculated scores for the i-th sample and the u-th preceding model. for:
[0065]
[0066] in It is a learnable weight matrix. It is a paranoid vector. It is a queryable query vector. d is the hidden layer output vector of the preceding model u, and d is the dimension of the hidden layer output vector.
[0067] Then the weights of the preceding model u were adjusted. The calculation is performed using the following formula:
[0068]
[0069] in For the number of preceding models, u and This indicates the corresponding preceding model number. This represents the sum of the weights of all models.
[0070] A new feature representation is obtained by weighting the output of the previous model:
[0071]
[0072] in, It is a mixed feature vector.
[0073] Step 1.4.2: Multi-gradient decision tree prediction: Mix the feature vectors As a new feature, multiple decision trees are trained using gradient descent.
[0074] Overall Predicted Score for:
[0075] in, It is the first A tree, The number of decision trees is denoted by , and each tree is constructed by minimizing the negative gradient direction of the loss function.
[0076] training objective function for:
[0077]
[0078] in It is the mean squared error loss function. For the true value of sample i, It is a regularization term used to control the complexity of the tree.
[0079] For each entity-concept prediction result, a prediction threshold is set, and only prediction concepts that pass the prediction threshold are adopted. In this embodiment, the prediction threshold is 0.90.
[0080] Step 2: A deep domain adversarial anomaly collaborative reasoning model based on a class of transfer learning is used to perform anomaly reasoning checks on the feature-extracted medical consumables data, conducting a preliminary detection to determine if the data is anomalous. Specifically, a feature-sharing encoder is used to extract cross-domain invariant features, a hyperpolyhedron decision boundary constrains the normal data distribution, and the maximum mean difference loss in the feature space aligns inter-domain differences. Simultaneously, adversarial training is used to prevent the discriminator from distinguishing domain origins, achieving anomaly scoring. The trained deep domain adversarial collaborative reasoning model is then used for preliminary screening. The deep domain adversarial anomaly collaborative reasoning model includes a feature domain-sharing encoder, a transfer domain single-class test classifier, a domain label discriminator, a hyperpolyhedron adaptation and MMD normalization integrator, with the overall structure as follows: Figure 3 As shown, the details are as follows:
[0081] Step 2.1: Generation of Multiple Transfer Data. For the source and target domain training data required for transfer learning, a stratified sampling method is constructed to divide the data into M+1 folds (1 source domain data and M target domain data) according to the category of medical consumables. The source domain data consists of standard consumable data, while the target domain data consists of a portion of pre-confirmed consumable data to be processed.
[0082] Step 2.2: Construct a feature domain shared encoder. This involves repeatedly mapping source domain data and target domain data to a common feature space to extract domain-invariant feature representations, obtaining the coordinates of the source and target domain data in the common feature space. This ensures that the distribution of data from the two domains in each group is as aligned as possible within the common feature space. The feature domain shared encoder includes a feature extraction module and a fully connected layer. The specific steps in this embodiment are as follows:
[0083] The characteristics of the mapping between source domain data and target domain data include entity features.
[0084] The goal of the feature domain shared encoder is to learn a mapping function that maps data to a D-dimensional common feature space through a fully connected layer, where D is the dimension of the data information;
[0085] Source domain data The data is composed of the form of D=7, and the M target domain data are ultimately represented as follows: The form is composed of the superscript S, which refers to the source domain, and the superscript mT, which refers to the m-th target domain T. Let these represent the number of parameters in the source domain and the m-th target domain, respectively. and These are the coordinates of the source domain data and the target domain data in the common feature space, respectively.
[0086] Step 2.3: Construct a single-class detection classifier for the transfer domain. The classifier is responsible for implementing the transfer of detection rules and constructing anomaly decision boundaries. A hyperpolyhedral model jointly optimized by two domains is used to represent the boundaries of normal data, and the transfer of detection rules is achieved through hyperpolyhedral spatial constraints. The hyperpolyhedral model specifically includes a source hyperpolyhedron and M target hyperpolyhedrons, each composed of K hyperplanes. The hyperpolyhedrons are used to describe the boundaries of normal data. The specific steps in this embodiment are as follows:
[0087] Feature Hyperpolyhedron Construction: In the common feature space, hyperpolyhedra are constructed for the source and target domains of each group to describe the boundaries of normal data. The source hyperpolyhedron and the M target hyperpolyhedra are each composed of K hyperplanes, each hyperplane... The formula is shown below:
[0088]
[0089] in It is the normal vector matrix, with the superscript T indicating the transpose operation. The normal vector matrix includes the normal vector matrix of the source hyperpolyhedron. and target hyperpolyhedron normal vector matrix k represents the number of the hyperplane in the hyperpolyhedron. The offset includes the source hyperpolygon normal matrix. and target hyperpolyhedron normal vector matrix The source hyperpoly is denoted as The m-th target hyperpoly is denoted as ;
[0090] Hyperpolyhedral space constraints: The empirical loss caused by hyperpolyhedral space constraints is evaluated using the following formula, which is shown below:
[0091]
[0092] in, It is a superpoly experience loss, which controls the compactness of all superpolyhedra; and These are the volume parameters of the source hyperpolyhedron and the target hyperpolyhedron, respectively. The formulas for these volume parameters are shown below:
[0093]
[0094] and The hyperpolyhedral loss from the source neighborhood and the hyperpolyhedral loss from the target neighborhood are respectively represented by the following formulas:
[0095]
[0096]
[0097] in, M is used as a penalty factor to balance the losses in the source and target domains. This represents the total number of parameters across all target domains.
[0098] The hyperpolyhedron spatial constraint measures the discrimination ability of hyperpolyhedra while calculating their volume, aiming to improve the discrimination ability of hyperpolyhedra while minimizing their volume.
[0099] Step 2.4: Construct a domain label discriminator to serve as a domain-invariant feature learned by the adversarial training engine. A two-layer discriminator network performs binary classification on the features, distinguishing whether they originate from the source or target domain. Simultaneously, cross-entropy loss is used to drive the domain-shared encoder to generate indistinguishable features. The specific steps in this embodiment are as follows:
[0100] A two-layer discriminator network is used to perform binary classification of features, and the calculation formula is as follows:
[0101]
[0102] in This indicates the output of the discriminator. This represents the coordinates of the transformed data in the common feature space, including the coordinates of the source domain data. and target domain data coordinates ; , , Here are the weights and biases of two discriminator layers, both of which are multilayer perceptron structures. It is the linearly modified activation function of the first-layer discriminator. is the sigmoid activation function for the output layer.
[0103] The cross-entropy loss caused by the current source polyhedron and the target hyperpolyhedron is calculated using the cross-entropy formula. The specific calculation formula is as follows:
[0104]
[0105] The domain discriminant loss represents the discriminant's ability to distinguish, and its value is positively correlated with the discriminant's ability to distinguish. The discriminant D is trained to distinguish between the source domain and the target domain, while driving the feature extractor to generate domain-invariant features.
[0106] Step 2.5: Construct a hyperpolyhedron adaptation and MMD normalization integrator to evaluate the feature space distribution differences and maximum mean difference (MMD) of the feature space between the target hyperpolyhedron and the source hyperpolyhedron in each group, and control a class of detection rules; the feature space distribution difference is calculated by randomly placing test points in the common feature space and calculating the distribution difference of the test points in the polyhedron, which measures the difference between the learning domain and the target domain in transfer learning; MMD measures the mean divergence of the two distributions in the reproducing kernel Hilbert space by providing a nonparametric metric; the specific implementation in this embodiment is as follows:
[0107] Feature space distribution differences: First, test points are randomly placed within the common feature space. If the test point is located within the source hyperpolyhedron or the target hyperpolyhedron, it is retained until the number of retained test points reaches a large target number N. A At this point, take N. A =1500, the number of test points falling inside the source polyhedron is recorded as . The number of test points falling inside the target hyperpolyhedron is denoted as . The number of test points that simultaneously fall inside both the source polyhedron and the target polyhedron is denoted as . The distributional differences are measured using the following formula:
[0108]
[0109] in, It is the hyperpolyhedral adaptation loss, which represents the difference in the set of mean values of each hyperpolyhedral domain in the total feature space; This represents the influence coefficient, which determines the impact of distribution differences on the total system loss.
[0110] The difference in feature space distribution is used to measure the classification difference between the source hyperpolyhedron and the target hyperpolyhedron, and to measure the difference between the learning domain and the target domain in transfer learning, so as to make the learned standard hyperpolyhedron more robust.
[0111] Maximum Mean Difference in Feature Space (MMD): MMD provides a nonparametric measure of the mean divergence between two distributions in a reproducing kernel Hilbert space (RKHS). The specific formula for the difference is as follows:
[0112]
[0113] in It is the maximum mean difference loss. Update parameters for the domain tag discriminator Minimize the inter-domain distance between any two sets in the regenerating kernel Hilbert space (RKHS); This represents the RHKS norm.
[0114] Step 2.6: Model parameter optimization. An alternating optimization strategy is adopted, which is divided into two sub-optimization components that are performed alternately. The optimization aims to minimize the hyperpolyhedral space constraints, feature space distribution differences, and maximum mean differences in the feature space, while maximizing the cross-entropy loss of the domain name label discriminator. Specifically, the model parameters are adjusted by calculating the loss of each module and the overall model. This mainly consists of total loss calculation and parameter optimization. In this embodiment, the implementation steps are as follows:
[0115] The model's total loss function is shown below:
[0116]
[0117] in, , , , This is a regularization parameter, responsible for controlling the influence of different attribute losses on the model's loss.
[0118] The main goal of model optimization is to minimize , , This results in compact features, overlapping hyperpolyhedra, and aligned distribution; maximizing... This prevents the domain discriminator from distinguishing between the source domain and the target domain.
[0119] An alternating optimization strategy is adopted, which is divided into two sub-optimization components that run alternately, using gradient descent and updating model parameters. The specific execution points are as follows:
[0120] Feature extraction and discrimination optimization: The parameters of the fixed feature domain shared encoder and the transfer domain uniclass test classifier are included. , , , Update the parameters of the domain name label discriminator and calculate the gradient of the parameters through backpropagation.
[0121] Hyper-polyhedral domain attribute optimization: Parameters of the fixed domain tag discriminator were updated. , , The gradient of the parameters is calculated through backpropagation.
[0122] The main adjustment is achieved by calculating the gradients of the hyperpolyhedral empirical loss, discriminator loss, distribution difference loss, and MMD loss and providing gradient changes.
[0123] Step 2.7: Based on the converged source domain hyperpolyhedral decision boundary, use the domain-invariant features and detection rules obtained through transfer learning to detect abnormal data of medical consumables; specifically, the detection rules generate anomaly scores for test samples by calculating the position of the medical consumable sample in the target feature space. A partial two-dimensional projection of the detection results is shown below. Figure 6 As shown in the figure, the shaded area represents the projection of the hyperpolyhedron onto a two-dimensional plane, and the detection rule formula is as follows:
[0124]
[0125]
[0126] in and These are the parameters of the k-th hyperplane in the final targeted hyperpoly after convergence and mean-squared processing. This indicates a cumulative multiplication operation.
[0127] Anomaly scores for the test samples are generated by calculating the position of the medical consumables sample in the target feature space (based on hyperpolyhedrons). This indicates that the model is biased towards considering the sample as a normal sample, and The larger the value, the higher the confidence level; when This indicates that the model is biased towards the sample being abnormal, and at this point, the medical consumable sample can be preliminarily identified as an abnormal sample.
[0128] Finally, the PR curve of the inference detection model on the test set is as follows: Figure 7 As shown, the model achieves a precision of around 90% and a recall of around 60% at the equilibrium point.
[0129] Step 3: A heterogeneous semantic knowledge graph information fusion model based on knowledge graphs and graph neural networks is used to repair and complete the features of medical consumables data detected as anomalous. A heterogeneous semantic knowledge graph is constructed and embedded into a graph neural network. One- and two-hop neighbor subgraphs are extracted centered on anchor triples. Triple-level and graph-level features are fused through a hierarchical graph transformation network to predict the probability distribution of missing entities. Data repair is then performed based on the probability distribution. The overall structure of the fusion model is as follows: Figure 4 As shown, the key points of model construction are as follows:
[0130] Step 3.1: Obtain standard medical consumables data to construct a knowledge graph of triple types; In this embodiment, a medical graph is constructed using standard medical consumables data obtained from local drug procurement centers, including data such as product registration certificate name, registration certificate number, manufacturer and its information, applicant and its information, specifications and model information, material and characteristic information, etc., and combined with the data features extracted by the medical consumables attribute feature cleaning and extraction system mentioned above, a knowledge graph of triple types is constructed.
[0131] Step 3.2: Context-level subgraph construction; Using the triples to be predicted as anchor triples, a context subgraph is constructed for each anchor triple to capture the semantic information of entities in different triples, thus solving the problem of ignoring the global graph structure or introducing noise in traditional methods; Using the subject entity of the anchor triple as the anchor point, other consumable information triples that share entities with the anchor point are selected as one-hop neighbors. Starting from each one-hop neighbor, the two-hop neighbors of the one-hop neighbor as the anchor point are further explored, and these triples are converted into an undirected subgraph, where the nodes of the undirected subgraph are triples, and the edges represent the entity relationships between consumable information triples; The specific implementation steps in this embodiment are as follows:
[0132] Define knowledge graph as ,in , , These represent the entity set, relation set, and triple set, respectively. Each consumable information triple is represented by... Defined in the form of Represents the source entity. Indicates the relationship between entities. Represent the target entity, each defined as Anchor triples construct a context subgraph, where The missing entity to be predicted is the subject entity (the entity that is not missing) of the anchor triple in this subgraph. As anchor point Select other consumable information triples that share the entity with the anchor point as one-hop neighbors. Starting from each one-hop neighbor entity, further explore its two-hop neighbors (2-hop) that are the anchor point. Transform these triples into an undirected subgraph where nodes are triples (not entities) and edges represent entity relationships between consumable information triples.
[0133] The specific construction process is divided into 1-hop one-hop neighbor selection, 2-hop multi-hop neighbor selection, and anchor subgraph construction, and their respective implementation processes are shown below:
[0134] Step 3.2.1: 1-hop Neighbor Selection: Identify all neighbor triplets of the anchor triplet from the relation graph, using the anchor entity as the reference. Find the set of triples centered at the center. Includes anchor entities The triples, which share a common reality with the anchor triples, represent the same entity among multiple triples.
[0135] To prevent an excessively large number of neighbor triples due to an overly large knowledge graph, a hyperparameter is set to limit the number of neighbor triples, and semantically closest triples are selected based on similarity. Given a set of triples, obtain the set of neighbor triples. The similarity is calculated using cosine similarity. These are the top K nearest semantic triples.
[0136] Step 3.2.2: 2-hop Neighbor Selection: For each selected 1-hop neighbor entity, further explore its 1-hop neighbors and filter those with relationships to the anchor point. Related paths.
[0137] For each selected 1-hop neighbor Get the target entity in the one-hop neighbor triplet. 1-hop neighbors For each from Starting Triples Move anchor point s to The path is regarded as a relation sequence The path embedding is obtained by using additive encoding of path semantics. Similarity is calculated by measuring the cosine similarity between path embedding and anchor point relationship, and the semantically most similar path is selected. Given a set of triples, we obtain a 2-hop neighbor triplet set. ,at this time = / 2 is used to control the size of the subgraph. These are the top M semantically closest triples.
[0138] Step 3.2.2: Merging Anchor Point Subgraphs: Finally, the subgraphs constructed from the anchor points are merged.
[0139] Step 3.3: Construct a hierarchical graph transformation network. Based on the obtained knowledge graph and undirected subgraph, fuse triplet-level and graph-level structural features. Through a hierarchical aggregation mechanism, fuse the neighbor information of anchor entities at different distances (1-hop and 2-hop) while maintaining computational efficiency.
[0140] Specifically, the triples are encoded using an encoder to generate embeddings, aggregating the embeddings of shared entities in different contexts: First, one-hop neighbor aggregation is performed, converting the anchor entity vector into a query vector and the one-hop neighbor entity vectors into key and value vectors. Then, the attention score between the anchor entity and each one-hop neighbor entity is calculated, and weighted aggregation is performed based on the node information and attention scores of the neighbor entities to obtain the aggregated single-hop information output. Next, multi-hop neighbor aggregation is calculated, starting from the anchor's one-hop neighbors, using the same method as one-hop neighbor aggregation to obtain the aggregated multi-hop information output for each one-hop neighbor. The average of all aggregated multi-hop information outputs is taken, and then the average of the original representation, the aggregated single-hop information output, and the aggregated multi-hop information output is concatenated to obtain a triple feature representation. In this embodiment, the specific implementation process is as follows:
[0141] Step 3.3.1: Build a triple-level transformer by replacing the entity to be predicted with [mask], generating embeddings from the triples using an encoder, which are then used by the subsequent model for recognition and differentiation; for anchor triples... The entity to be predicted Replace with [mask] tags, and form entity and relation embeddings using a standard positionless Transformer encoder. , , This indicates the embedding of source entities, relations, and target entities;
[0142] For each neighbor triplet ,in Represents all 1-hop neighbors, if the entity to be predicted... If it is the source entity, then we get , These are the embeddings of the relationship and the target entity, respectively.
[0143] Step 3.3.2: Hierarchical Graph Transformer Network: Aggregating Co-existing Entities Embedding in different contexts, the triple-level output is used as sequence input, and graph structure information is fused through an improved self-attention mechanism to generate a globally aware representation. The key points involved are as follows:
[0144] The embedding vectors of the common entities in the triples and context triples (including 1-hop neighbors and 2-hop neighbors) are merged into the input sequence. The detailed conversion formula is shown below:
[0145]
[0146] in It is the co-real entity embedding of the anchor triple. It is the co-present entity embedding of the y-th triple in the context triple set, where Y is the number of context triples and d is the embedding dimension.
[0147] 1-hop neighbor aggregation: Aggregates the interaction information between the anchor entity and its one-hop neighbors (1-hop), and merges the anchor entity vector... Convert to query vector Convert one-hop neighbor entity vectors into key vectors Sum value vector As shown below:
[0148]
[0149] in , , It is a learnable transformation matrix for query, key, and value, used to project the original entity vector onto the query vector, key vector, and value vector, respectively.
[0150] Then, the attention score between the anchor entity and each one-hop neighbor entity is calculated. The formula is as follows:
[0151]
[0152] in Describe the similarity between the anchor entity and its i-th one-hop neighbor entity. Key vectors The vectors in the i-th and j-th rows of the matrix, For path-aware bias. , The relationship of anchor triples. For the relationship of the i-th neighbor triple, It is a two-layer fully connected network.
[0153] By aggregating the information of 1-hop neighbor nodes, the aggregated single-hop information is output. :
[0154] ;
[0155] in Value vector The i-th row vector.
[0156] 2-hop neighbor aggregation: Aggregates the interaction information between the anchor entity and its multi-hop neighbors (2-hop). Using the same method as one-hop neighbor aggregation, it starts from the anchor's one-hop neighbor (1-hop) and gradually aggregates the 2-hop neighbors.
[0157] For each 1-hop entity, calculate its attention with the 2-hop entities to obtain the aggregated multi-hop information output of the i-th one-hop neighbor entity. ;
[0158] Then, the aggregated multi-hop information of all 1-hop entities is output. average value ( This yields the final aggregated multi-hop information output. ;
[0159] Cross-layer information aggregation: through link functions The original representation (the original context triple, By concatenating 1-hop and 2-hop representations, we obtain a triple feature representation. ;
[0160] Step 3.4: Introduce the relation embedding of anchor triples, dynamically fuse the triple feature representations through a gating mechanism to obtain the fused features, then calculate the similarity between the fused features and the candidate entities, and perform normalization to obtain the final repair result. The specific implementation steps in this embodiment are as follows:
[0161] Step 3.4.1: Introduce the relational embedding of anchor triples By dynamically fusing multi-level features through a gating mechanism, fused features are obtained. The specific formula is as follows:
[0162]
[0163]
[0164] in For gating weights, It is the sigmoid function. This represents vector concatenation. This indicates element-wise multiplication.
[0165] Step 3.4.2: Merge features With candidate entities Entity matching results are obtained by performing similarity calculations. The formula is as follows:
[0166]
[0167] in for transpose, Candidate entities Learnable bias terms, relational embedding Projected into the entity embedding space via a multilayer perceptron (MLP).
[0168] Step 3.4.3: Based on the entity matching results The final normalized probabilities are generated. Based on these probabilities, the candidate entity with the highest probability is selected and output as the repair result. The partial feature repair results are as follows: Figure 8 Show.
[0169] Step 4: Based on the type-two fuzzy recognition and analysis model, perform deep discrimination on the corrected abnormal products: Build a dynamic rule engine, dynamically update the rule base through rule dynamic growth, pruning, and fusion, improve interpretability through group-level, rule-level, and feature entropy three-order sparse constraints, output category confidence intervals, and make judgments based on the confidence interval data. For example... Figure 5 As shown, the key points for implementing the model are as follows:
[0170] Step 4.1: Based on the training set data, set G rule sets, each rule set including D interval type II fuzzy sets, where D is the total number of features; in this embodiment, the specific implementation is as follows:
[0171] by The training set data is in the form of , where Representing data information, This indicates whether the classification result is correct or not. Indicates the total number of training set data, in ,in, Let D represent the Dth data information feature, where D represents the dimension of the data feature, i.e., the total number of features. Let G be the set of rules used as the judgment criteria, where the gth rule... As shown below:
[0172]
[0173]
[0174] in, It is the interval type II fuzzy set of the D-th feature in the g-th rule, with a total of D features, where D=7. It contains the discriminant features of the rule, specifically the registration certificate number, registration certificate name, specification information, model information, material information, manufacturer information, and applicant information of the repaired medical consumables. arrive Together, they constitute the antecedent of rule g. This is the output of the g-th rule for the correct judgment class, representing the judgment score that there are no anomalies in the current consumable data. This represents the d-th result parameter of the g-th rule, with a value of 0 or 1, indicating whether the result is passed or not. The intercept is a constant.
[0175] Step 4.2: First, calculate the nonlinear correlation between features using a matrix to obtain the mutual information matrix between features, convert it into a similarity map, and then perform spectral clustering to obtain Z mutually exclusive feature groups. Next, calculate the joint membership degree and its upper and lower bounds in the mutually exclusive feature groups using a multivariate Gaussian membership function. Then, adjust the joint membership degree using an exponential gating function and its gating parameters. Based on the change of the gating parameters over time, calculate the update parameters. When the update parameters exceed a set threshold, feature reorganization is triggered, and rules are regenerated after updating the training data. The specific implementation in this embodiment is as follows:
[0176] high-dimensional feature space Divided into Z mutually exclusive feature groups Each group shares a single gating parameter. This is to reduce the impact of increased feature dimensions on model performance;
[0177] Step 4.2.1: Calculate features using matrix calculations The nonlinear correlation between them is used to replace the traditional linear correlation coefficient. The specific calculation formula is as follows:
[0178]
[0179] in, Representation of features The mutual information matrix, Represents the set of all features. It means that it has been removed. The subsequent feature set, Represents mutual information measurement. For joint probability distribution, This represents a marginal probability distribution.
[0180] Step 4.2.2: [The text abruptly ends here, likely due to an incomplete sentence or a formatting error Convert to similarity graph To meet the input requirements of spectral clustering, according to The degree to which the correlation approaches 1 is used to represent the degree of feature relevance.
[0181] Step 4.2.3: Map the original feature space to a low-dimensional spectral space. By preserving the topological structure, automatic feature grouping is achieved. The specific implementation process is shown below:
[0182] First, construct a degree matrix based on the similarity map. and Laplace matrix Then, it is standardized and decomposed into features, as shown in the following formula:
[0183]
[0184] in For the l-th eigenvalue, For the corresponding eigenvector, The standardized result matrix, The function is the eigenvalue decomposition function.
[0185] extract The first Z eigenvectors are obtained as follows: ,in The parameters are adaptive, and they also represent the Z-dimensional embedding space, ultimately forming spectral clustering: ,in For spectral clustering algorithms, the rows() function represents the expression for the feature matrix. Extract by row.
[0186] Step 4.2.4: Membership Adjustment Based on Shared Gating Parameters: The gating parameter sharing mechanism achieves efficient feature selection by sharing gating parameters among feature groups. The specific calculation process is shown below:
[0187] First, the features in the feature group are calculated using a multivariable Gaussian membership function. joint membership degree and its upper and lower boundaries and The calculation formula is as follows:
[0188]
[0189]
[0190]
[0191] in For data in feature groups The value on, Indicates that rule g is in the feature group The centroid vector on, Indicates that rule g is in the feature group Width vector on, For the Euclidean norm, and These are features g in the feature group. The upper and lower bound parameters are obtained through The calculation yielded, where = , = , The aspect ratio is calculated as follows: .
[0192] Subsequently, an exponential gating function was used to adjust its gating parameters. Applying this to membership, we obtain the gating-adjusted membership. :
[0193]
[0194] in The sharpening function varies over time, and the gate parameters are... Fine-tuning is performed based on the loss gradient descent.
[0195] Step 4.2.5: Adjust the grouping according to the changes in feature importance during the learning process to ensure that the grouping structure is always relevant to the current decision. The specific calculation principle is as follows:
[0196] Calculate the current update parameters
[0197] Represents the gating parameters at time t , These are the gate control parameter values before the last change. This indicates the time since the last change of the gating parameter. It means that when the reorganization condition is met: the current updated parameter is greater than the set threshold (set to 0.3 in this embodiment), feature reorganization is triggered, and the rules are regenerated after replacing the training data.
[0198] Step 4.3: Construct a rule dynamic generation module. After initializing each fuzzy rule, aggregate the interval membership values generated by the membership layer into the trigger intensity interval of each fuzzy rule, quantify the matching degree between the input sample and the fuzzy rule, and generate, trim and fuse the current rule.
[0199] First, the joint membership degree of the feature group and its upper and lower bounds are aggregated to obtain the rule triggering strength and its upper and lower bounds; the specific calculation formula is as follows:
[0200]
[0201]
[0202]
[0203] in, For cumulative multiplication, It is a feature group The joint membership degree on rule g, and These are the upper and lower bounds of the membership degree after gating adjustment. For input samples in feature groups The value on, and These are the upper and lower bounds of the rule trigger strength, respectively.
[0204] Rule set growth and splitting: Calculate the difference between the upper and lower bounds of the rule trigger strength and the activity of the rule trigger strength to determine whether to split the current rule; add random perturbation when splitting the rule;
[0205] The rule growth mechanism increases the capacity of the rule base by splitting existing rules to cover new or highly uncertain regions in the input space. The triggering conditions for rule splitting include the following logical judgments:
[0206]
[0207] in For input data, This represents the difference between the upper and lower bounds of the rule trigger strength. This represents the splitting threshold (default value 0.5), used to determine whether the uncertainty of the rule is high enough. This represents the activity threshold (default value 0.1), used to ensure that the rule is active.
[0208] When a set of rules meets the splitting condition, it will split into multiple new rules along the direction of high uncertainty. The specific operation is as follows:
[0209] First, for the k-th feature group, we examine which feature groups contribute the most uncertainty to the trigger strength calculation. Typically, the feature group contributing the most is selected for splitting.
[0210] Subsequently, for the selected feature group It splits along the centroid direction of this feature set. The centroid of the new rule... Through the original center of mass The formula is obtained by adding a perturbation:
[0211]
[0212] in It is the sample point that triggers the split (i.e., at this point) maximum), It is the offset coefficient (default is 0.3).
[0213] At the same time, adding random perturbations helps prevent the new rules from becoming too similar.
[0214] The width of the new rules Initialize to a certain percentage of the original rule width (default 80%), with the gating parameters remaining unchanged.
[0215] When the total number of rules G reaches the preset upper limit G max The splitting stops at that time.
[0216] Rule set self-pruning: The activity and contribution of rules are calculated; rules below a pruning threshold are considered candidate pruning rules. Then, the importance of these candidate rules is calculated, and the rule with the lowest importance is removed. Specifically, this is achieved through rule pruning. To remove inefficient or redundant rules, the candidate rules for pruning are selected based on the rule's activity and contribution, calculated using the following formula:
[0217]
[0218] in It represents the activity level of rule g at time step t, which is the trigger strength. The average value over the sample. It is a dynamically adjusted cropping threshold. It is the consequent parameter of the rule, equivalent to , This is the consequent parameter norm threshold (default is 0.01), the consequent parameter The initialization formula is as follows:
[0219]
[0220] in For the set of samples covered by the rules, Output the confidence score of the current pattern for sample i. This represents the classification result of whether sample i is actually correct or not.
[0221] The dynamic adjustment formula for the clipping threshold is shown below:
[0222]
[0223] in This is the initial threshold (default is 0.2). The decay rate is 0.01 (default value), and t is the time parameter.
[0224] After obtaining the candidate pruning rules, the importance of each rule is further evaluated to remove the rules with the lowest importance. The calculation formula is as follows:
[0225]
[0226] in It is the L1 norm of the consequent parameter of the rule. For input data.
[0227] Rule set fusion: Rules are fused based on their similarity and consistency. When generating new rules, if two rules are similar, a rule fusion mechanism is used to merge the two similar rules to generate a new fused rule, thereby optimizing the structure and size of the rule base. The specific decision-making steps are as follows:
[0228] First, use the KL divergence of the rule antecedent to analyze the rule. and Similarity between To make a judgment.
[0229] Next, the consistency of the consequents of the rule is verified, using the following formula:
[0230]
[0231] in Indicates consistent computation. This indicates the tolerance for consequent differences (default is 0.2). , Let be the consequent parameter vectors of rules i and j, representing the consequents of rules i and j respectively. .
[0232] Subsequently, rule similarity and consequent consistency are used to make a fusion decision, with the following fusion conditions:
[0233]
[0234] in This indicates the threshold for high similarity (default is 0.8). Indicates the behavior consistency threshold (default is 0.7).
[0235] When the fusion conditions are met, the rule antecedents are first fused, as shown in the following formula:
[0236]
[0237]
[0238] in , The weights of rules i and j are represented by the following formula: Where T is the total time. , Describes the centroid of the k-th feature in rules i and j. Let the centroid of the k-th feature in the new rule be... , Let represent the square of the width of the k-th feature in rules i and j. This represents the width of the k-th feature in the new rule.
[0239] The rule consequents are then fused, as shown in the following formula:
[0240]
[0241] in This indicates element-wise operations. and The rule is to square the element-wise consequent. Let be the consequent vectors of rules i and j, respectively. The function is a symbolic function. This is a small perturbation coefficient (default is 0.01). This is a consequence of the new rule.
[0242] Step 4.4: Third-order coefficient fusion learning: By fusing group-level sparsity, rule-level sparsity, and feature information constraints, multi-level feature selection is achieved to obtain fusion parameters. The core formula is shown below:
[0243]
[0244] The implementation of each part is shown below:
[0245] Sparse between groups Calculate the L0 norm. The key to achieving the balance coefficient is: , This is the smoothing factor (default is 50).
[0246] For feature-level information constraints, Here, Ent(·) is the balance coefficient, D is the number of features, and Ent(·) is the information entropy calculation.
[0247]
[0248]
[0249]
[0250] in This represents the characteristic mean and standard deviation, where B=10 is the fixed number of bins, and b is the bin index. This represents the number of samples in that binning interval. For the compartment numbered b, For indicator functions, .
[0251] It is a regular-level sparsity. For balance coefficient, The key calculation method for the L1 / 2 norm is as follows: , for The c-th parameter in the equation, where C is the total number of parameters.
[0252] Step 4.5: Transform the trigger intensity range into a deterministic confidence range, and then display the confidence level of each category through probabilistic output. The specific implementation method in this embodiment is as follows:
[0253] First, construct a type simplification layer to reduce the interval trigger strength generated by the rule layer. Transform into deterministic interval output This provides confidence intervals for each category in the final output layer. The type reduction layer uses an algorithm to calculate the left and right endpoints of each output component. and The calculation formula is as follows:
[0254]
[0255] in, express Values sorted in ascending order and Indicates and The corresponding trigger strength upper and lower bounds, with L and N as switching points, satisfy the following critical conditions, where and For G=L and G=N respectively :
[0256]
[0257] The final output layer will output the deterministic range of the type reduction layer. The model is transformed into a category probability distribution, and the probabilistic output intuitively displays the model's confidence level for each category, enabling multi-class decision-making. The entire model is optimized by the cross-entropy loss between the output probability distribution and the true label. The implementation process is as follows:
[0258] Calculate category output value : Trigger strength of the input rule layer and the result parameters of the rule set We then perform a weighted summation, using the following formula:
[0259]
[0260] in This indicates the basic tendency of rule g towards category C. This represents the weight of feature z in rule g on category C. Let z be the z-th feature of the input data.
[0261] By using the softmax function Convert to probability .
[0262] The form is .
[0263] Some test results are as follows Figure 9 As shown.
[0264] It is understood that the present invention has been described through some embodiments, and those skilled in the art will recognize that various changes or equivalent substitutions can be made to these features and embodiments without departing from the spirit and scope of the invention. Furthermore, under the teachings of the present invention, these features and embodiments can be modified to adapt to specific situations and materials without departing from the spirit and scope of the invention. Therefore, the present invention is not limited to the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application are within the protection scope of the present invention.
Claims
1. A type II fuzzy screening method for abnormal medical consumable data based on cross-domain collaboration and heterogeneous graphs, characterized in that, Includes the following steps: Step 1: Obtain the required medical consumables dataset and preprocess the data, construct positive sample data, negative sample data and target concept set, process the sample data through the pre-model and provide basic prediction results or features, calculate the weight of the pre-model based on the attention mechanism, and finally make the final prediction through the gradient decision tree to generate standardized entity-concept features. Step 2: Use a deep domain adversarial anomaly collaborative reasoning model based on a class of transfer learning to perform anomaly reasoning checks on the medical consumables data after feature extraction, and conduct preliminary detection to determine whether the data is anomalous. Specifically, a feature-sharing encoder is used to extract cross-domain invariant features, a hyperpolyhedral decision boundary is used to constrain the normal data distribution, and the maximum mean difference loss in the feature space is used to align inter-domain differences. At the same time, adversarial training is used to make the discriminator unable to distinguish the domain source, thereby achieving anomaly scoring. The trained deep domain adversarial collaborative reasoning model is then used for preliminary screening. Step 3: Based on the heterogeneous semantic knowledge graph information fusion model of knowledge graph and graph neural network, feature repair and completion are performed on the medical consumables features detected as abnormal data; a heterogeneous semantic knowledge graph is constructed and embedded into the graph neural network. The 1- and 2-hop neighbor subgraphs are extracted with anchor triples as the center. The triple-level and graph-level features are fused through the hierarchical graph transformation network to predict the probability distribution of missing entities. Data repair is performed according to the probability distribution. Step 4: Based on the type II fuzzy recognition and analysis model, perform in-depth discrimination on the corrected abnormal products: Build a dynamic rule engine, dynamically update the rule base through rule dynamic growth, pruning and fusion, improve interpretability through group-level, rule-level and feature entropy three-order sparse constraints, output category confidence intervals, and make judgments based on the confidence interval data.
2. The method for type II fuzzy screening of abnormal medical consumable data based on cross-domain collaboration and heterogeneous graphs according to claim 1, characterized in that, Step 1 is described in detail as follows: Step 1.1: Obtain the required training dataset and preprocess the data, which includes the medical consumables data to be processed and standard medical consumables product data; Step 1.2: Obtain basic positive samples, pair entities with standard concepts using standard labeled data, and then construct a positive sample dataset through a three-layer expansion strategy of data transfer expansion, symmetric expansion, and external knowledge expansion; filter out interfering erroneous concepts through semantic similarity retrieval to build a negative sample dataset; and generate a target concept set by filtering out standard concepts related to the input entities based on information retrieval technology. Step 1.3: Building the fusion architecture model, including the construction of the dual-stream transform attention model, the position-aware enhancement model, the dual-tower concept mapping model, and the implementation of the hybrid architecture; Write a permutation language model that samples a permutation of all input word lists in the dataset and uses only positions when predicting the target word. The previous content, Indicates the target words in the arrangement The position within the content stream; the two-stream transform attention model uses the content stream and query stream respectively to obtain the position of the content stream and query stream. The hidden state; The location-aware enhancement model addresses the location-content coupling problem and the relative location modeling problem by building a decoupled attention and enhanced mask decoder. The dual-tower concept mapping model independently encodes entities and concepts using two encoders with shared weights, and then compares their embeddings to generate semantic matching scores for entity-concept pairs. Step 1.4: Calculate and weight the weights of the three pre-models using an attention mechanism to obtain a mixed feature vector, and finally make the final prediction using a gradient decision tree.
3. The method for type II fuzzy screening of abnormal medical consumable data based on cross-domain collaboration and heterogeneous graphs according to claim 2, characterized in that, The deep domain adversarial anomaly collaborative reasoning model includes a feature domain shared encoder, a transfer domain single-class test classifier, a domain label discriminator, a hyperpolyhedron adaptor, and an MMD normalization integrator. Step 2 is as follows: Step 2.1: Generation of multiple transfer data; For the source domain and target domain training data required for transfer learning, a stratified sampling method is constructed to divide the data into M+1 folds according to the category of medical consumables, including 1 source domain data and M target domain data, where the source domain data is standard consumable data and the target domain data is consumable data to be processed; Step 2.2: Construct a feature domain shared encoder. By repeatedly mapping source domain data and target domain data to a common feature space, extract domain-invariant feature representations to obtain the coordinates of source domain data and target domain data in the common feature space. The feature domain shared encoder includes a feature extraction module and a fully connected layer; Step 2.3: Construct a transfer domain single-class test classifier, and use a hyperpolyhedron model jointly optimized by two domains to represent the boundary of normal data, and realize the transfer of detection rules through hyperpolyhedron spatial constraints; the hyperpolyhedron model specifically includes a source hyperpolyhedron and M target hyperpolyhedrons, each of which is composed of K hyperplanes. The hyperpolyhedron is used to describe the boundary of normal data. Step 2.4: Construct a domain label discriminator as an adversarial training engine. A two-layer discriminator network is used to perform binary classification of features to distinguish whether the features come from the source domain or the target domain. At the same time, cross-entropy loss is used to drive the domain-sharing encoder to generate indistinguishable features. Step 2.5: Construct a hyperpolyhedron adaptation and MMD normalization integrator, evaluate the feature space distribution differences and maximum mean difference (MMD) of the target hyperpolyhedron and source hyperpolyhedron in each group, and control a class of detection rules; The feature space distribution difference is measured by randomly placing test points in a common feature space and calculating the distribution difference of the test points on the polyhedron, which measures the difference between the learning domain and the target domain in transfer learning. MMD measures the mean divergence of two distributions in the reproducing kernel Hilbert space by providing a nonparametric metric. Step 2.6: Model parameter optimization. An alternating optimization strategy is adopted, which is divided into two sub-optimization components that are performed alternately to minimize the hyperpolyhedral space constraints, feature space distribution differences, and maximum mean differences in the feature space, while maximizing the cross-entropy loss of the domain name label discriminator. Step 2.7: Based on the converged source domain hyperpolyhedral decision boundary, use the domain-invariant features and detection rules obtained from transfer learning to detect abnormal data of medical consumables; the detection rules specifically generate anomaly scores for test samples by calculating the position of medical consumable samples in the target feature space.
4. The method for type II fuzzy screening of abnormal medical consumable data based on cross-domain collaboration and heterogeneous graphs according to claim 3, characterized in that, Step 3 is as follows: Step 3.1: Obtain standard medical consumables data and construct a knowledge graph of triple types; Step 3.2: Context-level subgraph construction; Using the triplet data to be predicted as anchor triplets, construct a context subgraph for each anchor triplet to be predicted; Using the topic entity of the anchor triplet as the anchor point, select other consumable information triplets that share entities with the anchor point as one-hop neighbors. Starting from each one-hop neighbor, further explore the two-hop neighbors that are the anchor points, and convert these triplets into an undirected subgraph, where the nodes of the undirected subgraph are triplets and the edges represent the entity relationships between consumable information triplets; Step 3.3: Construct a hierarchical graph transformation network. Based on the obtained knowledge graph and undirected subgraph, the information of the anchor entity's neighbors at different distances is fused through a hierarchical aggregation mechanism. The triples are used to generate embeddings using an encoder, and the embeddings of shared entities in different contexts are aggregated: First, one-hop neighbor aggregation is performed, the anchor entity vector is converted into a query vector, and the one-hop neighbor entity vector is converted into a key vector and a value vector. Then, the attention score between the anchor entity and each one-hop neighbor entity is calculated. Based on the node information and attention score of the neighbor entities, a weighted aggregation is performed to obtain the aggregated single-hop information output. Then, multi-hop neighbor aggregation is calculated. Starting from the one-hop neighbor of the anchor point, the same method as one-hop neighbor aggregation is used to obtain the aggregated multi-hop information output of each one-hop neighbor. The average value of all aggregated multi-hop information outputs is taken, and then the average value of the original representation, aggregated single-hop information output, and aggregated multi-hop information output is concatenated to obtain the triple feature representation. Step 3.4: Introduce the relation embedding of anchor triples, dynamically fuse the triple feature representation through a gating mechanism to obtain the fused feature, then calculate the similarity between the fused feature and the candidate entity, and normalize it to obtain the final repair result.
5. The method for type II fuzzy screening of abnormal medical consumable data based on cross-domain collaboration and heterogeneous graphs according to claim 4, characterized in that, Step 4 is as follows: Step 4.1: Based on the training set data, set G rule sets, each rule set including D interval type II fuzzy sets, where D is the total number of features; Step 4.2: First, calculate the nonlinear correlation between features using a matrix to obtain the mutual information matrix between features, convert it into a similarity map, and then perform spectral clustering to obtain Z mutually exclusive feature groups. Then, calculate the joint membership degree and upper and lower bounds of the joint membership degree of the features in the mutually exclusive feature groups using a multivariate Gaussian membership function. Subsequently, adjust the joint membership degree using an exponential gating function and its gating parameters. Calculate the update parameters based on the change of the gating parameters over time. When the update parameters exceed a set threshold, feature recombination is triggered, and rules are regenerated after updating the training data. Step 4.3: Construct a dynamic rule generation module to generate, trim, and merge the current rules; First, the joint membership degree of the feature group and its upper and lower bounds are aggregated to obtain the rule triggering strength and its upper and lower bounds. Rule set growth and splitting: Calculate the difference between the upper and lower bounds of the rule trigger strength and the activity of the rule trigger strength to determine whether to split the current rule; add random perturbation when splitting the rule; Rule set self-pruning: Calculate the activity and contribution of the rules, and use the rules that are less than the pruning threshold as candidate pruning rules. Then calculate the importance of the candidate pruning rules and remove the rules with the lowest importance. Rule set fusion: fusing rules based on their similarity and consistency; Step 4.4: Third-order coefficient fusion learning: Multi-level feature selection is achieved by fusing group-level sparsity, rule-level sparsity, and feature information constraints; Step 4.5: Transform the trigger intensity range into a deterministic confidence range, and then display the confidence level of each category through probabilistic output.
Citation Information
Patent Citations
Knowledge graph multi-hop reasoning method based on reinforcement learning
CN117217305A
Medical consumable audit-oriented monitoring evaluation method and system and storage medium
CN119004318A