Artificial intelligence-based precursor drug dynamic monitoring and early warning system
By collecting multi-source heterogeneous data, using the BERT model and deep learning model, we built a knowledge graph of precursor drugs, enterprises, personnel, and geographic space, solving the problems of data dispersion and insufficient supervision in existing technologies, achieving efficient fusion and dynamic monitoring of multimodal data, and improving the adaptability and early warning capabilities of supervision.
Patent Information
- Application Number
- CN202510808021.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-09-26
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The data of existing monitoring methods for precursor drugs are scattered, making it difficult to integrate multimodal data. Existing supervision relies on manual screening, making it difficult to discover complex correlations. The existing system cannot adapt to regulatory updates and changes in distribution routes, and lacks multi-dimensional anomaly evaluation values for early warning analysis.
Through the multi-source heterogeneous data collection and transmission module, the data preprocessing and knowledge graph construction module and the anomaly evaluation and early warning module, real-time data collection, processing and dynamic updating of the knowledge graph are realized. The BERT model and deep learning model are used for entity recognition and relationship calculation, and a knowledge graph of precursor drugs, enterprises, personnel and geographic space is constructed, and early warning is issued in combination with anomaly evaluation values.
It achieves efficient fusion and real-time monitoring of multimodal data, improves entity classification accuracy and relationship confidence, supports dynamic association, can trigger early warnings in a timely manner, and improves the adaptability and accuracy of supervision.
Smart Images

Figure CN120708844A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of public safety supervision technology, and specifically relates to an artificial intelligence-based dynamic monitoring and early warning system for precursor drugs. Background Art
[0002] With the increasing informatization of drug production, circulation and sales, precursor drugs, due to their special chemical properties, may become key raw materials for drug manufacturing when illegally used.
[0003] Although existing monitoring methods have achieved the supervision of precursor drugs to a certain extent, the data is relatively scattered. The data of production enterprises, logistics, sales and regulations are scattered in different systems, making it difficult to integrate multimodal data. Existing supervision relies on manual investigation, which makes it difficult to discover complex associations. Existing systems are mostly static rule matching and cannot adapt to dynamic scenarios such as regulatory updates and changes in distribution paths. It is difficult to build corresponding knowledge graphs and update them dynamically. They only focus on attribute values or relationship frequencies, ignore entity connection patterns and contextual semantics, and lack the construction of multi-dimensional anomaly evaluation values for early warning analysis.
[0004] Therefore, existing technologies have shortcomings in multimodal data fusion, dynamic updating of knowledge graphs, and multidimensional anomaly evaluation. There is an urgent need for an artificial intelligence-based dynamic monitoring and early warning system for precursor drugs. Summary of the Invention
[0005] The present invention aims to solve at least one of the technical problems existing in the prior art. To this end, the present invention proposes an artificial intelligence-based dynamic monitoring and early warning system for precursor drugs, which is used to solve the following technical problems:
[0006] Although existing monitoring methods have achieved the supervision of precursor drugs to a certain extent, the data is relatively scattered. The data of production enterprises, logistics, sales and regulations are scattered in different systems, making it difficult to integrate multimodal data. Existing supervision relies on manual investigation, which makes it difficult to discover complex associations. Existing systems are mostly static rule matching and cannot adapt to dynamic scenarios such as regulatory updates and changes in distribution paths. It is difficult to build corresponding knowledge graphs and update them dynamically. They only focus on attribute values or relationship frequencies, ignore entity connection patterns and contextual semantics, and lack the construction of multi-dimensional anomaly evaluation values for early warning analysis.
[0007] To solve the above problems, the present invention provides an artificial intelligence-based dynamic monitoring and early warning system for precursor drugs, which includes the following modules:
[0008] Multi-source heterogeneous data collection and transmission module: used by manufacturing enterprise ERP systems to push drug batch information in real time via APIs, logistics enterprise TMS systems to push GPS tracks in real time via Kafka message queues, pharmacy HIS systems to upload sales data daily via WebService interfaces, and drug administration bureaus to push regulatory documents via APIs;
[0009] Data preprocessing and knowledge graph construction module: This module processes multimodal data separately, defines entity types and relationship types, calculates entity category probabilities, analyzes the contextual information of entity pairs in the text, calculates the relationships and confidence levels between entity pairs, constructs a knowledge graph of precursor drugs, enterprises, personnel, and geographic space, and defines relationship rules and dynamically updates them.
[0010] The calculation of entity category probability generates a hidden state matrix through the BERT model, calculates the entity dictionary matching feature value in combination with the edit distance, introduces a part-of-speech tagging model, calculates the part-of-speech tagging feature value in combination with the weights of nouns and gerunds, concatenates the BERT vector with the mapped features, and calculates through a fully connected layer and Softmax;
[0011] Abnormal evaluation and early warning module: extract entity features, relationship features and attribute features respectively, construct entity abnormal evaluation value, relationship abnormal evaluation value and attribute abnormal evaluation value, add them up by weight to obtain comprehensive abnormal evaluation value, and trigger early warning based on dynamic threshold.
[0012] Preferably, the multi-source heterogeneous data acquisition and transmission module includes:
[0013] The manufacturer's ERP system pushes drug batch information in real time through the API. The data is encrypted and transmitted to the data platform via HTTPS and stored in a relational database.
[0014] The logistics company's TMS system pushes GPS tracks in real time through the Kafka message queue. The data is cleaned and stored in a time series database. Transaction records are uploaded daily via SFTP as CSV files, parsed, and stored in the operational data layer.
[0015] The pharmacy HIS system uploads sales data daily through the WebService interface. After verification, the data is linked to the drug dictionary and manufacturer information and stored in the public dimension layer.
[0016] The Drug Administration pushes the latest regulatory documents through the API. The data is encrypted and transmitted and stored in the application data layer, triggering dynamic updates of the knowledge graph.
[0017] Preferably, the multimodal data is processed separately to define entity types and relationship types, including:
[0018] Preprocess the collected data, specifically by directly mapping fields to structured data and unifying the field format; using the BERT model to extract entities and relationships from text for unstructured data, and identifying sensitive information based on a keyword library; and parsing CSV files uploaded via SFTP for semi-structured data, extracting fields and converting them to a unified format. Furthermore, data cleaning operations are performed, including deduplication, standardization, and anomaly correction.
[0019] Based on business needs and data characteristics, define the entity types and relationship types of the knowledge graph, calculate the entity recognition probability, and calculate the relationship and confidence between entity pairs by analyzing the contextual information of entity pairs in the text;
[0020] Define entity types. Core entities include precursor drugs, enterprises, personnel, and geographic space names. Extended entities include laws and policies. Define relationship types. Core relationships include precursor drugs-enterprise, enterprise-personnel, enterprise-geographic space, and precursor drugs-geographic space. Use SQL to associate data from multiple tables. Dynamic relationships include regulatory associations.
[0021] Preferably, the construction of a knowledge graph of precursor drugs-enterprises-personnel-geographic space, definition of relationship rules and dynamic updates include:
[0022] Use regular expressions to extract fixed-format fields, use the BERT model to identify entities in text, convert addresses to latitude and longitude through the Baidu API, link them to geospatial nodes, and distinguish entities with the same name based on context.
[0023] Define relationship rules: Production relationships are associated with the manufacturer ID and drug batch number in ERP data, and distribution path relationships are generated through GPS tracks in TMS to generate path nodes;
[0024] Use a deep learning model to label relational data. The training set comes from the labeled relations in the historical data. Use the training set to train the model and output labels.
[0025] Use Neo4j to store nodes and relationships, with core entities stored in the main graph and geospatial nodes and regulatory policies stored in the extended graph.
[0026] Preferably, the calculating entity category probability includes:
[0027] The input text A is processed by the BERT word segmenter to obtain the vocabulary sequence C = {c1, c2, ..., c w}, through the pre-trained BERT model, each vocabulary c i Mapped to a high-dimensional vector p i , get the hidden state matrix, specifically: P1={p1,p2,…,p w};
[0028] At the same time, the vocabulary c is calculated i With dictionary D C The normalized edit distance of each entity in , takes the maximum value as the corresponding entity dictionary matching feature value, specifically:
[0029]
[0030] Among them, f e (c i , a) is vocabulary c i The maximum similarity under category a, c i is the vocabulary in the input text, a is the entity category, D c is the entity dictionary corresponding to category c, d is D c An entity word in ED(c i , d) is c i The edit distance between d and f, i.e. the minimum number of edit operations, e (c i ) is the corresponding entity dictionary matching feature value, and A is the set of all entity categories;
[0031] The probability output of the part-of-speech tagging model is used to obtain the part-of-speech tagging feature value, specifically:
[0032] f p (c i )=K(p(c i )=n)+ω vn K(p(c i )=vn)
[0033] Among them, f p (c i ) is the feature value of part-of-speech tagging, p(c i ) is vocabulary c i The corresponding part-of-speech tags, K(p(c i )=n) is the model prediction c i is the probability of a noun, K(p(c i )=vn) is the model prediction c i is the probability of the gerund, ω vn is the weight coefficient of the gerund.
[0034] Preferably, the calculating entity category probability further includes:
[0035] The entity dictionary matching feature value and the part-of-speech tagging feature value are constructed to obtain the feature matrix F = [f e (c i ), f p (c i ),…,f e(c n ), f p (c n )], normalize the feature matrix and map the normalized features to the entity category set A, specifically:
[0036] P2=F norm W norm
[0037] Among them, P2 is the feature after mapping, F norm is the normalized feature, W norm is the mapping weight matrix;
[0038] The BERT vector P1 and the mapped feature P2 are concatenated into a mixed feature: P = [P1, P2]. The entity category probability of each word is calculated through the fully connected layer and Softmax, specifically:
[0039] H=Softmax(P·w h +b h )
[0040] Among them, H is the entity category probability, P is the mixed feature, w h and b h are the weights and biases of the classification layer.
[0041] Preferably, the calculating of the relationship and confidence between entity pairs by analyzing contextual information of entity pairs in text includes:
[0042] Extraction is performed based on a method that combines semantic rules and machine learning. For the identified entity pairs (d1, d2), the relationship and confidence between the entity pairs (d1, d2) are calculated by analyzing the contextual information of the entity pairs in the text. Specifically:
[0043] CL(d1, d2, r)=α1×M(d1, d2)+α2×P(r|context)
[0044] Among them, CL(d1, d2, r) is the confidence of the relationship r between the entity pair (d1, d2), r is the relationship type, which indicates the association between the entity pairs, M(d1, d2) is the entity matching degree calculated based on semantic rules, P(r|context) is the probability of the relationship type r appearing under given context conditions, context is the context information, and α1 and α2 are the corresponding weight coefficients.
[0045] Preferably, the abnormality evaluation and early warning module includes:
[0046] Extract features from the constructed knowledge graph of precursor drugs, enterprises, people, and geographic space, including entity features, relationship features, and attribute features;
[0047] The entity features include entity degree and entity attribute distribution; the relationship features include relationship frequency and relationship path entropy; the attribute features include attribute value deviation and attribute update frequency;
[0048] Construct entity anomaly evaluation value, relationship anomaly evaluation value and attribute anomaly evaluation value, and perform weighted addition to obtain the comprehensive anomaly evaluation value, specifically:
[0049] U=β1*u1+β2*u2+β3*u3
[0050] Among them, U is the comprehensive anomaly evaluation value, u1, u2 and u3 are the entity anomaly evaluation value, relationship anomaly evaluation value and attribute anomaly evaluation value respectively, β1, β2 and β3 are the corresponding weight coefficients respectively;
[0051] If the comprehensive abnormal evaluation value is greater than the preset threshold, an early warning is triggered; otherwise, continuous monitoring is performed.
[0052] Preferably, the constructing of entity anomaly evaluation values, relationship anomaly evaluation values and attribute anomaly evaluation values includes:
[0053] The entity anomaly evaluation value is obtained by collecting the degrees of all entities in the graph and calculating the mean and standard deviation of all entity degrees, specifically:
[0054]
[0055] Among them, u1 is the entity abnormality evaluation value, r s is the degree of an entity s, μ r is the mean of all entity degrees, σ r is the standard deviation of all entity degrees;
[0056] The relationship anomaly evaluation value is obtained by collecting the current relationship frequency, the current relationship path entropy, the mean and standard deviation of all relationship frequencies, and the mean and standard deviation of all relationship path entropies, specifically:
[0057]
[0058] Among them, u2 is the relationship abnormality evaluation value, γ is the weight coefficient, f now and h now is the current relationship frequency and relationship path entropy, μ f and μ h is the mean of all relationship frequencies and the mean of relationship path entropy, σ f and σ his the standard deviation of all relationship frequencies and the standard deviation of shutdown path entropy;
[0059] The attribute anomaly evaluation value shown is obtained by collecting the extreme values of the deviation of all entity attribute values, the mean and standard deviation of all attribute update frequencies, and is specifically:
[0060]
[0061] Among them, u3 is the attribute abnormality evaluation value, Δoff s is the attribute value deviation of an entity s, max(Δoff) is the maximum value of the attribute value deviation of all entities, is the weight coefficient, is the attribute update frequency of an entity s, μ F and σ F Update the mean and standard deviation of the frequencies for all attributes.
[0062] Beneficial effects of the present invention:
[0063] The present invention integrates ERP, TMS, HIS and drug administration data through API, Kafka, WebService and other protocols, supports HTTPS encrypted transmission and hierarchical storage, and ensures data integrity and timeliness;
[0064] This invention improves the accuracy of entity classification by combining BERT word segmentation, edit distance calculation and part-of-speech tagging, and calculates relationship confidence by analyzing contextual information. It supports dynamic association of production relations, circulation paths and regulations. Baidu API converts addresses into longitude and latitude to achieve the integration of geographic nodes and entity graphs. Regulatory documents trigger knowledge graph updates, and deep learning models annotate historical relationship data, improving graph adaptability.
[0065] The present invention extracts corresponding entity features, relationship features and attribute features based on the analysis of the constructed knowledge graph, constructs corresponding abnormality evaluation values respectively, and performs weighted addition to obtain a comprehensive abnormality evaluation value. When the comprehensive weighted evaluation value exceeds the threshold, an early warning is triggered for continuous monitoring and active intervention. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] Figure 1 It is a schematic diagram of the module flow of the present invention. DETAILED DESCRIPTION
[0067] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.
[0068] See also Figure 1 As shown, the present invention is an artificial intelligence-based dynamic monitoring and early warning system for precursor drugs, which includes the following modules:
[0069] Multi-source heterogeneous data collection and transmission module: used by manufacturing enterprise ERP systems to push drug batch information in real time via APIs, logistics enterprise TMS systems to push GPS tracks in real time via Kafka message queues, pharmacy HIS systems to upload sales data daily via WebService interfaces, and drug administration bureaus to push regulatory documents via APIs;
[0070] Data preprocessing and knowledge graph construction module: This module processes multimodal data separately, defines entity types and relationship types, calculates entity category probabilities, analyzes the contextual information of entity pairs in the text, calculates the relationships and confidence levels between entity pairs, constructs a knowledge graph of precursor drugs, enterprises, personnel, and geographic space, and defines relationship rules and dynamically updates them.
[0071] The calculation of entity category probability generates a hidden state matrix through the BERT model, calculates the entity dictionary matching feature value in combination with the edit distance, introduces a part-of-speech tagging model, calculates the part-of-speech tagging feature value in combination with the weights of nouns and gerunds, concatenates the BERT vector with the mapped features, and calculates through a fully connected layer and Softmax;
[0072] Abnormal evaluation and early warning module: extract entity features, relationship features and attribute features respectively, construct entity abnormal evaluation value, relationship abnormal evaluation value and attribute abnormal evaluation value, add them up by weight to obtain comprehensive abnormal evaluation value, and trigger early warning based on dynamic threshold.
[0073] In one embodiment of the present invention, the multi-source heterogeneous data acquisition and transmission module includes:
[0074] The manufacturer's ERP system pushes drug batch information in real time through the API. The data is encrypted and transmitted to the data platform via HTTPS and stored in a relational database.
[0075] The logistics company's TMS system pushes GPS tracks in real time through the Kafka message queue. The data is cleaned and stored in a time series database. Transaction records are uploaded daily via SFTP as CSV files, parsed, and stored in the operational data layer.
[0076] The pharmacy HIS system uploads sales data daily through the WebService interface. After verification, the data is linked to the drug dictionary and manufacturer information and stored in the public dimension layer.
[0077] The Drug Administration pushes the latest regulatory documents through the API. The data is encrypted and transmitted and stored in the application data layer, triggering dynamic updates of the knowledge graph.
[0078] Specifically, the API is connected to the ERP system of the manufacturing enterprise, and a push cycle is set, such as real-time / minute level. The ERP system pushes drug batch information (such as drug name, batch number, production date, and manufacturer ID) to the data platform through the HTTPS protocol. The received data is stored in a relational database, and the field mapping is a unified format; the logistics enterprise TMS system encapsulates GPS trajectory data (latitude and longitude, timestamp, vehicle ID) as a message, sends it to the Kafka topic, consumes Kafka messages, filters abnormal coordinates, fills in missing fields, and stores the cleaned data in a time series database, storing trajectory points in time series; the pharmacy HIS system uploads sales data (drug name, sales quantity, pharmacy ID, sales time) through the WebService interface every day, checks required fields, associates the drug dictionary, stores the data in the public dimension layer, and associates it with the drug dictionary and manufacturer information; the Drug Administration pushes regulatory documents (PDF / XML format) through the API, including metadata such as file name, release date, and effective date, and uses HTTPS or national encryption algorithm to encrypt and transmit files. The files are stored in the application data layer, and the file content is parsed and key terms are extracted.
[0079] In one embodiment of the present invention, the processing of multimodal data separately to define entity types and relationship types includes:
[0080] Preprocess the collected data, specifically by directly mapping fields to structured data and unifying the field format; using the BERT model to extract entities and relationships from text for unstructured data, and identifying sensitive information based on a keyword library; and parsing CSV files uploaded via SFTP for semi-structured data, extracting fields and converting them to a unified format. Furthermore, data cleaning operations are performed, including deduplication, standardization, and anomaly correction.
[0081] Based on business needs and data characteristics, define the entity types and relationship types of the knowledge graph, calculate the entity recognition probability, and calculate the relationship and confidence between entity pairs by analyzing the contextual information of entity pairs in the text;
[0082] Define entity types. Core entities include precursor drugs, enterprises, personnel, and geographic space names. Extended entities include laws and policies. Define relationship types. Core relationships include precursor drugs-enterprise, enterprise-personnel, enterprise-geographic space, and precursor drugs-geographic space. Use SQL to associate data from multiple tables. Dynamic relationships include regulatory associations.
[0083] Specifically, multimodal data is first preprocessed, fields are directly mapped for structured data, and the time format is unified; for unstructured data, the BERT model is used to extract entities and relationships from text, for example, entities (ephedrine), relationships (controlled by) and context (prohibited from circulation) are identified from regulatory documents; for semi-structured data, CSV files are parsed, fields are extracted and converted into JSON format; entity types and relationship types are defined, including: entity types are divided into core entities and extended entities, for example, core entities: precursor drugs (such as "ephedrine"), enterprises (such as "XX Pharmaceutical Factory"), personnel (such as "Zhang San"), geographic space (such as "Shanghai"), and extended entities: laws and policies (such as "Drug Control Regulations"); relationship types include core relationships and dynamic relationships, core relationships: production (enterprise → drug), employment (enterprise → personnel), location (enterprise → geographic space), circulation (drug → geographic space), and dynamic relationships: regulatory associations (such as "ephedrine is controlled by Drug Control Regulations").
[0084] In one embodiment of the present invention, the construction of a knowledge graph of precursor drugs-enterprises-personnel-geographic space, definition of relationship rules and dynamic update includes:
[0085] Use regular expressions to extract fixed-format fields, use the BERT model to identify entities in text, convert addresses to latitude and longitude through the Baidu API, link them to geospatial nodes, and distinguish entities with the same name based on context.
[0086] Define relationship rules: Production relationships are associated with the manufacturer ID and drug batch number in ERP data, and distribution path relationships are generated through GPS tracks in TMS to generate path nodes;
[0087] Use a deep learning model to label relational data. The training set comes from the labeled relations in the historical data. Use the training set to train the model and output labels.
[0088] Use Neo4j to store nodes and relationships, with core entities stored in the main graph and geospatial nodes and regulatory policies stored in the extended graph.
[0089] Specifically, use regular expressions to extract fixed format fields, for example, drug batch number: ABC12345, manufacturer: XX Pharmaceutical Factory, production date: 2023-01-01; regular expression: drug batch number: (\w+), manufacturer: ([\u4e00-\u9fa5]+), production date: (\d{4}-\d{2}-\d{2}), extract fixed fields, batch_id, producer_name, production_date; use Python's re library or Java's Pattern class to match data; among them, entities with the same name are distinguished by context, for example, "Beijing" in the text may refer to a geographic space or a company name, and the rule definition is: if keywords such as "located" and "address" appear before and after the entity, it is marked as a geographic space. If it is different from "production" and " If the data is associated with "employment", it is marked as an enterprise by adding rule judgment based on BERT output; relationship rules are defined, including production relationships (ERP data), by generating associations between enterprise IDs and drug batch numbers, and circulation path relationships (TMS data), and path nodes are generated through GPS trajectories; annotated relationships are collected from historical data, each data is (head entity, tail entity, relationship type), and a BiLSTM-CRF or Transformer-XL model is selected to convert text and relationship labels into model inputs. The model is trained using the training set, and the optimization objective is log-likelihood loss; new text data is input into the trained model, the predicted relationship labels are output, and the relationships are stored in Neo4j; core entities are stored in the main graph, including precursor drugs, enterprises, and personnel, and geographic spatial nodes and regulatory policy nodes are stored in the extended graph.
[0090] In one embodiment of the present invention, calculating the entity category probability includes:
[0091] The input text A is processed by the BERT word segmenter to obtain the vocabulary sequence C = {c1, c2, ..., c w}, through the pre-trained BERT model, each vocabulary c i Mapped to a high-dimensional vector p i , get the hidden state matrix, specifically: P1={p1,p2,…,p w};
[0092] At the same time, the vocabulary c is calculated i With dictionary D C The normalized edit distance of each entity in , takes the maximum value as the corresponding entity dictionary matching feature value, specifically:
[0093]
[0094] Among them, f e (c i, a) is vocabulary c i The maximum similarity under category a, c i is the vocabulary in the input text, a is the entity category, D c is the entity dictionary corresponding to category c, d is D c An entity word in ED(c i , d) is c i The edit distance between d and f, i.e. the minimum number of edit operations, e (c i ) is the corresponding entity dictionary matching feature value, and A is the set of all entity categories;
[0095] The probability output of the part-of-speech tagging model is used to obtain the part-of-speech tagging feature value, specifically:
[0096] f p (c i )=K(p(c i )=n)+ω vn K(p(c i )=vn)
[0097] Among them, f p (c i ) is the feature value of part-of-speech tagging, p(c i ) is vocabulary c i The corresponding part-of-speech tags, K(p(c i )=n) is the model prediction c i is the probability of a noun, K(p(c i )=vn) is the model prediction c i is the probability of the gerund, ω vn is the weight coefficient of the gerund.
[0098] Specifically, the pre-trained BERT model is used to obtain the context representation of each word and output the BERT vector of each word. For each word, the feature value related to the entity category is calculated through the feature function, including the entity dictionary matching feature value and the part-of-speech tagging feature value. The entity dictionary matching is defined as checking whether the word is in the predefined entity dictionary. The part-of-speech tagging feature is defined as obtaining the part-of-speech tag through the word segmentation tool. vn is the weight coefficient of the gerund, and the corresponding value is 0.5.
[0099] In one embodiment of the present invention, the calculating entity category probability further includes:
[0100] The entity dictionary matching feature value and the part-of-speech tagging feature value are constructed to obtain the feature matrix F = [f e (c i ), f p (c i),…,f e (c n ), f p (c n )], normalize the feature matrix and map the normalized features to the entity category set A, specifically:
[0101] P2=F norm W norm
[0102] Among them, P2 is the feature after mapping, F norm is the normalized feature, W norm is the mapping weight matrix;
[0103] The BERT vector P1 and the mapped feature P2 are concatenated into a mixed feature: P = [P1, P2]. The entity category probability of each word is calculated through the fully connected layer and Softmax, specifically:
[0104] H=Softmax(P·w h +b h )
[0105] Among them, H is the entity category probability, P is the mixed feature, w h and b h are the weights and biases of the classification layer.
[0106] Specifically, all possible entity categories are defined, the feature values of each word are normalized to ensure that the weights of different features are consistent, and the relevant feature values of each word and all entity categories are output; the hidden state of BERT is spliced with the manually designed features to form a hybrid feature, and the probability of each word belonging to each entity category is calculated through the fully connected layer and Softmax, and the entity category probability of each word is output.
[0107] In one embodiment of the present invention, calculating the relationship and confidence between entity pairs by analyzing contextual information of entity pairs in text includes:
[0108] Extraction is performed based on a method that combines semantic rules and machine learning. For the identified entity pairs (d1, d2), the relationship and confidence between the entity pairs (d1, d2) are calculated by analyzing the contextual information of the entity pairs in the text. Specifically:
[0109] CL(d1, d2, r)=α1×M(d1, d2)+α2×P(r|context)
[0110] Among them, CL(d1, d2, r) is the confidence of the relationship r between the entity pair (d1, d2), r is the relationship type, which indicates the association between the entity pairs, M(d1, d2) is the entity matching degree calculated based on semantic rules, P(r|context) is the probability of the relationship type r appearing under given context conditions, context is the context information, and α1 and α2 are the corresponding weight coefficients.
[0111] Specifically, a BERT model or other NLP tools is used to identify entity pairs in the text. For example, in the sentence "ephedrine is produced by XX Pharmaceutical Factory", the entity pair (ephedrine, XX Pharmaceutical Factory) is identified; contextual information related to the entity pair is extracted from the text, including sentence structure, keywords and semantic roles. For example, in the sentence "ephedrine is produced by XX Pharmaceutical Factory", the context information is "produced by...", semantic rules are defined, and the matching degree between entity pairs is calculated. The semantic rules are: if there is an explicit relation word between the entity pair (such as "produced" and "located in"), then M(d1, d2) = 1; if there is an implicit semantic relationship between the entity pairs, then M(d1, d2) = 1; if there is an implicit semantic relationship between the entity pairs, then M(d1, d2) = 1. If there is no clear association between entity pairs (such as "ephedrine is the core product of XX Pharmaceutical Factory"), then M(d1, d2) = 0.8; if there is no clear association between entity pairs, then M(d1, d2) = 0; use machine learning models (such as CRF, BiLSTM) to train the relationship type classifier, input entity pairs and context information, and output relationship types and their probabilities; calculate the confidence of the entity pair relationship according to the formula, where α1 and α2 are the corresponding weight coefficients, and the corresponding initial values are both 0.5. Use the validation set to adjust the weights. The goal is to maximize the accuracy and recall of relationship extraction. If the semantic rules are more reliable, increase α1; if the machine learning model is more reliable, increase α2.
[0112] In one embodiment of the present invention, the abnormality evaluation and early warning module includes:
[0113] Extract features from the constructed knowledge graph of precursor drugs, enterprises, people, and geographic space, including entity features, relationship features, and attribute features;
[0114] The entity features include entity degree and entity attribute distribution; the relationship features include relationship frequency and relationship path entropy; the attribute features include attribute value deviation and attribute update frequency;
[0115] Construct entity anomaly evaluation value, relationship anomaly evaluation value and attribute anomaly evaluation value, and perform weighted addition to obtain the comprehensive anomaly evaluation value, specifically:
[0116] U=β1*u1+β2*u2+β3*u3
[0117] Among them, U is the comprehensive anomaly evaluation value, u1, u2 and u3 are the entity anomaly evaluation value, relationship anomaly evaluation value and attribute anomaly evaluation value respectively, β1, β2 and β3 are the corresponding weight coefficients respectively;
[0118] If the comprehensive abnormal evaluation value is greater than the preset threshold, an early warning is triggered; otherwise, continuous monitoring is performed.
[0119] Specifically, the following features are extracted from the knowledge graph: entity features include entity degree and entity attribute distribution, relationship features include relationship frequency and relationship path entropy, and attribute features include attribute value deviation and attribute update frequency. The abnormal evaluation values of entities, relationships, and attributes are calculated respectively, and weighted addition is performed to obtain a comprehensive abnormal evaluation value. β1, β2, and β3 are corresponding weight coefficients, satisfying β1+β2+β3=1, where the value of β1 is 0.2, and the values of β2 and β3 are both 0.4; wherein, the preset threshold is set to μ+k·σ by calculating the distribution of the comprehensive abnormal evaluation value of historical data, where μ is the mean, σ is the standard deviation, and k is the adjustment coefficient. For example, using the statistical method, setting k=2 represents a 95% confidence interval; the threshold is recalculated regularly (e.g., monthly) to adapt to changes in data distribution. If a large number of false positives or negatives are monitored, k is adjusted.
[0120] In one embodiment of the present invention, constructing the entity anomaly evaluation value, the relationship anomaly evaluation value, and the attribute anomaly evaluation value includes:
[0121] The entity anomaly evaluation value is obtained by collecting the degrees of all entities in the graph and calculating the mean and standard deviation of all entity degrees, specifically:
[0122]
[0123] Among them, u1 is the entity abnormality evaluation value, r s is the degree of an entity s, μ r is the mean of all entity degrees, σ r is the standard deviation of all entity degrees;
[0124] The relationship anomaly evaluation value is obtained by collecting the current relationship frequency, the current relationship path entropy, the mean and standard deviation of all relationship frequencies, and the mean and standard deviation of all relationship path entropies, specifically:
[0125]
[0126] Among them, u2 is the relationship abnormality evaluation value, γ is the weight coefficient, f now and h now is the current relationship frequency and relationship path entropy, μ f and μ his the mean of all relationship frequencies and the mean of relationship path entropy, σ f and σ h is the standard deviation of all relationship frequencies and the standard deviation of shutdown path entropy;
[0127] The attribute anomaly evaluation value shown is obtained by collecting the extreme values of the deviation of all entity attribute values, the mean and standard deviation of all attribute update frequencies, and is specifically:
[0128]
[0129] Among them, u3 is the attribute abnormality evaluation value, Δoff s is the attribute value deviation of an entity s, max(Δoff) is the maximum value of the attribute value deviation of all entities, is the weight coefficient, is the attribute update frequency of an entity s, μ F and σ F Update the mean and standard deviation of the frequencies for all attributes.
[0130] Specifically, the importance of relationship features is weighted based on relationship frequency and relationship path entropy. If the relationship frequency can more directly reflect anomalies, such as high-frequency trading may indicate violations, a higher weight is given; if the relationship path entropy can better reflect complex associations, such as multi-path dependence may hide risks, a higher weight is given; γ is the weight of relationship frequency, and the corresponding value is 0.6, then the weight of relationship path entropy is 0.4; similarly, the importance of attribute features is weighted based on attribute value deviation and attribute update frequency. is the weight of the attribute value deviation, and the corresponding value is 0.7, so the weight of the attribute update frequency is 0.3.
[0131] The above embodiments are only used to illustrate the technical method of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical method of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical method of the present invention.
Claims
1. The AI-based dynamic monitoring and early warning system for precursor drugs is characterized by: Includes the following modules: Multi-source heterogeneous data collection and transmission module: used by manufacturing enterprise ERP systems to push drug batch information in real time via APIs, logistics enterprise TMS systems to push GPS tracks in real time via Kafka message queues, pharmacy HIS systems to upload sales data daily via WebService interfaces, and drug administration bureaus to push regulatory documents via APIs; Data preprocessing and knowledge graph construction module: This module processes multimodal data separately, defines entity types and relationship types, calculates entity category probabilities, analyzes the contextual information of entity pairs in the text, calculates the relationships and confidence levels between entity pairs, constructs a knowledge graph of precursor drugs, enterprises, personnel, and geographic space, and defines relationship rules and dynamically updates them. The calculation of entity category probability generates a hidden state matrix through the BERT model, calculates the entity dictionary matching feature value in combination with the edit distance, introduces a part-of-speech tagging model, calculates the part-of-speech tagging feature value in combination with the weights of nouns and gerunds, concatenates the BERT vector with the mapped features, and calculates through a fully connected layer and Softmax; Abnormal evaluation and early warning module: extract entity features, relationship features and attribute features respectively, construct entity abnormal evaluation value, relationship abnormal evaluation value and attribute abnormal evaluation value, add them up by weight to obtain comprehensive abnormal evaluation value, and trigger early warning based on dynamic threshold.
2. The artificial intelligence-based dynamic monitoring and early warning system for precursor drugs according to claim 1 is characterized in that: The multi-source heterogeneous data acquisition and transmission module includes: The manufacturer's ERP system pushes drug batch information in real time through the API. The data is encrypted and transmitted to the data platform via HTTPS and stored in a relational database. The logistics company's TMS system pushes GPS tracks in real time through the Kafka message queue. The data is cleaned and stored in a time series database. Transaction records are uploaded daily via SFTP as CSV files, parsed, and stored in the operational data layer. The pharmacy HIS system uploads sales data daily through the WebService interface. After verification, the data is linked to the drug dictionary and manufacturer information and stored in the public dimension layer. The Drug Administration pushes the latest regulatory documents through the API. The data is encrypted and transmitted and stored in the application data layer, triggering dynamic updates of the knowledge graph.
3. The artificial intelligence-based dynamic monitoring and early warning system for precursor drugs according to claim 1 is characterized in that: The multimodal data is processed separately to define entity types and relationship types, including: Preprocess the collected data, specifically by directly mapping fields to structured data and unifying the field format; using the BERT model to extract entities and relationships from text for unstructured data, and identifying sensitive information based on a keyword library; and parsing CSV files uploaded via SFTP for semi-structured data, extracting fields and converting them to a unified format. Furthermore, data cleaning operations are performed, including deduplication, standardization, and anomaly correction. Based on business needs and data characteristics, define the entity types and relationship types of the knowledge graph, calculate the entity recognition probability, and calculate the relationship and confidence between entity pairs by analyzing the contextual information of entity pairs in the text; Define entity types. Core entities include precursor drugs, enterprises, personnel, and geographic space names. Extended entities include laws and policies. Define relationship types. Core relationships include precursor drugs-enterprise, enterprise-personnel, enterprise-geographic space, and precursor drugs-geographic space. Use SQL to associate data from multiple tables. Dynamic relationships include regulatory associations.
4. The artificial intelligence-based dynamic monitoring and early warning system for precursor drugs according to claim 1 is characterized in that: The construction of the knowledge graph of precursor drugs-enterprises-personnel-geographic space, definition of relationship rules and dynamic updates include: Use regular expressions to extract fixed-format fields, use the BERT model to identify entities in text, convert addresses to latitude and longitude through the Baidu API, link them to geospatial nodes, and distinguish entities with the same name based on context. Define relationship rules: Production relationships are associated with the manufacturer ID and drug batch number in ERP data, and distribution path relationships are generated through GPS tracks in TMS to generate path nodes; Use a deep learning model to label relational data. The training set comes from the labeled relations in the historical data. Use the training set to train the model and output labels. Use Neo4j to store nodes and relationships, with core entities stored in the main graph and geospatial nodes and regulatory policies stored in the extended graph.
5. The artificial intelligence-based dynamic monitoring and early warning system for precursor drugs according to claim 1 is characterized in that: The calculating entity category probability includes: The input text A is processed by the BERT word segmenter to obtain the vocabulary sequence C = {c1, c2, ..., c w }, through the pre-trained BERT model, each vocabulary c i Mapped to a high-dimensional vector p i , get the hidden state matrix, specifically: P1={p1,p2,…,p w }; At the same time, the vocabulary c is calculated i With dictionary D C The normalized edit distance of each entity in , takes the maximum value as the corresponding entity dictionary matching feature value, specifically: Among them, f e (c i , a) is vocabulary c i The maximum similarity under category a, c i is the vocabulary in the input text, a is the entity category, D c is the entity dictionary corresponding to category c, d is D c An entity word in ED(c i , d) is c i The edit distance between d and f, i.e. the minimum number of edit operations, e (c i ) is the corresponding entity dictionary matching feature value, and A is the set of all entity categories; The probability output of the part-of-speech tagging model is used to obtain the part-of-speech tagging feature value, specifically: f p (c i )=K(p(c i )=n)+ω vn K(p(c i )=vn) Among them, f p (c i ) is the feature value of part-of-speech tagging, p(c i ) is vocabulary c i The corresponding part-of-speech tags, K(p(c i )=n) is the model prediction c i is the probability of a noun, K(p(c i )=vn) is the model prediction c i is the probability of the gerund, ω vn is the weight coefficient of the gerund.
6. The artificial intelligence-based dynamic monitoring and early warning system for precursor drugs according to claim 1 is characterized in that: The calculating entity category probability further includes: The entity dictionary matching feature value and part-of-speech tagging feature value are constructed to obtain the feature matrix F = [f e (c i ), f p (c i ),…,f e (c n ), f p (c n )], normalize the feature matrix and map the normalized features to the entity category set A, specifically: P2=F norm ·W norm Among them, P2 is the feature after mapping, F norm is the normalized feature, W norm is the mapping weight matrix; The BERT vector P1 and the mapped feature P2 are concatenated into a mixed feature: P = [P1, P2]. The entity category probability of each word is calculated through the fully connected layer and Softmax, specifically: H=Softmax(P·w h +b h ) Among them, H is the entity category probability, P is the mixed feature, w h and b h are the weights and biases of the classification layer.
7. The artificial intelligence-based dynamic monitoring and early warning system for precursor drugs according to claim 1 is characterized in that: The process of analyzing the contextual information of entity pairs in the text and calculating the relationship and confidence between entity pairs includes: Extraction is performed based on a method that combines semantic rules and machine learning. For the identified entity pairs (d1, d2), the relationship and confidence between the entity pairs (d1, d2) are calculated by analyzing the contextual information of the entity pairs in the text. Specifically: CL(d1, d2, r)=α1×M(d1, d2)+α2×P(r|context) Among them, CL(d1, d2, r) is the confidence of the relationship r between the entity pair (d1, d2), r is the relationship type, which indicates the association between the entity pairs, M(d1, d2) is the entity matching degree calculated based on semantic rules, P(r|context) is the probability of the relationship type r appearing under given context conditions, context is the context information, and α1 and α2 are the corresponding weight coefficients.
8. The artificial intelligence-based dynamic monitoring and early warning system for precursor drugs according to claim 1 is characterized in that: The abnormality evaluation and early warning module includes: Extract features from the constructed knowledge graph of precursor drugs, enterprises, people, and geographic space, including entity features, relationship features, and attribute features; The entity features include entity degree and entity attribute distribution; the relationship features include relationship frequency and relationship path entropy; the attribute features include attribute value deviation and attribute update frequency; Construct entity anomaly evaluation value, relationship anomaly evaluation value and attribute anomaly evaluation value, and perform weighted addition to obtain the comprehensive anomaly evaluation value, specifically: U=β1*u1+β2*u2+β3*u3 Among them, U is the comprehensive anomaly evaluation value, u1, u2 and u3 are the entity anomaly evaluation value, relationship anomaly evaluation value and attribute anomaly evaluation value respectively, β1, β2 and β3 are the corresponding weight coefficients respectively; If the comprehensive abnormal evaluation value is greater than the preset threshold, an early warning is triggered; otherwise, continuous monitoring is performed.
9. The artificial intelligence-based dynamic monitoring and early warning system for precursor drugs according to claim 8, characterized in that: The constructing of entity anomaly evaluation values, relationship anomaly evaluation values and attribute anomaly evaluation values includes: The entity anomaly evaluation value is obtained by collecting the degrees of all entities in the graph and calculating the mean and standard deviation of all entity degrees, specifically: Among them, u1 is the entity abnormality evaluation value, r s is the degree of an entity s, μ r is the mean of all entity degrees, σ r is the standard deviation of all entity degrees; The relationship anomaly evaluation value is obtained by collecting the current relationship frequency, the current relationship path entropy, the mean and standard deviation of all relationship frequencies, and the mean and standard deviation of all relationship path entropies, specifically: Among them, u2 is the relationship abnormality evaluation value, γ is the weight coefficient, f now and h now is the current relationship frequency and relationship path entropy, μ f and μ h is the mean of all relationship frequencies and the mean of relationship path entropy, σ f and σ h is the standard deviation of all relationship frequencies and the standard deviation of shutdown path entropy; The attribute anomaly evaluation value shown is obtained by collecting the extreme values of the deviation of all entity attribute values, the mean and standard deviation of all attribute update frequencies, and is specifically: Among them, u3 is the attribute abnormality evaluation value, Δoff s is the attribute value deviation of an entity s, max(Δoff) is the maximum value of the attribute value deviation of all entities, is the weight coefficient, is the attribute update frequency of an entity s, μ F and σ F Update the mean and standard deviation of the frequencies for all attributes.