Intelligent comparison method for drug keywords
By building a comprehensive drug keyword database and combining multiple matching technologies, the problem of incomplete keyword databases and low accuracy of accurate comparison methods for existing drug-related sensitive word detection methods is solved, and higher recall and detection efficiency are achieved.
Patent Information
- Application Number
- CN202510093697.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-06-13
AI Technical Summary
The existing keyword database of drug-related sensitive word detection methods is incomplete, the specificity is insufficient, the accurate comparison method is low, and manual intervention is required, resulting in low recall rate.
Using the intelligent comparison method of drug keywords, a comprehensive drug keyword database is constructed, combining text preprocessing, text binary classification model, fuzzy matching and precise matching, and using Levinstein distance algorithm and dictionary tree technology, fuzzy matching and precise matching results are fused to reduce error matching.
It improves the integrity of the drug keyword database, enhances the recall and accuracy of the intelligent comparison model, reduces manual intervention, and improves detection efficiency and accuracy.
Smart Images

Figure CN120146040A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of natural language processing, and specifically relates to the application of Internet information in the field of detecting drug-related sensitive words. Background Art
[0003] At present, there are few studies and achievements in the field of identifying the risk of drug-related content. The research related to drugs also focuses on the drug situation, management policies, and addiction research. On the other hand, many domestic mature intelligent detection tools have insufficient specificity in identifying drug-related content and cannot upgrade their algorithm models in a timely manner according to the changes in regulatory policies. In other aspects, industry practitioners mainly establish a drug keyword library based on the control list issued by China and use a rule-based matching algorithm to conduct manual or semi-automated compliance reviews on pre-released information.
[0004] The existing methods for detecting drug-related sensitive words have the following deficiencies:
[0005] First, the keyword libraries included in the existing methods and systems are incomplete, resulting in ineffective screening of some controlled drugs and low recall rates.
[0006] Second, the existing methods and systems have insufficient specificity in identifying drug-related content.
[0007] Third, the existing methods and systems often use precise comparison methods, with low accuracy, low comparison efficiency, and requiring manual intervention. Summary of the Invention
[0008] The purpose of the present invention is to provide an intelligent comparison method for drug keywords to help users quickly screen whether the pre-released information contains drug-related sensitive words and solve problems such as insufficient specificity and low accuracy of the existing methods.
[0009] An intelligent comparison method for drug keywords includes the following steps:
[0010] 1. Construct a drug keyword database: Collect controlled drug keywords, including Chinese names, English names, abbreviations, CAS numbers, and aliases, and construct a drug keyword database.
[0011] 2. Verify login: Provide user authorization and verification login functions, set a user login whitelist, and control user access permissions.
[0012] 3. Intelligent comparison model: After preprocessing and classifying the text, perform fuzzy matching and exact matching respectively, and fuse the results of fuzzy matching and exact matching.
[0013] 4. User feedback: Based on the intelligent comparison results, feedback on whether there is a drug-related risk is provided, and one of the following three results is output: no drug-related risk, drug-related risk (controlled drugs), drug-related risk (controlled precursor chemicals).
[0014] The text is respectively subjected to fuzzy matching and exact matching with the drug keyword database, and the results are fused, which can maximize the recall rate of the intelligent comparison model and improve the intelligent comparison performance.
[0015] Preferably, the intelligent comparison model includes the following steps:
[0016] 3.1 Text preprocessing includes text cleaning and text standardization. Text cleaning includes deleting useless tags and rewriting special symbols. Text standardization includes removing stop words, lemmatization, stemming, and text normalization.
[0017] Text preprocessing can standardize the text content and reduce the processing difficulty.
[0018] Preferably, the intelligent comparison model includes the following steps:
[0019] 3.2 The classification method is to use a text binary classification model to identify whether the text involves chemical categories. If so, enter fuzzy matching and exact matching. If not, it is considered not to involve chemical categories, and "no drug-related risk" is output, and the process ends.
[0020] The text binary classification model can filter out most incorrect matches of drug keywords. For example, LSD can represent lysergic acid diethylamide and can also represent limited slip differential.
[0021] Preferably, the method of fuzzy matching is
[0022] 3.3.1 Determine the language used in the text. If it is Chinese, use a Chinese word segmentation model. If it is English, use an English word segmentation model;
[0023] 3.3.2 After segmenting the text to form a word list, calculate the similarity between each word in the word list and the keywords in the drug keyword database;
[0024] 3.3.3 If the similarity is greater than or equal to the set value, it is considered that the fuzzy matching is successful, and the list of matching keywords is returned; if the similarity is less than the set value, it is considered that the fuzzy matching is unsuccessful, and an empty list is returned.
[0025] Preferably, in step 3.3.2, the Levenshtein distance algorithm is used to calculate the minimum number of single-character edits required to convert one word to another word to obtain the similarity between the two words.
[0026] Preferably, the method of exact matching is
[0027] 3.4.1 Construct a keyword character tree for storing and retrieving text content;
[0028] 3.4.2 Conduct precise matching for drug keywords including Chinese names, English names, and CAS numbers respectively. Starting from the root node of the character tree, search and match the drug keywords.
[0029] Match the CAS number separately to reduce incorrect matches of the model. Introduce a trie tree, which can minimize unnecessary string comparisons and improve the intelligent matching performance.
[0030] Preferably, the method of precise matching further includes the following steps:
[0031] 3.4.3 Establish character blacklist libraries for Chinese names, English names, and CAS numbers respectively. During precise matching, if the characters before and after the matching keyword belong to the character blacklist library, the matching fails and an empty list is returned; otherwise, the matching is successful and a list of matching keywords is returned.
[0032] There may be an inclusion relationship among the CAS numbers, Chinese names, and English names of chemicals. To avoid such incorrect matches, a character blacklist library is set up.
[0033] Preferably, the method of result fusion is
[0034] 3.5.1 Fusion the results of Chinese names, English names, and CAS numbers after precise matching, and delete redundant matching results;
[0035] 3.5.2 Take the union of the results of fuzzy matching and precise matching;
[0036] If the result set is empty, output "drug-free risk"; if the result set is not empty, according to the drug keyword database, add a risk classification to the matching results, that is, there is a drug-related risk (controlled drug) or there is a drug-related risk (controlled precursor chemicals), and finally output the matching results.
[0037] Reduce duplicate matching, combine fuzzy matching and precise matching, and maximize the recall rate of the model.
[0038] Due to the adoption of the above technical solutions, the present invention has the following beneficial effects:
[0039] 1. The included keyword library is more comprehensive, collecting the situation of controlled drugs in China and internationally controlled drugs, and expanding their keyword information.
[0040] 2. Introduce a text binary classification model, which can filter out most incorrect matches of drug keywords.
[0041] 3. Use a text preprocessing module, which can standardize the text content and reduce the processing difficulty.
[0042] 4. Introduce fuzzy matching and exact matching to maximize the recall rate of the model; match CAS numbers separately to reduce incorrect matches of the model.
[0043] 5. Introduce a trie tree to minimize unnecessary string comparisons and improve the performance of intelligent comparison. Description of the Drawings
[0044] The present invention will be further described below with reference to the drawings.
[0045] Figure 1 It is a flowchart of an intelligent comparison method for drug keywords.
[0046] Figure 2 It is a flowchart of an intelligent comparison model. Detailed Embodiments
[0047] An intelligent comparison method for drug keywords includes the following steps:
[0048] 1. Construct a drug keyword database: Collect controlled drug keywords, including Chinese names, English names, abbreviations, CAS numbers, and aliases, and store them in a MySQL database to construct a drug keyword database.
[0049] 2. Verify login: Provide user authorization and verification login functions, set a user login whitelist, and control user access rights.
[0050] 3. Intelligent comparison model:
[0051] 3.1 Text preprocessing: Preprocess the text content, including text cleaning, text standardization, etc. Text cleaning includes deleting useless tags, rewriting special symbols, etc. Text standardization includes removing stop words, lemmatization, stemming, text normalization, etc.
[0052] 3.2 For the input text content, use a text binary classification model to identify whether the text involves chemical categories. If so, enter fuzzy matching and exact matching. If not, it is considered not to involve chemical categories, output "no drug risk", and end.
[0053] 3.3 Fuzzy matching:
[0054] 3.3.1 Determine the language used in the text. If it is Chinese, use a Chinese word segmentation model. If it is English, use an English word segmentation model.
[0055] 3.3.2 After segmenting the text into a word list, use the Levenshtein distance algorithm to calculate the minimum number of single-character edits required to convert one word to another to obtain the similarity between two words.
[0056] For example, for the text content "we suppy a chemical named norfentanil, its basic infois...", the drug keyword norfentanyl is matched using the fuzzy matching algorithm, and the similarity between norfentanil and norfentanyl is 91%.
[0057] 3.3.3 If the similarity is greater than or equal to the set value of 90%, the fuzzy matching is considered successful, and the list of matching keywords is returned; if the similarity is less than the set value of 90%, the fuzzy matching is considered unsuccessful, and an empty list is returned.
[0058] 3.4 Exact matching:
[0059] 3.4.1 Construct a keyword character tree for storing and retrieving text content.
[0060] 3.4.2 Conduct precise matching for drug keywords including Chinese names, English names, and CAS numbers respectively. Starting from the root node of the character tree, search and match the drug keywords.
[0061] 3.4.3 Establish character blacklist libraries for Chinese names, English names, and CAS numbers respectively. During exact matching, if the characters before and after the matching keyword belong to the character blacklist library, the matching fails and an empty list is returned; otherwise, the matching is successful and the list of matching keywords is returned.
[0062] It should be noted that there may be an inclusion relationship between the CAS number, Chinese name, and English name of a chemical. When designing the algorithm, this incorrect matching needs to be avoided. The following situations exist:
[0063] 1: 58-08-2 is the CAS number of a certain drug. When the text content contains "181058-08-2", "58-08-2" will be matched, resulting in incorrect matching.
[0064] 2: "Methamphetamine" and "2-Fluoromethamphetamine" are both drug keywords. When the text content contains "2-Fluoromethamphetamine", both "Methamphetamine" and "2-Fluoromethamphetamine" will be matched, resulting in incorrect matching.
[0065] 3: "1-phenyl-2-propanone" is the English name of a precursor chemical controlled in China, and "1-Bromo-1-phenyl-2-propanone" is the English name of a non-controlled chemical in China. When the text content contains "1-Bromo-1-phenyl-2-propanone", "1-phenyl-2-propanone" will be matched, resulting in incorrect matching.
[0066] In view of the above situation, character blacklist libraries for CAS numbers, Chinese names, and English names are set respectively. If the characters before and after the keyword match belong to the character blacklist library, the match fails.
[0067] 3.5 Result Fusion
[0068] 3.5.1 Merge the results of the Chinese name, English name, and CAS number after exact matching, and delete the redundant matching results.
[0069] The result of character tree matching is (start_index, end_index), and text[start_index:end_index + 1] is used to represent the keyword text that is matched. It is necessary to delete the redundant matching results.
[0070] For example, if "methamphetamine" and "2-fluoromethylamphetamine" are matched at the same character position, only the matching result of "2-fluoromethylamphetamine" is retained.
[0071] The matching results before merging are
[0072] Match = [(4, 6), (2, 6), (0, 6), (2, 4), (3, 9), (3, 4), (4, 9), (4, 6), (10, 12), (13, 14)], and the matching results after merging are Result = [(0, 6), (3, 9), (10, 12), (13, 14)].
[0073] 3.5.2 Take the union of the results of fuzzy matching and the results of exact matching.
[0074] If the results of fuzzy matching are ["methamphetamine", "58-08-2"], and the results of exact matching are ["methamphetamine", "1-phenyl-2-propanone"], then the results are ["methamphetamine", "1-phenyl-2-propanone", "58-08-2"].
[0075] 3.5.3 According to the drug keyword database, add risk classifications to the matching results, and finally output the matching results. If the results are empty, return "No drug-related risk"; if the results are not empty, output "Drug-related risk (controlled drug)" or "Drug-related risk (controlled precursor chemicals)" according to the risk type.
[0076] 4. User Feedback: Based on the intelligent comparison results, feedback whether there is a drug-related risk, and output one of the three results: No risk, Drug-related risk (controlled drug), Drug-related risk (controlled precursor chemicals).
[0077] The above are only specific embodiments of the present invention, but the technical features of the present invention are not limited thereto. Any simple changes, equivalent replacements or modifications made based on the present invention to solve substantially the same technical problems and achieve substantially the same technical effects are all covered by the protection scope of the present invention.
Claims
1. A drug keyword intelligent comparison method, characterized in that: The steps include: 1) Construct a drug keyword database: collect controlled drug keywords, including Chinese name, English name, abbreviation, CAS number and alias, and construct a drug keyword database; 2) Verify login: provide user authorization and login verification functions, set user login whitelist, and control user access rights; 3) Intelligent matching model: After preprocessing and classifying the text, fuzzy matching and exact matching are performed respectively, and the results of fuzzy matching and exact matching are merged; 4) User feedback: Based on the intelligent comparison results, feedback is provided on whether there is any drug-related risk.
2. According to claim 1, a drug keyword intelligent comparison method is characterized by: The intelligent comparison model includes the following steps: 3.1 Text preprocessing includes text cleaning and text standardization. Text cleaning includes deleting useless tags and rewriting special symbols. Text standardization includes removing stop words, word form restoration, stemming, and text normalization.
3. According to claim 1, a drug keyword intelligent comparison method is characterized by: The intelligent comparison model includes the following steps: 3.2 The classification method is to use a text binary classification model to identify whether the text involves chemical categories. If so, it will enter fuzzy matching and exact matching; if not, it is considered that it does not involve chemical categories, and the output is "no toxic risk", and the process ends.
4. The intelligent drug keyword comparison method according to claim 1 is characterized by: The fuzzy matching method is 3.3.1 to determine the language used by the text. If it is Chinese, use the Chinese word segmentation model; if it is English, use the English word segmentation model; 3.3.2 Segment the text into words and form a vocabulary, and calculate the similarity between each word in the vocabulary and the keywords in the drug keyword library; 3.3.3 If the similarity is greater than or equal to the set value, the fuzzy match is considered successful and a list of matching keywords is returned; If the similarity is less than the set value, the fuzzy match is considered unsuccessful and an empty list is returned.
5. According to claim 4, a drug keyword intelligent comparison method is characterized by: In step 3.1.2, the Levenshtein distance algorithm is used to calculate the minimum single-character edits required to convert one word into another word, and obtain the similarity between the two words.
6. According to claim 1, a drug keyword intelligent comparison method is characterized by: The exact matching method is 3.4.1 to build a keyword character tree for storing and retrieving text content; 3.4.2 Carry out precise matching for drug keywords including Chinese name, English name and CAS number respectively, starting from the root node of the character tree, and search and match the drug keywords.
7. According to claim 6, a drug keyword intelligent comparison method is characterized by: The exact matching method also includes the following steps: 3.4.3 Establish character blacklist libraries for Chinese names, English names and CAS numbers respectively. When performing an exact match, if the characters before and after the matching keyword belong to the character blacklist library, the match fails and an empty list is returned; Otherwise, the match is successful and a list of matching keywords is returned.
8. According to claim 1, a drug keyword intelligent comparison method is characterized by: The method for result fusion is as follows: 3.5.1: fusion of the results of the Chinese name, English name and CAS number after exact matching, and deletion of redundant matching results; 3.5.2 Take the union of the fuzzy matching result and the exact matching result; 3.5.3 If the result set is empty, then output "no drug-related risk"; if the result set is not empty, then add risk classification to the matching results based on the drug keyword database, that is, there are controlled drugs with drug-related risks or there are controlled precursor chemicals with drug-related risks, and finally output the matching results.