A pharmacological similarity detection method and system combined with neural network

By combining BioBERT and BERT encoders to extract semantic features of drug names, and using SimHash values ​​and dynamic token encryption schemes, the misidentification and security issues in pharmacological similarity detection are resolved, achieving high-precision and high-reliability drug similarity detection.

CN120470338BActive Publication Date: 2025-10-03FUJIAN THINKWIN BIG DATA APPLICATION SERVICE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510962589.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-14
Publication Date
2025-10-03
Estimated Expiration
2045-07-14

AI Technical Summary

Technical Problem

Existing pharmacological similarity detection methods cannot effectively identify the semantic similarity of drug names, and there is a risk of key leakage and replay attacks, resulting in insufficient detection accuracy and reliability.

Method used

A pharmacological similarity detection method combined with neural networks is adopted. The semantic features of drug names are extracted through BioBERT and BERT encoders, SimHash values ​​are used for double detection, and dynamic tokens and multiple encryption schemes are used for authentication and data transmission to improve detection accuracy and reliability.

Benefits of technology

It greatly improves the ability to recognize semantic similarities between drug names, reduces the probability of false detection, effectively prevents session key reuse and replay attacks, ensures the security of data transmission and storage, and improves the accuracy and reliability of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120470338B_ABST
    Figure CN120470338B_ABST
Patent Text Reader

Abstract

The present invention provides a pharmacological similarity detection method and system in the field of biomedical information technology combined with a neural network. The method includes: step S1, obtaining a large number of drug names to construct a data set to train the created pharmacological similarity detection model; step S2, the server creates a drug database, allocates an account number and a key to each detection client and stores it in a password table, pre-sets the key into the detection client, and sets an authentication interface; step S3, obtains a similarity detection request sent by the detection client through the authentication interface and performs authentication; step S4, respectively detects the drug name to be detected carried by the similarity detection request through the pharmacological similarity detection model and the drug database, and obtains the first and second similarity detection sub-results; step S5, generates a similarity detection report based on the first and second similarity detection sub-results. The advantage of the present invention is that it greatly improves the accuracy and reliability of pharmacological similarity detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of biomedical information technology, and in particular to a pharmacological similarity detection method and system combined with a neural network. Background Art

[0002] In the field of clinical application of antimicrobial drugs, accurate identification of pharmacological similarity is of great value, and its application value can be systematically developed as follows: 1. Optimizing individualized treatment plans for patients: By establishing a pharmacological similarity assessment matrix, clinicians can make accurate alternative choices among similar drugs based on the patient's pathogen sensitivity results, organ function status and allergy history; for example, among β-lactam antibiotics, by comparing PK parameters such as drug plasma half-life and tissue penetration, carbapenems that do not require dose adjustment can be optimized for patients with renal insufficiency; it is particularly noteworthy that pharmacological similarity analysis can effectively avoid the risk of cross-resistance. For example, for ESBLs-producing Enterobacteriaceae, by identifying cephalosporins with similar β-lactamase stability, the excessive use of carbapenems can be avoided while ensuring efficacy. 2. Establish a scientific medication management system: At the medical institution level, pharmacological similarity maps provide a molecular-level classification basis for developing a hierarchical management system for antimicrobial drugs. By classifying drugs with similar mechanisms of action and resistance profiles into the same category through cluster analysis, the "same-class substitution priority" management strategy can be effectively implemented. For example, third-generation cephalosporins and monocyclic β-lactams can be assigned different management levels to reduce the irrational use of broad-spectrum drugs at the source. 3. Innovate technical approaches for drug resistance prevention and control: By constructing an antimicrobial drug similarity-resistance evolution prediction model, the impact of the use of specific drugs on the resistance rate of similar drugs can be prospectively assessed. 4. Drive innovation in precision medicine practice: Integrating genomics and pharmacological similarity databases can achieve a transition from "empirical medication" to "molecular targeted therapy."

[0003] Traditionally, pharmacological similarity assessment involves accessing drug management databases and performing string matching or keyword comparison on drug names. This static feature matching method is unable to identify semantic similarities between drug names (e.g., the equivalence between "cefotaxime sodium" and "ceftriaxone"), resulting in an error rate as high as 35%. Traditionally, similarity assessment algorithms have been used, but these methods are unable to handle the nested structure of specialized terms within drug names. For example, the accuracy of identifying the association between "meropenem for injection" and "meropenem" is less than 40%. Furthermore, traditional drug management databases utilize static key management mechanisms, which pose the risk of session key reuse and replay attacks. If a key leak allows a man-in-the-middle attack, resulting in malicious tampering with drug names, the confidence level of similarity detection results will decrease exponentially.

[0004] Therefore, how to provide a pharmacological similarity detection method and system combined with neural networks to improve the accuracy and reliability of pharmacological similarity detection has become a technical problem that needs to be solved urgently. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide a pharmacological similarity detection method and system combined with a neural network to improve the accuracy and reliability of pharmacological similarity detection.

[0006] In a first aspect, the present invention provides a pharmacological similarity detection method combined with a neural network, comprising the following steps:

[0007] Step S1: The server creates a pharmacological similarity detection model based on the input module, the feature extraction module, the feature fusion module, the similarity calculation module, and the output module, and sets a loss function of the pharmacological similarity detection model;

[0008] Step S2: The server obtains a large number of drug names including common names and trade names, pre-processes and labels each of the drug names, and then constructs a data set;

[0009] Step S3: The server trains a pharmacological similarity detection model based on the data set and the loss function, and deploys the trained pharmacological similarity detection model;

[0010] Step S4: The server creates a drug database and a password table for storing drug data, assigns an account and a key to each detection client and stores them in the password table, pre-installs each key into the corresponding detection client, and sets an authentication interface;

[0011] Step S5: The server obtains the similarity detection request sent by the detection client through the authentication interface, which carries the encrypted message, dynamic token and account number;

[0012] Step S6: After authenticating the similarity detection request, the server releases the calling authority of the pharmacological similarity detection model and the drug database;

[0013] Step S7: The server inputs the name of the drug to be tested carried in the similarity detection request into the pharmacological similarity detection model to obtain a first similarity detection sub-result;

[0014] Step S8: The server calculates the SimHash value of the drug name to be tested, and performs a match from the drug database based on the SimHash value to obtain a second similarity detection sub-result;

[0015] Step S9: The server generates a similarity detection report based on the first similarity detection sub-result and the second similarity detection sub-result, encrypts the similarity detection report into an encrypted report, pushes the encrypted report to the detection client, and generates and stores a detection log;

[0016] In step S1, the input module is used to input the drug name including the common name and the trade name, and pre-process the input drug name;

[0017] The feature extraction module is composed of a common name encoder and a product name encoder; the common name encoder is built based on BioBERT and is used to extract pharmacological semantic features from common names; the product name encoder is built based on BERT and is used to extract product semantic features from product names;

[0018] The feature fusion module is used to combine the pharmacological semantic features and the product semantic features, and then perform dimensionality reduction through a fully connected layer to obtain fused features;

[0019] The similarity calculation module is used to calculate the cosine similarity between the fused features, and use the Sigmoid function to map the cosine similarity to the similarity probability;

[0020] The output module is used to output a first similarity detection sub-result according to the similarity probability;

[0021] The loss function adopts the binary cross entropy loss function, and the formula is: ;

[0022] Among them, L represents the loss value of the loss function; N represents the total number of samples of drug names; represents the true similarity label of the i-th drug name; represents the similarity probability of the i-th drug name.

[0023] Furthermore, the step S2 is specifically as follows:

[0024] The server obtains a large number of drug names including generic names and trade names, preprocesses each of the drug names including at least word segmentation, removal of stop words, removal of auxiliary words, and removal of conjunctions, and annotates each of the preprocessed drug names with pharmacological similarities to construct a data set;

[0025] The step S3 is specifically as follows:

[0026] The server groups the data set based on drug type, splits each group into a training subset, a validation subset, and a test subset in a ratio of 7:2:1, and combines each of the training subsets, validation subsets, and test subsets to obtain a training set, a validation set, and a test set;

[0027] The pharmacological similarity detection model is trained using the training set, and during the training process, the loss function is optimized using a stochastic gradient descent method until a loss value of the loss function is less than a preset loss threshold;

[0028] The detection accuracy is calculated using the validation set to determine whether the detection accuracy is greater than a preset accuracy threshold. If not, the verification fails and the training set is expanded to continue training. If so, the verification passes, and:

[0029] Calculate the F1 score using the test set and determine whether the F1 score is greater than a preset F1 threshold. If not, the test fails and the training set is expanded to continue training. If so, the test passes and the pharmacological similarity detection model is subjected to knowledge distillation and deployed via a Docker container.

[0030] In step S4, the drug data includes at least the drug code, generic name, trade name, SimHash value of the generic name, SimHash value of the trade name, drug type, chemical information, pharmacological information and clinical information; the key is generated based on the AES256 algorithm; and the authentication interface is an API interface.

[0031] Furthermore, in step S5, the encrypted message generation process is specifically as follows: obtaining the current timestamp, intercepting the time parameter from the timestamp based on a preset interception rule, concatenating the preset key and the time parameter to obtain a first dynamic key, and calling the first dynamic key through the SM4 algorithm to encrypt the name of the drug to be tested twice to obtain an encrypted message;

[0032] The dynamic token generation process is as follows: the account number and timestamp are spliced ​​to obtain spliced ​​data, the spliced ​​data is encrypted by calling the key through the SM4 algorithm to obtain a second dynamic key, a one-time random number is added to the end of the second dynamic key to obtain a third dynamic key, and the third dynamic key is base64 encoded to obtain a dynamic token;

[0033] The step S6 is specifically as follows:

[0034] The server parses the similarity detection request to obtain an encrypted message, a dynamic token, and an account number, and matches a key from a password table using the account number;

[0035] The server performs base64 decoding on the dynamic token to obtain a third dynamic key, parses the third dynamic key to obtain a second dynamic key and a one-time random number, and determines whether the one-time random number already exists in a preset random number management table. If so, authentication fails and a notification of authentication failure is fed back; if not, the one-time random number is stored in the random number management table and:

[0036] The key is called by the SM4 algorithm to decrypt the second dynamic key to obtain spliced ​​data, the spliced ​​data is parsed to obtain the account number and timestamp, and a time validity check is performed based on the timestamp. If the timeout is exceeded, the authentication fails and an authentication failure notification is fed back; if the timeout is not exceeded, the authentication succeeds, and the access rights to the pharmacological similarity detection model and the drug database are released, and:

[0037] Based on the preset interception rules, the time parameter is intercepted from the timestamp, the key and the time parameter are spliced ​​to obtain the first dynamic key, and the first dynamic key is called through the SM4 algorithm to decrypt the encrypted message twice to obtain the name of the drug to be tested.

[0038] Furthermore, the step S7 is specifically as follows:

[0039] The server creates a detection queue for storing the name of the drug to be tested, the detection status, the process lock status, the number of attempts, and the creation time, stores the name of the drug to be tested carried in the similarity detection request in the detection queue, and updates and initializes the corresponding detection status, process lock status, number of attempts, and creation time; the detection status is not detected, executing the first detection, executing the second detection, execution success, or execution failure; the first detection is detected by a pharmacological similarity detection model; the second detection is detected by calculating a SimHash value;

[0040] The server creates a producer-consumer model including a producer and a consumer. The producer selects a preset number of drug names to be tested from the test queue. The consumer inputs the selected drug names to be tested into the pharmacological similarity detection model to obtain the corresponding first similarity detection sub-result, and the test queue is updated synchronously.

[0041] The step S8 is specifically as follows:

[0042] The server selects a preset number of drug names to be tested from the test queue through the producer, performs word segmentation on each of the drug names to be tested to obtain drug nouns, obtains a proportional weight based on the ratio of the length of each drug noun to the length of the drug name to be tested, counts the number of occurrences of the drug noun in the preset drug list, and calculates the comprehensive weight of each drug noun based on the proportional weight and the number of occurrences: ;

[0043] Among them, C represents the comprehensive weight; M represents the number of occurrences; represents the proportion weight of the i-th drug noun; n represents the total number of drug nouns contained in the names of the drugs to be tested;

[0044] The SimHash value of the drug name to be tested is calculated based on the comprehensive weight and the drug noun, the SimHash value in the matching drug database is traversed based on the SimHash value, and the second similarity detection sub-result is generated based on the code distance calculated based on the SimHash value.

[0045] Furthermore, the step S9 is specifically as follows:

[0046] The server generates a similarity detection report including names of pharmacologically similar drugs based on the first similarity detection sub-result and the duplicate items of the second similarity detection sub-result;

[0047] The server performs HMAC calculation on the similarity detection report to obtain a MAC value, encrypts the similarity detection report and the MAC value using the key to obtain primary encrypted data, cyclically shifts each character of the primary encrypted data to the left by 5 characters to obtain secondary encrypted data, encrypts the secondary encrypted data using the 3DES algorithm to obtain an encrypted report, and pushes the encrypted report to the detection client in real time via the TLS protocol;

[0048] The server generates a detection log based on the detection time, the similarity detection report, and the similarity detection request, performs MD5 calculation on the detection log to obtain an MD5 value, encrypts the detection log and the MD5 value using the key to obtain first encrypted data, circularly shifts each character of the first encrypted data to the right by 7 characters to obtain second encrypted data, encrypts the second encrypted data using the RC6 algorithm to obtain an encrypted log, stores the encrypted log, and performs distributed backup.

[0049] In a second aspect, the present invention provides a pharmacological similarity detection system combined with a neural network, comprising the following modules:

[0050] A pharmacological similarity detection model creation module is used for the server to create a pharmacological similarity detection model based on the input module, feature extraction module, feature fusion module, similarity calculation module and output module, and set the loss function of the pharmacological similarity detection model;

[0051] The data set construction module is used for the server to obtain a large number of drug names including common names and trade names, and to construct a data set after preprocessing and labeling each of the drug names;

[0052] A pharmacological similarity detection model training module is used for the server to train the pharmacological similarity detection model based on the data set and the loss function, and deploy the trained pharmacological similarity detection model;

[0053] A drug database and password table creation module is used for the server to create a drug database and a password table for storing drug data, assign an account and a key to each detection client and store them in the password table, pre-set each key into the corresponding detection client, and set an authentication interface;

[0054] A similarity detection request receiving module is used for the server to obtain the similarity detection request sent by the detection client through the authentication interface, which carries the encrypted message, dynamic token and account number;

[0055] An authentication module, configured to release the calling authority of the pharmacological similarity detection model and the drug database after the server authenticates the similarity detection request;

[0056] A first similarity detection sub-result generating module is configured to input the name of the drug to be tested carried in the similarity detection request into a pharmacological similarity detection model on the server to obtain a first similarity detection sub-result;

[0057] A second similarity detection sub-result generation module is used for the server to calculate the SimHash value of the name of the drug to be tested, and match it from the drug database based on the SimHash value to obtain a second similarity detection sub-result;

[0058] a similarity detection report generation module, configured to generate a similarity detection report based on the first similarity detection sub-result and the second similarity detection sub-result, encrypt the similarity detection report into an encrypted report, push the encrypted report to the detection client, and generate and store a detection log;

[0059] In the pharmacological similarity detection model creation module, the input module is used to input the drug name including the common name and the trade name, and pre-process the input drug name;

[0060] The feature extraction module is composed of a common name encoder and a product name encoder; the common name encoder is built based on BioBERT and is used to extract pharmacological semantic features from common names; the product name encoder is built based on BERT and is used to extract product semantic features from product names;

[0061] The feature fusion module is used to combine the pharmacological semantic features and the product semantic features, and then perform dimensionality reduction through a fully connected layer to obtain fused features;

[0062] The similarity calculation module is used to calculate the cosine similarity between the fused features, and use the Sigmoid function to map the cosine similarity to the similarity probability;

[0063] The output module is used to output a first similarity detection sub-result according to the similarity probability;

[0064] The loss function adopts the binary cross entropy loss function, and the formula is: ;

[0065] Among them, L represents the loss value of the loss function; N represents the total number of samples of drug names; represents the true similarity label of the i-th drug name; represents the similarity probability of the i-th drug name.

[0066] Furthermore, the dataset construction module is specifically used to:

[0067] The server obtains a large number of drug names including generic names and trade names, preprocesses each of the drug names including at least word segmentation, removal of stop words, removal of auxiliary words, and removal of conjunctions, and annotates each of the preprocessed drug names with pharmacological similarities to construct a data set;

[0068] The pharmacological similarity detection model training module is specifically used for:

[0069] The server groups the data set based on drug type, splits each group into a training subset, a validation subset, and a test subset in a ratio of 7:2:1, and combines each of the training subsets, validation subsets, and test subsets to obtain a training set, a validation set, and a test set;

[0070] The pharmacological similarity detection model is trained using the training set, and during the training process, the loss function is optimized using a stochastic gradient descent method until a loss value of the loss function is less than a preset loss threshold;

[0071] The detection accuracy is calculated using the validation set to determine whether the detection accuracy is greater than a preset accuracy threshold. If not, the verification fails and the training set is expanded to continue training. If so, the verification passes, and:

[0072] Calculate the F1 score using the test set and determine whether the F1 score is greater than a preset F1 threshold. If not, the test fails and the training set is expanded to continue training. If so, the test passes and the pharmacological similarity detection model is subjected to knowledge distillation and deployed via a Docker container.

[0073] In the drug database and password table creation module, the drug data includes at least the drug code, generic name, trade name, SimHash value of the generic name, SimHash value of the trade name, drug type, chemical information, pharmacological information and clinical information; the key is generated based on the AES256 algorithm; and the authentication interface is an API interface.

[0074] Furthermore, in the similarity detection request receiving module, the encrypted message generation process is specifically as follows: obtaining a current timestamp, intercepting a time parameter from the timestamp based on a preset interception rule, concatenating the preset key and the time parameter to obtain a first dynamic key, and calling the first dynamic key through the SM4 algorithm to encrypt the name of the drug to be tested twice to obtain an encrypted message;

[0075] The dynamic token generation process is as follows: the account number and timestamp are spliced ​​to obtain spliced ​​data, the spliced ​​data is encrypted by calling the key through the SM4 algorithm to obtain a second dynamic key, a one-time random number is added to the end of the second dynamic key to obtain a third dynamic key, and the third dynamic key is base64 encoded to obtain a dynamic token;

[0076] The authentication module is specifically used for:

[0077] The server parses the similarity detection request to obtain an encrypted message, a dynamic token, and an account number, and matches a key from a password table using the account number;

[0078] The server performs base64 decoding on the dynamic token to obtain a third dynamic key, parses the third dynamic key to obtain a second dynamic key and a one-time random number, and determines whether the one-time random number already exists in a preset random number management table. If so, authentication fails and a notification of authentication failure is fed back; if not, the one-time random number is stored in the random number management table and:

[0079] The key is called by the SM4 algorithm to decrypt the second dynamic key to obtain spliced ​​data, the spliced ​​data is parsed to obtain the account number and timestamp, and a time validity check is performed based on the timestamp. If the timeout is exceeded, the authentication fails and an authentication failure notification is fed back; if the timeout is not exceeded, the authentication succeeds, and the access rights to the pharmacological similarity detection model and the drug database are released, and:

[0080] Based on the preset interception rules, the time parameter is intercepted from the timestamp, the key and the time parameter are spliced ​​to obtain the first dynamic key, and the first dynamic key is called through the SM4 algorithm to decrypt the encrypted message twice to obtain the name of the drug to be tested.

[0081] Furthermore, the first similarity detection sub-result generating module is specifically configured to:

[0082] The server creates a detection queue for storing the name of the drug to be tested, the detection status, the process lock status, the number of attempts, and the creation time, stores the name of the drug to be tested carried in the similarity detection request in the detection queue, and updates and initializes the corresponding detection status, process lock status, number of attempts, and creation time; the detection status is not detected, executing the first detection, executing the second detection, execution success, or execution failure; the first detection is detected by a pharmacological similarity detection model; the second detection is detected by calculating a SimHash value;

[0083] The server creates a producer-consumer model including a producer and a consumer. The producer selects a preset number of drug names to be tested from the test queue. The consumer inputs the selected drug names to be tested into the pharmacological similarity detection model to obtain the corresponding first similarity detection sub-result, and the test queue is updated synchronously.

[0084] The second similarity detection sub-result generation module is specifically used to:

[0085] The server selects a preset number of drug names to be tested from the test queue through the producer, performs word segmentation on each of the drug names to be tested to obtain drug nouns, obtains a proportional weight based on the ratio of the length of each drug noun to the length of the drug name to be tested, counts the number of occurrences of the drug noun in the preset drug list, and calculates the comprehensive weight of each drug noun based on the proportional weight and the number of occurrences: ;

[0086] Among them, C represents the comprehensive weight; M represents the number of occurrences; represents the proportion weight of the i-th drug noun; n represents the total number of drug nouns contained in the names of the drugs to be tested;

[0087] The SimHash value of the drug name to be tested is calculated based on the comprehensive weight and the drug noun, the SimHash value in the matching drug database is traversed based on the SimHash value, and the second similarity detection sub-result is generated based on the code distance calculated based on the SimHash value.

[0088] Furthermore, the similarity detection report generation module is specifically used to:

[0089] The server generates a similarity detection report including names of pharmacologically similar drugs based on the first similarity detection sub-result and the duplicate items of the second similarity detection sub-result;

[0090] The server performs HMAC calculation on the similarity detection report to obtain a MAC value, encrypts the similarity detection report and the MAC value using the key to obtain primary encrypted data, cyclically shifts each character of the primary encrypted data to the left by 5 characters to obtain secondary encrypted data, encrypts the secondary encrypted data using the 3DES algorithm to obtain an encrypted report, and pushes the encrypted report to the detection client in real time via the TLS protocol;

[0091] The server generates a detection log based on the detection time, the similarity detection report, and the similarity detection request, performs MD5 calculation on the detection log to obtain an MD5 value, encrypts the detection log and the MD5 value using the key to obtain first encrypted data, circularly shifts each character of the first encrypted data to the right by 7 characters to obtain second encrypted data, encrypts the second encrypted data using the RC6 algorithm to obtain an encrypted log, stores the encrypted log, and performs distributed backup.

[0092] The advantages of the present invention are:

[0093] 1. Create a pharmacological similarity detection model based on the input module, feature extraction module, feature fusion module, similarity calculation module and output module through the server, and set the loss function of the pharmacological similarity detection model; then obtain a large number of drug names including generic names and trade names, preprocess and annotate each drug name to build a data set, train the pharmacological similarity detection model based on the data set and loss function, and deploy the trained pharmacological similarity detection model; then create a drug database and password table, assign an account and key to each detection client and store them in the password table, pre-set each key into the corresponding detection client, and set an authentication interface; then the server obtains the similarity detection request sent by the detection client through the authentication interface, which carries an encrypted message, a dynamic token and an account, authenticates the similarity detection request, releases the calling authority of the pharmacological similarity detection model and the drug database, inputs the name of the drug to be detected carried by the similarity detection request into the pharmacological similarity detection model to obtain the first similarity detection sub-result, calculates the SimHash value of the drug name to be detected, and calculates the SimHash value based on Si The mHash value is matched against the drug database to obtain a second similarity detection sub-result. A similarity detection report is generated based on the first and second similarity detection sub-results. The similarity detection report is encrypted and pushed to the detection client, and a detection log is generated and stored. That is, the pharmacological similarity detection process combines the pharmacological similarity detection model and SimHash value. The pharmacological similarity detection model greatly improves the feature extraction capability by extracting and fusing pharmacological semantic features and product semantic features. It can not only effectively identify the semantic similarity between drug names, but also effectively handle the nested structure of professional terms in drug names. Combined with SimHash value for dual detection, the probability of false detection is greatly reduced. In addition, the similarity detection request carries an encrypted message, a dynamic token, and an account number. Authentication through the dynamic token can effectively prevent the risk of session key reuse and replay attacks. Transmitting data through encrypted messages also prevents data from being stolen in plain text. Combined with the encrypted push of the similarity detection report and the encrypted storage of the detection log, the accuracy and reliability of pharmacological similarity detection are ultimately greatly improved.

[0094] 2. By setting the loss function of the pharmacological similarity detection model to a binary cross-entropy loss function, the difference between the first similarity detection sub-result output by the pharmacological similarity detection model and the true label can be effectively measured, thereby optimizing the model performance and improving the training efficiency and stability of the model.

[0095] 3. By preprocessing the names of each drug, including at least word segmentation, removal of stop words, removal of auxiliary words, and removal of conjunctions, the quality of the dataset is effectively improved, thereby effectively improving the training effect of the pharmacological similarity detection model.

[0096] 4. Group the data set by drug type, and split each group into training subset, validation subset, and test subset in a ratio of 7:2:1. Combine each training subset, validation subset, and test subset to obtain the training set, validation set, and test set, respectively. Make sure that the training set, validation set, and test set all contain the drug names of each drug type, so as to avoid the influence of data set division bias on subsequent training results.

[0097] 5. By using the stochastic gradient descent method to optimize the loss function during the training process of the pharmacological similarity detection model, the training efficiency is effectively improved, the convergence speed is accelerated, the memory usage is reduced, and it has good adaptability and scalability.

[0098] 6. By setting the key to be generated based on the AES256 algorithm with a 256-bit key length, the amount of calculation required for cracking increases exponentially, which can provide extremely high security and effectively resist brute force attacks.

[0099] 7. The first dynamic key is obtained by concatenating the key and the time parameter, and the first dynamic key is called through the SM4 algorithm to encrypt the name of the drug to be tested twice to obtain an encrypted message; that is, the encryption process of the name of the drug to be tested combines the key, time parameter, SM4 algorithm, and two encryptions, which greatly increases the difficulty of cracking, and thus greatly improves the security of the transmission of the name of the drug to be tested.

[0100] 8. The account number and timestamp are spliced ​​together to obtain spliced ​​data, and the spliced ​​data is encrypted by calling the key through the SM4 algorithm to obtain the second dynamic key. A one-time random number is added to the end of the second dynamic key to obtain the third dynamic key, and the third dynamic key is base64 encoded to obtain the dynamic token; that is, the generation of the dynamic token combines the timestamp, SM4 algorithm, one-time random number, and base64 encoding. The dynamic token can be subsequently subjected to multi-dimensional verification, thereby greatly improving the security of authentication.

[0101] 9. By creating a detection queue for storing the names of drugs to be tested, detection status, process lock status, number of attempts, and creation time, a producer-consumer model is created. The producer selects a preset number of drug names to be tested from the detection queue, and the consumer performs pharmacological similarity testing on the selected drug names to be tested. This can effectively solve the data sharing and synchronization problems in a multi-threaded environment, thereby effectively improving concurrency performance and resource utilization.

[0102] 10. By selecting the duplicates of the first similarity detection sub-result and the second similarity detection sub-result, a similarity detection report containing the names of pharmacologically similar drugs is generated, which greatly improves the confidence of the similarity detection report.

[0103] 11. Perform HMAC calculation on the similarity detection report to obtain a MAC value, encrypt the similarity detection report and the MAC value with a key to obtain first-level encrypted data, shift each character of the first-level encrypted data 5 characters to the left in a circular manner to obtain second-level encrypted data, encrypt the second-level encrypted data with the 3DES algorithm to obtain an encrypted report, and push the encrypted report to the detection client in real time via the TLS protocol. That is, at least five security measures (HMAC calculation, key, character shift, 3DES algorithm, TLS protocol) are combined in the push process of the similarity detection report to prevent the plain text from being stolen and tampered with during the push process of the similarity detection report, thereby greatly improving the security of the push of the similarity detection report.

[0104] 12. Perform MD5 calculation on the detection log to obtain an MD5 value, encrypt the detection log and the MD5 value using a key to obtain first encrypted data, shift each character of the first encrypted data rightward by 7 characters to obtain second encrypted data, encrypt the second encrypted data using the RC6 algorithm to obtain an encrypted log, store the encrypted log, and perform distributed backup. This means that at least five security measures (MD5 calculation, key, character shift, RC6 algorithm, and distributed backup) are combined in the detection log storage process, greatly improving the security of detection log storage and facilitating later traceability.

[0105] 13. Security is further improved by adopting different encryption schemes in each link of similarity detection request authentication, similarity detection report push and detection log storage.

[0106] 14. By using BioBERT (a pre-trained model in the biomedical field) to process common names and combining it with standard BERT to process product names, efficient extraction of pharmacological semantic features and commercial features is achieved; and the dual encoder structure solves the problem of incompatibility of cross-domain text features, effectively improving feature relevance compared to a single model, thereby greatly improving the accuracy of pharmacological similarity detection.

[0107] 15. The concatenated features are nonlinearly reduced in dimension through the fully connected layer, effectively retaining key discriminant information and effectively reducing computational complexity compared to the traditional attention mechanism.

[0108] 16. The comprehensive weight is calculated by proportional weight and number of occurrences, and then the SimHash value is calculated based on the comprehensive weight, which greatly reduces the misjudgment rate of pharmacological similarity detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0109] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0110] Figure 1 The present invention is a flowchart of a pharmacological similarity detection method combined with a neural network.

[0111] Figure 2 It is a structural schematic diagram of a pharmacological similarity detection system combined with a neural network of the present invention. DETAILED DESCRIPTION

[0112] The technical solution in the embodiments of the present application has the following overall idea: pharmacological similarity detection is performed in combination with a pharmacological similarity detection model and a SimHash value. The pharmacological similarity detection model greatly improves the feature extraction capability by extracting and fusing pharmacological semantic features and commodity semantic features. It can not only effectively identify the semantic similarity between drug names, but also effectively process the nested structure of professional terms in drug names. Double detection is performed in combination with the SimHash value, which greatly reduces the probability of false detection. The similarity detection request carries an encrypted message, a dynamic token, and an account number. Authentication through a dynamic token can effectively prevent the risk of session key reuse and replay attack. Transmitting data through encrypted messages also prevents data from being stolen in plain text. Combined with the encrypted push of the similarity detection report and the encrypted storage of the detection log, the accuracy and reliability of pharmacological similarity detection can be improved.

[0113] Please refer to Figures 1 to 2 As shown, a preferred embodiment of the present invention is a pharmacological similarity detection method combined with a neural network, comprising the following steps:

[0114] Step S1: The server creates a pharmacological similarity detection model based on the input module, the feature extraction module, the feature fusion module, the similarity calculation module, and the output module, and sets a loss function of the pharmacological similarity detection model;

[0115] Step S2: The server obtains a large number of drug names including common names and trade names, pre-processes and labels each of the drug names, and then constructs a data set;

[0116] Step S3: The server trains a pharmacological similarity detection model based on the data set and the loss function, and deploys the trained pharmacological similarity detection model;

[0117] Step S4: The server creates a drug database and a password table for storing drug data, assigns an account and a key to each detection client and stores them in the password table, pre-installs each key into the corresponding detection client, and sets an authentication interface;

[0118] Step S5: The server obtains the similarity detection request sent by the detection client through the authentication interface, which carries the encrypted message, dynamic token and account number;

[0119] Step S6: After authenticating the similarity detection request, the server releases the calling authority of the pharmacological similarity detection model and the drug database;

[0120] Step S7: The server inputs the name of the drug to be tested carried in the similarity detection request into the pharmacological similarity detection model to obtain a first similarity detection sub-result;

[0121] Step S8: The server calculates the SimHash value of the drug name to be tested, and performs a match from the drug database based on the SimHash value to obtain a second similarity detection sub-result;

[0122] Step S9: The server generates a similarity detection report based on the first similarity detection sub-result and the second similarity detection sub-result, encrypts the similarity detection report into an encrypted report, pushes the encrypted report to the detection client, and generates and stores a detection log;

[0123] In step S1, the input module is used to input the drug name including the common name and the trade name, and pre-process the input drug name;

[0124] The feature extraction module is composed of a common name encoder and a product name encoder; the common name encoder is built based on BioBERT and is used to extract pharmacological semantic features from common names; the product name encoder is built based on BERT and is used to extract product semantic features from product names;

[0125] By using BioBERT (a pre-trained model in the biomedical field) to process common names and combining it with standard BERT to process product names, efficient extraction of pharmacological semantic features and commercial features is achieved; and the dual encoder structure solves the problem of incompatibility of cross-domain text features, effectively improving feature correlation compared to a single model, thereby greatly improving the accuracy of pharmacological similarity detection.

[0126] The feature fusion module is used to combine the pharmacological semantic features and the product semantic features, and then perform dimensionality reduction through a fully connected layer to obtain fused features;

[0127] The concatenated features are nonlinearly reduced in dimension through the fully connected layer, effectively retaining key discriminant information and significantly reducing computational complexity compared to the traditional attention mechanism.

[0128] The similarity calculation module is used to calculate the cosine similarity between the fused features, and use the Sigmoid function to map the cosine similarity to the similarity probability;

[0129] The output module is used to output a first similarity detection sub-result according to the similarity probability;

[0130] The loss function adopts the binary cross entropy loss function, and the formula is: ;

[0131] Among them, L represents the loss value of the loss function; N represents the total number of samples of drug names; represents the true similarity label of the i-th drug name; represents the similarity probability of the i-th drug name.

[0132] By setting the loss function of the pharmacological similarity detection model to a binary cross-entropy loss function, the difference between the first similarity detection sub-result output by the pharmacological similarity detection model and the true label can be effectively measured, the model performance can be optimized, and the training efficiency and stability of the model can be improved.

[0133] The step S2 is specifically as follows:

[0134] The server obtains a large number of drug names including generic names and trade names, preprocesses each of the drug names including at least word segmentation, removal of stop words, removal of auxiliary words, and removal of conjunctions, and annotates each of the preprocessed drug names with pharmacological similarities to construct a data set;

[0135] By preprocessing the names of each drug, including at least word segmentation, removal of stop words, removal of auxiliary words, and removal of conjunctions, the quality of the dataset is effectively improved, thereby effectively improving the training effect of the pharmacological similarity detection model.

[0136] The step S3 is specifically as follows:

[0137] The server groups the data set based on drug type, splits each group into a training subset, a validation subset, and a test subset in a ratio of 7:2:1, and combines each of the training subsets, validation subsets, and test subsets to obtain a training set, a validation set, and a test set;

[0138] The data set was grouped by drug type, and each group was split into training subsets, validation subsets, and test subsets in a ratio of 7:2:1. The training subsets, validation subsets, and test subsets were combined to obtain the training set, validation set, and test set, respectively. The training set, validation set, and test set all contained the names of drugs of each drug type, avoiding the influence of data set division bias on subsequent training effects.

[0139] The pharmacological similarity detection model is trained using the training set, and during the training process, the loss function is optimized using a stochastic gradient descent method until a loss value of the loss function is less than a preset loss threshold;

[0140] The detection accuracy is calculated using the validation set to determine whether the detection accuracy is greater than a preset accuracy threshold. If not, the verification fails and the training set is expanded to continue training. If so, the verification passes, and:

[0141] Calculate the F1 score using the test set and determine whether the F1 score is greater than a preset F1 threshold. If not, the test fails and the training set is expanded to continue training. If so, the test passes and the pharmacological similarity detection model is subjected to knowledge distillation and deployed via a Docker container.

[0142] By using the stochastic gradient descent method to optimize the loss function during the training process of the pharmacological similarity detection model, the training efficiency is effectively improved, the convergence speed is accelerated, the memory usage is reduced, and it has good adaptability and scalability.

[0143] In step S4, the drug data includes at least the drug code, generic name, trade name, SimHash value of the generic name, SimHash value of the trade name, drug type, chemical information, pharmacological information and clinical information; the key is generated based on the AES256 algorithm; and the authentication interface is an API interface.

[0144] By setting the key to be generated based on the AES256 algorithm with a 256-bit key length, the amount of computation required for cracking increases exponentially, providing extremely high security and effectively resisting brute force attacks.

[0145] In step S5, the encrypted message generation process is specifically as follows: obtaining the current timestamp, intercepting the time parameter from the timestamp based on a preset interception rule, concatenating the preset key and the time parameter to obtain a first dynamic key, and calling the first dynamic key through the SM4 algorithm to encrypt the name of the drug to be tested twice to obtain an encrypted message;

[0146] The first dynamic key is obtained by concatenating the key and the time parameter, and the first dynamic key is called through the SM4 algorithm to encrypt the name of the drug to be tested twice to obtain an encrypted message; that is, the encryption process of the name of the drug to be tested combines the key, time parameter, SM4 algorithm, and two encryptions, which greatly increases the difficulty of cracking, and thus greatly improves the security of the transmission of the name of the drug to be tested.

[0147] The dynamic token generation process is as follows: the account number and timestamp are spliced ​​to obtain spliced ​​data, the spliced ​​data is encrypted by calling the key through the SM4 algorithm to obtain a second dynamic key, a one-time random number is added to the end of the second dynamic key to obtain a third dynamic key, and the third dynamic key is base64 encoded to obtain a dynamic token;

[0148] The account number and timestamp are spliced ​​to obtain spliced ​​data, the spliced ​​data is encrypted by calling the key through the SM4 algorithm to obtain the second dynamic key, a one-time random number is added to the end of the second dynamic key to obtain the third dynamic key, and the third dynamic key is base64 encoded to obtain the dynamic token; that is, the generation of the dynamic token combines the timestamp, SM4 algorithm, one-time random number, and base64 encoding, and the dynamic token can be subsequently subjected to multi-dimensional verification, thereby greatly improving the security of authentication.

[0149] The step S6 is specifically as follows:

[0150] The server parses the similarity detection request to obtain an encrypted message, a dynamic token, and an account number, and matches a key from a password table using the account number;

[0151] The server performs base64 decoding on the dynamic token to obtain a third dynamic key, parses the third dynamic key to obtain a second dynamic key and a one-time random number, and determines whether the one-time random number already exists in a preset random number management table. If so, authentication fails and a notification of authentication failure is fed back; if not, the one-time random number is stored in the random number management table and:

[0152] The key is called by the SM4 algorithm to decrypt the second dynamic key to obtain spliced ​​data, the spliced ​​data is parsed to obtain the account number and timestamp, and a time validity check is performed based on the timestamp. If the timeout is exceeded, the authentication fails and an authentication failure notification is fed back; if the timeout is not exceeded, the authentication succeeds, and the access rights to the pharmacological similarity detection model and the drug database are released, and:

[0153] Based on the preset interception rules, the time parameter is intercepted from the timestamp, the key and the time parameter are spliced ​​to obtain the first dynamic key, and the first dynamic key is called through the SM4 algorithm to decrypt the encrypted message twice to obtain the name of the drug to be tested.

[0154] The step S7 is specifically as follows:

[0155] The server creates a detection queue for storing the name of the drug to be tested, the detection status, the process lock status, the number of attempts, and the creation time, stores the name of the drug to be tested carried in the similarity detection request in the detection queue, and updates and initializes the corresponding detection status, process lock status, number of attempts, and creation time; the detection status is not detected, executing the first detection, executing the second detection, execution success, or execution failure; the first detection is performed by a pharmacological similarity detection model; the second detection is performed by calculating the SimHash value; when the number of attempts reaches a preset threshold, the abnormal process automatically rolls back and switches to a backup node;

[0156] The server creates a producer-consumer model including a producer and a consumer. The producer selects a preset number of names of drugs to be tested from the detection queue, and the consumer inputs the selected names of drugs to be tested into the pharmacological similarity detection model to obtain the corresponding first similarity detection sub-result, and synchronously updates the detection queue. When the consumer's process is executing, a process lock is formed based on the local IP+uuid (i.e., a process lock is formed by combining the IP address and the universally unique identifier). In combination with the select * from table1 for update skip locked (this SQL statement is used to implement an efficient row-level locking mechanism in database concurrent access scenarios, and is particularly suitable for high-concurrency queue processing systems, SELECT * FROM table1 is the basic query part, FOR UPDATE is the lock modifier, and SKIP LOCKED is skip lock) row lock method to prevent multiple threaded tasks from operating on the same data at the same time.

[0157] By creating a detection queue for storing the names of drugs to be tested, detection status, process lock status, number of attempts and creation time, a producer-consumer model is created. The producer selects a preset number of drug names to be tested from the detection queue, and the consumer performs pharmacological similarity testing on the selected drug names to be tested. This can effectively solve the data sharing and synchronization problems in a multi-threaded environment, thereby effectively improving concurrency performance and resource utilization.

[0158] The step S8 is specifically as follows:

[0159] The server selects a preset number of drug names to be tested from the test queue through the producer, performs word segmentation on each of the drug names to be tested to obtain drug nouns, obtains a proportional weight based on the ratio of the length of each drug noun to the length of the drug name to be tested, counts the number of occurrences of the drug noun in the preset drug list, and calculates the comprehensive weight of each drug noun based on the proportional weight and the number of occurrences: ;

[0160] Among them, C represents the comprehensive weight; M represents the number of occurrences; represents the proportional weight of the i-th drug noun; n represents the total number of drug nouns contained in the names of the drugs to be tested; when the corresponding drug noun does not exist in the drug list, the comprehensive weight of the corresponding drug noun is set to the minimum value, for example, to 0.01;

[0161] The SimHash value of the drug name to be tested is calculated based on the comprehensive weight and the drug noun, the SimHash value in the matching drug database is traversed based on the SimHash value, and the second similarity detection sub-result is generated based on the code distance calculated based on the SimHash value.

[0162] The comprehensive weight is calculated by proportional weight and number of occurrences, and the SimHash value is calculated based on the comprehensive weight, which greatly reduces the misjudgment rate of pharmacological similarity detection.

[0163] The step S9 is specifically as follows:

[0164] The server generates a similarity detection report including names of pharmacologically similar drugs based on the first similarity detection sub-result and the duplicate items of the second similarity detection sub-result;

[0165] By selecting duplicates of the first similarity detection sub-result and the second similarity detection sub-result, a similarity detection report containing the names of pharmacologically similar drugs is generated, which greatly improves the confidence of the similarity detection report.

[0166] The server performs HMAC calculation on the similarity detection report to obtain a MAC value, encrypts the similarity detection report and the MAC value using the key to obtain primary encrypted data, cyclically shifts each character of the primary encrypted data to the left by 5 characters to obtain secondary encrypted data, encrypts the secondary encrypted data using the 3DES algorithm to obtain an encrypted report, and pushes the encrypted report to the detection client in real time via the TLS protocol;

[0167] The MAC value is obtained by performing HMAC calculation on the similarity detection report, and the similarity detection report and MAC value are encrypted with a key to obtain the first-level encrypted data. Each character of the first-level encrypted data is circularly shifted to the left by 5 characters to obtain the second-level encrypted data. The second-level encrypted data is encrypted using the 3DES algorithm to obtain an encrypted report, and the encrypted report is pushed to the detection client in real time via the TLS protocol. That is, at least five security measures (HMAC calculation, key, character shift, 3DES algorithm, TLS protocol) are combined in the process of pushing the similarity detection report to prevent the plaintext from being stolen and tampered with during the push process of the similarity detection report, thereby greatly improving the security of the push of the similarity detection report.

[0168] The server generates a detection log based on the detection time, the similarity detection report, and the similarity detection request, performs MD5 calculation on the detection log to obtain an MD5 value, encrypts the detection log and the MD5 value using the key to obtain first encrypted data, circularly shifts each character of the first encrypted data to the right by 7 characters to obtain second encrypted data, encrypts the second encrypted data using the RC6 algorithm to obtain an encrypted log, stores the encrypted log, and performs distributed backup.

[0169] The MD5 value is obtained by performing MD5 calculation on the detection log. The detection log and the MD5 value are encrypted using a key to obtain the first encrypted data. Each character of the first encrypted data is circularly shifted right by 7 characters to obtain the second encrypted data. The second encrypted data is encrypted using the RC6 algorithm to obtain an encrypted log. The encrypted log is stored and distributedly backed up. In other words, the detection log storage process combines at least five security measures (MD5 calculation, key, character shift, RC6 algorithm, and distributed backup), which greatly improves the security of detection log storage and facilitates later traceability.

[0170] Security is further improved by adopting different encryption schemes in each link of similarity detection request authentication, similarity detection report push and detection log storage.

[0171] A preferred embodiment of the pharmacological similarity detection system combined with a neural network of the present invention includes the following modules:

[0172] A pharmacological similarity detection model creation module is used for the server to create a pharmacological similarity detection model based on the input module, feature extraction module, feature fusion module, similarity calculation module and output module, and set the loss function of the pharmacological similarity detection model;

[0173] The data set construction module is used for the server to obtain a large number of drug names including common names and trade names, and to construct a data set after preprocessing and labeling each of the drug names;

[0174] A pharmacological similarity detection model training module is used for the server to train the pharmacological similarity detection model based on the data set and the loss function, and deploy the trained pharmacological similarity detection model;

[0175] A drug database and password table creation module is used for the server to create a drug database and a password table for storing drug data, assign an account and a key to each detection client and store them in the password table, pre-set each key into the corresponding detection client, and set an authentication interface;

[0176] A similarity detection request receiving module is used for the server to obtain the similarity detection request sent by the detection client through the authentication interface, which carries the encrypted message, dynamic token and account number;

[0177] An authentication module, configured to release the calling authority of the pharmacological similarity detection model and the drug database after the server authenticates the similarity detection request;

[0178] A first similarity detection sub-result generating module is configured to input the name of the drug to be tested carried in the similarity detection request into a pharmacological similarity detection model on the server to obtain a first similarity detection sub-result;

[0179] A second similarity detection sub-result generation module is used for the server to calculate the SimHash value of the name of the drug to be tested, and match it from the drug database based on the SimHash value to obtain a second similarity detection sub-result;

[0180] a similarity detection report generation module, configured to generate a similarity detection report based on the first similarity detection sub-result and the second similarity detection sub-result, encrypt the similarity detection report into an encrypted report, push the encrypted report to the detection client, and generate and store a detection log;

[0181] In the pharmacological similarity detection model creation module, the input module is used to input the drug name including the common name and the trade name, and pre-process the input drug name;

[0182] The feature extraction module is composed of a common name encoder and a product name encoder; the common name encoder is built based on BioBERT and is used to extract pharmacological semantic features from common names; the product name encoder is built based on BERT and is used to extract product semantic features from product names;

[0183] By using BioBERT (a pre-trained model in the biomedical field) to process common names and combining it with standard BERT to process product names, efficient extraction of pharmacological semantic features and commercial features is achieved; and the dual encoder structure solves the problem of incompatibility of cross-domain text features, effectively improving feature correlation compared to a single model, thereby greatly improving the accuracy of pharmacological similarity detection.

[0184] The feature fusion module is used to combine the pharmacological semantic features and the product semantic features, and then perform dimensionality reduction through a fully connected layer to obtain fused features;

[0185] The concatenated features are nonlinearly reduced in dimension through the fully connected layer, effectively retaining key discriminant information and significantly reducing computational complexity compared to the traditional attention mechanism.

[0186] The similarity calculation module is used to calculate the cosine similarity between the fused features, and use the Sigmoid function to map the cosine similarity to the similarity probability;

[0187] By setting the loss function of the pharmacological similarity detection model to a binary cross-entropy loss function, the difference between the first similarity detection sub-result output by the pharmacological similarity detection model and the true label can be effectively measured, the model performance can be optimized, and the training efficiency and stability of the model can be improved.

[0188] The output module is used to output a first similarity detection sub-result according to the similarity probability;

[0189] The loss function adopts the binary cross entropy loss function, and the formula is: ;

[0190] Among them, L represents the loss value of the loss function; N represents the total number of samples of drug names; represents the true similarity label of the i-th drug name; represents the similarity probability of the i-th drug name.

[0191] The dataset construction module is specifically used for:

[0192] The server obtains a large number of drug names including generic names and trade names, preprocesses each of the drug names including at least word segmentation, removal of stop words, removal of auxiliary words, and removal of conjunctions, and annotates each of the preprocessed drug names with pharmacological similarities to construct a data set;

[0193] By preprocessing the names of each drug, including at least word segmentation, removal of stop words, removal of auxiliary words, and removal of conjunctions, the quality of the dataset is effectively improved, thereby effectively improving the training effect of the pharmacological similarity detection model.

[0194] The pharmacological similarity detection model training module is specifically used for:

[0195] The server groups the data set based on drug type, splits each group into a training subset, a validation subset, and a test subset in a ratio of 7:2:1, and combines each of the training subsets, validation subsets, and test subsets to obtain a training set, a validation set, and a test set;

[0196] The data set was grouped by drug type, and each group was split into training subsets, validation subsets, and test subsets in a ratio of 7:2:1. The training subsets, validation subsets, and test subsets were combined to obtain the training set, validation set, and test set, respectively. The training set, validation set, and test set all contained the names of drugs of each drug type, avoiding the influence of data set division bias on subsequent training effects.

[0197] The pharmacological similarity detection model is trained using the training set, and during the training process, the loss function is optimized using a stochastic gradient descent method until a loss value of the loss function is less than a preset loss threshold;

[0198] The detection accuracy is calculated using the validation set to determine whether the detection accuracy is greater than a preset accuracy threshold. If not, the verification fails and the training set is expanded to continue training. If so, the verification passes, and:

[0199] Calculate the F1 score using the test set and determine whether the F1 score is greater than a preset F1 threshold. If not, the test fails and the training set is expanded to continue training. If so, the test passes and the pharmacological similarity detection model is subjected to knowledge distillation and deployed via a Docker container.

[0200] By using the stochastic gradient descent method to optimize the loss function during the training process of the pharmacological similarity detection model, the training efficiency is effectively improved, the convergence speed is accelerated, the memory usage is reduced, and it has good adaptability and scalability.

[0201] In the drug database and password table creation module, the drug data includes at least the drug code, generic name, trade name, SimHash value of the generic name, SimHash value of the trade name, drug type, chemical information, pharmacological information and clinical information; the key is generated based on the AES256 algorithm; and the authentication interface is an API interface.

[0202] By setting the key to be generated based on the AES256 algorithm with a 256-bit key length, the amount of computation required for cracking increases exponentially, providing extremely high security and effectively resisting brute force attacks.

[0203] In the similarity detection request receiving module, the encrypted message generation process is specifically as follows: obtaining the current timestamp, intercepting the time parameter from the timestamp based on a preset interception rule, splicing the preset key and the time parameter to obtain a first dynamic key, and calling the first dynamic key through the SM4 algorithm to encrypt the name of the drug to be tested twice to obtain an encrypted message;

[0204] The first dynamic key is obtained by concatenating the key and the time parameter, and the first dynamic key is called through the SM4 algorithm to encrypt the name of the drug to be tested twice to obtain an encrypted message; that is, the encryption process of the name of the drug to be tested combines the key, time parameter, SM4 algorithm, and two encryptions, which greatly increases the difficulty of cracking, and thus greatly improves the security of the transmission of the name of the drug to be tested.

[0205] The dynamic token generation process is as follows: the account number and timestamp are spliced ​​to obtain spliced ​​data, the spliced ​​data is encrypted by calling the key through the SM4 algorithm to obtain a second dynamic key, a one-time random number is added to the end of the second dynamic key to obtain a third dynamic key, and the third dynamic key is base64 encoded to obtain a dynamic token;

[0206] The account number and timestamp are spliced ​​to obtain spliced ​​data, the spliced ​​data is encrypted by calling the key through the SM4 algorithm to obtain the second dynamic key, a one-time random number is added to the end of the second dynamic key to obtain the third dynamic key, and the third dynamic key is base64 encoded to obtain the dynamic token; that is, the generation of the dynamic token combines the timestamp, SM4 algorithm, one-time random number, and base64 encoding, and the dynamic token can be subsequently subjected to multi-dimensional verification, thereby greatly improving the security of authentication.

[0207] The authentication module is specifically used for:

[0208] The server parses the similarity detection request to obtain an encrypted message, a dynamic token, and an account number, and matches a key from a password table using the account number;

[0209] The server performs base64 decoding on the dynamic token to obtain a third dynamic key, parses the third dynamic key to obtain a second dynamic key and a one-time random number, and determines whether the one-time random number already exists in a preset random number management table. If so, authentication fails and a notification of authentication failure is fed back; if not, the one-time random number is stored in the random number management table and:

[0210] The key is called by the SM4 algorithm to decrypt the second dynamic key to obtain spliced ​​data, the spliced ​​data is parsed to obtain the account number and timestamp, and a time validity check is performed based on the timestamp. If the timeout is exceeded, the authentication fails and an authentication failure notification is fed back; if the timeout is not exceeded, the authentication succeeds, and the access rights to the pharmacological similarity detection model and the drug database are released, and:

[0211] Based on the preset interception rules, the time parameter is intercepted from the timestamp, the key and the time parameter are spliced ​​to obtain the first dynamic key, and the first dynamic key is called through the SM4 algorithm to decrypt the encrypted message twice to obtain the name of the drug to be tested.

[0212] The first similarity detection sub-result generation module is specifically used to:

[0213] The server creates a detection queue for storing the name of the drug to be tested, the detection status, the process lock status, the number of attempts, and the creation time, stores the name of the drug to be tested carried in the similarity detection request in the detection queue, and updates and initializes the corresponding detection status, process lock status, number of attempts, and creation time; the detection status is not detected, executing the first detection, executing the second detection, execution success, or execution failure; the first detection is performed by a pharmacological similarity detection model; the second detection is performed by calculating the SimHash value; when the number of attempts reaches a preset threshold, the abnormal process automatically rolls back and switches to a backup node;

[0214] The server creates a producer-consumer model including a producer and a consumer. The producer selects a preset number of names of drugs to be tested from the detection queue, and the consumer inputs the selected names of drugs to be tested into the pharmacological similarity detection model to obtain the corresponding first similarity detection sub-result, and synchronously updates the detection queue. When the consumer's process is executing, a process lock is formed based on the local IP+uuid (i.e., a process lock is formed by combining the IP address and the universally unique identifier). In combination with the select * from table1 for update skip locked (this SQL statement is used to implement an efficient row-level locking mechanism in database concurrent access scenarios, and is particularly suitable for high-concurrency queue processing systems, SELECT * FROM table1 is the basic query part, FOR UPDATE is the lock modifier, and SKIP LOCKED is skip lock) row lock method to prevent multiple threaded tasks from operating on the same data at the same time.

[0215] By creating a detection queue for storing the names of drugs to be tested, detection status, process lock status, number of attempts and creation time, a producer-consumer model is created. The producer selects a preset number of drug names to be tested from the detection queue, and the consumer performs pharmacological similarity testing on the selected drug names to be tested. This can effectively solve the data sharing and synchronization problems in a multi-threaded environment, thereby effectively improving concurrency performance and resource utilization.

[0216] The second similarity detection sub-result generation module is specifically used to:

[0217] The server selects a preset number of drug names to be tested from the test queue through the producer, performs word segmentation on each of the drug names to be tested to obtain drug nouns, obtains a proportional weight based on the ratio of the length of each drug noun to the length of the drug name to be tested, counts the number of occurrences of the drug noun in the preset drug list, and calculates the comprehensive weight of each drug noun based on the proportional weight and the number of occurrences: ;

[0218] Among them, C represents the comprehensive weight; M represents the number of occurrences; represents the proportional weight of the i-th drug noun; n represents the total number of drug nouns contained in the names of the drugs to be tested; when the corresponding drug noun does not exist in the drug list, the comprehensive weight of the corresponding drug noun is set to the minimum value, for example, to 0.01;

[0219] The SimHash value of the drug name to be tested is calculated based on the comprehensive weight and the drug noun, the SimHash value in the matching drug database is traversed based on the SimHash value, and the second similarity detection sub-result is generated based on the code distance calculated based on the SimHash value.

[0220] The comprehensive weight is calculated by proportional weight and number of occurrences, and the SimHash value is calculated based on the comprehensive weight, which greatly reduces the misjudgment rate of pharmacological similarity detection.

[0221] The similarity detection report generation module is specifically used to:

[0222] The server generates a similarity detection report including names of pharmacologically similar drugs based on the first similarity detection sub-result and the duplicate items of the second similarity detection sub-result;

[0223] By selecting duplicates of the first similarity detection sub-result and the second similarity detection sub-result, a similarity detection report containing the names of pharmacologically similar drugs is generated, which greatly improves the confidence of the similarity detection report.

[0224] The server performs HMAC calculation on the similarity detection report to obtain a MAC value, encrypts the similarity detection report and the MAC value using the key to obtain primary encrypted data, cyclically shifts each character of the primary encrypted data to the left by 5 characters to obtain secondary encrypted data, encrypts the secondary encrypted data using the 3DES algorithm to obtain an encrypted report, and pushes the encrypted report to the detection client in real time via the TLS protocol;

[0225] The MAC value is obtained by performing HMAC calculation on the similarity detection report, and the similarity detection report and MAC value are encrypted with a key to obtain the first-level encrypted data. Each character of the first-level encrypted data is circularly shifted to the left by 5 characters to obtain the second-level encrypted data. The second-level encrypted data is encrypted using the 3DES algorithm to obtain an encrypted report, and the encrypted report is pushed to the detection client in real time via the TLS protocol. That is, at least five security measures (HMAC calculation, key, character shift, 3DES algorithm, TLS protocol) are combined in the process of pushing the similarity detection report to prevent the plaintext from being stolen and tampered with during the push process of the similarity detection report, thereby greatly improving the security of the push of the similarity detection report.

[0226] The server generates a detection log based on the detection time, the similarity detection report, and the similarity detection request, performs MD5 calculation on the detection log to obtain an MD5 value, encrypts the detection log and the MD5 value using the key to obtain first encrypted data, circularly shifts each character of the first encrypted data to the right by 7 characters to obtain second encrypted data, encrypts the second encrypted data using the RC6 algorithm to obtain an encrypted log, stores the encrypted log, and performs distributed backup.

[0227] The MD5 value is obtained by performing MD5 calculation on the detection log. The detection log and the MD5 value are encrypted using a key to obtain the first encrypted data. Each character of the first encrypted data is circularly shifted right by 7 characters to obtain the second encrypted data. The second encrypted data is encrypted using the RC6 algorithm to obtain an encrypted log. The encrypted log is stored and distributedly backed up. In other words, the detection log storage process combines at least five security measures (MD5 calculation, key, character shift, RC6 algorithm, and distributed backup), which greatly improves the security of detection log storage and facilitates later traceability.

[0228] Security is further improved by adopting different encryption schemes in each link of similarity detection request authentication, similarity detection report push and detection log storage.

[0229] Although the specific embodiments of the present invention are described above, those skilled in the art should understand that the specific embodiments described are merely illustrative and are not intended to limit the scope of the present invention. Equivalent modifications and changes made by those skilled in the art in accordance with the spirit of the present invention should be included within the scope of protection of the claims of the present invention.

Claims

1. A pharmacological similarity detection method combined with a neural network, characterized by: The steps include: Step S1: The server creates a pharmacological similarity detection model based on the input module, the feature extraction module, the feature fusion module, the similarity calculation module, and the output module, and sets a loss function of the pharmacological similarity detection model; The feature extraction module is composed of a common name encoder and a product name encoder; the common name encoder is built based on BioBERT and is used to extract pharmacological semantic features from common names; the product name encoder is built based on BERT and is used to extract product semantic features from product names; The feature fusion module is used to combine the pharmacological semantic features and the product semantic features, and then perform dimensionality reduction through a fully connected layer to obtain fused features; Step S2: The server obtains a large number of drug names including common names and trade names, pre-processes and labels each of the drug names, and then constructs a data set; Step S3: The server trains a pharmacological similarity detection model based on the data set and the loss function, and deploys the trained pharmacological similarity detection model; Step S4: The server creates a drug database and a password table for storing drug data, assigns an account and a key to each detection client and stores them in the password table, pre-installs each key into the corresponding detection client, and sets an authentication interface; Step S5: The server obtains the similarity detection request sent by the detection client through the authentication interface, which carries the encrypted message, dynamic token and account number; The encrypted message generation process is specifically as follows: obtaining the current timestamp, intercepting the time parameter from the timestamp based on a preset interception rule, concatenating the preset key and the time parameter to obtain a first dynamic key, and using the SM4 algorithm to call the first dynamic key to encrypt the name of the drug to be tested twice to obtain an encrypted message; The dynamic token generation process is as follows: the account number and timestamp are spliced ​​to obtain spliced ​​data, the spliced ​​data is encrypted by calling the key through the SM4 algorithm to obtain a second dynamic key, a one-time random number is added to the end of the second dynamic key to obtain a third dynamic key, and the third dynamic key is base64 encoded to obtain a dynamic token; Step S6: After authenticating the similarity detection request, the server releases the calling authority of the pharmacological similarity detection model and the drug database; Step S7: The server inputs the name of the drug to be tested carried in the similarity detection request into the pharmacological similarity detection model to obtain a first similarity detection sub-result; Step S8: The server calculates the SimHash value of the drug name to be tested, and performs a match from the drug database based on the SimHash value to obtain a second similarity detection sub-result; Step S9: The server generates a similarity detection report based on the first similarity detection sub-result and the second similarity detection sub-result, encrypts the similarity detection report into an encrypted report, pushes the encrypted report to the detection client, and generates and stores a detection log.

2. The pharmacological similarity detection method combined with a neural network according to claim 1, characterized in that: In step S1, the input module is used to input the drug name including the common name and the trade name, and pre-process the input drug name; The similarity calculation module is used to calculate the cosine similarity between the fused features, and use the Sigmoid function to map the cosine similarity to the similarity probability; The output module is used to output a first similarity detection sub-result according to the similarity probability; The loss function adopts the binary cross entropy loss function, and the formula is: ; Among them, L represents the loss value of the loss function; N represents the total number of samples of drug names; represents the true similarity label of the i-th drug name; represents the similarity probability of the i-th drug name; The step S2 is specifically as follows: The server obtains a large number of drug names including generic names and trade names, preprocesses each of the drug names including at least word segmentation, removal of stop words, removal of auxiliary words, and removal of conjunctions, and annotates each of the preprocessed drug names with pharmacological similarities to construct a data set; The step S3 is specifically as follows: The server groups the data set based on drug type, splits each group into a training subset, a validation subset, and a test subset in a ratio of 7:2:1, and combines each of the training subsets, validation subsets, and test subsets to obtain a training set, a validation set, and a test set; The pharmacological similarity detection model is trained using the training set, and during the training process, the loss function is optimized using a stochastic gradient descent method until a loss value of the loss function is less than a preset loss threshold; The detection accuracy is calculated using the validation set to determine whether the detection accuracy is greater than a preset accuracy threshold. If not, the verification fails and the training set is expanded to continue training. If so, the verification passes, and: Calculate the F1 score using the test set and determine whether the F1 score is greater than a preset F1 threshold. If not, the test fails and the training set is expanded to continue training. If so, the test passes and the pharmacological similarity detection model is subjected to knowledge distillation and deployed via a Docker container. In step S4, the drug data includes at least the drug code, generic name, trade name, SimHash value of the generic name, SimHash value of the trade name, drug type, chemical information, pharmacological information and clinical information; the key is generated based on the AES256 algorithm; and the authentication interface is an API interface.

3. The pharmacological similarity detection method combined with a neural network as claimed in claim 1, characterized in that: The step S6 is specifically as follows: The server parses the similarity detection request to obtain an encrypted message, a dynamic token, and an account number, and matches a key from a password table using the account number; The server performs base64 decoding on the dynamic token to obtain a third dynamic key, parses the third dynamic key to obtain a second dynamic key and a one-time random number, and determines whether the one-time random number already exists in a preset random number management table. If so, authentication fails and a notification of authentication failure is fed back; if not, the one-time random number is stored in the random number management table and: The key is called by the SM4 algorithm to decrypt the second dynamic key to obtain spliced ​​data, the spliced ​​data is parsed to obtain the account number and timestamp, and a time validity check is performed based on the timestamp. If a timeout occurs, the authentication fails and an authentication failure notification is fed back; If the timeout is not reached, the authentication is successful, and the access rights to the pharmacological similarity detection model and the drug database are released, and: Based on the preset interception rules, the time parameter is intercepted from the timestamp, the key and the time parameter are spliced ​​to obtain the first dynamic key, and the first dynamic key is called through the SM4 algorithm to decrypt the encrypted message twice to obtain the name of the drug to be tested.

4. The pharmacological similarity detection method combined with a neural network as claimed in claim 1, characterized in that: The step S7 is specifically as follows: The server creates a detection queue for storing the name of the drug to be tested, the detection status, the process lock status, the number of attempts, and the creation time, stores the name of the drug to be tested carried in the similarity detection request in the detection queue, and updates and initializes the corresponding detection status, process lock status, number of attempts, and creation time; the detection status is not detected, executing the first detection, executing the second detection, execution success, or execution failure; the first detection is detected by a pharmacological similarity detection model; the second detection is detected by calculating a SimHash value; The server creates a producer-consumer model including a producer and a consumer. The producer selects a preset number of drug names to be tested from the test queue. The consumer inputs the selected drug names to be tested into the pharmacological similarity detection model to obtain the corresponding first similarity detection sub-result, and the test queue is updated synchronously. The step S8 is specifically as follows: The server selects a preset number of drug names to be tested from the test queue through the producer, performs word segmentation on each of the drug names to be tested to obtain drug nouns, obtains a proportional weight based on the ratio of the length of each drug noun to the length of the drug name to be tested, counts the number of occurrences of the drug noun in the preset drug list, and calculates the comprehensive weight of each drug noun based on the proportional weight and the number of occurrences: ; Among them, C represents the comprehensive weight; M represents the number of occurrences; represents the proportion weight of the i-th drug noun; n represents the total number of drug nouns contained in the names of the drugs to be tested; The SimHash value of the drug name to be tested is calculated based on the comprehensive weight and the drug noun, the SimHash value in the matching drug database is traversed based on the SimHash value, and the second similarity detection sub-result is generated based on the code distance calculated based on the SimHash value.

5. The pharmacological similarity detection method combined with a neural network as claimed in claim 1, characterized in that: The step S9 is specifically as follows: The server generates a similarity detection report including names of pharmacologically similar drugs based on the first similarity detection sub-result and the duplicate items of the second similarity detection sub-result; The server performs HMAC calculation on the similarity detection report to obtain a MAC value, encrypts the similarity detection report and the MAC value using the key to obtain primary encrypted data, cyclically shifts each character of the primary encrypted data to the left by 5 characters to obtain secondary encrypted data, encrypts the secondary encrypted data using the 3DES algorithm to obtain an encrypted report, and pushes the encrypted report to the detection client in real time via the TLS protocol; The server generates a detection log based on the detection time, the similarity detection report, and the similarity detection request, performs MD5 calculation on the detection log to obtain an MD5 value, encrypts the detection log and the MD5 value using the key to obtain first encrypted data, circularly shifts each character of the first encrypted data to the right by 7 characters to obtain second encrypted data, encrypts the second encrypted data using the RC6 algorithm to obtain an encrypted log, stores the encrypted log, and performs distributed backup.

6. A pharmacological similarity detection system incorporating a neural network, characterized by: Includes the following modules: A pharmacological similarity detection model creation module is used for the server to create a pharmacological similarity detection model based on the input module, feature extraction module, feature fusion module, similarity calculation module and output module, and set the loss function of the pharmacological similarity detection model; The feature extraction module is composed of a common name encoder and a product name encoder; the common name encoder is built based on BioBERT and is used to extract pharmacological semantic features from common names; the product name encoder is built based on BERT and is used to extract product semantic features from product names; The feature fusion module is used to combine the pharmacological semantic features and the product semantic features, and then perform dimensionality reduction through a fully connected layer to obtain fused features; The data set construction module is used for the server to obtain a large number of drug names including common names and trade names, and to construct a data set after preprocessing and labeling each of the drug names; A pharmacological similarity detection model training module is used for the server to train the pharmacological similarity detection model based on the data set and the loss function, and deploy the trained pharmacological similarity detection model; A drug database and password table creation module is used for the server to create a drug database and a password table for storing drug data, assign an account and a key to each detection client and store them in the password table, pre-set each key into the corresponding detection client, and set an authentication interface; A similarity detection request receiving module is used for the server to obtain the similarity detection request sent by the detection client through the authentication interface, which carries the encrypted message, dynamic token and account number; The encrypted message generation process is specifically as follows: obtaining the current timestamp, intercepting the time parameter from the timestamp based on a preset interception rule, concatenating the preset key and the time parameter to obtain a first dynamic key, and using the SM4 algorithm to call the first dynamic key to encrypt the name of the drug to be tested twice to obtain an encrypted message; The dynamic token generation process is as follows: the account number and timestamp are spliced ​​to obtain spliced ​​data, the spliced ​​data is encrypted by calling the key through the SM4 algorithm to obtain a second dynamic key, a one-time random number is added to the end of the second dynamic key to obtain a third dynamic key, and the third dynamic key is base64 encoded to obtain a dynamic token; An authentication module, configured to release the calling authority of the pharmacological similarity detection model and the drug database after the server authenticates the similarity detection request; A first similarity detection sub-result generating module is configured to input the name of the drug to be tested carried in the similarity detection request into a pharmacological similarity detection model on the server to obtain a first similarity detection sub-result; A second similarity detection sub-result generation module is used for the server to calculate the SimHash value of the name of the drug to be tested, and match it from the drug database based on the SimHash value to obtain a second similarity detection sub-result; A similarity detection report generation module is used for the server to generate a similarity detection report based on the first similarity detection sub-result and the second similarity detection sub-result, encrypt the similarity detection report into an encrypted report, push the encrypted report to the detection client, and generate and store a detection log.

7. The pharmacological similarity detection system incorporating a neural network as claimed in claim 6, characterized in that: In the pharmacological similarity detection model creation module, the input module is used to input the drug name including the common name and the trade name, and pre-process the input drug name; The similarity calculation module is used to calculate the cosine similarity between the fused features, and use the Sigmoid function to map the cosine similarity to the similarity probability; The output module is used to output a first similarity detection sub-result according to the similarity probability; The loss function adopts the binary cross entropy loss function, and the formula is: ; Among them, L represents the loss value of the loss function; N represents the total number of samples of drug names; represents the true similarity label of the i-th drug name; represents the similarity probability of the i-th drug name; The dataset construction module is specifically used for: The server obtains a large number of drug names including generic names and trade names, preprocesses each of the drug names including at least word segmentation, removal of stop words, removal of auxiliary words, and removal of conjunctions, and annotates each of the preprocessed drug names with pharmacological similarities to construct a data set; The pharmacological similarity detection model training module is specifically used for: The server groups the data set based on drug type, splits each group into a training subset, a validation subset, and a test subset in a ratio of 7:2:1, and combines each of the training subsets, validation subsets, and test subsets to obtain a training set, a validation set, and a test set; The pharmacological similarity detection model is trained using the training set, and during the training process, the loss function is optimized using a stochastic gradient descent method until a loss value of the loss function is less than a preset loss threshold; The detection accuracy is calculated using the validation set to determine whether the detection accuracy is greater than a preset accuracy threshold. If not, the verification fails and the training set is expanded to continue training. If so, the verification passes, and: Calculate the F1 score using the test set and determine whether the F1 score is greater than a preset F1 threshold. If not, the test fails and the training set is expanded to continue training. If so, the test passes and the pharmacological similarity detection model is subjected to knowledge distillation and deployed via a Docker container. In the drug database and password table creation module, the drug data includes at least the drug code, generic name, trade name, SimHash value of the generic name, SimHash value of the trade name, drug type, chemical information, pharmacological information and clinical information; the key is generated based on the AES256 algorithm; and the authentication interface is an API interface.

8. The pharmacological similarity detection system incorporating a neural network as claimed in claim 6, characterized in that: The authentication module is specifically used for: The server parses the similarity detection request to obtain an encrypted message, a dynamic token, and an account number, and matches a key from a password table using the account number; The server performs base64 decoding on the dynamic token to obtain a third dynamic key, parses the third dynamic key to obtain a second dynamic key and a one-time random number, and determines whether the one-time random number already exists in a preset random number management table. If so, authentication fails and a notification of authentication failure is fed back; if not, the one-time random number is stored in the random number management table and: The key is called by the SM4 algorithm to decrypt the second dynamic key to obtain spliced ​​data, the spliced ​​data is parsed to obtain the account number and timestamp, and a time validity check is performed based on the timestamp. If a timeout occurs, the authentication fails and an authentication failure notification is fed back; If the timeout is not reached, the authentication is successful, and the access rights to the pharmacological similarity detection model and the drug database are released, and: Based on the preset interception rules, the time parameter is intercepted from the timestamp, the key and the time parameter are spliced ​​to obtain the first dynamic key, and the first dynamic key is called through the SM4 algorithm to decrypt the encrypted message twice to obtain the name of the drug to be tested.

9. The pharmacological similarity detection system incorporating a neural network as claimed in claim 6, characterized in that: The first similarity detection sub-result generation module is specifically used to: The server creates a detection queue for storing the name of the drug to be tested, the detection status, the process lock status, the number of attempts, and the creation time, stores the name of the drug to be tested carried in the similarity detection request in the detection queue, and updates and initializes the corresponding detection status, process lock status, number of attempts, and creation time; the detection status is not detected, executing the first detection, executing the second detection, execution success, or execution failure; the first detection is detected by a pharmacological similarity detection model; the second detection is detected by calculating a SimHash value; The server creates a producer-consumer model including a producer and a consumer. The producer selects a preset number of drug names to be tested from the test queue. The consumer inputs the selected drug names to be tested into the pharmacological similarity detection model to obtain the corresponding first similarity detection sub-result, and the test queue is updated synchronously. The second similarity detection sub-result generation module is specifically used to: The server selects a preset number of drug names to be tested from the test queue through the producer, performs word segmentation on each of the drug names to be tested to obtain drug nouns, obtains a proportional weight based on the ratio of the length of each drug noun to the length of the drug name to be tested, counts the number of occurrences of the drug noun in the preset drug list, and calculates the comprehensive weight of each drug noun based on the proportional weight and the number of occurrences: ; Among them, C represents the comprehensive weight; M represents the number of occurrences; represents the proportion weight of the i-th drug noun; n represents the total number of drug nouns contained in the names of the drugs to be tested; The SimHash value of the drug name to be tested is calculated based on the comprehensive weight and the drug noun, the SimHash value in the matching drug database is traversed based on the SimHash value, and the second similarity detection sub-result is generated based on the code distance calculated based on the SimHash value.

10. The pharmacological similarity detection system combined with a neural network as claimed in claim 6, characterized in that: The similarity detection report generation module is specifically used to: The server generates a similarity detection report including names of pharmacologically similar drugs based on the first similarity detection sub-result and the duplicate items of the second similarity detection sub-result; The server performs HMAC calculation on the similarity detection report to obtain a MAC value, encrypts the similarity detection report and the MAC value using the key to obtain primary encrypted data, cyclically shifts each character of the primary encrypted data to the left by 5 characters to obtain secondary encrypted data, encrypts the secondary encrypted data using the 3DES algorithm to obtain an encrypted report, and pushes the encrypted report to the detection client in real time via the TLS protocol; The server generates a detection log based on the detection time, the similarity detection report, and the similarity detection request, performs MD5 calculation on the detection log to obtain an MD5 value, encrypts the detection log and the MD5 value using the key to obtain first encrypted data, circularly shifts each character of the first encrypted data to the right by 7 characters to obtain second encrypted data, encrypts the second encrypted data using the RC6 algorithm to obtain an encrypted log, stores the encrypted log, and performs distributed backup.

Citation Information

Patent Citations

  • Drug indexing method and drug retrieval method and system

    CN111198887A

  • Semantic extraction method and device based on artificial intelligence, electronic equipment and medium

    CN114077841A