New Psychoactive Substances Identification and Rapid Comparison Method and System Based on Big Data
Through a new psychoactive substance identification and rapid comparison system based on big data, the problem that traditional detection methods take a long time and are difficult to quickly deal with new drugs is solved, and the rapid and efficient identification and comparison of new drugs is achieved, and the response ability of public safety supervision is improved.
Patent Information
- Application Number
- CN202510097608.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-01-22
AI Technical Summary
Traditional new psychoactive substances (NPS) detection and identification methods rely on manual empirical analysis, which takes time and is difficult to quickly deal with the complex chemical structures and rapidly growing types of new drugs, resulting in regulatory authorities being in a passive situation in dealing with NPS.
A new psychoactive substance recognition and rapid comparison system based on big data is adopted, and the rapid identification and comparison of new drugs can be achieved through the automatic data acquisition module, HBase-based data storage module, NPS molecular structure analysis module, NPS intelligent comparison module and NPS result feedback and early warning module.
It significantly shortens the time from preliminary detection to final confirmation, improves the speed and accuracy of new drug identification, reduces artificial errors, can cope with the continuous increase in NPS types, improves the response speed and processing capabilities of regulatory authorities, and ensures public safety.
Smart Images

Figure CN119517223B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of identification of new addictive chemical substances, and in particular to a method and system for identifying and rapidly comparing new psychoactive substances based on big data. Background Art
[0002] With the rapid increase in the types of new psychoactive substances (NPS), the global drug market has shown unprecedented complexity and diversity. These substances usually change their chemical structures to evade existing legal controls and detection technologies, enabling them to circulate in the market without being detected in a timely manner. Traditional NPS detection and identification methods mainly rely on sophisticated instrument detection in laboratories. After detection, the results are compared with the substances listed in the control list to determine whether they match. However, the structural analysis step in this process highly depends on manual experience analysis. Especially when the detection results do not exactly match the known substances in the control list, the appraisers need to conduct manual comparison based on the similarity of chemical structures to determine whether the substance is related to some listed substances. This manual analysis process not only takes a long time but also requires high professional knowledge of the appraisers and is often difficult to complete quickly.
[0003] In the current manual comparison process, due to the complex chemical structures and a vast number of possible combinations, the analysis of each sample takes several hours or even several days. This means that a relatively long time period is required from the initial detection to the final confirmation and submission to the regulatory authorities. During this period, the detected new psychoactive substances may have already flowed into the market, posing a serious threat to public health and social safety. Facing these problems, the lag and dependence of traditional detection methods put the regulatory authorities in a passive position when dealing with NPS, making it difficult to quickly respond to emergencies and reduce public risks.
[0004] Finally, due to the complex chemical structure of new psychoactive substances (NPS) and the ease of generating a large number of new derivatives through minor modifications, the types of NPS show a geometric growth trend. Manufacturers often circumvent existing legal regulations by fine-tuning the molecular structure. This flexibility results in each chemical variation potentially forming a new substance, increasing the detection difficulty and the need for scheduling. However, traditional database architectures rely on the concept of regular updates and linear expansion, making it difficult to keep up with the rapid increase in the types of NPS. As the data volume increases, the query speed and processing efficiency of traditional databases in structural similarity comparison also decrease significantly, unable to meet the requirements of rapid and efficient identification. Therefore, a new database concept capable of adapting to the explosive growth of data is needed, which supports efficient horizontal expansion and can complete the comparison and query of a large number of substances within seconds to cope with the increasingly complex chemical changes and growth trends of NPS, helping laboratories and regulatory authorities to promptly discover and handle potential new addictive substances. Summary of the Invention
[0005] Aiming at the above-mentioned shortcomings of the existing technology, the present invention proposes a method and system for identifying and rapidly comparing new psychoactive substances based on big data, which can achieve rapid, efficient, and automated identification and comparison of new psychoactive substances and can cope with the possible geometric growth of new psychoactive substances in the future.
[0006] The system for identifying and rapidly comparing new psychoactive substances based on big data disclosed by the present invention includes a data automatic collection module, an NPS data storage module based on HBase, an NPS molecular structure analysis module, an NPS intelligent comparison module, an NPS result feedback and warning module, and an application terminal;
[0007] The data automatic collection module, the NPS data storage module based on HBase, the NPS molecular structure analysis module, the NPS intelligent comparison module, and the NPS result feedback and warning module are all deployed on a distributed server in a microservices architecture, and data interaction is realized between each module through an API interface;
[0008] Among them:
[0009] The data automatic collection module is used to collect in real time the result data detected by laboratory instruments and equipment and the updated data in the specified network scheduling list;
[0010] The NPS data storage module based on HBase is used to receive the data sent by the data automatic collection module, clean the data and convert it into a key-value pair format containing the molecular structure, and store it in the Hbase database;
[0011] The NPS molecular structure analysis module is used to convert and analyze the data in the key-value pair format containing the molecular structure, obtain the molecular fingerprint of the new psychoactive substance, and return it to the NPS data storage module based on HBase;
[0012] The NPS intelligent comparison module is used to compare the molecular fingerprint of the substance to be detected one by one with the molecular fingerprint data of the listed controlled substances stored in the MySQL database, and send the comparison result to the NPS result feedback and warning module;
[0013] The NPS result feedback and warning module is used to generate warning information for the comparison results exceeding the threshold through threshold comparison, send it to the application terminal for the user to judge, and at the same time receive the judgment result of the warning information from the application terminal, and dynamically adjust the threshold according to the judgment result;
[0014] The application terminal is used to provide a graphical interface for the user, facilitating the user to initiate a comparison request to the NPS intelligent comparison module and display and feedback the user's judgment result.
[0015] Furthermore, the data automatic collection module includes:
[0016] The instrument equipment data collection unit is used to automatically collect the newly generated test result data of the laboratory instrument equipment;
[0017] The network data collection unit is used to automatically obtain the updated NPS data in the listed controlled list published on the Internet through a web crawler, and the collected NPS data flows into the distributed server unidirectionally through the boundary platform.
[0018] Furthermore, during the storage process of the NPS data storage module based on HBase, first, the NPS molecular structure analysis module analyzes the data in the key-value pair format containing the molecular structure, and judges whether the analyzed data is in the listed controlled directory;
[0019] If so, assign the mark "1" to the analyzed data, indicating "confirmed as a new psychoactive substance", and transfer it to the MySQL database;
[0020] If not, send the analyzed data to the NPS intelligent comparison module, and judge whether the analyzed data should apply for listing control according to the feedback results of the NPS intelligent comparison module and the NPS result feedback and warning module; if so, assign the mark "1" to the analyzed data, and transfer it to the MySQL database, if not, assign the mark "0" to it, indicating "not confirmed as a new psychoactive substance", and store it in the Hbase database.
[0021] Furthermore, the NPS molecular structure analysis module includes:
[0022] A molecular structure conversion unit is used to convert key-value pair data containing molecular structures into SMILES structures, extract the structural features of the molecules from the SMILES structures, perform hash processing on the structural features to generate bit vectors, and use the bit vectors as molecular fingerprints.
[0023] A new key-value pair generation unit is used to generate a pair of new key-value pairs based on the SMILES structure and return the new key-value pairs to the NPS data storage module based on HBase for storage in the Hbase database.
[0024] Among them, the key is the molecular fingerprint, and the value is a byte array of a unique two-dimensional matrix character describing the structural features of the molecule.
[0025] Furthermore, the NPS intelligent comparison module includes:
[0026] A request data format judgment unit is used to receive a comparison request sent by an application terminal and judge whether the data format in the comparison request is a molecular fingerprint format.
[0027] If so, it is sent to the comparison unit; if not, the comparison request is forwarded to the NPS molecular structure parsing module, and the data in the comparison request is converted into a molecular fingerprint by the NPS molecular structure parsing module and then returned to the NPS intelligent comparison module.
[0028] The comparison unit, based on the Tanimoto coefficient, compares the received molecular fingerprints with the data stored in the MySQL database one by one, outputs a value between 0 and 1, and sends the value as the comparison result to the NPS result feedback and warning module.
[0029] Furthermore, the NPS result feedback and warning module includes:
[0030] A comparison result receiving unit is used to receive the comparison result sent by the NPS intelligent comparison module.
[0031] A threshold comparison unit is used to compare the value of the comparison result with a preset threshold, generate a warning message for comparison results greater than or equal to the threshold, and send the warning message to the application terminal.
[0032] A threshold self-adjustment unit is used to receive the judgment result of the application terminal on the warning message, learn according to the judgment result, and automatically adjust the preset threshold.
[0033] The method for identifying and quickly comparing new psychoactive substances based on big data is implemented according to the system for identifying and quickly comparing new psychoactive substances based on big data, and it includes the following implementation steps:
[0034] Step 1: Build a distributed server;
[0035] Step 2: The user configures the basic parameters of each microservice;
[0036] Step 3: The data automatic acquisition module automatically obtains the detection result data of laboratory instruments and equipment and the updated data in the specified network control list, and then the NPS molecular structure analysis module analyzes these data into the molecular fingerprint format and stores the data through the HBase-based data storage module;
[0037] Step 4: The user initiates a comparison request through the application terminal. The NPS intelligent comparison module receives the comparison request, compares the similarity of the molecular fingerprints in the comparison request data, and sends the comparison result to the NPS result feedback and warning module;
[0038] The NPS result feedback and warning module compares the comparison result with a preset threshold, generates a warning message for the comparison result greater than or equal to the threshold, and sends it to the application terminal;
[0039] After receiving the warning message, the user makes a judgment on the application terminal and returns the judgment result to the NPS result feedback and warning module to readjust the threshold.
[0040] Further, the specific steps of Step 3 include:
[0041] Step 3.1: The network data acquisition unit monitors the specified network control list on the Internet in real time. If it finds that there is data update, it crawls the latest data and uploads it to the HBase-based NPS data storage module;
[0042] When storing in the HBase-based NPS data storage module, if the column family of the current data already exists, the current data is updated, and the original data is stored in the old version at the same time; if the current data column family does not exist, a new row is added and the data is marked; then the new data containing the molecular structure is passed to the NPS molecular structure analysis module;
[0043] Step 3.2: After receiving the data containing the molecular structure, the NPS molecular structure analysis module converts it into a SMILES structure body, then extracts the structural features of the molecule through the SMILES structure body, performs hash processing on the structural features of the molecule to obtain the molecular fingerprint, and then returns the molecular fingerprint to the HBase-based NPS data storage module.
[0044] Further, the specific steps of Step 4 include:
[0045] Step 4.1: The user logs in to this system through the application terminal. If the current molecular fingerprint is a new substance, a comparison request is initiated;
[0046] Step 4.2: The application terminal sends the comparison request to the NPS intelligent comparison module. After receiving the comparison request, the NPS intelligent comparison module determines the data format of the comparison request. If the format is a non-molecular fingerprint format, the comparison request is forwarded to the NPS molecular structure analysis module.
[0047] The NPS molecular structure analysis module converts the comparison request into a molecular fingerprint format and returns it to the NPS intelligent comparison module. After receiving the molecular fingerprint, the NPS intelligent comparison module compares the structural features of the vector string molecules one by one with the bit vectors already existing in the MySQL database, and sends the comparison results to the NPS result feedback and warning module;
[0048] Step 4.3: The NPS result feedback and warning module compares the received comparison result with the preset threshold, generates warning information for the result greater than or equal to the threshold, and sends the warning information to the application terminal;
[0049] Step 4.4: The user views the warning information through the application terminal, and at the same time initiates a three-dimensional molecular formula viewing request to the HBase-based NPS data storage module according to the relevant substance ID in the warning information. The user again determines whether the new substance needs to be listed based on the three-dimensional molecular formula viewing result, and feeds back the judgment result to the NPS result feedback and warning module. The NPS result feedback and warning module learns according to the user's judgment result and automatically adjusts the preset threshold;
[0050] Step 4.5: The NPS result feedback and early warning module generates a corresponding application report based on the application listing information determined by the user, and sends the relevant information of the new substance to the NPS data storage module based on HBase, and modifies the corresponding column attributes in the Hbase database according to the relevant substance ID, changing its identifier to "1".
[0051] Furthermore, in step 4.4, the specific steps of automatically adjusting the threshold include:
[0052] Step 4.4.1: Set the threshold;
[0053] Step 4.4.2: Calculate the similarity between compound pairs using molecular fingerprints;
[0054] Step 4.4.3: Use cross validation to test the accuracy and recall under different thresholds and select the optimal threshold;
[0055] Step 4.4.4: During the comparison process of the NPS intelligent comparison module, the threshold is dynamically adjusted according to the judgment result fed back by the application terminal to ensure the comparison quality;
[0056] Step 4.4.5: Verify whether the adjusted threshold can improve the comparison accuracy and obtain the verification result;
[0057] Step 4.4.6: Finally determine the adjusted threshold according to the verification result.
[0058] Therefore, the present invention adopts the above-mentioned method and system for identifying and rapidly comparing new psychoactive substances based on big data, and has the following beneficial effects:
[0059] First, the present invention can greatly improve the speed and accuracy of identifying new psychoactive substances (NPS). Different from the traditional detection method that relies on manual comparison, the system compares the detection results with a vast amount of known substances through an automated process. This can greatly shorten the time from preliminary detection to final confirmation, avoid the risk of new drugs circulating in the market, and thus effectively ensure public safety.
[0060] Second, based on the multi-version feature of HBase, this system can provide accurate decision support based on historical data, chemical structure, and similarity analysis. Whenever new detection results enter the system, the system will automatically analyze the molecular structure of the substance, compare it with the existing data, and generate a label. This data-driven approach not only reduces human error but also can discover potential similarities that are difficult to detect manually, thus helping the detection personnel to more accurately identify new drugs.
[0061] Third, the present invention uses HBase as the infrastructure for data storage, which can effectively cope with the continuous increase in the types of NPS and support horizontal expansion. This means that as the types of NPS substances surge, the system can handle larger-scale data by simply adding column families and columns without large-scale reconstruction of the system.
[0062] Fourth, when a new substance is identified and marked as a potential new psychoactive substance, this system will automatically generate a warning message to remind the regulatory authorities to take appropriate measures. The warning mechanism can effectively improve the response speed and processing ability by linking with the information systems of other regulatory authorities, and avoid new drugs from flowing into the market and causing harm.
[0063] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Description of the Drawings
[0064] Figure 1 It is the architecture diagram of the system proposed by the present invention.
[0065] Figure 2 It is the reference EI mass spectrum and characteristic mass spectral fragments of etomidate.
[0066] Figure 3 It is the reference EI mass spectrum and characteristic mass spectral fragments of metomidate.
[0067] Figure 4 Reference EI mass spectrum and characteristic mass spectral fragments of Isopropoxate.
[0068] Figure 5 Similarity of Methomidate to Etomidate using MACCS molecular fingerprints.
[0069] Figure 6 Similarity of Isopropoxate to Etomidate using MACCS molecular fingerprints.
[0070] Figure 7 Two-dimensional structure comparison diagram of Methomidate and Etomidate.
[0071] Figure 8 Three-dimensional structure comparison diagram of Methomidate and Etomidate.
[0072] Figure 9 Similarity results and frequency distribution histograms (except Khat) of 46 NPS in 2024 compared with controlled drugs in previous years. Detailed implementation manners
[0073] In the description of the present invention, it should also be noted that unless otherwise clearly specified and limited, these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. In addition, it should be understood that after reading the content taught by the present invention, those skilled in the art make various changes or modifications to the present invention, and these equivalent forms also fall within the scope defined by the appended claims of this application.
[0074] The present invention proposes a method and system for the identification and rapid comparison of new psychoactive substances based on big data. Based on advanced cheminformatics and data science technologies, the data obtained by precision detection instruments is directly compared with an ever-updating NPS database. Through the NPS intelligent comparison module, the closest matching results are quickly screened out, and similarity analysis and detailed chemical structure information are provided within seconds. This can not only significantly reduce the manual intervention time, but also discover some potential similarities that are difficult to identify manually through the application of machine learning and big data analysis, accelerating the identification and supervision process of suspected substances. At the same time, the automated characteristics of the system can greatly improve the processing efficiency, cope with the challenge of the continuous increase in the types of NPS in the future, and provide solid technical support for public safety supervision.
[0075] 1. New psychoactive substance identification and rapid comparison system based on big data:
[0076] As Figure 1As shown in the figure, the system includes a data automatic acquisition module, an NPS data storage module based on HBase, an NPS molecular structure analysis module, an NPS intelligent comparison module, an NPS result feedback and warning module, and an application terminal. The data automatic acquisition module, the NPS data storage module based on HBase, the NPS molecular structure analysis module, the NPS intelligent comparison module, and the NPS result feedback and warning module are all deployed on a distributed server in a microservices architecture. Each service communicates through a local area network and realizes data interaction between modules through an API interface.
[0077] Among them, any version of Nginx is used for the distributed server. The fingerprint data of the listed molecules is stored in the Mysql 5.7 database. The overall database uses HBase 2.4.9, and the file storage system uses Hadoop 3.3.0. The backend programming languages are JAVA and Python. The system architecture combines Spring Boot 2 and Vue 2.0.
[0078] The following is an introduction to each module in the system:
[0079] (1) Data automatic acquisition module:
[0080] This module is used to collect the detection result data of experimental instrument and equipment in real time and the network listed inventory data. It includes:
[0081] The instrument and equipment data acquisition unit is used to automatically collect the newly generated mass spectrometry fragments and detection result data of experimental instrument and equipment.
[0082] The network data acquisition unit is used to automatically obtain the updated NPS data in the listed inventory published on the Internet through a web crawler. The collected NPS data flows unidirectionally into the distributed server through the border platform of the public security network to ensure the security and consistency of the data. The listed inventory is specified by configuring the website address.
[0083] Specifically, the experimental instrument and equipment are mainly liquid chromatography-mass spectrometry or gas chromatography-mass spectrometry. The detection data formats of different instrument manufacturers are different, but the data all contain information such as ion characteristic fragments, retention time, and peak area. Generally, the values of 4 characteristic ion fragments are extracted as representatives to form the data format shown in Table 1:
[0084] Table 1 Detection data format of experimental instrument and equipment:
[0085] ;
[0086] After processing the data shown in Table 1, the data format shown in Table 2 is obtained.
[0087] Table 2 Format of the detected data of the processed experimental instruments and equipment:
[0088] ;
[0089] The data structure of Table 2 is stored using the Hbase database. Three column families are defined, namely mz_info, area_info, and rt_info. These three column families are defined as basic column families, and these three column families are the mass-to-charge ratio, peak area, and retention time respectively. If other attributes are needed in the future, only new column families need to be added, which can fully meet the extended requirements of subsequent data changes and give full play to the flexible expansion characteristics of HBase.
[0090] After the detection of the experimental instruments and equipment is completed, data files in a specified file format will be generated. These data files are stored in the corresponding storage directory, and this directory is synchronized to the distributed server. The instrument data acquisition unit obtains the new data in this directory in real time through the Server Message Block (SMB) protocol.
[0091] As can be seen from the above, through this module, the system can not only obtain experimental data in real time, but also dynamically update the list information of the column tubes, providing complete data support for subsequent comparison and analysis.
[0092] (2) NPS Data Storage Module Based on HBase:
[0093] This module is used to receive the data sent by the data automatic acquisition module, clean and transform it, convert the data into key-value pair format, and store it in the preset column family or the added column family of HBase.
[0094] During the storage process, first, the NPS molecular structure parsing module parses the converted data, and then judges whether the parsed data is in the listed tube directory. If so, assign the mark "1" to this data, indicating "confirmed as a new psychoactive substance", and transfer it to the MySQL database; if not, send the parsed data to the NPS intelligent comparison module, and the user judges whether it is a new psychoactive substance based on the feedback results of the NPS intelligent comparison module and the NPS result feedback and warning module. If so, assign the mark "1" to this data, otherwise assign the mark "0" to it, indicating "not confirmed as a new psychoactive substance", and store it in the Hbase database;
[0095] Since the column families of HBase support infinite expansion, the system can flexibly adapt to different data formats without paying attention to the specific structure of the data. At the same time, it reserves expansion space for possible future added mark types, ensuring that the system can achieve horizontal expansion by adding columns to meet the ever-changing data requirements.
[0096] In summary, the present invention stores based on Hase data, which is a data storage method based on the concept of big data. It can process both structured and unstructured data simultaneously, and can also scale horizontally to cope with future data growth and adapt to changing data requirements.
[0097] (3) NPS Molecular Structure Analysis Module:
[0098] This module is used to analyze the molecular formula of new psychoactive substances to obtain the molecular fingerprints of the new psychoactive substances;
[0099] It includes the following units:
[0100] The molecular structure conversion unit is used to convert the key-value pairs containing molecular structures collected by the data automatic collection module into SMILES (Simplified Molecular Input Line Entry System) structures, then extract the structural features of the molecules from the SMILES structures, and perform hashing processing on the structural features to generate bit vectors, and use the bit vectors as molecular fingerprints;
[0101] The new key-value pair generation unit is used to generate a pair of new key-value pairs according to the SMILES structure, where the key is the molecular fingerprint and the value is a byte array of unique two-dimensional matrix characters describing the structural features of the molecule, and return the new key-value pairs to the NPS data storage module based on HBase and store them in the Hbase database.
[0102] (4) NPS Intelligent Comparison Module:
[0103] This module is used to compare the molecular fingerprints with the data stored in MySQL one by one and send the comparison results to the NPS result feedback and warning module; it includes the following units:
[0104] The request data format judgment unit is used to judge whether the format of the request data is the molecular fingerprint format when receiving the comparison request sent by the application terminal. If not, forward the request data to the NPS molecular structure analysis module, and after converting the request data into a molecular fingerprint through the NPS molecular structure analysis module, return it to the NPS intelligent comparison module;
[0105] The comparison unit is based on the Tanimoto coefficient to compare the values of the received molecular fingerprints with the listed controlled molecular fingerprint data stored in MySQL one by one, and finally output a value between 0 and 1 and send the result to the NPS result feedback and warning module;
[0106] The working principle of the comparison unit is as follows: Each bit corresponds to a molecular fragment. Assuming that there must be many common fragments between similar molecules, then molecules with similar fingerprints are very likely to be similar in 2D structure. Therefore, in the present invention, the Tanimoto coefficient is used to evaluate the quantization comparison result of the bit string, and finally a value between 0 and 1 is obtained.
[0107] (5)NPS result feedback and warning module:
[0108] This module is used to receive the comparison result sent by the NPS intelligent comparison module in real time, compare the value of the comparison result with a preset threshold, generate a warning message for the comparison result exceeding the threshold, and send the warning message to the application terminal at the same time. At the same time, it receives the judgment result of the application terminal for the warning, can learn according to the judgment result, and automatically adjust the threshold;
[0109] The warning message at least includes relevant substance IDs, such as the substance to be detected, the existing substances, etc. After receiving the warning message, the application terminal queries in the MySQL database through the relevant substance IDs in the warning message, and generates a three-dimensional molecular structure body in a visual manner for the queried molecular structure, which is convenient for the staff to further analyze and judge.
[0110] (6)Application terminal:
[0111] The application terminal includes an all-in-one touch computer, a tablet computer, a notebook computer, a desktop computer, etc., and needs to have a network connection function. Its main function is to provide a graphical interface for the user, which is convenient for the user to initiate a comparison request to the NPS intelligent comparison module and display the processing result.
[0112] 2. Big data-based new psychoactive substance identification and rapid comparison method:
[0113] Based on the big data-based new psychoactive substance identification and rapid comparison system of the present invention, a big data-based new psychoactive substance identification and rapid comparison method is realized, which includes the following implementation steps:
[0114] Step 1: Build a distributed server;
[0115] Step 1.1: Design the roles and node numbers of the distributed storage file system according to its own data volume. In the present invention, two NameNode nodes are set, one of which is Active and the other is StandBy. According to the current data volume, 5 DataNode nodes are designed;
[0116] Step 1.2: Install Hbase, place the role of HMaster on the same node as NameNode, and place the role of HRegionServer on the same node as DataNode;
[0117] Step 1.3: Install zookeeper, which is used for the master-slave election of NameNode and HMaster, and stores the metadata of HRegionServer of HBase at the same time;
[0118] Step 1.4: Install the MySQL database;
[0119] Step 1.5: Install and configure the environment for the data automatic collection module, the NPS data storage module based on HBase, the NPS molecular structure parsing module, the NPS intelligent comparison module, and the NPS result feedback and warning module as microservices.
[0120] Step 2: Initialize;
[0121] Step 2.1: The user sets the number of versions to be saved according to the actual application;
[0122] Step 2.2: The user sets the Internet website address to be automatically obtained;
[0123] Step 2.3: The user sets the IP address of the instrument and equipment to be automatically obtained and the storage directory corresponding to the experimental detection data.
[0124] Step 2.4: The user sets the threshold in the NPS result feedback and warning module;
[0125] Step 3: Automatically obtain NPS data and parse NPS data through the data automatic collection module and the NPS molecular structure parsing module respectively;
[0126] Step 3.1: The network data collection unit monitors the specified Internet website in real time, compares the data update date captured with the system's last version date. If an update is found, it crawls the latest data and uploads it to the NPS data storage module based on HBase for storage; when storing, if the column family of the current data already exists, the current data is updated, and the original data is stored as an old version; if the current data column family does not exist, a new row is added, and the data is marked according to the identifier of the automatic collection end; finally, the data containing the molecular formula of the new version is passed to the NPS molecular structure parsing module;
[0127] Step 3.2: After receiving the data, the NPS molecular structure analysis module judges it. If it is a molecular formula, it converts it into a SMILES structure according to the molecular formula, abstracts the molecular features through the SMILES structure to obtain a molecular fingerprint, and then returns the molecular fingerprint to the NPS data storage module based on HBase;
[0128] Step 4: Compare the similarity of the molecular fingerprints through the NPS intelligent comparison module;
[0129] Step 4.1: The user logs in to the system through the application terminal. If the current molecular fingerprint is a new substance, a similarity comparison request is initiated;
[0130] Step 4.2: The application terminal sends the comparison request to the NPS intelligent comparison module. After receiving the comparison request, the NPS intelligent comparison module judges the data format of the request. If the format is not the molecular fingerprint format, the request is forwarded to the NPS molecular structure analysis module; the NPS molecular structure analysis module converts the molecular structure into the molecular fingerprint format and returns it to the NPS intelligent comparison module. After receiving the molecular fingerprint, the NPS intelligent comparison module connects to MySQL and compares the molecular features of the vector strings one by one with the existing bit vectors in MySQL, and sends the comparison result to the NPS result feedback and warning module;
[0131] Step 4.3: After receiving the comparison result, the NPS result feedback and warning module compares the received comparison result with the preset threshold, generates a warning message for the result greater than or equal to the threshold, and feeds back the warning message to the application terminal;
[0132] Step 4.4: The user views the warning message through the application terminal, and at the same time initiates a three-dimensional molecular formula viewing request to the data storage module based on HBase according to the relevant substance ID in the warning message. The user judges again whether the substance needs to apply for scheduling according to the three-dimensional molecular formula viewing result, and feeds back the judgment result to the NPS result feedback and warning module, and learns according to the user's judgment result to automatically adjust the preset threshold;
[0133] Step 4.5: The NPS result feedback and warning module generates a corresponding application report according to the scheduling application information determined by the user, and at the same time sends the information of the substance to the data storage module based on HBase. The data storage module based on HBase modifies the identifier of the data to "1" by modifying its corresponding column attribute, and stores each version of the application report.
[0134] Specifically, the specific steps for automatically adjusting the threshold in Step 4.4 are as follows:
[0135] Step 1: Based on experience or literature, set an initial threshold (for example, a Tanimoto coefficient > 0.7 is considered that compounds are similar), and this threshold refers to the preset threshold;
[0136] Step 2: Calculate the similarity between compound pairs using molecular fingerprints (such as MACCS, ECFP) or other descriptors (the present invention uses the Tanimoto coefficient);
[0137] Step 3: Use cross-validation to test the accuracy (Precision) and recall rate (Recall) under different thresholds, and select the optimal threshold;
[0138] Step 4: During the comparison process of the NPS intelligent comparison module, dynamically adjust the threshold according to the judgment results feedback by the application terminal using the reward mechanism in the prior art to ensure the comparison quality;
[0139] Step 5: Use a standard data set or manually labeled results to verify whether the adjusted threshold can improve the comparison accuracy;
[0140] Step 6: Determine the best threshold according to the verification results of Step 5.
[0141] The specific process of dynamically adjusting the threshold using the reward mechanism is as follows:
[0142] First, collect all the judgment results (approved or not approved) of users within a day from the application terminal and record the confidence levels corresponding to the judgment results;
[0143] Secondly, analyze the judgment results according to the confidence levels, and count the approval rate and error rate of the judgment results;
[0144] Finally, judge whether there are too many false positives according to the approval rate. If so, increase the threshold; if not, decrease the threshold; judge whether there are too many missed reports according to the error rate. If so, decrease the threshold; if not, increase the threshold;
[0145] Among them, the specific threshold adjustment formula is:
[0146] New threshold = Current threshold Δ;
[0147] Among them, Δ represents the threshold increment
[0148] Example:
[0149] The present invention proposes a new psychoactive substance identification and rapid comparison method and system based on big data. In this example, the rapid comparison of the controlled narcotic and psychotropic drugs medetomidine and isopropylparaben with etomidate is used to verify the performance of the present invention.
[0150] The standard EI mass spectrum of the controlled narcotic drug etomidate and the characteristic fragments 244, 105, 79, and 77 correspond to specific chemical structure fragments (such as Figure 2 ) are preset in the system.
[0151] When judging an unknown new imidate substitute (such as medetomidine and isopamidate controlled on July 1, 2024), you first need to collect the EI mass spectrum of the sample. This system will automatically output the possible structural fragments ( Figure 3 , Figure 4 ) and prompts the closest substance structure, i.e. the preset spectrum of etomidate ( Figure 2 ). Researchers can quickly infer and troubleshoot the structure based on the system-assisted output results combined with the original EI mass spectrum of the sample. Although the EI mass spectra of both medetomide and ipropagate can both reveal structural fragments such as 198, 105, 77, and 79, the molecular ion peaks at the high-mass end are quite different: medetomide is 230, etomidate is 244, and ipropagate is 258, with a fixed difference between the two. , that is (CH2)n, where n is a positive integer. The system will prompt that it may be a homologue. Researchers can easily infer the preliminary structure of the suspicious substance based on the prompted homologue information. At this time, other analytical methods need to be combined to further verify and confirm the results.
[0152] After confirming the structure of the suspicious substance, the system can convert the chemical structure into a SMILES structural expression that the computer can quickly identify and compare similarities through the NPS molecular structure analysis module. The system then uses this system to compare the similarity of the suspicious substances (SMILES structures of medetomidine and ipropafenac) with the previously controlled drug library in SMILES format, and quantitatively evaluates the similarity of new abuse substances from the perspective of chemical structure through different molecular fingerprints, thereby assessing the risk of abuse.
[0153] Taking the most commonly used MACCS molecular fingerprint as an example, the similarities between medetomidine and ipropamate and etomidate are 86.11% and 81.58%, respectively ( Figure 5 , Figure 6 ), the system can simultaneously output two-dimensional and three-dimensional structure comparison diagrams ( Figure 7 , Figure 8 ) Assist researchers to further confirm. The NPS result feedback and early warning module compares the current similarity recognition result with the preset threshold. In this example, the threshold is determined based on the lowest value of the similarity comparison between the 46 newly listed NPS on July 1, 2024 and the controlled drug catalogs of previous years, which is 72% ( Figure 9), The similarity between methomidate and isopropylparaben and etomidate is relatively high, both higher than the preset threshold, indicating the risk of substitution and abuse, which requires further attention. The NPS result feedback and early warning module feeds back the recognition result to the application terminal. Researchers determine again whether the substance should be listed for control application through the prompt information on the application terminal, and return the judgment result to the NPS result feedback and early warning module. In this module, the original preset threshold is adjusted (increased or decreased) based on the reward mechanism to achieve dynamic threshold adjustment.
[0154] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that they can still modify or equivalently replace the technical solutions of the present invention, and these modifications or equivalent replacements cannot make the modified technical solutions deviate from the spirit and scope of the technical solutions of the present invention.
Claims
1. A new psychoactive substance identification and rapid comparison system based on big data, characterized in that: It includes automatic data collection module, HBase-based NPS data storage module, NPS molecular structure analysis module, NPS intelligent comparison module, NPS result feedback and early warning module and application terminal; The automatic data collection module, HBase-based NPS data storage module, NPS molecular structure analysis module, NPS intelligent comparison module, NPS result feedback and early warning module are all deployed on a distributed server with a microservice architecture, and data interaction is achieved between the modules through an API interface; in: Automatic data collection module, used to collect the result data detected by laboratory instruments and equipment and the updated data in the designated network control list in real time; The HBase-based NPS data storage module is used to receive data sent by the automatic data collection module, clean the data and convert it into a key-value pair format containing a molecular structure, and store it in the Hbase database; The NPS molecular structure analysis module is used to convert and analyze the data in the key-value pair format containing the molecular structure, obtain the molecular fingerprint of the new psychoactive substance, and return it to the NPS data storage module based on HBase; The NPS intelligent comparison module is used to compare the molecular fingerprint of the substance to be tested with the listed molecular fingerprint data stored in the MySQL database, and send the comparison results to the NPS result feedback and early warning module; The NPS result feedback and warning module is used to generate warning information for comparison results that exceed the threshold through threshold comparison, and send it to the application terminal for user judgment. At the same time, it receives the judgment result of the application terminal on the warning information and dynamically adjusts the threshold according to the judgment result; The application terminal is used to provide a graphical interface for users to initiate comparison requests to the NPS intelligent comparison module and to display and feedback the user's judgment results; The automatic data collection module comprises: Instrument and equipment data acquisition unit, used to automatically collect the test result data newly generated by laboratory instruments and equipment; The network data collection unit is used to automatically obtain updated NPS data from the control list published on the Internet through a network crawler. The collected NPS data flows unidirectionally into the distributed server through the border platform; During the storage process, the HBase-based NPS data storage module first parses the data in the key-value pair format containing the molecular structure through the NPS molecular structure parsing module, and determines whether the parsed data is in the listed directory; If yes, the parsed data is assigned a flag "1" to indicate "confirmed as a new psychoactive substance" and is transferred to the MySQL database; If not, the parsed data is sent to the NPS intelligent comparison module, and the feedback results of the NPS intelligent comparison module and the NPS result feedback and early warning module are used to determine whether the parsed data should be subject to listing application; if so, the parsed data is assigned a mark of "1" and is transferred to the MySQL database; if not, the parsed data is assigned a mark of "0", indicating "not confirmed as a new psychoactive substance", and is stored in the Hbase database; The NPS result feedback and early warning module includes: A comparison result receiving unit receives the comparison result sent by the NPS intelligent comparison module; A threshold comparison unit compares the value of the comparison result with a preset threshold, generates warning information for comparison results that are greater than or equal to the threshold, and sends the warning information to the application terminal; The threshold self-adjusting unit receives the judgment result of the application terminal on the warning information, and learns according to the judgment result to automatically adjust the preset threshold; The specific process of dynamically adjusting the threshold according to the judgment results is as follows: first, collect all the judgment results of users within one day from the application terminal and record the confidence of the corresponding judgment results; second, analyze the judgment results according to the confidence, and count the recognition rate and error rate of the judgment results; finally, judge whether there are too many false alarms according to the recognition rate, if so, increase the threshold; if not, reduce the threshold; judge whether there are too many missed alarms according to the error rate, if so, reduce the threshold; if not, increase the threshold.
2. The new psychoactive substance identification and rapid comparison system based on big data as claimed in claim 1, characterized in that: The NPS molecular structure analysis module includes: A molecular structure conversion unit is used to convert the key-value pair data containing the molecular structure into a SMILES structure, then extract the structural features of the molecule from the SMILES structure, and perform hash processing on the structural features to generate a bit vector, and use the bit vector as the molecular fingerprint; A new key-value pair generation unit is used to generate a new key-value pair according to the SMILES structure, and return the new key-value pair to the HBase-based NPS data storage module and store it in the HBase database; The key is the molecular fingerprint, and the value is a byte array of unique two-dimensional matrix characters that describe the structural characteristics of the molecule.
3. The new psychoactive substance identification and rapid comparison system based on big data as claimed in claim 2, characterized in that: The NPS intelligent comparison module includes: A request data format determination unit, used for receiving a comparison request sent by an application terminal, and determining whether the data format in the comparison request is a molecular fingerprint format; If yes, it is sent to the comparison unit; if no, the comparison request is forwarded to the NPS molecular structure analysis module, which converts the data in the comparison request into a molecular fingerprint and returns it to the NPS intelligent comparison module; The comparison unit, based on the Tanimoto coefficient, compares the received molecular fingerprint with the data stored in the MySQL database one by one, outputs a value between 0 and 1, and sends the value as the comparison result to the NPS result feedback and early warning module.
4. A method for identifying and quickly comparing new psychoactive substances based on big data, characterized in that: The system according to any one of claims 1 to 3 is implemented, comprising the following implementation steps: Step 1: Build a distributed server; Step 2: The user configures the basic parameters of each microservice; Step 3: Automatically obtain the test result data of laboratory instruments and equipment and the updated data in the designated network control list through the data automatic acquisition module, and then parse these data into molecular fingerprint format through the NPS molecular structure analysis module, and store the data through the HBase-based data storage module; Step 4: The user initiates a comparison request through the application terminal. The NPS intelligent comparison module receives the comparison request, performs a similarity comparison on the molecular fingerprints in the comparison request data, and sends the comparison result to the NPS result feedback and warning module; The NPS result feedback and warning module compares the comparison results with the preset threshold, and generates warning information for the comparison results that are greater than or equal to the threshold, and sends it to the application terminal; After receiving the warning information, the user makes a judgment on the application terminal and returns the judgment result to the NPS result feedback and warning module to readjust the threshold.
5. The method for identifying and quickly comparing new psychoactive substances based on big data as claimed in claim 4, characterized in that: The specific steps of step 3 include: Step 3.1: The network data collection unit monitors the designated network control list on the Internet in real time. If any data update is found, the latest data is crawled and uploaded to the NPS data storage module based on HBase; When the HBase-based NPS data storage module stores data, if the column family of the current data already exists, the current data is updated and the original data is stored as the old version; if the column family of the current data does not exist, a new row is added and the data is marked; then the new data containing the molecular structure is passed to the NPS molecular structure analysis module; Step 3.2: After receiving the data containing the molecular structure, the NPS molecular structure analysis module converts it into a SMILES structure, extracts the structural features of the molecule through the SMILES structure, performs hash processing on the structural features of the molecule to obtain the molecular fingerprint, and then returns the molecular fingerprint to the NPS data storage module based on HBase.
6. The method for identifying and quickly comparing new psychoactive substances based on big data as claimed in claim 4, characterized in that: The specific steps of step 4 include: Step 4.1: The user logs in to the system through the application terminal. If the current molecular fingerprint is a new substance, a comparison request is initiated; Step 4.2: The application terminal sends the comparison request to the NPS intelligent comparison module. After receiving the comparison request, the NPS intelligent comparison module determines the data format of the comparison request. If the format is a non-molecular fingerprint format, the comparison request is forwarded to the NPS molecular structure analysis module. The NPS molecular structure analysis module converts the comparison request into a molecular fingerprint format and returns it to the NPS intelligent comparison module. After receiving the molecular fingerprint, the NPS intelligent comparison module compares the structural features of the vector string molecules one by one with the bit vectors already existing in the MySQL database, and sends the comparison results to the NPS result feedback and warning module; Step 4.3: The NPS result feedback and warning module compares the received comparison result with the preset threshold, generates warning information for the result greater than or equal to the threshold, and sends the warning information to the application terminal; Step 4.4: The user views the warning information through the application terminal, and at the same time initiates a three-dimensional molecular formula viewing request to the HBase-based NPS data storage module according to the relevant substance ID in the warning information. The user again determines whether the new substance needs to be listed based on the three-dimensional molecular formula viewing result, and feeds back the judgment result to the NPS result feedback and warning module. The NPS result feedback and warning module learns according to the user's judgment result and automatically adjusts the preset threshold; Step 4.5: The NPS result feedback and early warning module generates a corresponding application report based on the application listing information determined by the user, and sends the relevant information of the new substance to the NPS data storage module based on HBase, and modifies the corresponding column attributes in the Hbase database according to the relevant substance ID, changing its identifier to "1".
7. The method for identifying and quickly comparing new psychoactive substances based on big data as claimed in claim 6, characterized in that: In step 4.4, the specific steps of automatically adjusting the threshold include: Step 4.4.1: Set the threshold; Step 4.4.2: Calculate the similarity between compound pairs using molecular fingerprints; Step 4.4.3: Use cross validation to test the accuracy and recall under different thresholds and select the optimal threshold; Step 4.4.4: During the comparison process of the NPS intelligent comparison module, the threshold is dynamically adjusted according to the judgment result fed back by the application terminal to ensure the comparison quality; Step 4.4.5: Verify whether the adjusted threshold can improve the comparison accuracy and obtain the verification result; Step 4.4.6: Finalize the adjusted threshold value based on the verification results.
Citation Information
Patent Citations
Law enforcement site drug rapid detection system and drug detector thereof
CN114544581A