Drug type identification method and system based on sensor array
By acquiring drug gas response data through a sensor array, extracting steady-state and dynamic features, and combining them with classification models to identify drug types, it solves the low-cost and rapid identification needs of grassroots law enforcement agencies and achieves efficient and accurate drug type identification.
Patent Information
- Application Number
- CN202510868646.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-09-23
AI Technical Summary
Existing technologies are unable to provide a low-cost, fast and effective solution for drug type identification, especially for grassroots law enforcement agencies, which find it difficult to accurately identify multiple drugs.
A sensor array is used to obtain the sensor response data time series. By extracting steady-state features and dynamic features, a classification model is used to identify the type of drug, including steady-state features as the ratio of the response data average value and dynamic features as the cumulative value of the response data. Combined with temperature and humidity compensation and optimized weight processing, the identification accuracy is improved.
It achieves low-cost, fast and accurate identification of multiple drugs. The cost is much lower than existing technologies, the calculation amount is small, the operation is simple, the identification speed is increased by more than 50%, and the accuracy rate reaches 96%.
Smart Images

Figure CN120687919A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of drug detection, and in particular to a sensor array-based drug type identification method and system. Background Art
[0002] In recent years, the types of drugs have increased, which has brought great challenges and difficulties to the drug control work of relevant departments. In particular, grassroots law enforcement departments are in urgent need of low-cost drug identification solutions that can quickly and effectively identify the types of drugs and precursor chemicals.
[0003] Related technologies offer methods for drug identification, including gas chromatography-mass spectrometry (GC-MS), liquid chromatography-mass spectrometry (LC-MS), infrared spectroscopy, and chemical colorimetry. These methods require expensive, specialized equipment and specialized operators, making them costly and time-consuming. Chemical colorimetry suffers from shortcomings such as low sensitivity, susceptibility to environmental interference, and an inability to accurately identify new drugs. Consequently, related technologies fail to provide a low-cost, rapid, and effective drug identification solution that meets the needs of grassroots law enforcement agencies. Summary of the Invention
[0004] The present application aims to at least solve the technical problems existing in the prior art and provide a method and system for drug type identification based on a sensor array.
[0005] In the first aspect, the present application provides a method for identifying drug types based on a sensor array, which is applied to a sensor array configured in a reaction chamber and reacting with the gas of the item to be tested, the method comprising: obtaining a response data time series of at least three sensors in the sensor array, the response data time series of the sensor being obtained based on an original response data time series of a preset reaction time output by the sensor after the gas of the item to be tested is passed into the reaction chamber; extracting a steady-state feature and one or more dynamic features corresponding to the sensor based on the response data time series of the sensor, wherein the steady-state feature is the ratio of the average value of the response data in the last data segment of the response data time series to the average value of the response data in the starting data segment, and the one or more dynamic features include at least the cumulative value of the response data in a predefined response area in the response data time series; and obtaining a drug type identification result of the item to be tested using a classification model based on the steady-state features and one or more dynamic features corresponding to at least three sensors.
[0006] In the second aspect, the present application provides a drug type identification system, comprising: a reaction gas chamber, provided with a gas inlet allowing the gas of the item to be tested to flow in, and internally configured with a sensor array that reacts with the gas of the item to be tested; a processor, configured to be electrically connected to at least three sensors in the acquisition sensor array, and obtain a drug type identification result of the item to be tested in accordance with a drug type identification method based on a sensor array as described in the first aspect of the present application.
[0007] Beneficial technical effects of this application:
[0008] 1. This application introduces a gas sample to be tested into a reaction chamber. A sensor array disposed within the reaction chamber senses the gas sample and outputs at least three raw response data time series. A response data time series is obtained based on the raw response data time series. Based on the response data time series, steady-state features and one or more dynamic features corresponding to the sensors are extracted. This application features very low hardware cost and ease of operation.
[0009] 2. The steady-state feature can avoid the impact on the accuracy of drug identification caused by the inconsistent response speed of the sensors in the sensor array to the volatile gas of drugs. The calculation is simple and fast, saving computing resources.
[0010] 3. The characteristic of volatile drug gases that distinguishes them from common gases is that the sensor responds slowly. Therefore, the cumulative value of the response data in the predefined response area is used as a dynamic feature. This facilitates extraction and also increases the differentiation from common gases. Extracting only the cumulative value of the response data in the predefined response area can reduce the amount of calculation and further improve the calculation speed.
[0011] 4. Based on the steady-state characteristics corresponding to at least three sensors and one or more dynamic characteristics, a classification model is used to obtain drug type identification results, which can realize the identification of multiple drug types.
[0012] Therefore, the cost of the technical solution of the present application is extremely low, far lower than gas chromatography-mass spectrometry (GC-MS), liquid chromatography-mass spectrometry (LC-MS), infrared spectroscopy, chemical colorimetry and other solutions, and it is easy to operate, requires little calculation, and can quickly and accurately identify a variety of drugs. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 1 is a flow chart of a method for drug identification based on a sensor array in a preferred embodiment of the present invention;
[0014] Figure 2 is a schematic diagram of a sensor array in an example of the present invention;
[0015] Figure 3 This is a schematic diagram of the classification model training process in an example;
[0016] Figure 4 is the confusion matrix diagram of the classification model;
[0017] Figure 5 This is a schematic diagram of the layout of a drug type identification system in a preferred embodiment of the present invention;
[0018] Figure 6 This is a system architecture diagram of a drug type identification system in an example. DETAILED DESCRIPTION
[0019] The following describes embodiments of the present invention in detail. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended only to explain the present invention and are not to be construed as limiting the present invention.
[0020] In the description of the present invention, it should be understood that the terms "longitudinal", "transverse", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", etc., indicating the orientation or position relationship, are based on the orientation or position relationship shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as limiting the present invention.
[0021] In the description of the present invention, unless otherwise specified and limited, it should be noted that the terms "installed", "connected" and "connected" should be understood in a broad sense. For example, it can be a mechanical connection or an electrical connection, or it can be the internal communication between two components. It can be a direct connection or an indirect connection through an intermediate medium. For ordinary technicians in this field, the specific meanings of the above terms can be understood according to the specific circumstances.
[0022] The execution subject of the drug type identification method based on the sensor array provided by the present invention includes but is not limited to at least one of the electronic devices that can be configured to execute the method provided by the embodiment of the present application, such as a server, a terminal, and a processor. In other words, the drug type identification method based on the sensor array 2 provided can be executed by software or hardware installed in a terminal device or a server device. The server includes but is not limited to: a single server, a server cluster, a cloud server or a cloud server cluster, etc. The server can be an independent server, or it can be a cloud server that provides cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content distribution networks (CDN), as well as basic cloud computing services such as big data and artificial intelligence platforms. The processor can be a processing chip system such as ARM, FPGA, single-chip microcomputer, CPU, etc.
[0023] The present invention provides a method for identifying drug types based on a sensor array, which is applied to a reaction chamber 1 equipped with a sensor array 2 that reacts with the gas of an object to be tested. Types of drugs include opium, heroin, methamphetamine (ice), morphine, marijuana, cocaine, as well as narcotic drugs and psychotropic drugs that can cause addiction. Most drugs exist in solid form, and some exist in liquid or gaseous form. In the field of drug identification, free volatile gas of drugs (volatile drugs) is usually collected as drug gas. In order to accelerate volatilization, the drugs can also be heated (especially for drugs that are difficult to evaporate freely) to obtain a higher concentration of drug gas. The drug gas is introduced into the reaction chamber 1 to react with the sensor array 2. Different types of drugs have different characteristics of the response data time series of the sensor array 2. The sensor array 2 includes a plurality of sensors 21 arranged in an array, and the sensor 21 is not limited to a semiconductor gas sensor or a PID photoionization sensor.
[0024] In a preferred embodiment, please refer to Figure 1 , the method comprising:
[0025] Step S1, obtaining the response data time series of at least three sensors 21 in the sensor array 2, wherein the response data time series of the sensor 21 is obtained based on the original response data time series of the preset reaction time output by the sensor 21 after the gas of the test object is passed into the reaction gas chamber 1.
[0026] In this embodiment, the preset reaction time and the flow rate of the gas of the test item into the reaction chamber can be set based on experience. The preset reaction time is generally set to 3 to 10 minutes, and the flow rate is generally set to 0.5 L / min to 2 L / min. The sensor array 2 has at least two sensors 21. The time series of raw response data output by at least three sensors 21 after the gas of the test item is passed into the reaction chamber 1 for the preset reaction time can be defined as a set of data. The time series of raw response data of sensor 21 includes multiple raw response data arranged in chronological order according to the collection time, or includes multiple raw response data arranged in chronological order according to the collection time, and the collection time corresponding to each raw response data. The gas of the test item is the gas freely volatilized from the test item or the gas generated by heating the gas.
[0027] For example, Figure 2 As shown, sensor array 2 has eight sensors 21. First, clean air is introduced into reaction chamber 1, where sensor array 2 resides, at a flow rate of 1 L / min for a reaction time of 5 minutes, allowing the baseline values of sensors 21 to reach a stable value. Then, gas from the object to be detected is introduced into reaction chamber 1, where sensor array 2 resides, at a flow rate of 1 L / min for a reaction time of 5 minutes. During this process, the data acquisition module collects the raw response data from the eight sensors 21 to generate a set of data. The data acquisition module's acquisition frequency can be set based on experience and hardware performance, and is not limited to 3 Hz.
[0028] In this embodiment, the raw response data time series is preprocessed to obtain the response data time series. Preferably, the preprocessing process includes performing a default check on the raw response data according to the data acquisition standard, padding any missing data bits with zeros, and checking the raw response data for outliers, resetting any data bits with outliers to zero. This ensures the integrity and consistency of the data format of the raw response data.
[0029] Step S2: extracting a steady-state feature and one or more dynamic features corresponding to the sensor 21 based on the response data time series of the sensor 21, wherein the steady-state feature is the ratio of the average value of the response data in the last data segment to the average value of the response data in the starting data segment in the response data time series, and the one or more dynamic features include at least the cumulative value of the response data in the predefined response area in the response data time series.
[0030] In this embodiment, each response data time series includes multiple data items. In the above example, it includes 1800 data items. The starting data segment is the continuous multiple data items at the beginning of the response data time series, and the ending data segment is the continuous multiple data items at the end of the response data time series. The number of data items in the starting data segment and the ending data segment can be equal or unequal. For example, the number of data items in the starting data segment and the ending data segment is 20.
[0031] The steady-state characteristic E corresponding to a certain sensor 21 is expressed as:
[0032]
[0033] Here, M2 represents the average response data within the last data segment of the response data time series of a particular sensor 21, and M1 represents the average response data within the first data segment of the response data time series of a particular sensor 21. The ratio calculation method for steady-state feature E is relatively simple and can quickly extract the steady-state features of the response curve corresponding to the response data time series. This saves a significant amount of computing resources and can increase the computational speed by dozens of times compared to other calculation methods.
[0034] In this embodiment, the response data time series of the sensor 21 can be represented as a two-bit array including the response data value and the acquisition sequence number, for example (i, R i ), i represents the serial number of each data, R i Representing the response data value of each data, the response data time series can generate a corresponding response curve. The predefined response area is used to represent the data set of the interval that meets the response condition in the response data time series, or the predefined response area is used to represent the curve domain that meets the response condition in the response curve. The response condition is set according to the actual situation. For example, after the sensor 21 responds to the toxic gas, the numerical value of the response data is shown as a value that continuously decreases and then tends to a stable value, the response condition is set to a slope less than 0; for example, after the sensor 21 responds to the toxic gas, the numerical value of the response data is shown as a value that continuously increases and then tends to a stable value, the response condition is set to a slope greater than 0.
[0035] In this embodiment, the first dynamic feature is defined as the cumulative response data value of a predefined response region within the response data time series. The cumulative response data value is the curve domain integral value, i.e., the area between the curve domain and the horizontal axis. The calculation method for the first dynamic feature is relatively simple, conserving computing resources. Compared to other feature extraction methods, it is more flexible in application, capable of integrating calculations with most feature values and adapting to classification model training. Utilizing only response data within the interval that meets the response criteria can effectively accelerate drug identification. In actual data verification, the speed of operation has increased by over 50%.
[0036] Step S3: obtaining a drug type identification result of the object to be tested using a classification model based on the steady-state features and one or more dynamic features corresponding to at least three sensors 21.
[0037] In this embodiment, the drug classification results output multiple confidence levels for the drug types. For example, five drug types are selected: heroin, methamphetamine, ecstasy, marijuana, and air. The classification model is a pre-trained machine learning model, such as an MLP multi-layer perceptron. To reduce training workload, the classification model can also be a support vector machine.
[0038] In a preferred embodiment, the one or more dynamic features further include an average value of the first-order differential value of the response data within a predefined response area in the response data time series.
[0039] In this embodiment, the average value of the first-order differential values of the response data within a predefined response region in the response data time series is defined as the second dynamic feature. The method for determining the predefined response region can be referred to in the previous embodiment and will not be repeated here. The average value of the first-order differential values of the response data within the predefined response region can also be considered as the average value of the slopes of each data point within the predefined response region. The predefined response region is extracted from the response data time series, and the slope of each data point within the predefined response region is calculated. The average value of the slopes of the data points within the predefined response region is calculated and used as the second dynamic feature.
[0040] In this embodiment, since drugs are different from general inorganic compounds, the slow response speed of the sensor array 2 in the response stage to drug gas makes it easier to extract the integral and first-order differential characteristics of its response curve (i.e., the first dynamic characteristic and the second dynamic characteristic). At the same time, the characteristic performance is more distinguishable from that of common gases. Different types of drugs also have obvious differences in the performance of steady-state characteristics and one or more dynamic characteristics, thereby improving the accuracy of drug type identification.
[0041] For example, when the sensor array 2 has 8 sensors 21, the 8 sensors 21 correspond to 8 response data time series. Based on each response data time series, three features, namely steady-state feature, first dynamic feature and second dynamic feature, are obtained. The sensor array 2 obtains a 24-dimensional feature data.
[0042] In a preferred embodiment, in step S1, obtaining the response data time series of at least three sensors 21 in the sensor array 2 includes:
[0043] Step S11 , obtaining a time series of original response data of a preset reaction time output by at least three sensors 21 after the gas of the object to be tested is introduced into the reaction gas chamber 1 .
[0044] Step S12 : using the temperature and humidity of the reaction chamber 1 to compensate the original response data time series to obtain the response data time series of the sensor 21 .
[0045] In this embodiment, the temperature and humidity of reaction chamber 1 can be detected by temperature and humidity sensor 3 disposed within reaction chamber 1. Temperature and humidity sensor 3 can be independent temperature and humidity sensors, or it can be an integrated temperature and humidity sensor. Preferably, to improve the accuracy of drug identification, temperature and humidity sensor 3 is disposed within sensor array 2.
[0046] In this embodiment, the temperature and humidity of the reaction chamber 1 are used to compensate the original response data time series, which can significantly reduce the impact of temperature and humidity on the accuracy of drug identification. Preferably, the compensation process is:
[0047] Polynomial fitting is performed based on a time series set of original response data at different temperatures and humidities. In the polynomial function f, temperature and humidity are used as independent variables, and the response data offset is used as the dependent variable. The polynomial is preferably, but not limited to, a linear equation of two variables. The polynomial fitting method is preferably, but not limited to, a least squares method or a neural network mapping method (temperature and humidity are input features, and the response data offset is an output feature).
[0048] Substitute the current temperature T and humidity H into the polynomial function f to obtain the current response data offset f(T,H). Fuse the current response data offset f(T,H) with each original response data R in the original response data time series to obtain the corrected response data R. c , R c =R+f(T,H). The response data offset f(T,H) can be positive or negative.
[0049] In a preferred embodiment, in step S3, a classification model is used to obtain a drug identification result of the test object based on the steady-state features and one or more dynamic features corresponding to at least three sensors 21, including:
[0050] In step S31, principal component analysis is performed on the steady-state features and one or more dynamic features corresponding to at least three sensors 21 to obtain multiple primary features. For example, principal component analysis is performed on the steady-state features, first dynamic features, and second dynamic features corresponding to eight sensors 21, totaling 24 features. K feature vectors are selected as primary features, where k is a positive integer.
[0051] In step S32, a plurality of main features are input into a classification model to obtain a drug type identification result of the object to be tested.
[0052] In this embodiment, the principal component analysis method can be used to select multiple key main features to improve the accuracy of drug type identification.
[0053] In a preferred embodiment, drugs usually exist in the form of salt compounds when in solid form, and some of them are relatively difficult to volatilize. When the temperature is suitable, they will volatilize organic gas molecules such as organic amines and benzene ring derivatives. These gas molecules are aggregated and unevenly distributed at room temperature. Therefore, in the process of the reaction between the gas and the sensor array 2, there will be a certain randomness that causes a certain sensor 21 in the sensor array 2 to respond first; because the drug gas is input into the reaction gas chamber 1 at a certain rate, and the input direction may be different each time. The above reasons result in the response data amplitude and starting response time of different sensors 21 in the sensor array 2 being different each time. The 24-dimensional features obtained when different batches of the same type of drug gas are input into the reaction gas chamber 1 are different. This difference may interfere with the classification model and reduce the accuracy of drug type identification. Therefore, to solve the above problem, preferably, in step S2, based on the steady-state features corresponding to at least three sensors 21 and one or more dynamic features, the classification model is used to obtain the drug type identification result of the test object, including:
[0054] Step S21, determining the optimization weight of sensor 21 based on the start response time of the response data time series of sensor 21 and / or the cumulative value of the response data of the predefined response area, that is, determining the optimization weight of the sensor 21 based on the start response time of the response data time series of each sensor 21 and / or the corresponding first dynamic feature.
[0055] Step S22 : using the optimized weight of the sensor 21 to weight the steady-state feature and one or more dynamic features corresponding to the sensor 21 , respectively, to obtain the optimized steady-state feature and one or more optimized dynamic features corresponding to the sensor 21 .
[0056] Specifically, for a certain sensor 21, the optimization weight of the sensor 21 is multiplied by the steady-state characteristic, the first dynamic characteristic, and the second dynamic characteristic corresponding to the sensor 21 to obtain the optimized steady-state characteristic, the first optimized dynamic characteristic, and the second optimized dynamic characteristic corresponding to the sensor 21.
[0057] Step S23 , obtaining a drug type identification result of the object to be tested using a classification model based on the optimized steady-state features and one or more optimized dynamic features corresponding to at least three sensors 21 .
[0058] In this embodiment, accurate features are amplified by optimizing weights, reducing the adverse effects of airflow disturbances, differences in drug gas input directions, differences in input rates, and impurity gases in drug gases on drug type identification results, which is beneficial to improving the accuracy of drug type identification and significantly improving the generalization ability of drug type identification methods.
[0059] In this embodiment, further preferably, in step S22, determining the optimization weight of the sensor 21 according to the start response time of the response data time series of the sensor 21 and / or the cumulative value of the response data of the predefined response area includes:
[0060] The optimization weight of the sensor 21 is set to be negatively correlated with the order value of the start response time of the response data time series of the sensor 21 in the start response time series;
[0061] and / or, setting the optimization weight of the sensor 21 to be negatively correlated with the order value of the response data cumulative value of the predefined response area of the response data time series of the sensor in the response cumulative value sequence;
[0062] The start response time sequence is obtained by sorting the start response times of the response data time series of at least three sensors 21 in chronological order, and the response cumulative value sequence is obtained by sorting the response data cumulative values of the predefined response areas of the response data time series of at least three sensors 21 in descending order.
[0063] In this embodiment, the negative correlation can be any negative correlation function, such as a linear negative correlation function, an inverse proportional function, an exponential decay function, etc. The optimization weight of sensor 21 is set using the above method. The earlier the start response time of sensor 21, the greater its optimization weight, indicating that the characteristic data of sensor 21 is more core and important. The later the start response time, the less important the characteristic data of sensor 21. The larger the first dynamic characteristic of sensor 21, the greater its optimization weight, indicating that sensor 21 reacts with more toxic gas and its output characteristic has more reference value. Conversely, the lower the reference value of its output characteristic.
[0064] In a preferred embodiment of step S21, determining the optimization weight of the sensor 21 specifically includes:
[0065] Step A1: sort the start response times of the response data time series of at least three sensors 21 in order to obtain a start response time series.
[0066] Obtain the acquisition time of the first data point that meets the response condition in the response data time series of sensor 21, and then obtain the start response time of the response data time series. Sorting the start response times of the M sensors 21 in order to obtain the start response time series {ts1,…ts m ,…ts M}, M represents a positive integer greater than or equal to 1, ts1,…ts m ,…ts M Indicates the start response time of M response data time series. m represents the start response time ts m The sequence value is also a positive integer, and 1≤m≤M.
[0067] Step A2: Calculate the ratio of the adjustment coefficient K to the sequence value of the start response time of the response data time series of sensor 21 in the start response time series, and use this ratio as the optimization weight of sensor 21. The adjustment coefficient is greater than 0. Preferably, the value range of K can be set to 50 to 150 based on experience, preferably 100, so that the optimized feature value will not lose useful detailed information as much as possible. The sequence value of the start response time in the start response time series is a positive integer. For the sth sensor 21, let the start response time ts of its response data time series be m The order value in the start response time sequence is m, then the optimization weight ρ of the sth sensor 21 is s Expressed as:
[0068] In this embodiment, the earlier the start response time of sensor 21 is, the greater its optimization weight is, indicating that the feature data of sensor 21 is more core and important; the later the start response time is, the less important the feature data of sensor 21 is.
[0069] In another preferred embodiment of step S21, determining the optimization weight of the sensor 21 includes:
[0070] Step B1 : sorting the response data cumulative values of the predefined response areas of the response data time series of at least three sensors 21 from large to small to obtain a response cumulative value sequence.
[0071] The response data cumulative value of the predefined response area of the acquired response data time series is the first dynamic feature corresponding to the sensor 21. The first dynamic features of the M sensors 21 are sorted from large to small to obtain a response cumulative value sequence {d1,…d j ,…d M}, M represents a positive integer greater than or equal to 1, d1,…d j ,…d Mrepresents the cumulative value of the response data of the predefined response area of the M response data time series. j represents the cumulative value of the response data d j The order value in the response cumulative value sequence is also a positive integer, and 1≤j≤M.
[0072] Step B2: Calculate the ratio of the adjustment coefficient K to the order value of the response data accumulation value of the sensor 21 in the response accumulation value sequence, and use the ratio as the optimization weight of the sensor 21, wherein the adjustment coefficient is greater than 0 and the order value of the response data accumulation value in the response accumulation value sequence is a positive integer. For the sth sensor 21, let its corresponding first dynamic feature d j When the order value of the response cumulative value sequence is j, the optimization weight ρ of the sth sensor 21 is s Expressed as:
[0073] In this embodiment, the larger the first dynamic characteristic of sensor 21 is, the greater the optimization weight it obtains, indicating that the sensor 21 reacts to a larger amount of toxic gas and the characteristics it outputs are more valuable for reference. Conversely, the reference value of the characteristics it outputs is lower.
[0074] In another preferred embodiment of step S21, determining the optimization weight of the sensor 21 includes:
[0075] Step C1, sorting the start response times of the response data time series of at least three sensors 21 in order to obtain a start response time series;
[0076] The response data cumulative values of the predefined response areas of the response data time series of the at least three sensors 21 are sorted from large to small to obtain a response cumulative value sequence.
[0077] Obtain the acquisition time of the first data point that meets the response condition in the response data time series of sensor 21, and then obtain the start response time of the response data time series. Sorting the start response times of the M sensors 21 in order to obtain the start response time series {ts1,…ts m ,…ts M}, M represents a positive integer greater than or equal to 1, ts1,…ts m ,…ts M Indicates the start response time of M response data time series. m represents the start response time ts m The sequence value is also a positive integer, and 1≤m≤M.
[0078] The response data cumulative value of the predefined response area of the acquired response data time series is the first dynamic feature corresponding to the sensor 21. The first dynamic features of the M sensors 21 are sorted from large to small to obtain a response cumulative value sequence {d1,…dj ,…d M}, where M represents a positive integer greater than or equal to 1, and d1, … d j ,…d M represent the cumulative response values of the predefined response regions of the M response data time series. j represents the sequential value of the cumulative response value d j in the sequence of cumulative response values, which is also a positive integer, and 1 ≤ j ≤ M.
[0079] Step C2, calculate the product of the sequential value of the start response time of the response data time series of sensor 21 in the start response time series and the sequential value of the cumulative response value of sensor 21 in the cumulative response value series to obtain the sequential value product.
[0080] For the s-th sensor 21, assume that the sequential value of the start response time of its response data time series in the start response time series is m, and assume that the sequential value of its cumulative response value in the cumulative response value series is j. Then the sequential value product of the s-th sensor 21 is m * j.
[0081] Step C3, calculate the ratio of the adjustment coefficient K to the sequential value product, and use this ratio as the optimization weight of sensor 21, where the adjustment coefficient is greater than 0, the sequential value of the start response time in the start response time series is a positive integer, and the sequential value of the cumulative response value in the cumulative response value series is a positive integer.
[0082] For the optimization weight ρ of the s-th sensor 21 s is expressed as:
[0083] In the three preferred embodiments of the above step S21, feature adaptive optimization is achieved.
[0084] Further preferably, the adjustment system K is greater than 1 times the square of the number of sensors 21 in sensor array 2 and less than 2 times the square of the number of sensors 21 in sensor array 2, so that the optimization weight of sensor 21 is always a coefficient greater than 1 during the calculation process. The value of K can be adaptively set according to the number of sensors 21 in sensor array 2, which can ensure that the optimized feature value does not lose useful detail information. Exemplarily, when there are 8 sensors 21 in sensor array 2, 64 < K < 128, and preferably, K = 100.
[0085] In a preferred implementation manner, when the feature value of sensor array 2 is optimized by the optimization weight, in step S3, based on the optimized steady-state features corresponding to at least three sensors 21 and more than one optimized dynamic feature, the drug type identification result of the待测物品 (to-be-detected item) is obtained by using the classification model, including:
[0086] Performing principal component analysis on the optimized steady-state features and one or more optimized dynamic features corresponding to at least three sensors 21 to obtain a plurality of main features;
[0087] Multiple main features are input into the classification model to obtain the drug type identification results of the tested items.
[0088] In this embodiment, multiple main features are extracted from the optimized features, and the multiple main features are used to identify the types of drugs, which can improve the identification accuracy.
[0089] In order to prove that the classification model of this application has high classification accuracy, the specific training and verification process of the classification model is introduced below:
[0090] Step 1: Obtain the original sample data set and build a classification model.
[0091] Assume that the sensor array 2 is provided with Figure 2 The system shows eight sensors 21 of different models and one temperature and humidity sensor 3. The raw response data time series and temperature and humidity values from the multiple sensors 21 obtained after the introduction of drug gases are used as raw sample data. The preset drug gas types include air, heroin, methamphetamine, ecstasy, and marijuana. Ten sets of response data time series were collected for each drug gas type, totaling 50 raw sample data points. A support vector machine classification model was constructed.
[0092] Step 2: Preprocess the original sample data set to obtain a sample data set, and mark the true category label for each sample data.
[0093] The 50 data sets in the original sample dataset are preprocessed to obtain 50 time series of response data, thus forming the sample dataset. The class label setting rules are: air: "0"; heroin: "1"; methamphetamine: "2"; ecstasy: "3"; marijuana: "4". A true class label is set for each data point in the sample dataset.
[0094] Step 3: Extract the characteristic values of each sensor in each sample data, namely the steady-state feature, the first dynamic feature, and the second dynamic feature, to obtain the 24-dimensional feature corresponding to each sample data. This 24-dimensional feature is used as a model sample, and the true classification label of the sample data is assigned to the model sample. In this way, a model sample set consisting of 50 24-dimensional model samples is obtained. Then, principal component analysis is performed on the model sample set:
[0095] First, we need to calculate the average value of each characteristic value, such as the steady-state characteristic or the first dynamic characteristic or the second dynamic characteristic. The calculation formula is:
[0096]
[0097] Among them, v is the average value of each eigenvalue, n is the number of samples in the model sample set, here the value of n is 50, x i is the eigenvalue of the i-th model sample. Then use each eigenvalue to subtract the eigenvalue mean v to get a new 50*24 dimensional matrix. This data matrix is defined as matrix X * , calculate the matrix X through matrix operations * The covariance matrix C of is calculated as follows:
[0098]
[0099] After obtaining the covariance matrix, it is necessary to perform eigendecomposition on the covariance matrix and construct the characteristic equation:
[0100] det(C-λI)=0
[0101] Among them, λ is the eigenvalue, I is the 24*24 unit matrix, and 24 eigenvalues λ can be obtained by solving the characteristic equation i (i∈[1,24]). For each eigenvalue λ i , construct the linear equation system u i (C-λ i I)=0, solve the equations to get each λ i The corresponding eigenvector u i .
[0102] Then use the sorting algorithm to sort the eigenvalues to get 24 eigenvalues arranged from large to small, and get their corresponding eigenvectors u i sequence.
[0103] Then calculate the cumulative variance contribution rate. The calculation formula for the cumulative variance contribution rate is:
[0104]
[0105] When the cumulative variance contribution rate is greater than 0.8, select the k value at this time. Take the first k eigenvectors u i Analyze as the main features. Then calculate the data matrix X * The product with the eigenvector U obtains a feature matrix with reduced data dimension.
[0106] Step 4: Use the feature matrix after reducing the data dimension to train the support vector machine to obtain the drug type identification results of each model sample.
[0107] Step 5: Verify the classification accuracy of the classification model.
[0108] To verify the classification model, the response data of the gas to be detected must be collected in the same manner as above, and then the classification model must be called by a computer program for identification.
[0109] After verification, the accuracy of the verification results needs to be statistically analyzed. When the accuracy reaches above 95%, the model training is considered complete. If the accuracy is less than 95%, the training parameters need to be readjusted for training, that is, return to step 1 to step 4.
[0110] Step 6: Performance and advantages of the method and system.
[0111] Figure 4 The confusion matrix diagram of the classification model is shown. During the verification, five sets of data were collected for each category of gas under different environments for testing, and the accuracy rate of drug identification has reached 96%.
[0112] It has been verified that the present invention has the following advantages:
[0113] In terms of the speed of identification in actual applications, the total time from single gas data collection to giving identification results does not exceed 15 minutes, which is more than ten times shorter than the current conventional drug identification technology.
[0114] In terms of ease of operation, the method and system of the present invention can be operated after simple training of users after model training is completed, and the required learning cost is almost negligible.
[0115] In terms of cost and consumption, the method and system of the present invention can be replicated in batches after the model training is completed, and the cost of the entire equipment is much lower than technical methods such as gas chromatography-mass spectrometry (GC-MS), liquid chromatography-mass spectrometry (LC-MS), infrared spectroscopy, and chemical colorimetry.
[0116] The present invention also provides a drug type identification system. Figure 5 The system architecture diagram of the system is shown. Figure 6 A schematic diagram showing the layout of the system includes:
[0117] Reaction chamber 1 is equipped with a gas inlet for allowing the flow of gas from the test article, and houses a sensor array 2 that reacts with the gas from the test article. Sensor array 2 comprises a base 22 and multiple sensors 21 mounted on base 22. Preferably, a temperature and humidity sensor 3 is also mounted between the multiple sensors 21. Sensor array 2 is not limited to being located at the bottom of reaction chamber 1; the gas inlet can be located above or to either side of sensor array 2. To improve measurement accuracy, the sensors 21 in sensor array 2 can be of different models and manufacturers.
[0118] The processor is configured to be electrically connected to at least three sensors 21 in the acquisition sensor array 2, and obtain the drug type identification result of the object to be tested according to the above-mentioned drug type identification method based on the sensor array 2.
[0119] In this embodiment, referring to Figure 5 The processor includes a data acquisition module and a classification model analysis module. The data acquisition module is used to collect raw response data time series from at least three sensors 21. The algorithm analysis module is connected to the data acquisition module and obtains a drug type identification result of the test object according to the sensor array-based drug type identification method provided by the present invention.
[0120] Throughout this specification, reference to terms such as "one embodiment," "some embodiments," "example," "specific example," "one implementation," "a preferred implementation," or "some examples" means that the specific features, structures, materials, or characteristics described in conjunction with that embodiment or example are included in at least one embodiment or example of the present invention. In this specification, schematic representations of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.
[0121] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to the embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the claims and their equivalents.
Claims
1. A method for drug identification based on a sensor array, which is applied to a reaction chamber equipped with a sensor array that reacts with the gas of the object to be tested, characterized in that: The method comprises: Acquiring a time series of response data from at least three sensors in the sensor array, the time series of the sensor response data being obtained based on a time series of original response data output by the sensors with a preset reaction time after the gas of the article to be tested is passed into the reaction gas chamber; Extracting a steady-state feature and one or more dynamic features corresponding to the sensor based on a time series of response data of the sensor, wherein the steady-state feature is a ratio of an average value of response data in a terminal data segment to an average value of response data in a starting data segment in the time series of the response data, and the one or more dynamic features include at least a cumulative value of response data in a predefined response region in the time series of the response data; A classification model is used based on steady-state features and one or more dynamic features corresponding to at least three sensors to obtain drug type identification results of the object to be tested.
2. The method for drug identification based on a sensor array according to claim 1, wherein: The one or more dynamic features further include an average value of first-order differential values of the response data within a predefined response region in the response data time series.
3. The method for drug identification based on a sensor array according to claim 1, wherein: The step of obtaining a time series of response data of at least three sensors in the sensor array includes: Obtaining a time series of original response data of a preset reaction time output by at least three sensors after the gas of the test object is passed into the reaction gas chamber; The temperature and humidity of the reaction chamber are used to compensate the original response data time series to obtain the sensor response data time series.
4. The method for drug identification based on a sensor array according to claim 1, wherein: The method of obtaining a drug type identification result of the test object using a classification model based on the steady-state features and one or more dynamic features corresponding to at least three sensors includes: Performing principal component analysis on steady-state features and one or more dynamic features corresponding to at least three sensors to obtain multiple main features; Multiple main features are input into the classification model to obtain the drug type identification results of the tested items.
5. A drug type identification method based on a sensor array as claimed in claim 1, 2 or 3, characterized in that: The method of obtaining a drug type identification result of the test object using a classification model based on the steady-state features and one or more dynamic features corresponding to at least three sensors includes: determining an optimization weight of the sensor according to a start response time of a response data time series of the sensor and / or a cumulative value of response data in a predefined response area; The optimized weight of the sensor is used to weight the steady-state characteristics and one or more dynamic characteristics corresponding to the sensor to obtain the optimized steady-state characteristics and one or more optimized dynamic characteristics corresponding to the sensor; Based on the optimized steady-state features and one or more optimized dynamic features corresponding to at least three sensors, a classification model is used to obtain the drug type identification result of the object to be tested.
6. The method for drug identification based on a sensor array according to claim 5, wherein: The step of determining the optimization weight of the sensor according to the start response time of the sensor response data time series and / or the cumulative value of the response data in the predefined response area includes: The optimization weight of the sensor is set to be negatively correlated with the order value of the start response time of the response data time series of the sensor in the start response time series; and / or, setting the optimization weight of the sensor to be negatively correlated with the order value of the response data cumulative value of a predefined response area of the response data time series of the sensor in the response cumulative value sequence; The start response time sequence is obtained by sorting the start response times of the response data time series of at least three sensors in chronological order, and the response cumulative value sequence is obtained by sorting the response data cumulative values of the predefined response areas of the response data time series of at least three sensors in descending order.
7. The method for drug identification based on a sensor array according to claim 6, wherein: Determine the optimization weights of sensors, including: Calculate the ratio of the adjustment coefficient to the sequence value of the start response time of the response data time series of the sensor in the start response time series, and use the ratio as the optimization weight of the sensor, wherein the adjustment coefficient is greater than 0, and the sequence value of the start response time in the start response time series is a positive integer.
8. The method for drug identification based on a sensor array according to claim 6, wherein: Determine the optimization weights of sensors, including: The ratio of the adjustment coefficient to the order value of the response data cumulative value of the sensor in the response cumulative value sequence is calculated, and the ratio is used as the optimization weight of the sensor, wherein the adjustment coefficient is greater than 0, and the order value of the response data cumulative value in the response cumulative value sequence is a positive integer.
9. The method for drug identification based on a sensor array according to claim 6, wherein: Determine the optimization weights of sensors, including: Calculating the product of the sequence value of the start response time of the response data time series of the sensor in the start response time series and the sequence value of the response data cumulative value of the sensor in the response cumulative value series to obtain a sequence value product value; Calculate the ratio of the adjustment coefficient to the product value of the sequence value, and use the ratio as the optimization weight of the sensor, wherein the adjustment coefficient is greater than 0, the sequence value of the start response time in the start response time sequence is a positive integer, and the sequence value of the response data cumulative value in the response cumulative value sequence is a positive integer.
10. A drug identification system, characterized in that: include: A reaction gas chamber is provided with a gas inlet for allowing the gas of the object to be tested to flow in, and a sensor array is disposed inside thereof to react with the gas of the object to be tested; The processor is configured to be electrically connected to at least three sensors in the acquisition sensor array and obtain a drug type identification result of the object to be tested according to a drug type identification method based on a sensor array according to any one of claims 1-9.