Drug Identification Method, Device, Equipment and Medium Based on Spectral Correlation Dimension Ascension

Through the method based on spectral correlation dimensionality upgrade, the problem of traditional drug detection technology being difficult to portable and fast detection is solved, and high-accurate drug recognition is achieved, which improves the robustness of the recognition.

CN119810575BActive Publication Date: 2025-06-24NATIONAL INSTITUTE OF METROLOGY CHINA +3
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510289705.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-06-24
Estimated Expiration
2045-03-12

AI Technical Summary

Technical Problem

Traditional drug detection technology is difficult to meet the needs of portable and rapid detection, and the multi-spectral fusion technology cannot exceed 90% of the drug identification accuracy.

Method used

Using a method based on spectral correlation dimensionality-up, the near-infrared and mid-infrared spectral data of drug samples were obtained, pre-processing, feature spectral screening, correlation analysis and heat map conversion were performed, and the pre-trained recognition model was used for identification.

Benefits of technology

The accuracy and robustness of drug identification are improved, the model input characteristics are optimized, the negative impact of dopants is reduced, and the recognition accuracy is increased to 94.44%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119810575B_ABST
    Figure CN119810575B_ABST
Patent Text Reader

Abstract

The present invention provides a drug recognition method, device, equipment and medium based on spectral correlation dimension elevation, which relates to the field of drug recognition. The method includes: obtaining spectral data of a drug sample and preprocessing it using at least two preprocessing methods to obtain target spectral data, where the target spectral data includes near-infrared data and mid-infrared spectral data; using a spectral variable selection method to screen the target spectral data to obtain characteristic spectral data; constructing a spectral matrix based on the wavelengths of the characteristic spectral data to obtain a near-infrared matrix and a mid-infrared matrix, performing a correlation analysis on the near-infrared matrix and the mid-infrared matrix to obtain a correlation coefficient matrix, converting the correlation coefficient matrix into a heat map, and using a pre-trained recognition model to recognize the heat map to obtain a recognition result. By screening characteristic variables and performing a correlation analysis, the present invention elevates one-dimensional data to two-dimensional, highlights the spectral information of drugs, and the recognition accuracy of the established model reaches more than 94%.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of drug identification, and more particularly, to a drug identification method, device, equipment and medium based on spectral correlation dimensionality elevation. Background Art

[0002] Although traditional drug detection technologies have high accuracy, they are limited by expensive detection equipment, the reserve of professional technical personnel, and the large floor area of the instruments, etc., and cannot meet the portable requirements. They can often only be used for laboratory detection, and the operation is relatively cumbersome and time-consuming, making it difficult to meet the needs of rapid on-site detection in anti-drug combat. Mid-infrared spectroscopy (MIR) and near-infrared spectroscopy (NIR) technologies have shown great application potential in rapid on-site drug detection due to their characteristics such as fast detection speed, simple operation, non-invasiveness, and simultaneous analysis of multiple components. Especially, they have unique advantages in terms of non-destruction of samples and protection of the integrity of evidence. Multispectral fusion technology is a method that comprehensively utilizes spectral information of different types or the same type. By integrating and optimizing multiple spectral information, it makes up for problems such as insufficient characteristic information reflected by a single spectrum, and realizes the complementary advantages between multiple spectra. However, the existing multispectral fusion technologies cannot exceed 90% in the recognition accuracy of drugs, and the recognition success rate still needs to be improved. Summary of the Invention

[0003] The purpose of the present invention is to provide a drug identification method, device, equipment and medium based on spectral correlation dimensionality elevation to improve the above problems. To achieve the above purpose, the technical solutions adopted by the present invention are as follows:

[0004] In a first aspect, the present application provides a drug identification method based on spectral correlation dimensionality elevation, including:

[0005] Obtain the spectral data of a drug sample and preprocess it using at least two preprocessing methods to obtain target spectral data, where the target spectral data includes near-infrared data and mid-infrared spectral data;

[0006] Use a spectral variable selection method to screen the target spectral data to obtain characteristic spectral data;

[0007] Construct a spectral matrix based on the wavelengths of the characteristic spectral data to obtain a near-infrared matrix and a mid-infrared matrix, perform correlation analysis on the near-infrared matrix and the mid-infrared matrix to obtain a correlation coefficient matrix, convert the correlation coefficient matrix into a heat map, and use a pre-trained recognition model to recognize the heat map to obtain a recognition result.

[0008] In a second aspect, the present application also provides a drug identification device based on spectral correlation dimensionality elevation, including:

[0009] A preprocessing module for obtaining spectral data of a drug sample and preprocessing it using at least two preprocessing methods to obtain target spectral data, where the target spectral data includes near-infrared data and mid-infrared spectral data;

[0010] A screening module for screening the target spectral data using a spectral variable selection method to obtain characteristic spectral data;

[0011] An identification module for constructing a spectral matrix based on the wavelengths of the characteristic spectral data to obtain a near-infrared matrix and a mid-infrared matrix, performing a correlation analysis on the near-infrared matrix and the mid-infrared matrix to obtain a correlation coefficient matrix, converting the correlation coefficient matrix into a heat map, and using a pre-trained identification model to identify the heat map to obtain an identification result.

[0012] In a third aspect, the present application also provides a drug identification device based on spectral correlation dimensionality elevation, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the drug identification method based on spectral correlation dimensionality elevation as described above are implemented.

[0013] In a fourth aspect, the present application also provides a readable storage medium, characterized in that a computer program is stored on the readable storage medium, and when the computer program is executed by a processor, the steps of the drug identification method based on spectral correlation dimensionality elevation as described above are implemented.

[0014] The beneficial effects of the present invention are as follows:

[0015] By performing feature variable selection on spectral data, the present invention screens out the main feature information therein, combines the characteristic spectra under different spectral preprocessing conditions, finally performs dimensionality elevation on the data through correlation analysis, and makes a heat map, optimizing the model input features, reducing the negative impact brought by impurities in the actual samples, and effectively improving the accuracy and robustness of identification.

[0016] Other features and advantages of the present invention will be described in the subsequent specification, and, in part, will become apparent from the specification or be understood by implementing the embodiments of the present invention. Description of the Drawings

[0017] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for use in the embodiments. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0018] Figure 1Flow chart of the drug identification method based on spectral correlation dimension elevation in the embodiments of the present application;

[0019] Figure 2 NIR spectral data (A) and MIR spectral data (B) collected for methamphetamine samples in the embodiments of the present application;

[0020] Figure 3 NIR spectral data (A) and MIR spectral data (B) collected for heroin samples in the embodiments of the present application;

[0021] Figure 4 VIP screening results of NIR spectrum (A) and MIR spectrum (B) under unpretreated conditions in the embodiments of the present application;

[0022] Figure 5 MIV screening results of NIR spectrum (A) and MIR spectrum (B) under unpretreated conditions in the embodiments of the present application;

[0023] Figure 6 Structure diagram of the drug identification device based on spectral correlation dimension elevation in the embodiments of the present application;

[0024] Figure 7 Structure diagram of the drug identification device based on spectral correlation dimension elevation in the embodiments of the present application.

[0025] Markings in the figure: 100 - Pretreatment module; 200 - Screening module; 210 - First construction unit; 220 - First calculation unit; 230 - Selection unit; 310 - Third construction unit; 320 - Fourth calculation unit; 330 - Second identification unit; 300 - Identification module; 800 - Drug identification device based on spectral correlation dimension elevation; 801 - Processor; 802 - Memory; 803 - Multimedia component; 804 - I / O interface; 805 - Communication component. Detailed implementation manners

[0026] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Usually, the components of the embodiments of the present invention described and illustrated herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed present invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0027] It should be noted that similar reference numerals and letters refer to similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. Embodiment 1

[0028] Refer to Figure 1 , this application provides a drug recognition method based on spectral correlation dimension elevation, including steps S100, S200, and S300;

[0029] S100. Obtain the spectral data of the drug sample and preprocess it using at least two preprocessing methods to obtain target spectral data, where the target spectral data includes near-infrared data and mid-infrared spectral data;

[0030] In this embodiment, 24 groups of heroin samples and 30 groups of methamphetamine sample data are collected. The calibration set and the validation set are randomly divided in a ratio of 2:1. The calibration set samples are used for model construction, and the validation set samples are used to measure the external validation performance of the model. At the same time, near-infrared (NIR) data and mid-infrared (MIR) data of the above samples are collected, as Figure 2 and Figure 3 shown. The NIR image and the MIR image are curves formed by connecting multiple points. Each point corresponds to a spectral data, which includes wavelength data (abscissa) and the intensity corresponding to that wavelength (ordinate);

[0031] Since the infrared spectrum of the sample to be measured not only contains rich chemical structure information but also contains external redundant data, such as instrument noise, sample background, or stray light, etc. Without losing effective spectral data, eliminating part of the redundant data can improve the prediction performance of the model to a certain extent. Therefore, the spectral preprocessing method aims to eliminate this part of the redundant data. Commonly used spectral preprocessing methods include normalization, smoothing, derivative, standard normal variate transformation, multiplicative scatter correction, etc.

[0032] In this embodiment, a total of four preprocessing methods are used to preprocess the NIR data and the MIR data respectively, specifically including:

[0033] Vector normalization: The purpose of normalization is to unify the scale of the spectral data to prevent the large difference in spectral data within different wavenumbers from having a negative optimization on the model, including area normalization, maximum-minimum normalization, and vector normalization (VN), etc. VN is the most commonly used normalization method. The calculation formula for the spectral intensity x is as follows:

[0034] ;

[0035] ;

[0036] where \(m\) is the number of points (wavelength points) of the spectral curve, \(k = 1, 2,\cdots,m\). is the spectral intensity. is the average spectral intensity. is the normalized spectral intensity. is the \(k\)-th point.

[0037] Savitzkg - Golag Smoothing: Smoothing is one of the most commonly used methods to eliminate background noise and helps to eliminate random errors in spectral data. SG (Savitzkg - Golag) smoothing assumes that the background noise is zero uniform random white noise and achieves the smoothing effect by applying a set of sliding windows and polynomial fitting to the data. The spectral data at the center of the window has a larger weight, while the spectral data at the edge of the window has a smaller weight. The deconvolution - type smoothing by this method can better retain the original shape and characteristics of the spectral data and is applicable to spectral data with obvious peaks and valleys. The calculation formula is as follows:

[0038] First, for SG smoothing, the smoothing window width needs to be selected, such as \(2w + 1\). Then each window consists of an odd number of points. Taking a certain point in the spectrum as the center, the average value of the measured values of the \(w\) points before and after is calculated as the smoothed value. The specific formula is as follows:

[0039] ;

[0040] ;

[0041] where \(h\) i is the smoothing coefficient (which needs to query the SG smoothing coefficient table according to different window widths), \(k = 1, 2,\cdots,m\); \(H\) is the normalization factor; represents the smoothing of the \(k\)-th point and is also denoted by ; represents the value of the data point offset by \(i\) positions relative to .

[0042] Derivation: The first derivative (1st Derivative, 1st D) and the second derivative (2nd Derivative, 2nd D) of the spectrum are two commonly used preprocessing methods for baseline correction of the spectrum. 1st D usually calculates the slope or rate of change of spectral data using numerical differentiation methods to highlight the spikes or valleys in the spectral data and amplify the characteristic signals of the spectrum. 2nd D usually calculates the curvature or the rate of change of the rate of change of spectral data by differentiating the first derivative again to highlight the broad peaks and plateaus in the spectral data. The Savitzky-Golay smoothing can also be used for the derivative calculation of spectral data. The smoothing coefficient in SG smoothing can be replaced with a derivative coefficient, and the SG derivative coefficient table can be queried according to different window widths. SG derivatives are usually more stable and accurate than traditional numerical differentiation methods. While highlighting the spectral signal characteristics, they can also effectively eliminate the influence of baseline drift on data analysis and modeling.

[0043] Standard normal variate transformation: The standard normal variate transformation (SNV) is often used in spectral preprocessing to eliminate the scattering effects caused by factors such as differences in optical path and particle size. SNV makes all spectral data conform to a normal distribution with a mean of 0 and a standard deviation of 1 by subtracting the average value of all wavelength points from each spectral data point and then dividing by the standard deviation of the spectrum. The specific formula is as follows:

[0044] ;

[0045] where, is the spectral intensity after standard normal variate transformation; m is the number of points on the spectral curve, k = 1, 2, ……, m; is the spectral intensity, is the average spectral intensity, is the spectral intensity at the k-th point.

[0046] Based on the technical concept of this application, the optional spectral preprocessing methods are not only the above four, but can also be replaced with techniques such as multiplicative scatter correction.

[0047] S200. Use a spectral variable selection method to screen the target spectral data to obtain characteristic spectral data;

[0048] The infrared spectral data processed in step S100 usually contains a large amount of chemical structure information, but not all spectral data has a high correlation with the property to be measured. Therefore, the role of the spectral variable selection method is to screen out the characteristic spectra with a high correlation with the property to be measured from all spectral data, aiming to reduce the spectral dimension and improve the prediction performance of the model.

[0049] In this embodiment, two spectral variable selection methods are preferably used. One is the variable importance in projection (VIP) method, and the other is the mean influence value (MIV) method.

[0050] Variable importance projection method: Based on the target spectral data, a partial least squares regression model (PLS) is established, and the model parameters are recorded. The model parameters include the loadings, scores, and principal components of each target spectral data, and the model parameters are all represented as vectors.

[0051] Based on the model parameters, the variable importance projection method is used to calculate the VIP value of each target spectral data; when the VIP value is very large, it means that the spectral data has a strong explanatory ability for the dependent variable. Usually, the wavelengths with VIP greater than 1 are selected as the characteristic data.

[0052] Select the target spectral data with VIP value greater than 1 as the characteristic spectral data.

[0053] ;

[0054] Among them, h is the number of PLS principal components; w k is the weight vector of the k-th principal component, w jk is the weight vector of the j-th data in the k-th principal component; t k is the score vector of the k-th principal component, t k T is the transpose of t k ; q k is the loading vector of the k-th principal component; m is the number of wavelengths, that is, the number of data; j represents the j-th data point, j = 1, 2,..., m; is the sum of the squares of the weight vectors of the k-th principal component ; is the VIP value of the j-th data .

[0055] According to the VIP values of each spectral parameter, the spectral parameters are screened to obtain the characteristic spectral data.

[0056] Mean impact value: MIV is a new method for screening characteristic variables based on the importance degree of input variables to output variables. The calculation method is as follows: Multiple training samples are constructed according to the target spectral data. Each training sample includes multiple spectral parameters, and a neural network recognition model is constructed based on the training samples.

[0057] The training samples are represented by P, and there are a total of z training samples. Each training sample includes n spectral parameters, and one training sample is the data of the points on a spectral curve; the output parameter corresponding to the training sample is represented by A, that is, P is the spectral data and A is the predicted value (drug type).

[0058] ;

[0059] ;

[0060] Among them, X represents the spectral parameter, and Y represents the predicted value;

[0061] Perform the same numerical increase and decrease transformation on the i-th spectral parameter in each training sample to obtain multiple sets of prediction samples, where each set of prediction samples contains two transformed samples obtained by respectively increasing and decreasing the same training sample by a numerical value (increasing or decreasing by 10%); that is:

[0062] ;

[0063] is the set of transformed samples after the change of the i-th spectral parameter of z training samples;

[0064] Input the transformed samples in each set of prediction samples into the neural network recognition model respectively to obtain the first prediction result and the second prediction result; among them, the prediction result of increasing by 10% is denoted as A i (1), and the prediction result of decreasing by 10% is denoted as A i (2);

[0065] Perform a difference operation on the first prediction result and the second prediction result to obtain the influence value of the i-th spectral parameter on the model prediction result;

[0066] ;

[0067] I v, i is the influence value of the i-th spectral parameter on the model prediction result;

[0068] Calculate the influence value of the i-th spectral parameter on the model prediction result in all training samples, and average the influence values according to the number of training samples to obtain the average influence value of the i-th spectral parameter;

[0069] ;

[0070] Among them, is the influence value of the i-th spectral parameter of the j-th training sample, is the average influence value of the i-th spectral parameter; z is the number of training samples;

[0071] Thus, the average influence value of each spectral parameter can be obtained;

[0072] Screen the spectral parameters according to the average influence value of each spectral parameter to obtain the characteristic spectral data.

[0073] After calculating the spectral data using VIP or MIV, in addition to determining the characteristic spectral data by screening the VIP value and the MIV value by setting a threshold, the spectra can also be sorted from large to small according to the degree of spectral importance (i.e., VIP value and MIV value), and the spectral parameters with the highest ranking, such as the first 200 spectral parameters, can be selected as the characteristic spectral data.

[0074] It should be noted that the first 200 spectral parameters were screened out based on experience, and the number of screened parameters can be appropriately adjusted based on the actual model prediction effect.

[0075] It should be noted that the spectral variable selection method is not limited to VIP and MIV, but can also be other methods, such as competitive adaptive weighted sampling, Monte Carlo uninformative variable elimination algorithm, etc. When using different spectral variable screening methods to screen data, the accuracy of the final recognition model will be slightly different.

[0076] S300. Construct a spectral matrix according to the wavelength of the characteristic spectral data to obtain a near-infrared matrix and a mid-infrared matrix, perform correlation analysis on the near-infrared matrix and the mid-infrared matrix to obtain a correlation coefficient matrix, convert the correlation coefficient matrix into a heat map, and use a pre-trained recognition model to recognize the heat map to obtain a recognition result.

[0077] The wavelength of the characteristic spectrum is used as the column of the spectrum matrix, the preprocessing method is used as the row of the spectrum matrix, and the characteristic spectrum data corresponding to different wavelengths and different preprocessing methods are used as the elements of the spectrum matrix to construct a spectrum matrix, wherein the spectrum matrix includes a near-infrared matrix and a mid-infrared matrix;

[0078] For example, this application selects 200 characteristic spectral data (corresponding to 200 wavelength points) in both the near-infrared moment data and the mid-infrared data, and adopts four preprocessing methods, then the size of the near-infrared matrix and the mid-infrared matrix is ​​4×200;

[0079] Performing Pearson correlation coefficient analysis on each column of the near infrared matrix and the mid infrared matrix to obtain a correlation coefficient matrix;

[0080] For example, the near infrared matrix and the mid-infrared matrix each have 200 columns, and the size of the correlation coefficient matrix is ​​200×200;

[0081] The correlation coefficient matrix is ​​converted into a heat map, and the pre-trained Bayes-CNN model is used to identify the heat map to obtain the recognition result.

[0082] Select a suitable plotting library for heatmap plotting and color each cell in the matrix. The input of the pre-trained Bayes-CNN model is the heatmap training set. The heatmap training set is obtained by sequentially preprocessing, screening characteristic spectral data, constructing a correlation matrix, and performing transformation on spectral samples through the above method of this application; the output of the model is the drug type.

[0083] Based on the technical concept of this application, the correlation analysis technology that can be adopted is not only Pearson correlation analysis, but can also be replaced by Spearman correlation analysis, Kendall correlation analysis, and so on.

[0084] In this embodiment, multiple models are constructed according to different spectral variable screening methods and different correlation analysis methods, and the performance of different Bayes-CNN models is tested respectively. The results are shown in Table 1;

[0085] Table 1 Performance of correlation analysis heatmaps of different Bayes-CNN models

[0086]

[0087] In the table, TP is the true positive, FN is the false negative, TN is the true negative, FP is the false positive, and ACC is the accuracy rate.

[0088] It can be seen that the data processing method of feature screening through VIP and dimensionality increase through Pearson correlation analysis enables the model to have a higher recognition accuracy rate.

[0089] In addition, this application also establishes other spectral recognition models and tests the performance of the models for comparison with the model of the present invention, specifically as follows:

[0090] Single spectral recognition model:

[0091] The single spectral model uses NIR and MIR data to establish partial least squares discriminant (PLS-DA) models respectively. To eliminate the noise in the NIR and MIR spectral data and avoid large differences in the spectral absorption intensity between NIR and MIR, it is necessary to perform spectral preprocessing on NIR and MIR, such as SG first derivative (SG 1 st D), SG smoothing, and vector normalization (VN). After performing the above preprocessing methods on the spectra, a single spectral model is established using calibration set samples, and the model is used to predict and externally validate the validation set samples. The external validation prediction results are shown in Table 2. Under the condition of SG 1 st D, the single spectral MIR-PLS-DA model has the highest recognition accuracy rate for methamphetamine and heroin samples, which is 83.33%.

[0092] Table 2 Performance of single spectral PLS-DA models under different preprocessing methods

[0093]

[0094] Low - level and medium - level data fusion recognition models:

[0095] After vector splicing the pre - processed spectra, a low - level data fusion (LLDF) - PLS - DA model was established, and the performance results are shown in Table 3. Further, after VIP spectral variable screening of the NIR and MIR spectral data of drug samples. Taking the un - pre - processed condition as an example, the results of VIP screening of NIR and MIR spectral data are as Figure 4 shown. The spectral variables with VIP values greater than 1 were used as characteristic variables, and a VIP - medium - level data fusion (MLDF) - PLS - DA model was established by vector splicing. Similarly, the results of MIV screening of NIR and MIR spectral data are as Figure 5 shown. Taking the average value of MIV as the screening threshold, the spectral variables greater than the average value of MIV were selected as characteristic variables to establish an MIV - MLDF - PLS - DA model. The results show that under the SG 1 st D pre - processing condition, the VIP - MLDF - PLS - DA model established through VIP screening has the highest recognition accuracy for methamphetamine and heroin, reaching 88.89%, and can accurately identify heroin samples. The reason why the MIV - MLDF - PLS - DA model has slightly worse performance may be that the spectral variables screened by the MIV algorithm have nothing to do with the characteristic information of methamphetamine and heroin, resulting in poor prediction performance of the model. At the same time, it should be noted that whether using the VIP or MIV algorithm, or even other algorithms such as CARS (competitive adaptive reweighted sampling), the recognition success rate of the model depends on the correlation between the characteristic spectrum and the property to be measured. If the property to be measured changes, the most suitable spectral variable selection algorithm may be different, which requires detailed model comparison to screen out the spectral variable selection algorithms suitable for different properties to be measured.

[0096] Table 3 Performance of LLDF - PLS - DA and VIP - MLDF - PLS - DA models under different pre - processing methods

[0097]

[0098] It can be seen that the recognition accuracy of methamphetamine and heroin by the method provided in this application is 94.44%. Under various pre - processing conditions, the recognition accuracies of other recognition models cannot reach 90%. Through comparison, it can be known that the improvement range of the recognition success rate of the drug recognition method based on spectral correlation dimension elevation in this application is between 5.55% - 27.77%. Example 2

[0099] SeeFigure 6 , this application also provides a drug recognition device based on spectral correlation dimensionality increase, including:

[0100] A preprocessing module 100, configured to obtain spectral data of a drug sample and preprocess it using at least two preprocessing methods to obtain target spectral data, where the target spectral data includes near-infrared data and mid-infrared spectral data;

[0101] A screening module 200, configured to screen the target spectral data using a spectral variable selection method to obtain characteristic spectral data;

[0102] An identification module 300, configured to construct a spectral matrix based on the wavelengths of the characteristic spectral data to obtain a near-infrared matrix and a mid-infrared matrix, perform correlation analysis on the near-infrared matrix and the mid-infrared matrix to obtain a correlation coefficient matrix, convert the correlation coefficient matrix into a heat map, and use a pre-trained identification model to identify the heat map to obtain an identification result.

[0103] As an optional implementation manner, the screening module includes:

[0104] A first construction unit 210, configured to establish a partial least squares regression model based on the target spectral data and record model parameters, where the model parameters include the load, score, and principal component of each target spectral data;

[0105] A first calculation unit 220, configured to calculate the VIP value of each target spectral data using the variable importance projection method based on the model parameters;

[0106] A selection unit 230, configured to select the target spectral data with a VIP value greater than 1 as the characteristic spectral data.

[0107] As another optional implementation manner, the screening module includes:

[0108] A second construction unit, configured to construct multiple training samples according to the target spectral data, where each training sample includes multiple spectral parameters, and construct a neural network identification model based on the training samples;

[0109] A transformation unit, configured to perform the same numerical increase and decrease transformation on the i-th spectral parameter in each training sample to obtain multiple groups of prediction samples, where each group of prediction samples includes two transformation samples obtained by respectively performing numerical increase and decrease transformations on the same training sample;

[0110] A first identification unit, configured to input the transformation samples in each group of prediction samples into the neural network identification model respectively to obtain a first prediction result and a second prediction result;

[0111] A second calculation unit, configured to perform a subtraction operation on the first prediction result and the second prediction result to obtain the influence value of the i-th spectral parameter on the model prediction result;

[0112] A third calculation unit, configured to calculate the influence value of the i-th spectral parameter on the model prediction result in all training samples, and average the influence values according to the number of training samples to obtain the average influence value of the i-th spectral parameter;

[0113] A screening unit, configured to screen the spectral parameters according to the average influence value of each spectral parameter to obtain characteristic spectral data.

[0114] Either of the above two screening modules can be used.

[0115] As an optional implementation manner, the recognition module includes:

[0116] A third construction unit 310, configured to use the wavelengths of the characteristic spectra as the columns of the spectral matrix, the preprocessing methods as the rows of the spectral matrix, and the characteristic spectral data corresponding to different wavelengths and different preprocessing methods as the elements of the spectral matrix, and construct a spectral matrix, where the spectral matrix includes a near-infrared matrix and a mid-infrared matrix;

[0117] A fourth calculation unit 320, configured to perform Pearson correlation coefficient analysis between each column of the near-infrared matrix and the mid-infrared matrix to obtain a correlation coefficient matrix;

[0118] A second recognition unit 330, configured to convert the correlation coefficient matrix into a heat map, and use a pre-trained Bayes-CNN model to recognize the heat map to obtain a recognition result. Embodiment 3

[0119] Corresponding to the above method embodiment, in this embodiment, a drug recognition device based on spectral correlation dimensionality increase is further provided. A drug recognition device based on spectral correlation dimensionality increase described below can be mutually corresponding and referred to with the drug recognition method based on spectral correlation dimensionality increase described above.

[0120] Figure 7 is a block diagram of a drug recognition device 800 based on spectral correlation dimensionality increase shown according to an exemplary embodiment. As Figure 7As shown, the drug identification device 800 based on spectral correlation dimensionality elevation includes a processor 801 and a memory 802. The drug identification device 800 based on spectral correlation dimensionality elevation may also include one or more of a multimedia component 803, an input / output (I / O) interface 804, and a communication component 805. Among them, the processor 801 is used to control the overall operation of the drug identification device 800 based on spectral correlation dimensionality elevation to complete all or part of the steps in the above-mentioned drug identification method based on spectral correlation dimensionality elevation. The memory 802 is used to store various types of data to support the operation of the drug identification device 800 based on spectral correlation dimensionality elevation. These data may include, for example, commands for any application or method operating on the drug identification device 800 based on spectral correlation dimensionality elevation, as well as application-related data, such as contact data, messages sent and received, pictures, audio, video, and so on. The memory 802 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disc.

[0121] The multimedia component 803 may include a screen and an audio component. Among them, the screen may be, for example, a touch screen, and the audio component is used to output and / or input audio signals. For example, the audio component may include a microphone, and the microphone is used to receive external audio signals.

[0122] The received audio signal can be further stored in the memory 802 or sent via the communication component 805. The audio component also includes at least one speaker for outputting the audio signal. The I / O interface 804 provides an interface between the processor 801 and other interface modules, and the other interface modules may be a keyboard, a mouse, buttons, etc. These buttons can be virtual buttons or physical buttons. The communication component 805 is used for the drug recognition device 800 based on spectral correlation dimension elevation to communicate with other devices in a wired or wireless manner. Wireless communication, such as Wi-Fi, Bluetooth, Near Field Communication (NFC), 2G, 3G, or 4G, or a combination of one or more of them. Accordingly, the communication component 805 may include: a Wi-Fi module, a Bluetooth module, and an NFC module.

[0123] In an exemplary embodiment, the device 800 for mutual signature and verification of digital files can be implemented by one or more Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), Field Programmable Gate Arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components, and is used to execute the above-mentioned drug recognition method based on spectral correlation dimension elevation.

[0124] In another exemplary embodiment, a computer-readable storage medium including program commands is further provided. When the program commands are executed by a processor, the steps of the above-mentioned drug recognition method based on spectral correlation dimension elevation are implemented. For example, the computer-readable storage medium may be the above-mentioned memory 802 including program commands, and the above-mentioned program commands can be executed by the processor 801 of the drug recognition device 800 based on spectral correlation dimension elevation to complete the above-mentioned drug recognition method based on spectral correlation dimension elevation.

[0125] Embodiment 4

[0126] Corresponding to the above method embodiment of drug recognition based on spectral correlation dimension elevation, in this embodiment, a readable storage medium is further provided, and a readable storage medium described below can be correspondingly referred to the above-mentioned drug recognition method based on spectral correlation dimension elevation.

[0127] A readable storage medium stores a computer program thereon. When the computer program is executed by a processor, the steps of the above-mentioned embodiment of the drug recognition method based on spectral correlation dimension elevation are implemented.

[0128] Specifically, the readable storage medium can be various readable storage media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc that can store program codes.

[0129] It should be noted that in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element.

[0130] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.

Claims

1. A drug identification method based on spectral correlation dimension upgrading, characterized in that: include: Acquire spectral data of drug samples and preprocess them using at least two preprocessing methods to obtain target spectral data, wherein the target spectral data includes near-infrared data and mid-infrared spectral data; Using a spectral variable selection method to screen the target spectral data to obtain characteristic spectral data; A spectral matrix is ​​constructed according to the wavelength of the characteristic spectral data to obtain a near-infrared matrix and a mid-infrared matrix, the near-infrared matrix and the mid-infrared matrix are subjected to correlation analysis to obtain a correlation coefficient matrix, the correlation coefficient matrix is ​​converted into a heat map, and the heat map is recognized using a pre-trained recognition model to obtain a recognition result, including: The wavelength of the characteristic spectrum is used as the column of the spectrum matrix, the preprocessing method is used as the row of the spectrum matrix, and the characteristic spectrum data corresponding to different wavelengths and different preprocessing methods are used as the elements of the spectrum matrix to construct a spectrum matrix, wherein the spectrum matrix includes a near-infrared matrix and a mid-infrared matrix; Performing a Pearson correlation coefficient analysis on each column of the near infrared matrix and the mid infrared matrix to obtain a correlation coefficient matrix; The correlation coefficient matrix is ​​converted into a heat map, and the pre-trained Bayes-CNN model is used to identify the heat map to obtain the recognition result.

2. The drug identification method based on spectral correlation dimension upgrading according to claim 1 is characterized in that: The method of selecting a spectral variable to select the target spectral data to obtain characteristic spectral data includes: A partial least squares regression model is established based on the target spectral data, and model parameters are recorded, wherein the model parameters include the load, score and principal component of each target spectral data; Based on the model parameters, the VIP value of each target spectral data is calculated using the variable importance projection method; The target spectral data with a VIP value greater than 1 are selected as the characteristic spectral data.

3. The drug identification method based on spectral correlation dimension upgrading according to claim 1 is characterized in that: The method of selecting a spectral variable to select the target spectral data to obtain characteristic spectral data includes: Constructing a plurality of training samples according to the target spectral data, each of the training samples including a plurality of spectral parameters, and constructing a neural network recognition model based on the training samples; Perform the same numerical increase and decrease transformation on the i-th spectral parameter in each training sample to obtain multiple groups of prediction samples, wherein each group of prediction samples includes two transformed samples obtained by respectively increasing and decreasing the numerical value of the same training sample; Inputting the transformed samples in each group of prediction samples into the neural network recognition model respectively to obtain a first prediction result and a second prediction result; Perform a difference operation on the first prediction result and the second prediction result to obtain the influence value of the i-th spectral parameter on the model prediction result; Calculate the influence of the i-th spectral parameter on the model prediction results in all training samples, and average the influence values ​​according to the number of training samples to obtain the average influence value of the i-th spectral parameter; The spectral parameters are screened according to the average influence value of each spectral parameter to obtain characteristic spectral data.

4. A drug identification device based on spectral correlation dimension upgrading, characterized in that: include: A preprocessing module, used for acquiring spectral data of drug samples and preprocessing them using at least two preprocessing methods to obtain target spectral data, wherein the target spectral data includes near-infrared data and mid-infrared spectral data; A screening module, used for screening the target spectral data by using a spectral variable selection method to obtain characteristic spectral data; A recognition module is used to construct a spectrum matrix according to the wavelength of the characteristic spectrum data to obtain a near-infrared matrix and a mid-infrared matrix, perform correlation analysis on the near-infrared matrix and the mid-infrared matrix to obtain a correlation coefficient matrix, convert the correlation coefficient matrix into a heat map, and use a pre-trained recognition model to recognize the heat map to obtain a recognition result; The identification module comprises: A third construction unit is used to construct a spectrum matrix using the wavelength of the characteristic spectrum as the column of the spectrum matrix, using the preprocessing method as the row of the spectrum matrix, and using the characteristic spectrum data corresponding to different wavelengths and different preprocessing methods as the elements of the spectrum matrix, wherein the spectrum matrix includes a near-infrared matrix and a mid-infrared matrix; A fourth calculation unit is used to perform a Pearson correlation coefficient analysis between each column of the near infrared matrix and the mid-infrared matrix to obtain a correlation coefficient matrix; The second recognition unit is used to convert the correlation coefficient matrix into a heat map, and use the pre-trained Bayes-CNN model to recognize the heat map to obtain a recognition result.

5. The drug identification device based on spectral correlation dimension upgrading according to claim 4 is characterized in that: The screening module comprises: A first construction unit is used to establish a partial least squares regression model based on the target spectral data and record model parameters, wherein the model parameters include a load, a score and a principal component of each target spectral data; A first calculation unit is used to calculate the VIP value of each target spectral data by using a variable importance projection method based on the model parameters; The selection unit is used to select target spectrum data with a VIP value greater than 1 as characteristic spectrum data.

6. The drug identification device based on spectral correlation dimension upgrading according to claim 4 is characterized in that: The screening module comprises: A second construction unit is used to construct a plurality of training samples according to the target spectral data, each of the training samples includes a plurality of spectral parameters, and to construct a neural network recognition model based on the training samples; A transformation unit is used to perform the same numerical increase and decrease transformation on the i-th spectral parameter in each training sample to obtain multiple groups of prediction samples, wherein each group of prediction samples includes two transformed samples obtained by respectively increasing and decreasing the numerical value of the same training sample; A first recognition unit is used to input the transformed samples in each group of prediction samples into the neural network recognition model to obtain a first prediction result and a second prediction result; The second calculation unit is used to perform a difference operation on the first prediction result and the second prediction result to obtain an influence value of the i-th spectral parameter on the model prediction result; The third calculation unit is used to calculate the influence value of the i-th spectral parameter in all training samples on the model prediction result, and average the influence values ​​according to the number of training samples to obtain the average influence value of the i-th spectral parameter; The screening unit is used to screen the spectral parameters according to the average influence value of each spectral parameter to obtain characteristic spectral data.

7. A drug identification device based on spectral correlation dimension upgrading, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the drug identification method based on spectral correlation dimensionality upgrading as described in any one of claims 1 to 3 are implemented.

8. A readable storage medium, characterized in that: The readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the drug identification method based on spectral correlation dimensionality upgrading as claimed in any one of claims 1 to 3.

Citation Information

Patent Citations

  • Associated effective characteristic spectrum section selection method and oil product index content rapid detection method

    CN115372309A

  • Bill identification method and device based on machine vision

    CN119068504A