Mobile terminal lightweight drug taking auxiliary detection method and system based on facial image

By designing a mobile lightweight drug detection method based on facial images, using lightweight neural network module and model conversion technology, the convenience and efficiency of traditional drug detection methods are solved, and fast and accurate drug detection is achieved on the mobile terminal.

CN120032407APending Publication Date: 2025-05-23GUANGZHOU INSTITUTE OF TECHNOLOY XIDIAN UNIVERSITY

Patent Information

Application Number
CN202411709829.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-11-27
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

Traditional drug use detection methods have problems such as low convenience, long detection cycle and high cost. The existing drug use detection methods based on facial images either require close-range matching collection or require pre-installed software to perform specific signal extraction, resulting in inconvenience and accuracy that are difficult to guarantee. At the same time, ordinary neural network models have high requirements for computing resources and are difficult to apply on mobile terminals.

Method used

A lightweight drug addict assisted detection method and system on mobile terminal based on facial images was designed. After using a camera to collect facial images and perform preprocessing operations, a lightweight neural network module was designed to build a lightweight drug addict detection network, and transplant it to the mobile terminal through model conversion technology.

Benefits of technology

It realizes rapid drug use detection without high coordination, reduces detection costs, improves detection efficiency, and is successfully deployed to the mobile terminal, meeting the computing resource limitations of the mobile terminal.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032407A_ABST
    Figure CN120032407A_ABST
Patent Text Reader

Abstract

The invention discloses a mobile terminal lightweight drug taking auxiliary detection method and system based on a face image, and the method comprises the following steps: S1, collecting a face image of a detected person through a camera, and carrying out the preprocessing operation of face detection, background removal and face alignment; s2, designing a lightweight neural network module, and constructing a lightweight drug taking detection network based on the module; s3, acquiring a face data set of a drug addict and training a drug addict detection network; s4, extracting face features by using the drug taking detection network and obtaining a classification result and a certainty degree; and S5, converting and transplanting the drug taking detection network model to a mobile terminal, and constructing a mobile terminal application. According to the method, the drug addicts are detected based on the facial images, the defects that a traditional biochemical sample detection method is high in matching requirement, long in detection period, difficult to screen in a large range and the like are overcome, meanwhile, deep learning and lightweight thoughts are used, the algorithm performance is effectively improved, and the method is suitable for mobile terminal application scenes, portable application scenes and other application scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of digital image processing and pattern recognition, and in particular relates to a mobile terminal lightweight drug abuse auxiliary detection method and system based on facial images. Background Art

[0002] Traditional drug detection methods mainly include various biochemical test-based technologies, that is, by collecting urine, blood or hair from the person being tested, using biological detection methods and related equipment for detection. These detection methods can accurately detect whether or not drug use has occurred, thereby assisting the police in identifying drug users and taking corresponding measures. However, traditional forms of drug detection technology have serious deficiencies in terms of convenience, detection cycle, and economy: (1) Traditional drug detection methods require a high degree of cooperation from the person being tested, which brings inconvenience to the detection; (2) Traditional drug detection methods require a long time and cycle, and cannot quickly obtain detection results; (3) Traditional drug detection methods are relatively expensive, so they cannot be used for large-scale screening and can only be implemented on a small number of key drug users.

[0003] Therefore, intelligent visual drug detection technology based on facial images of the person being tested has become an important supplement to traditional drug detection technology. At present, deep neural networks can extract and analyze facial features very well. It is completely feasible and efficient to use deep learning and facial feature extraction to distinguish drug users from normal people (drug users have obvious features such as sunken eye sockets, dilated pupils, and mental depression). Vision-based drug detection technology has strong real-time performance, does not require long waiting times, and can completely get rid of the requirement for high user cooperation (such as through surveillance cameras). At the same time, vision-based drug detection technology is inexpensive and does not require the purchase of a large amount of biochemical consumables. More specifically, for its ease of use, facial image drug detection algorithms and software can be deployed to mobile terminals (such as various handheld devices). This application method can help relevant personnel use and exert their effectiveness more conveniently, and maximize the efficiency of drug control.

[0004] On the other hand, the existing methods for drug detection or judgment based on facial images either require the cooperation of suspected drug users, such as the need to collect pupil images of suspected drug users at close range, see patents CN103186765A and CN112603256A. This measure causes many inconveniences in actual deployment and use, and cannot be operated at a long distance; or some pre-software is required to extract certain specific signals from the human body. For example, the pulse signal is extracted using imaging photoplethysmography IPPG software, see patent CN111870235A. The accuracy of this idea depends on the accuracy of the pre-software, the performance is not necessarily guaranteed, and it requires additional costs and consumes more computing resources. How to get rid of the defects of drug detection methods based on pupil images and IPPG pulse signals is one of the current research directions for drug detection based on facial images.

[0005] In addition, many current neural network models are designed and trained for PCs. Although the prediction accuracy of the models has been improved, they have deeper layers and larger parameters, and require higher computing resources, making them unsuitable for mobile applications with lower computing power. First, the storage of more network parameters will take up more space, so these large models cannot be loaded on mobile devices with smaller memory. Second, more parameters also mean a greater amount of computing, which leads to insufficient processor performance, making it impossible for the network model to respond in a timely manner, making it difficult to achieve real-time performance. Therefore, lightweighting the drug detection model is an important issue that needs to be addressed when its network model is ported to mobile devices.

[0006] Therefore, in order to implement drug detection algorithms and technologies based on facial images on mobile devices, we need to streamline the network model. The two common methods are model compression and lightweight model design. Model compression includes parameter quantization, pruning and quantization, which aim to reduce the storage and computing overhead of the model. Lightweight model design is to reduce the complexity of the model while maintaining performance by designing a simpler network structure or introducing specially designed lightweight modules.

[0007] After searching, the patent with announcement number CN 118506413A published a method for face recognition of drug users. The method uses a camera to collect facial images of potential drug users, uses an auxiliary network to extract the identity information of the person, and fuses the face with drug taking characteristics and attribute characteristics to determine whether the person in the image is a drug user. However, this method cannot perform model conversion. At the same time, since this method does not build a lightweight neural network, its effect speed delay is relatively high, which is not conducive to mobile terminal deployment.

[0008] In general, the difficulties of current auxiliary methods for drug detection are: (1) Traditional drug detection methods based on biochemical tests have difficulties such as high cooperation requirements, long detection time period, and high cost; (2) Current drug detection methods based on facial images either require close-range cooperative acquisition, which makes deployment and use inconvenient, or require front-end software to extract specific signals, which makes use indirect, accuracy is difficult to guarantee, introduces additional costs, and has high computational overhead; (3) The current structure of ordinary neural networks has many more network parameters and requires larger memory storage space and higher computing power requirements, making it difficult to directly apply to auxiliary methods for drug detection based on facial images; (4) Current ordinary neural networks are more suitable for deployment on PCs, and their usage scenarios are often limited to indoor environments.

[0009] The significance of solving the above problems and defects is: by researching and designing drug detection technology based on image and visual intelligence, we can get rid of the defects of traditional biochemical detection technology such as low convenience, poor timeliness and low economy. At the same time, by designing lightweight network modules and corresponding lightweight drug detection neural networks, we can effectively improve the operating efficiency of drug detection technology based on image and visual intelligence, so as to successfully deploy it on mobile terminals. Summary of the invention

[0010] In view of the above problems, the present invention provides a lightweight drug abuse auxiliary detection method and system for a mobile terminal based on facial images.

[0011] The technical solution of the present invention is: A lightweight method for detecting and identifying drug addicts on a mobile terminal based on facial vision specifically includes the following steps: S1. Use a camera to collect the facial image of the person being tested, and perform preprocessing operations such as face detection, background removal, and face alignment; S2. Design a lightweight neural network module and build a lightweight drug detection network based on the module; S3, obtain the face data set of drug addicts and train the drug detection network; S4, using the drug use detection network to extract facial features and obtain classification results and confidence; S5. Convert and transplant the drug detection network model to the mobile terminal, build a mobile application and display the results.

[0012] As a preferred solution of the present invention: the face detection operation in S1 is performed using a representative cascade face detector MTCNN, which is specifically composed of three sub-networks PNet, RNet, and ONet. The face detector first scales an input face image into multiple face images in equal proportions, and inputs them one by one into the PNet network to obtain a three-dimensional feature map containing an offset rate and a confidence level; then, the coordinates of the prediction box in the original image are calculated according to the confidence level and a set threshold to obtain the prediction box; next, non-maximum suppression (NMS) processing is performed on all prediction boxes, and redundant pre-selected boxes are removed by the degree of overlap (i.e., intersection-over-union (IOU)), and the corresponding positions in the original image are captured and input into RNet for processing.

[0013] After that, the above-retained prediction box is processed by non-maximum suppression, and the redundant pre-selected boxes are removed by overlapping IOU, and the corresponding position in the original image is intercepted and input into ONet for processing. The retained prediction box is processed by non-maximum suppression again, and the redundant pre-selected boxes are removed by overlapping IOU to obtain the final target prediction box. Finally, the predicted face position in the image is drawn according to the prediction box coordinates.

[0014] As a preferred solution of the present invention: the face cropping operation in S1 needs to be cropped according to the face border detected by the face detector, and after obtaining the face detection frame, the image is cut and saved according to the four-point coordinates of the detection frame to obtain a face image with the background removed.

[0015] As a preferred solution of the present invention: the face alignment operation in S1 uses affine transformation to perform alignment, and before face alignment, it is necessary to detect facial key points and analyze and annotate the facial image.

[0016] As a preferred solution of the present invention: the annotated area in the analysis and annotation of the face image in S1 includes the eyes, eyebrows, nose, mouth and edge contours of the person, so as to obtain the key points (landmarks) of the face image.

[0017] As a preferred solution of the present invention: the face alignment operation in S1 is performed using an affine transformation, which converts the spatial coordinates of the face image before alignment to the spatial coordinates of the face image after alignment through a series of operations such as translation and rotation. The operation has a linear relationship, that is, it does not change the relative shape of the face image before and after alignment, nor does it change the degree of straightness of the lines. The affine transformation is expressed in the form of matrix multiplication. When determining the affine matrix, the mapping relationship of at least three non-collinear points (i.e., the aforementioned face key points) is required. According to the linear equations corresponding to these non-collinear points, the least squares method is used to solve the values ​​of each element of the affine matrix.

[0018] Specifically, the affine transformation process of face alignment is expressed as, (1) in, , , , , , is the matrix element value we need to determine, corresponding to the rotation parameter, translation parameter and scaling parameter; , These are the coordinates of the binocular center key points and the lip center key points we obtained; , It is the preset key point mapping coordinate point (i.e. the coordinate after alignment).

[0019] As a preferred solution of the present invention: the lightweight network module in S2 is based on group convolution and separable convolution, and is constructed by drawing on the idea of ​​​​first reducing the dimension and then performing feature processing on the network module in SqueezeNet. First, the group convolution method is used to divide the input features into multiple groups according to the path, and then convolution operations are performed in each group respectively; for the convolution operation in each group, the ordinary convolution is matrix decomposed and replaced with depthwise separable convolution; finally, the 1×1 convolution operation and the depthwise convolution operation of the depthwise separable convolution are replaced in sequence, so as to achieve the operation of first performing feature dimensionality reduction and then performing feature extraction, thereby realizing the lightweight design of the entire model.

[0020] As a preferred solution of the present invention: the face dataset of drug addicts obtained in S3 is, on the one hand, a drug detection dataset is constructed using face images of drug addicts provided by relevant institutions, and the drug detection dataset is divided into a training set and a test set for training and testing the model; on the other hand, a Python crawler tool is used to crawl public drug addict images from public websites to construct a test dataset, which is used for testing the generalization ability of the model across libraries.

[0021] As a preferred solution of the present invention: the neural network training in S3 is trained using the FocalLoss loss function, and a binary classification fully connected layer is used to classify whether the person has taken drugs or not (i.e., drug use and non-drug use are each classified into one category).

[0022] As a preferred solution of the present invention: the facial feature extraction operation in S4 uses a lightweight feature extraction module designed based on the MobileNet network structure as the main body, combined with a lightweight drug addict feature recognition and extraction network constructed with a residual structure to extract facial related features.

[0023] As a preferred solution of the present invention: the lightweight neural network model conversion and transplantation and deployment to the mobile terminal in S5 is performed using ncnn. ncnn is a high-performance neural network forward computing framework highly optimized for mobile phones, which can support most common convolutional neural network models.

[0024] The beneficial effects of the present invention are: 1. Based on deep neural networks and drawing on face recognition technology, the present invention designs a neural network for drug detection based on facial images. It does not require high cooperation from the detected person, can greatly improve the efficiency of drug control and prevention, and help maintain social order; 2. It has a high effect speed. With the construction of lightweight neural networks, it is possible to realize intelligent detection and identification of drug addicts with lower parameters and faster response time and complete deployment on mobile terminals. Therefore, the present invention designs a drug addict detection and identification network based on lightweight networks, and studies lightweight methods for models to meet the transplantation needs of mobile terminals; 3. The model conversion is convenient and conducive to mobile terminal deployment. After completing the network training on the PC, the network model is first converted into an intermediate model file using the onnx library, and then the model file is converted into a mobile terminal callable format using the ncnn library, so that it can be transplanted, deployed and used on the mobile terminal. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 It is a schematic diagram of a lightweight drug-assisted detection method for a mobile terminal based on facial images provided by an embodiment of the present invention; Figure 2 This is a training and testing flow chart of a lightweight drug detection model for a mobile terminal based on facial images provided by an embodiment of the present invention; Figure 3 is a schematic diagram of a preprocessing operation flow of a data set provided by an embodiment of the present invention; Figure 4 Schematic diagram of the structure of the LMDI lightweight module provided by an embodiment of the present invention; Figure 5 It is a structural diagram of a lightweight drug detection network DIFace provided by an embodiment of the present invention; Figure 6 It is a flow chart of a method for converting and deploying a drug abuse detection model on an Android mobile terminal provided by an embodiment of the present invention; Figure 7 It is a schematic diagram of the framework structure of an Android application end of a lightweight drug detection auxiliary method based on facial images provided by an embodiment of the present invention; Figure 8This is a flow chart of a lightweight drug abuse auxiliary detection method based on facial images on a mobile terminal provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0026] The embodiments of the present invention will be further described below in conjunction with the accompanying drawings.

[0027] Embodiment 1: The embodiment of the present invention provides a lightweight drug abuse auxiliary detection method based on facial images on a mobile terminal, comprising the following steps: Step 1: lightweight neural network module construction; Specifically, in the step 1 of the embodiment of the present invention, the lightweight model design first uses the grouped convolution method to divide the input features into multiple groups according to the path, and then performs convolution operations in each group respectively; for the convolution operation in each group, the ordinary convolution is replaced by matrix decomposition with depthwise separable convolution; finally, the 1×1 convolution operation and the depthwise convolution operation of the depthwise separable convolution are replaced in sequence, so as to achieve feature dimensionality reduction before feature extraction operation, thereby realizing the design of the entire lightweight model; Step 2: Build a lightweight drug detection network; Specifically, the lightweight drug detection network of step 2 of the embodiment of the present invention is built based on the lightweight module of step 1 and the MobileNet network architecture; the network uses the lightweight module in step 1 as a feature extraction module, and integrates the lightweight module into the MobileNet network structure to realize the construction of a lightweight drug detection network; in order to reduce the number of parameters of the neural network, the lightweight drug detection network replaces the depth separable convolution module in the MobileNet network with the lightweight module of step 1; Step 3: lightweight neural network model format conversion; Specifically, step 2 of the embodiment of the present invention uses onnx and ncnn to realize the conversion of the Android mobile terminal model; the trained model is converted and deployed to the Android mobile terminal for model deployment and application, so as to realize the construction of a handheld mobile terminal application system for auxiliary face recognition of drug addicts; Step 4: Android application development and model migration and deployment; Specifically, step three of the embodiment of the present invention provides a mobile terminal platform, which uses the front camera or rear camera of an Android phone to collect image information in a real scene, and then performs face detection on the collected image to determine whether it contains a face. If it does, the image is processed and input into a drug addict detection and auxiliary recognition network for inference and recognition, and finally the detection result is displayed on the interface; Step 5: face detection and capture; Specifically, in step 4 of the embodiment of the present invention, the face part in the person image is detected by using opencv and dlib, and the detected face frame is intercepted and retained to obtain a face image after removing the background; Step 6: facial key point detection; Specifically, step 5 of the embodiment of the present invention performs facial key point detection on the captured facial image. The present invention uses a 68-point face detection model, which will mark 68 key points of the facial parts. The key points include eyebrows, eye sockets, nose, mouth and the overall contour of the face, a total of 68 facial key points; Step 7: Face alignment; Specifically, in step 6 of the embodiment of the present invention, since different people present different actions and postures in the face image, the posture of the face image is corrected, which is conducive to the neural network to extract more important features, and the face image is corrected by using affine transformation. The key points used in the affine transformation are obtained based on the face key point detection in the previous step, and the required face image is made to present a unified frontal posture through the face alignment operation; Step 8: facial feature extraction; Specifically, step seven of the embodiment of the present invention gradually extracts and analyzes features through a lightweight module in a neural network, and the residual structure can well help improve the accuracy of the network; Step 9: Feature analysis and result display; Specifically, step eight of the embodiment of the present invention uses softmax to perform feature classification based on the network features extracted from the front part of the network, and sets the classification results between 0 and 1, which is the probability predicted by the neural network for each classification result, and selects the classification result with the largest probability as the classification output result of the model; after the image classification results are obtained, they will be displayed in the mobile application interface to facilitate the acquisition of real-time analysis information.

[0028] The embodiment of the present invention provides a mobile terminal auxiliary detection method for drug addicts, which effectively solves the problems of low detection and recognition efficiency, poor accuracy, high complexity of ordinary visual models, and difficulty in mobile terminal adaptation in traditional drug addict methods. The method can be used in mobile terminal real-time monitoring, public safety management, drug rehabilitation center management, mobile law enforcement and other occasions.

[0029] Embodiment 2: In view of the shortcomings of current drug detection, such as the need for high cooperation from the person being tested, the inability to obtain test results quickly, and the high cost that makes large-scale screening difficult, this implementation uses facial images of drug users and related deep learning technology to distinguish between drug users and normal people. Therefore, please refer to Figure 1 and Figure 2 .

[0030] This embodiment provides a model lightweight method for drug use detection on a mobile terminal based on facial images, the method comprising the following steps: Step 1: Use the facial images of drug users provided by relevant institutions to build a drug detection dataset, and divide the drug detection dataset into a training set and a test set for model training and testing. At the same time, use Python crawler tools to crawl public drug user images from public websites to build a test dataset for cross-database generalization test of the model. Specifically, this embodiment uses drug addict images provided by relevant institutions to construct a drug detection dataset (PsDrug) for model training and testing; and uses an online drug addict dataset (WebDrug) constructed from drug addict photos publicly available on the Internet to test the generalization ability of the model across libraries; the sample number ratio of drug addict face images to non-drug addict face images in the two datasets is approximately 1:3, that is, the drug addict data is less than the non-drug addict face data, which is consistent with the scenario mode where there are fewer drug addicts in reality; Among them, the PsDrug dataset contains the corresponding 400 provided by relevant institutions 1037 drug addict face images and 3100 non-drug addict face images of multiple drug addicts; this dataset contains images of drug addicts of multiple age groups and different drug-taking durations; the WebDrug dataset contains 57 drug addict face images and 180 non-drug addict face images; this dataset is obtained by crawling drug addict images from public websites using Python crawler tools, and the drug addict images with better quality are retained after data cleaning; the images in this dataset are taken in an unrestricted environment, with non-uniform formats and diverse postures. This embodiment uses this dataset as a cross-database test dataset for model generalization ability; Step 2, preprocessing of the dataset: perform face detection on the images in the PsDrug dataset and the WebDrug dataset to remove useless background information, then perform face key point detection and face alignment to make the face images in the dataset uniform in posture; Figure 3 Step 2 of the illustrated embodiment includes: Face detection and capture: For the collected face images of the detected persons, the cascaded face detector MTCNN is used for face detection. The face detector consists of three sub-networks: PNet, RNet, and ONet. First, a face image to be detected is scaled into multiple images in equal proportions and input into the PNet network one by one to obtain a three-dimensional feature map containing the offset rate and confidence. At the same time, the coordinates of the prediction box in the original image are calculated according to the confidence and the set threshold to obtain the prediction box; Perform non-maximum suppression (NMS) processing on all prediction boxes, remove redundant pre-selected boxes through overlap IOU, and capture the corresponding positions in the original image and input them into RNet for processing; Perform non-maximum suppression (NMS) on the prediction boxes retained after the above processing, remove redundant pre-selected boxes through the overlap IOU, and capture the corresponding position in the original image and input it into ONet for processing; Perform non-maximum suppression (NMS) on the prediction box retained after the above processing, remove the redundant pre-selected boxes through the overlap IOU, and obtain the final target prediction box; Draw the predicted face position in the image based on the prediction box coordinates.

[0031] This embodiment intercepts and retains the face part in the image according to the vertex coordinates of the face rectangular frame to obtain a face image after removing the background; Face key point detection: Face key point detection is performed on the captured face image to prepare for the subsequent face alignment. This embodiment uses a 68-point face detection model, which marks 68 key points of the face. The key points include eyebrows, eye sockets, nose, mouth and the overall contour of the face. Based on these key points, this embodiment determines the position of the key parts of the face for the subsequent face alignment operation; Face alignment: In order to reduce the impact of different facial posture angles on subsequent drug use detection, this embodiment uses affine transformation to correct the posture of the face image (i.e., face alignment). According to the three key points obtained by facial key point detection, such as the center coordinates of the left and right eyes and the center position coordinates of the bottom of the nose, a corresponding affine matrix is ​​constructed to perform face alignment operations, so that the face images in the data set present a unified frontal posture.

[0032] Specifically, in the face symmetry pre-training process in step 2 of this embodiment, the face alignment operation is performed using an affine transformation, which converts the spatial coordinates of the face image before alignment to the spatial coordinates of the face image after alignment through a series of translation, rotation, scaling and other operations. The operation has a linear relationship, that is, it will not change the relative shape of the face images before and after alignment, nor will it change the straightness of the lines. The affine transformation is represented in the form of matrix multiplication. When determining the affine matrix, the mapping relationship of at least three non-collinear points (i.e., the aforementioned face key points) is required. According to the linear equations corresponding to these non-collinear points, the values ​​of each element of the affine matrix are solved.

[0033] Specifically, the affine transformation process of face alignment is expressed as, (1) in, , , , , , is the matrix element value we need to determine, corresponding to the rotation parameter, translation parameter and scaling parameter; , These are the coordinates of the binocular center key points and the lip center key points we obtained; , It is the preset key point mapping coordinate point (i.e. the aligned coordinate point). Through the three sets of coordinates of the binocular center key point and the lip center key point, six equations are constructed, and all element values ​​of the affine matrix are solved through the six equations. When any face image is input, the input image is aligned by multiplying the above-determined affine matrix with the coordinates of each pixel of the input face image; Step 3, construct a lightweight module LMDI for drug addict identification, such as Figure 4 As shown, step 3 of this embodiment includes: The lightweight module for drug addict identification is constructed based on grouped convolution and separable convolution, and draws on the idea of ​​dimensionality reduction followed by feature processing in SqueezeNet. In the first stage of lightweight model design, the grouped convolution method is used to reduce the number of model parameters. The grouped convolution method first divides the input features into N groups according to the number of pathways, and then performs convolution operations on each group of features after division to obtain each group of features. Finally, the obtained groups of features are fused to obtain the final output of the layer.

[0034] In the second stage of lightweight model design, separable convolution is used to decompose the convolution kernel to reduce the calculation parameters and amount of calculation; after the first step of feature grouping, depth-wise separable convolution is used to replace ordinary convolution. The separable convolution is divided into spatially separable convolution and depth-wise separable convolution. In this embodiment, depth-wise separable convolution is used. The depth-wise separable convolution decomposes the convolution kernel into depth-wise convolution and point-wise convolution, where the depth-wise convolution performs convolution operations on each path, and the point-wise convolution uses a 1×1 size convolution kernel for convolution operations.

[0035] In the third stage of lightweight model design, the idea of ​​SqueezeNet to reduce dimension first and then perform convolution is used to further lightweight the model; the depthwise separable convolution in the second step uses the relevant theory of matrix decomposition to achieve the reduction of parameters, and it performs depthwise convolution and then pointwise convolution. This embodiment uses the same matrix decomposition method but changes the order of convolution, that is, first using a 1×1 convolution kernel to perform feature dimensionality reduction to reduce the number of parameters, and then using depthwise convolution to extract and integrate features. Because feature dimensionality reduction first reduces the number of feature paths, it reduces the number of parameters and calculation time of the overall calculation.

[0036] Step 4, construct a lightweight drug detection network DIFace; Specifically, see Figure 5The lightweight drug detection network DIFace is built based on the lightweight module LMDI and the MobileNet network architecture. The network uses LMDI as the feature extraction module and integrates the lightweight module into the MobileNet network structure to realize the network construction. In order to reduce the number of parameters of the neural network, the DIFace network replaces the deep separable convolution module in the MobileNet network with the LMDI lightweight module. For the input of features, group convolution is first used to group multiple features before performing feature processing and extraction operations. For network modules whose output feature dimensions are lower than the input feature dimensions, a 1×1 convolution kernel is first used for dimensionality reduction operations, and then feature extraction operations are performed, which can effectively reduce the amount of calculation in the network reasoning process. At the same time, the network structure uses a residual module to perform identity mapping between the input features and the output features to improve the accuracy of reasoning prediction.

[0037] The structural diagram of the DIFace network consists of multiple LMDI lightweight feature extraction modules. The input image is preprocessed and then input into the network. The input features are output through multiple residual feature extraction modules with additional identity mapping, followed by pooling layers, activation function layers and fully connected layers to achieve the final extraction and classification of features.

[0038] The input of the DIFace network structure designed by the present invention is an image of size 3×224×224. First, the image is preprocessed to perform face detection on the image to remove useless background information, and then a face alignment preprocessing operation is performed to unify the postures of the face images in the data set to facilitate training. After the preprocessing operation is completed, the image is input into the network, and the relevant face features of the image are extracted through the feature extraction layers of multiple lightweight modules LMDI. Finally, the final classification result is output through a pooling layer, a fully connected layer and a softmax function.

[0039] The parameter design of each layer in the lightweight drug detection network DIFace during the verification process of this embodiment is specifically shown in Table 1. The padding method during the convolution process is 0 padding.

[0040] Table 1 Parameter design of each layer of lightweight drug detection network (DIFace) Specifically, the experiment in this embodiment uses accuracy and F1 score as performance evaluation indicators. The exact definitions of accuracy and F1 score are as follows: (1) Accuracy (ACC) refers to the proportion of correctly classified samples in the model prediction to all predicted samples. The accuracy calculation formula is as follows.

[0041] (1) Among them, TP: true positive, the number of samples that are actually positive examples and predicted as positive examples; FP: false positive, the number of samples that are actually negative but predicted to be positive; TN: true negative, the number of samples that are actually negative and predicted as negative; FN: false negative, the number of samples that are actually positive but predicted as negative.

[0042] The accuracy rate allows us to intuitively understand the performance of the model. However, when the positive and negative samples in the data set are unbalanced, the samples that account for the majority of the data set will become the main factor affecting the results. The trained neural network model may be more inclined to predict the samples as the majority samples in the training set.

[0043] (2) The F1 score is a trade-off between accuracy and recall, taking into account the impact of the accuracy and recall of the classification model. The F1 score ranges from 0 to 1. The larger the value, the better the model performance. It is often used in scenarios where positive and negative samples are unbalanced. The calculation formula is as follows: (2) (3) (4) To sum up, this embodiment uses the lightweight module LMDI for feature extraction, that is, group convolution plus depth-separable convolution module plus dimensionality reduction, and embeds the lightweight module into the introduction of the MobileNet network structure, which effectively reduces the amount of calculation in the network reasoning process and realizes the lightweight model for identifying drug addicts.

[0044] The experiment designed in this embodiment is demonstrated from the following aspects: (1) To illustrate the effectiveness of the lightweight module LMDI in the lightweight drug detection network DIFace of the present invention, the effectiveness of each strategy is verified by ablation experiments. Table 2 shows the results of comparing the theoretical number of forward reasoning calculations of the network in the neural network constructed with different lightweight modules. Table 3 shows the results of comparing the actual running time spent by the network in the neural network constructed with different lightweight modules to process an input face image in a mobile application.

[0045] Table 2 Comparison of theoretical calculation times Table 3 Comparison of actual running time (2) In order to prove the effectiveness of the lightweight module LMDI in the lightweight drug detection network DIFace of the present invention, the effectiveness of each strategy was verified by ablation experiments, and the F1 value and accuracy rate in the reasoning process of each neural network were calculated. Table 4 shows the F1 value and accuracy rate (ACC) test results of the ablation experiment on the neural network composed of various convolutional modules on the relevant data set (PsDrug), and Table 5 shows the F1 value and accuracy rate (ACC) test results of the ablation experiment on the neural network composed of various convolutional modules on the network data set (WebDrug).

[0046] Table 4 F1 value and accuracy results of ablation experiments on relevant datasets (PsDrug) Table 5 F1 value and accuracy results of ablation experiment on the network dataset (WebDrug) (3) In order to prove that the efficiency of the lightweight drug detection network DIFace of the present invention is better than that of traditional face recognition and other mainstream deep learning methods, a comparative experiment was conducted on traditional face recognition and other mainstream deep learning methods, and the F1 value and accuracy index of each method were calculated. Table 6 shows the F1 value and accuracy (ACC) test results of the comparative experiment on the relevant data set, and Table 7 shows the F1 value and accuracy (ACC) test results of the comparative experiment on the network data set.

[0047] Table 6 Comparative experiment F1 value and accuracy results on related data sets (PsDrug) Table 7 Comparative experiment F1 value and accuracy results on the network dataset (WebDrug) It can be seen from the results of ablation experiments and comparative experiments that the DIFace neural network constructed by the LMDI lightweight module in the embodiment of the present invention has good improvements in parameter quantity, calculation times and running time. The theoretical number of multiplication and addition of the network is 424 million times, and the actual running time is 34ms; the F1 value of the lightweight network on the relevant data set is 0.862, and the accuracy is 0.906; the F1 value of the neural network on the network data set is 0.742, and the accuracy is 0.822.

[0048] Due to the uneven data quality of the network data set, it is relatively difficult for the model to predict the data set, and the prediction accuracy is reduced; due to the imbalance of samples, the accuracy is more inclined to multi-sample prediction, resulting in a higher accuracy. The F1 value can better measure the network performance. From the ablation experiment, it can be seen that the lightweight module used in this embodiment is lower than the existing grouped convolution and depth-separable convolution methods in terms of computational complexity, and the overall reasoning running time of the network is also relatively low. In the comparative experiment, this embodiment compares the accuracy of the existing lightweight model and the lightweight model designed in this embodiment on the drug addict identification data set, and verifies that the lightweight network designed in this embodiment reduces the number of parameters and computational complexity while taking into account the performance of reasoning and prediction.

[0049] The embodiment of the present invention provides a lightweight drug abuse auxiliary detection method based on facial images on a mobile terminal, which can solve the shortcomings of traditional detection technologies in detection convenience and result acquisition cycle; this embodiment implements a lightweight drug abuse feature recognition model, which reduces the model complexity while maintaining or even improving the model performance.

[0050] Embodiment 3: Based on the above Example 2, see Figure 6 , Figure 6 The embodiment of the present invention provides a method for converting and deploying a drug detection model on an Android mobile terminal. The specific process of converting and deploying a drug detection model on a mobile terminal is as follows: Step 1: Save the structure and weight parameters of the trained model; Step 2: Use the onnx method to convert the model into an .onnx file and test the usability of the model; Step 3: Use onnxsim to simplify the generated .onnx file; Step 4: Import the third-party library ncnn and convert the intermediate file onnx to the ncnn file to obtain the .param and .bin files. The param file is the network structure file, and the bin file is the network weight parameter file; Step 5: Generate .id.h file and .mem.h file using param file and bin file in VisualStudio environment; Step 6: Use the Android Studio platform to develop mobile applications; Step 7: Deploy the ncnn file obtained after the above conversion to a mobile application to use the neural network model.

[0051] The embodiment of the present invention provides a method for converting and deploying a drug detection network model on an Android mobile terminal, which can convert the above-mentioned lightweight drug detection network DIFace model and deploy it on an Android mobile terminal. The present invention solves the complexity problem of deploying a neural network model on an Android mobile terminal and can be widely used in various occasions where a neural network model needs to be run on a mobile device.

[0052] Embodiment 4: Handheld devices can make it easier for the present invention to be used in outdoor areas. In order to make it more convenient for relevant personnel to use, the present invention designs and develops an Android mobile terminal drug addict identification application system. Figure 7 , Figure 7 This is a schematic diagram of the Android application framework structure of a lightweight drug detection auxiliary method based on facial images provided by an embodiment of the present invention. The embodiment of the present invention provides an Android application framework of a lightweight drug detection auxiliary method based on facial images. The specific functions and structural modules of the application are as follows: The drug detection mobile application system realizes functions such as camera device detection and activation, camera image acquisition method selection, face detection, drug addict identification, and hardware reasoning method selection. By using the deployment and application method of the mobile terminal deep learning model, the lightweight drug addict identification model is deployed and applied, and multiple modules such as image acquisition module and image display module are used to jointly build the application system.

[0053] The main modules of the application platform include loading and using deep learning models, detecting, opening and switching modules of camera devices, image interface display modules and detection mode selection modules. First, the front camera or rear camera of the Android phone is used to collect image information in real scenes, and then the collected images are subjected to face detection to determine whether they contain faces. If they do, the images are processed and input into the drug addict detection and recognition network for inference and recognition, and finally the detection results are displayed on the interface.

[0054] Android application development uses the AndroidStudio platform, combined with third-party libraries such as ncnn and opencv, to realize image acquisition using the Android phone camera, and can select the model to be loaded and used. In terms of mode selection, you can choose CPU processing or GPU processing to input images into the neural network for reasoning.

[0055] After completing the development of the application on the AndroidStudio platform, you can install it on the mobile phone to collect facial images. The front and rear cameras of the mobile phone can be used for facial image collection. After completing the image collection task through the camera, the back-end program will input the collected image into the network model for reasoning and mark the face detection box and prediction results on the display interface.

[0056] An Android application-side framework for lightweight drug detection on a mobile terminal based on facial images provided in an embodiment of the present invention solves the problem that traditional methods are inconvenient to use outdoors and that recognition accuracy is low due to limitations on computing power and memory on the mobile terminal. The framework can be widely used in the identification and monitoring of drug users by relevant departments during actual law enforcement.

[0057] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When it is implemented in whole or in part in the form of a computer program product, the computer program product includes one or more computer instructions. When the computer program instructions are loaded or executed on a computer, the process or function described in the embodiment of the present invention is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from one website, computer, server or data center to another website, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL) or wireless (e.g., infrared, wireless, microwave, etc.)). The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state hard disk Solid State Disk (SSD)), etc.

[0058] Data source statement: The use of all data complies with the "Data Security Law of the People's Republic of China", "Personal Information Protection Law" and relevant international regulations. We hereby declare.

[0059] The above-mentioned embodiments only express the specific implementation of the present invention, and the description thereof is relatively specific and detailed, but it cannot be understood as limiting the scope of the present invention. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the present invention, which all belong to the protection scope of the present invention.

Claims

1. A lightweight drug detection method based on facial images on a mobile terminal, characterized in that: The following steps are involved: S1. Use a camera to collect a face image of the person being detected, and perform preprocessing operations such as face detection, background removal, and face alignment; S2. Design a lightweight neural network module and build a lightweight drug detection network based on the module; S3, obtain the face data set of drug addicts and train the drug detection network; S4, using the drug use detection network to extract facial features and obtain classification results and confidence; S5. Convert and transplant the drug detection network model to the mobile terminal, build a mobile application, and display the results.

2. The lightweight drug detection method based on facial images on a mobile terminal as claimed in claim 1, characterized in that: The face detection operation in step S1 is performed using a cascaded MTCNN face detector; The face detector is composed of sub-networks PNet, RNet and ONet. The face image is scaled into multiple images and input into the cascade network one by one to obtain the face detection frame; Calibrate the detection frame based on its top, bottom, left, and right coordinates. Finally, the calibrated face image is cropped to obtain a face image with the background removed.

3. The lightweight drug detection method based on facial images on a mobile terminal as claimed in claim 1, characterized in that: The face alignment operation in step S1 is implemented using affine transformation, and three face key points are used for alignment processing; Define the key points of the eyes, eyebrows, nose, mouth and edge contours, then use an automatic algorithm to extract the facial key points, and finally align the face by constructing an affine matrix.

4. The lightweight drug detection method based on facial images on a mobile terminal as claimed in claim 1, characterized in that: The lightweight neural network module in step S2 uses matrix decomposition and convolution kernel decomposition technology to reduce the number of network module parameters and the running speed, and builds a complete lightweight drug detection network based on the lightweight network module.

5. The lightweight drug detection method based on facial images on a mobile terminal as claimed in claim 4, characterized in that: The lightweight neural network module is constructed by using group convolution and separable convolution, and the idea of ​​first reducing the dimension and then performing feature processing; First, the grouped convolution method is used to divide the input features into multiple groups according to the path, and convolution operations are performed on each group separately; Secondly, the matrix decomposition of ordinary convolution is replaced by depth-wise separable convolution; The 1×1 convolution operation and the depthwise convolution operation of the depthwise separable convolution are replaced in order again to achieve feature dimensionality reduction before feature extraction, thus realizing the design of a lightweight model.

6. The lightweight drug detection method based on facial images on a mobile terminal as claimed in claim 4, characterized in that: The network module parameter amount is reduced by decomposing the convolution kernel into depth-separable convolutions for feature extraction through matrix decomposition, and using grouped convolution to achieve it. In addition, the network module parameter amount is further reduced by first reducing the dimension of the feature and then performing convolution.

7. The lightweight drug detection method based on facial images on a mobile terminal as claimed in claim 4, characterized in that: The lightweight drug detection network is based on the MobileNet network architecture and is built in combination with a lightweight neural network module (LMDI); The drug detection network uses LMDI as a feature extraction module, and integrates this lightweight module into the MobileNet network structure to achieve network construction.

8. The lightweight drug detection method based on facial images on a mobile terminal as claimed in claim 1, characterized in that: The neural network training in step S3 is trained using the FocalLoss loss function, a binary classification fully connected layer is used to classify whether the person has taken drugs or not, and the performance of the network is evaluated using two indicators, accuracy ACC and F1 score.

9. The lightweight drug detection method based on facial images on a mobile terminal as claimed in claim 1, characterized in that: The mobile application method in step S5 is to convert the neural network model trained on the PC side and build a mobile application to realize the transplantation and deployment of the model.

10. The lightweight drug detection method based on facial images on a mobile terminal as claimed in claim 6, characterized in that: The model conversion is performed using onnx and ncnn methods; Use the onnx method to convert the model to an .onnx file and test the usability of the model; Use ncnn and convert the intermediate file .onnx to .ncnn file to get .param and .bin files; Transplant and deploy the generated .param and .bin files to the mobile terminal and run them.

Citation Information

Patent Citations

  • Method of identifying drug addict in pupil identification mode though mobile phone

    CN103186765A

  • Drug addict screening method based on IPPG

    CN111870235A

  • Pupil size-based high-precision non-contact virus-related detection method and detection system

    CN112603256A

  • Face recognition method for drug addicts

    CN118506413A

Cited By

  • Vision health monitoring method based on shooting, tracking and comparison of mobile terminal equipment

    CN120599201A