Training of User Behavior Recognition Model, Facial Recognition Method, System and Device
By using the training method of user behavior recognition model in facial recognition technology, the multi-dimensional features of timing sensor data are extracted using the multi-head attention mechanism and interactive attention mechanism, and the problem of insufficient facial recognition accuracy in the prior art is solved, especially on mobile devices, and higher facial recognition accuracy is achieved.
Patent Information
- Application Number
- CN202111286929.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-02
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2041-11-02
AI Technical Summary
There are shortcomings in the accuracy of existing facial recognition technology, especially in mobile devices that lack multiple image sensors. There are limitations on the acquisition and processing of depth sensor data, which affects the accuracy of facial recognition.
A training method for user behavior recognition model is proposed. By obtaining labelless and labeled timing sensor data, preprocessing and retraining, the user behavior recognition model is generated. This model uses multi-head attention mechanism and interactive attention mechanism to extract multi-dimensional features of timing sensor data to improve the accuracy of behavioral category recognition.
By combining the multi-dimensional characteristics of timing sensor data, the accuracy of facial recognition model understanding of user behavior is improved and the accuracy of facial recognition is improved, especially in the absence of multiple image sensors.
Smart Images

Figure CN114220136B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of face recognition technology, and specifically relates to a method for training a user behavior recognition model, a method for face recognition, a system for face recognition, a device for training a user behavior recognition model, a device for face recognition, an electronic device, and a computer storage medium. Background Art
[0002] In recent years, with the development of deep learning technology, face recognition systems have been widely used in fields such as finance, security, transportation, and education. While face recognition systems are widely popularized, many security problems have emerged. For example, the privacy security, transmission security, and storage security of face data, as well as various prosthetic attacks on face recognition systems. In particular, due to the manufacture of facial prosthetics of legitimate users through various media such as electronic photos, printed photos, and recorded videos to attack face recognition systems, prosthetic attacks pose higher requirements for the security guarantee of face recognition, which also forces the face recognition technology to be further improved.
[0003] Existing face recognition technologies mostly use multiple optical sensing cameras to extract the essential difference features between real and fake faces, so as to better cope with the threats of various presentation attack means to face recognition systems. The sensors currently used in face recognition methods mainly include: visible light cameras, near-infrared cameras, depth cameras, thermal cameras, and multispectral cameras. These sensors can capture or enhance certain human physiological information, facial texture information, and geometric shape information, etc., which are key features for face recognition methods.
[0004] However, most mobile devices do not have multiple image sensors. Therefore, it becomes less practical to use multi-view image sensors for anti-counterfeiting detection, which affects the accuracy of face recognition. Moreover, the acquisition and processing of depth sensor data have limitations, which is also not conducive to improving the accuracy of face recognition.
[0005] Therefore, how to improve the accuracy of face recognition has become an urgent problem to be solved by those skilled in the art. Summary of the Invention
[0006] Embodiments of this application provide a method for training a user behavior recognition model to solve the problem of how to improve the accuracy of face recognition in the prior art. Embodiments of this application also provide a method for face recognition, a system for face recognition, a device for training a user behavior recognition model, a device for face recognition, an electronic device, and a computer storage medium.
[0007] Embodiments of this application provide a method for training a user behavior recognition model, including:
[0008] Obtain the data of the time-series sensor without labels as the original data;
[0009] Preprocess the original data to obtain time-domain data, frequency-domain data, and statistic data;
[0010] Input the time-domain data, frequency-domain data, and statistic data into the model respectively for pre-training to obtain a pre-trained model;
[0011] Obtain the time-domain data of the time-series sensor with labels, and obtain the corresponding frequency-domain data and statistic data;
[0012] Input the time-domain data, frequency-domain data, and statistic data of the time-series sensor with labels into the pre-trained model respectively for re-training to obtain the user behavior recognition model.
[0013] Optionally, the step of inputting the time-domain data, frequency-domain data, and statistic data into the model respectively for pre-training to obtain a pre-trained model includes:
[0014] Extract the time-domain features, frequency-domain features, and statistic features corresponding to the time-domain data, frequency-domain data, and statistic data;
[0015] Aggregate the time-domain features, frequency-domain features, and statistic features as training data.
[0016] Optionally, after extracting the time-domain features, frequency-domain features, and statistic features corresponding to the time-domain data, frequency-domain data, and statistic data, it further includes:
[0017] Obtain the similarity features and difference features among the time-domain features, frequency-domain features, and statistic features;
[0018] Aggregate the similarity features and difference features as training data.
[0019] Optionally, the step of inputting the time-domain data, frequency-domain data, and statistic data into the model respectively for pre-training to obtain a pre-trained model includes:
[0020] Input the time-domain data, frequency-domain data, and statistic data into different processing units of the model in parallel.
[0021] Optionally, the model is a transformer model, and the processing unit is the encoder of the transformer.
[0022] Optionally, before inputting the time-domain data, frequency-domain data, and statistic data into the model respectively, it further includes: preprocessing the time-domain data, its frequency-domain data, and statistic data;
[0023] Correspondingly, before inputting the time-domain data, frequency-domain data, and statistical data of the tagged time-series sensor into the pre-trained model respectively, it further includes: preprocessing the time-domain data, frequency-domain data, and statistical data of the tagged time-series sensor.
[0024] Optionally, the data of the untagged time-series sensor and the time-domain data of the tagged time-series sensor respectively include timestamp information.
[0025] An embodiment of the present application further provides a method for face recognition, including:
[0026] Obtaining the face information of the object to be recognized through an image sensor, where the face information at least includes the motion information of the face or a part of the face;
[0027] Inputting the face information into an image recognition model to obtain a first recognition result including the identity information of the object to be recognized and whether it is a live body;
[0028] While obtaining the face information of the object to be recognized, monitoring the object to be recognized through a time-series sensor to obtain the monitoring data of the time-series sensor;
[0029] Inputting the monitoring data into the user behavior recognition model generated by the above-mentioned training method of the user behavior recognition model to obtain a second recognition result including the behavior category of the object to be recognized;
[0030] Obtaining the identity information of the object to be recognized and a third recognition result for determining whether the object to be recognized is a live body according to the first recognition result and the second recognition result.
[0031] An embodiment of the present application further provides a system for face recognition, including: an image recognition model, a time-series sensor, a decision model, and a user behavior recognition model;
[0032] The image sensor is used to obtain the face information of the object to be recognized and send the face information of the object to be recognized to the image recognition model; the face information at least includes the motion information of the face or a part of the face;
[0033] The image recognition model is used to receive the face information of the object to be recognized, obtain a first recognition result including the identity information of the object to be recognized and whether it is a live body according to the face information of the object to be recognized, and send the first recognition result including the identity information of the object to be recognized and whether it is a live body to the decision model;
[0034] The timing sensor is used to obtain the monitoring data of the timing sensor by monitoring the object to be recognized while obtaining the facial information of the object to be recognized, and input the monitoring data of the timing sensor into the user behavior recognition model;
[0035] The user behavior recognition model is used to receive the monitoring data of the timing sensor, obtain a second recognition result including the behavior category of the object to be recognized according to the monitoring data, and send the second recognition result to the decision model;
[0036] The decision model is used to receive the first recognition result and the second recognition result, and obtain the identity information of the object to be recognized and a third recognition result for determining whether the object to be recognized is a living body according to the first recognition result and the second recognition result.
[0037] The embodiment of the present application further provides a training device for a user behavior recognition model, including:
[0038] An unlabeled data acquisition unit, configured to acquire the data of the timing sensor without labels as the original data;
[0039] A preprocessing unit, configured to preprocess the original data to obtain time-domain data, frequency-domain data, and statistic data;
[0040] A pre-training unit, configured to input the time-domain data, frequency-domain data, and statistic data into the model respectively for pre-training to obtain a pre-trained model;
[0041] A labeled data acquisition unit, configured to acquire the time-domain data of the timing sensor with labels, and obtain the corresponding frequency-domain data and statistic data;
[0042] A re-training unit, configured to input the time-domain data, frequency-domain data, and statistic data of the timing sensor with labels into the pre-trained model respectively for re-training to obtain the user behavior recognition model.
[0043] The embodiment of the present application further provides a device for face recognition, including:
[0044] A facial information acquisition unit, configured to acquire the facial information of the object to be recognized through an image sensor, where the facial information at least includes the motion information of the face or a part of the face;
[0045] A first recognition result acquisition unit, configured to input the facial information into an image recognition model to obtain a first recognition result including the identity information of the object to be recognized and whether it is a living body;
[0046] A monitoring data acquisition unit is configured to monitor the object to be recognized through a time series sensor while acquiring the facial information of the object to be recognized, so as to obtain the monitoring data of the time series sensor;
[0047] A second recognition result acquisition unit is configured to input the monitoring data into the user behavior recognition model generated by the training method of the user behavior recognition model described above, so as to obtain a second recognition result including the behavior category of the object to be recognized;
[0048] A recognition result determination unit is configured to obtain the identity information of the object to be recognized and a third recognition result for determining whether the object to be recognized is a live body according to the first recognition result and the second recognition result.
[0049] An embodiment of the present application further provides an electronic device, where the electronic device includes: a processor; a memory for storing a computer program, and the computer program is run by the processor to execute the method described above.
[0050] An embodiment of the present application further provides a computer storage medium, where the computer storage medium stores a computer program, and the computer program is run by the processor to execute the method described above.
[0051] Compared with the prior art, the present application has the following advantages:
[0052] An embodiment of the present application provides a training method for a user behavior recognition model, including: acquiring data of a time series sensor without labels as original data; preprocessing the original data to obtain time domain data, frequency domain data, and statistic data; respectively inputting the time domain data, frequency domain data, and statistic data into a model for pre-training to obtain a pre-trained model; acquiring time domain data of a time series sensor with labels, and obtaining corresponding frequency domain data and statistic data; respectively inputting the time domain data, frequency domain data, and statistic data of the time series sensor with labels into the pre-trained model for re-training to obtain the user behavior recognition model. In the embodiment of the present application, the data of the time series sensor is input into the model. Based on the multi-head attention mechanism of the model, when extracting the features of the data of the time series sensor, problems such as the long-term dependence of the traditional neural network and its inability to capture long-distance features due to its own sequential attributes can be solved, and at the same time, the extraction time of the features of the data of the time series sensor is shortened. In addition, based on the interactive attention mechanism of the model, multi-dimensional features of the data of the time series sensor can be interacted, that is, multi-dimensional feature extraction, fusing time domain features, frequency domain features, and statistic features, which can not only incorporate prior knowledge but also automatically extract multi-dimensional complementary features through the model, so that the obtained features can be expressed more comprehensively, thereby improving the understanding accuracy of the model for the data of the time series sensor, and further improving the accuracy of behavior category recognition.
[0053] An embodiment of the present application provides a method for face recognition. The face information of an object to be recognized is obtained through an image sensor, and the face information at least includes the motion information of the face or a part of the face. The face information is input into an image recognition model to obtain a first recognition result including the identity information of the object to be recognized and whether it is a live body. While obtaining the face information of the object to be recognized, the object to be recognized is monitored through a time series sensor to obtain the monitoring data of the time series sensor. The monitoring data is input into a user behavior recognition model to obtain a second recognition result including the behavior category of the object to be recognized. A third recognition result of the identity information of the object to be recognized and determining whether the object to be recognized is a live body is obtained according to the first recognition result and the second recognition result. In the embodiment of the present application, the face information of the user is obtained through the image sensor, and the first recognition result is correspondingly obtained. While monitoring the face information, the monitoring data of the time series sensor is collected, and the monitoring data of the time series sensor is input into the user behavior recognition model to obtain the second recognition result. Thus, the identity information of the object to be recognized and whether the object to be recognized is a live body are obtained according to the first recognition result and the second recognition result. Based on the monitoring data of the time series sensor during normal user authentication, there are obvious differences from the time series sensor information during a prosthesis attack system. And relying on the second recognition result output by the user behavior recognition model corresponding to the monitoring data of the time series sensor and combining it with the first recognition result, the accuracy of face recognition is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 is a schematic diagram of an application scenario provided by the present application.
[0055] Figure 2 is a flowchart of a training method for a user behavior recognition model provided by the first embodiment of the present application.
[0056] Figure 3 is a schematic diagram of a user behavior recognition model provided by the first embodiment of the present application.
[0057] Figure 4 is a flowchart of a method for face recognition provided by the second embodiment of the present application.
[0058] Figure 5 is a schematic diagram of a face recognition system provided by the third embodiment of the present application.
[0059] Figure 6 is a schematic diagram of a training device for a user behavior recognition model provided by the fourth embodiment of the present application.
[0060] Figure 7 is a schematic diagram of a device for face recognition provided by the fifth embodiment of the present application.
[0061] Figure 8Schematic diagram of the electronic device provided in the sixth embodiment of the present application. Detailed implementation manners
[0062] In the following description, many specific details are set forth in order to provide a thorough understanding of the embodiments of the present application. However, the embodiments of the present application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the connotation of the embodiments of the present application. Therefore, the embodiments of the present application are not limited by the specific implementations disclosed below.
[0063] In order to enable those skilled in the art to better understand the solution of the present application, the following describes in detail the specific application scenarios of the embodiments based on the method for face recognition provided by the present application. As Figure 1 shown, it is a schematic diagram of the application scenario provided by the present application.
[0064] This application scenario is a scenario where a user unlocks a terminal. The terminal can be various types of devices such as a mobile phone, a computer, a tablet computer, etc. When the user needs to unlock the terminal, the user places the terminal in a suitable position by hand, and the image sensor of the terminal captures the user's face and obtains the user's face information, which at least includes the motion information of the face or part of the face. Among them, the motion information of the face refers to all the motion information of the face that the user makes according to the instructions of the terminal, and the motion information refers to the information of the face making actions. For example, making facial expressions, shaking the head, etc. The motion information of part of the face refers to the motion information of part of the face that the user makes according to the instructions of the terminal, and the motion information refers to the information of the face making actions. For example, blinking, opening the mouth, etc.
[0065] After obtaining the user's face information, a first recognition result including the identity information of the object to be recognized and whether it is a live body corresponding to the user's face information can be obtained through an image recognition model. Among them, the identity information of the object to be recognized is used to represent the identity of the object to be recognized, for example, it is XX user. It should be noted that the first recognition result corresponding to the user's face information can determine whether the object to be recognized is a live body under the corresponding detection conditions. For example, the detection condition is set to detect a picture. When the terminal instructs the object to be recognized to make corresponding actions such as shaking the head and smiling, the object to be detected in the picture cannot complete them, and then it can be determined whether the object to be recognized is a live body through this first recognition result.
[0066] In this scenario, while the image sensor of the terminal acquires the facial information of the user, the timing sensor of the terminal also detects the user's face and obtains the monitoring data of the timing sensor accordingly. Among them, the monitoring data of the timing sensor at least includes gyroscope monitoring data and linear accelerator monitoring data. The monitoring data of the timing sensor is the data monitored correspondingly when the user controls the terminal to move with the hand. For example, when the user uses the terminal to take a picture of the face to obtain the facial information, the user's hand shakes, so that the monitoring data of the timing sensor is the data monitored when the terminal moves.
[0067] After obtaining the monitoring data of the timing sensor, the monitoring data of the timing sensor can be input into the user behavior recognition model, so that the user behavior recognition model obtains a second recognition result including the behavior category of the object to be recognized according to the monitoring data of the timing sensor. Among them, the second recognition result obtained according to the user behavior recognition model is coordinated with the first recognition result. It can be a supplement to the first recognition result or assist the image recognition model to further determine whether the object to be recognized is a live body. For example, when verifying in the form of video recording, there are actions such as shaking the head and smiling in the video recording. At this time, the first recognition result cannot correctly determine whether the object to be recognized is a live body. However, the behavior classification information determined by the second recognition result has a specific relationship and specific law with the user's actions such as shaking the head and smiling. Therefore, while obtaining the identity information of the object to be recognized, the recognition result of whether the object to be recognized is a live body can be obtained. That is, the identity information of the object to be recognized and the recognition result of whether the object to be recognized is a live body are obtained according to the first recognition result and the second recognition result. After obtaining the identity information of the object to be recognized and the recognition result of whether the object to be recognized is a live body, the obtained facial recognition result of the object to be recognized is sent to the terminal.
[0068] Among them, the user behavior recognition model is obtained through training in the following way: acquiring the data of the timing sensor without labels as the original data; preprocessing the original data to obtain time-domain data, frequency-domain data, and statistic data; inputting the time-domain data, frequency-domain data, and statistic data into the model respectively for pre-training to obtain a pre-trained model. Then, acquiring the time-domain data of the timing sensor with labels and obtaining the corresponding frequency-domain data and statistic data, and inputting the time-domain data, frequency-domain data, and statistic data of the timing sensor with labels into the pre-trained model respectively for re-training to obtain the user behavior recognition model. For the further training process of the user behavior recognition model, refer to the specific description of the subsequent embodiments.
[0069] Input the monitoring data of the time series sensor into the user behavior recognition model. Based on the multi-head attention mechanism of the transformer encoding layer of the user behavior recognition model, when extracting the features of the monitoring data, problems such as the long-term dependence of traditional neural networks and the inability to capture long-distance features due to their own sequential attributes can be solved, and at the same time, the extraction time of the features of the monitoring data is shortened. In addition, based on the interactive attention mechanism of the user behavior recognition model, the multi-dimensional features of the monitoring data can be interacted, thereby improving the understanding accuracy of the user behavior recognition model for the monitoring data.
[0070] In this scenario, the facial information of the user is obtained through an image sensor, and a first recognition result including the identity information of the object to be recognized and whether it is a live body is correspondingly obtained. At the same time, the monitoring data corresponding to the facial information monitored by the time series sensor is obtained, and the monitoring data of the time series sensor is input into the user behavior recognition model to obtain a second recognition result. Thus, the identity information of the object to be recognized and the recognition result of whether the object to be recognized is a live body are obtained according to the first recognition result and the second recognition result. The monitoring data of the time series sensor based on normal user authentication is significantly different from the time series sensor information during a prosthetic attack system. By relying on the second recognition result corresponding to the monitoring data of the time series sensor output by the user behavior recognition model and combining it with the first recognition result, the accuracy of facial recognition is improved.
[0071] It should be noted that the specific limitation on the application scenario of the method for facial recognition in the embodiments of the present application is only one embodiment of the application scenario of the method for facial recognition provided by the present application. The purpose of providing this application scenario embodiment is to facilitate the understanding of the method for facial recognition provided by the present application, rather than to limit the method for facial recognition provided by the present application. The embodiments of the present application have other application scenarios for facial recognition, which will not be elaborated here one by one.
[0072] Corresponding to the above scenario, the first embodiment of the present application provides a training method for a user behavior recognition model, as Figure 2 shown, Figure 2 is a flowchart of a training method for a user behavior recognition model provided by the first embodiment of the present application. The method includes the following steps:
[0073] Step S201, obtain the data of the time series sensor without labels as the original data.
[0074] In this step, the data of the untagged time-series sensor can be obtained from the database or from the historical data obtained by the time-series sensor. There are many ways to obtain the data of the untagged time-series sensor, and the first embodiment of this application does not make specific restrictions on this. The data of the untagged time-series sensor includes timestamp information. The timestamp information is used to inform the model when each data is input. Based on the large amount of data of the untagged time-series sensor, the model can be trained using a large-scale data, which can effectively prevent the problem of model overfitting caused by a small amount of data.
[0075] Step S202: Preprocess the original data to obtain time-domain data, frequency-domain data, and statistic data.
[0076] After obtaining the data of the untagged time-series sensor, the data of the untagged time-series sensor can be used as the original data, and the original data is preprocessed to obtain frequency-domain data and statistic data.
[0077] Among them, preprocessing the original data to obtain frequency-domain data includes: transforming the original data to obtain the corresponding frequency-domain data. The way to transform the original data can be Fourier transform or wavelet transform, etc.
[0078] Among them, preprocessing the original data to obtain statistic data includes: obtaining statistic data from the original data according to the preset expert experience. The statistic data includes first- and second-order statistics such as extreme values, means, variances, and higher-order statistics. Compared with the existing algorithms that only use time-domain data to obtain the features of the sensor, this application combines frequency-domain data and statistic data to describe the features of the sensor. The frequency-domain data describes the frequency and amplitude of the data itself from another angle, while the statistic data can integrate expert experience and extract features with prior knowledge through the calculation of statistics, such as first- and second-order statistics such as extreme values, means, variances, and higher-order statistics, etc., so that the behavior features of the sensor can be described from more dimensions.
[0079] Step S203: Input the time-domain data, frequency-domain data, and statistic data into the model respectively for pre-training to obtain a pre-trained model.
[0080] In this step, the model can be a Transformer model. After preprocessing the original data to obtain frequency-domain data and statistical data, it is necessary to preprocess the time-domain data, its frequency-domain data, and statistical data. The preprocessing is, for example, normalization processing, and it is embedded as a one-dimensional token by means of linear mapping, etc. The preprocessing result is used as the input of the Transformer model. The Transformer model includes an encoding component and a decoding component. The encoding component can include multiple parallel processing units, that is, multiple encoders, and each encoder can include multiple layers of Transformer blocks. In this embodiment, the step of separately inputting the time-domain data, frequency-domain data, and statistical data into the model includes: separately and parallelly inputting the preprocessing results of the time-domain data, frequency-domain data, and statistical data into three different encoders of the model.
[0081] Specifically, as Figure 3 shown, for example, input the preprocessed time-domain data, frequency-domain data, and statistical data into Encoder 1, Encoder 2, and Encoder 3 respectively. In the aforementioned three encoders, each encoder can include N stacked Transformer blocks, such as 6 layers. Each Transformer block can include a multi-domain fusion attention layer and a feed-forward neural network layer. In each encoder, the data vectors of the aforementioned time-domain data, frequency-domain data, and statistical data are respectively defined as Q (query vector), K (key vector), and V (value vector). After being input into the multi-layer multi-domain fusion attention layer and the feed-forward neural network layer, a higher-dimensional feature representation containing the differences of the three domains is output. In the aforementioned three encoders, the Q, K, and V vectors respectively defined by the time-domain data, frequency-domain data, and statistical data are different. For example, when input into the first encoder, the time-domain data is defined as the V vector, the frequency-domain data is defined as the Q vector, and the number of statistical items is defined as the K vector. The data in the aforementioned three domains are defined as different vectors, and features are extracted through the three encoders respectively.
[0082] After obtaining the higher-dimensional features, the feature representations of each encoder are fused and then input into the decoder. The decoder includes multiple multi-head attention layer-feed-forward neural network layers. In the decoder, a certain number of elements can be randomly masked, and the model is allowed to predict the masked elements, so as to train the model's prediction ability.
[0083] It should be noted that the amount of data of the untagged time series sensors is large. Using the data of a large number of untagged time series sensors for self-supervised pre-training can enable the model to obtain the implicit information of a larger amount of data to prevent the model from overfitting. In addition, the essence of the multi-head attention mechanism is the calculation of multiple independent self-attention mechanisms, and the final concatenation serves as an integration function, which also prevents overfitting to a certain extent.
[0084] In this embodiment, through the above method, the time-domain data, frequency-domain data, and statistic data are respectively input into the model for pre-training to obtain a pre-trained model, including: extracting the time-domain features, frequency-domain features, and statistic features corresponding to the time-domain data, frequency-domain data, and statistic data; aggregating the time-domain features, frequency-domain features, and statistic features as training data. In this embodiment, after extracting the time-domain features, frequency-domain features, and statistic features corresponding to the time-domain data, frequency-domain data, and statistic data, it further includes: obtaining the similarity features and difference features between the time-domain features, frequency-domain features, and statistic features; aggregating the similarity features and difference features as training data.
[0085] In this embodiment, the transformer attention mechanism is used to process the time-domain data, frequency-domain data, and statistic data, calculate the similarity features between the time-domain data, frequency-domain data, and statistic data, and further obtain the difference features between the time-domain data, frequency-domain data, and statistic data, solving the problem of ignoring the interaction between dimensions in the direct fusion of multi-dimensional features, improving the model's understanding and generalization ability of the data semantics of time series sensors, and thus enhancing the classification ability of the model.
[0086] Step S204, obtain the time-domain data of the tagged time series sensors, and obtain the corresponding frequency-domain data and statistic data.
[0087] In this step, the time-domain data of the tagged time series sensors can be obtained through a database or from the historical data obtained by the time series sensors. There are many ways to obtain the time-domain data of the tagged time series sensors, and the first embodiment of this application does not make specific limitations on this. The time-domain data of the tagged time series sensors includes timestamp information. The timestamp information is used to inform the model when each data is input and the input order of each data.
[0088] After obtaining the time-domain data of the tagged time series sensors, the time-domain data of the tagged time series sensors can be preprocessed to obtain the frequency-domain data and statistic data.
[0089] Among them, the time-domain data of the tagged time-series sensor is preprocessed to obtain frequency-domain data, including: transforming the time-domain data of the tagged time-series sensor to obtain the corresponding frequency-domain data. The method of transforming the time-domain data of the tagged time-series sensor can be Fourier transform or wavelet transform, etc.
[0090] Among them, the time-domain data of the tagged time-series sensor is preprocessed to obtain statistic data, including: obtaining statistic data from the time-domain data of the tagged time-series sensor according to the preset expert experience. The statistic data includes first- and second-order statistics such as extreme values, means, variances, and higher-order statistics. Compared with the existing algorithms that only use time-domain data to obtain the features of the sensor, this application combines frequency-domain data and statistic data to describe the features of the sensor. The frequency-domain data describes the frequency and amplitude of the data itself from another perspective, while the statistic data can integrate expert experience and extract features with prior knowledge through the calculation of statistics, such as first- and second-order statistics like extreme values, means, variances, and higher-order statistics, etc., so as to describe the behavioral characteristics of the sensor from more dimensions.
[0091] Step S205: Input the time-domain data, frequency-domain data, and statistic data of the tagged time-series sensor into the pre-trained model respectively for retraining to obtain the user behavior recognition model.
[0092] After preprocessing the time-domain data of the tagged time-series sensor to obtain frequency-domain data and statistic data, it is necessary to preprocess the time-domain data, frequency-domain data, and statistic data of the tagged time-series sensor. The preprocessing is, for example, normalization processing, and it is embedded as a one-dimensional token by means of linear mapping, etc. The preprocessing result is used as the input of the Transformer model. In this step, the model can be a Transformer model. The Transformer model includes an encoding component and a decoding component. The encoding component can include multiple parallel processing units, that is, multiple encoders. Each encoder can include multiple layers of Transformer blocks. In this embodiment, the inputting the time-domain data, frequency-domain data, and statistic data into the pre-trained model respectively includes: inputting the preprocessing results of the time-domain data, frequency-domain data, and statistic data into three different encoders of the pre-trained model in parallel.
[0093] Specifically, such as Figure 3As shown, for example, the preprocessed time-domain data, frequency-domain data, and statistical data are respectively input into the encoder 1, encoder 2, and encoder 3. In the above three encoders, each encoder may include N stacked transformer blocks, such as 6 layers. Each transformer block may include a multi-domain fusion attention layer and a feed-forward neural network layer. In each encoder, the data vectors of the above time-domain data, frequency-domain data, and statistical data are respectively defined as Q (query vector), K (key vector), and V (value vector). After being input into the multi-layer multi-domain fusion attention layer and the feed-forward neural network layer, a higher-dimensional feature representation containing three-domain difference features is output. In the above three encoders, the Q, K, and V vectors respectively defined for the time-domain data, frequency-domain data, and statistical data are different. For example, when input into the first encoder, the time-domain data is defined as the V vector, the frequency-domain data is defined as the Q vector, and the statistical item number is defined as the K vector. The data in the above three domains are defined as different vectors and the features are extracted through three encoders respectively.
[0094] After obtaining the higher-dimensional features, the features of each encoder are input into the decoder after representation fusion. The decoder includes multiple multi-head attention layer-feed-forward neural network layers. In the decoder, a certain number of elements can be randomly masked, and the model is allowed to predict the masked elements to train the model's prediction ability.
[0095] It should be noted that the essence of the multi-head attention mechanism is the calculation of multiple independent self-attention mechanisms. The final splicing serves as an integrated role and also prevents overfitting to a certain extent.
[0096] In this embodiment, through the above method, the time-domain data, frequency-domain data, and statistical data are respectively input into the pre-trained model for retraining to obtain a user behavior recognition model, including: extracting time-domain features, frequency-domain features, and statistical features corresponding to the time-domain data, frequency-domain data, and statistical data; aggregating the time-domain features, frequency-domain features, and statistical features as training data. In this embodiment, after extracting the time-domain features, frequency-domain features, and statistical features corresponding to the time-domain data, frequency-domain data, and statistical data, it further includes: obtaining the similarity features and difference features between the time-domain features, frequency-domain features, and statistical features, and aggregating the similarity features and difference features as training data.
[0097] In this embodiment, the Transformer attention mechanism is used to process time-domain data, frequency-domain data, and statistical data, calculate the similarity features between the time-domain data, frequency-domain data, and statistical data, and further obtain the difference features between the time-domain data, frequency-domain data, and statistical data, solving the problem of ignoring the interaction between dimensions in direct fusion of multi-dimensional features, improving the model's understanding and generalization ability of the data semantics of time-series sensors, and thus enhancing the classification ability of the model.
[0098] The first embodiment of this application provides a training method for a user behavior recognition model, including: obtaining the data of a time-series sensor without labels as the original data; preprocessing the original data to obtain time-domain data, frequency-domain data, and statistical data; respectively inputting the time-domain data, frequency-domain data, and statistical data into the model for pre-training to obtain a pre-trained model; obtaining the time-domain data of a time-series sensor with labels, and obtaining the corresponding frequency-domain data and statistical data; respectively inputting the time-domain data, frequency-domain data, and statistical data of the time-series sensor with labels into the pre-trained model for re-training to obtain the user behavior recognition model. In the first embodiment of this application, the data of the time-series sensor is input into the model. Based on the multi-head attention mechanism of the model, when extracting the features of the data of the time-series sensor, problems such as the long-term dependence of traditional neural networks and the inability to capture long-distance features due to their own sequential attributes can be solved, and at the same time, the extraction time of the features of the data of the time-series sensor is shortened. In addition, based on the interactive attention mechanism of the model, the multi-dimensional features of the data of the time-series sensor can be interacted, that is, multi-dimensional feature extraction, integrating time-domain features, frequency-domain features, and statistical features, which can not only incorporate prior knowledge but also automatically extract multi-dimensional complementary features through the model, so that the obtained features can be expressed more comprehensively, thereby improving the understanding accuracy of the model for the data of the time-series sensor, and further improving the accuracy of behavior category recognition.
[0099] The second embodiment of this application provides a method for face recognition, as Figure 4 shown Figure 4 is a flowchart of a method for face recognition provided by the second embodiment of this application. The method includes the following steps:
[0100] Step S301, obtaining the facial information of the object to be recognized through an image sensor, where the facial information includes at least the movement information of the face or a part of the face.
[0101] In this step, an image sensor is set in a terminal, which can be various types of devices such as a mobile phone, a computer, a tablet, etc. The image sensor of the terminal captures the face of the object to be recognized and obtains the face information of the object to be recognized, which at least includes the motion information of the face or a part of the face. Among them, the motion information of the face refers to all the motion information of the face of the object to be recognized according to the instructions of the terminal, and the motion information refers to the information of the face making actions. For example, making facial expressions, shaking the head, etc. The motion information of a part of the face refers to the motion information of a part of the face of the object to be recognized according to the instructions of the terminal, and the motion information refers to the information of the face making actions. For example, blinking, opening the mouth, etc.
[0102] Step S302: Input the face information into an image recognition model to obtain a first recognition result including the identity information of the object to be recognized.
[0103] After obtaining the face information of the user, a first recognition result including the identity information of the object to be recognized and whether it is a live body corresponding to the face information of the user can be obtained through the image recognition model. Among them, the identity information of the object to be recognized is used to represent the identity of the object to be recognized. For example, it is XX user. Among them, this first recognition result corresponding to the face information of the user can determine whether the object to be recognized is a live body under the corresponding detection conditions. For example, the detection condition is set to detect a picture. When the terminal instructs the object to be recognized to make corresponding actions such as shaking the head and smiling, the object to be detected in the picture cannot complete them, and then it can be determined whether the object to be recognized is a live body through this first recognition result.
[0104] Step S303: While obtaining the face information of the object to be recognized, monitor the object to be recognized through a timing sensor to obtain the monitoring data of the timing sensor.
[0105] While the image sensor of the terminal obtains the face information of the object to be recognized, the timing sensor of the terminal also detects the face of the object to be recognized and correspondingly obtains the monitoring data of the timing sensor. Among them, the monitoring data of the timing sensor at least includes gyroscope monitoring data and linear accelerator monitoring data. The monitoring data of the timing sensor is the data corresponding to the monitoring when the object to be recognized controls the terminal to move. For example, when the object to be recognized uses the terminal to capture the face to obtain face information, there is a phenomenon of shaking of the hand of the object to be recognized, so that the monitoring data of the timing sensor is the data corresponding to the monitoring when the terminal moves.
[0106] Step S304: Input the monitoring data into a user behavior recognition model to obtain a second recognition result including the behavior category of the object to be recognized.
[0107] After obtaining the monitoring data of the time series sensor, the monitoring data of the time series sensor can be input into the user behavior recognition model, so that the user behavior recognition model can obtain a second recognition result including the behavior category of the object to be recognized according to the monitoring data of the time series sensor. Among them, the second recognition result obtained according to the user behavior recognition model is coordinated with the first recognition result, which can be a supplement to the first recognition result or assist the image recognition model to further determine whether the object to be recognized is a live body. For example, when verifying in the form of a video recording, there are actions such as shaking the head and smiling in the video recording. At this time, the first recognition result cannot correctly determine whether the object to be recognized is a live body. However, the behavior classification information determined by the second recognition result has a specific relationship and specific law with the user's actions such as shaking the head and smiling, so that it can correctly determine whether the object to be recognized is a live body. As in step S305.
[0108] Step S305, obtain the identity information of the object to be recognized and a third recognition result for determining whether the object to be recognized is a live body according to the first recognition result and the second recognition result.
[0109] In this step, by combining the first recognition result and the second recognition result, the identity information of the object to be recognized and whether the object to be recognized is a live body can be correspondingly obtained, thereby improving the recognition accuracy.
[0110] The second embodiment of the present application provides a method for face recognition. The face information of the object to be recognized is obtained through an image sensor, and the face information at least includes the motion information of the face or a part of the face; the face information is input into the image recognition model to obtain a first recognition result including the identity information of the object to be recognized and whether it is a live body; while obtaining the face information of the object to be recognized, the object to be recognized is monitored through a time series sensor to obtain the monitoring data of the time series sensor; the monitoring data is input into the user behavior recognition model to obtain a second recognition result including the behavior category of the object to be recognized; the identity information of the object to be recognized and a third recognition result for determining whether the object to be recognized is a live body are obtained according to the first recognition result and the second recognition result. In the embodiment of the present application, the face information of the user is obtained through the image sensor, and the first recognition result is correspondingly obtained. While monitoring the face information, the monitoring data of the time series sensor is collected, and the monitoring data of the time series sensor is input into the user behavior recognition model to obtain the second recognition result, so as to determine whether the object to be recognized is a live body according to the first recognition result and the second recognition result. Based on the monitoring data of the time series sensor during normal user authentication is significantly different from the time series sensor information during a prosthesis attack system, and relying on the second recognition result output by the user behavior recognition model corresponding to the monitoring data of the time series sensor is combined with the first recognition result, thereby improving the accuracy of face recognition.
[0111] The third embodiment of the present application provides a system for face recognition, as Figure 5 shown Figure 5 is a schematic diagram of a system for face recognition provided by the third embodiment of the present application.
[0112] The third embodiment of the present application provides a system 400 for face recognition, including: an image sensor 401, a timing sensor 402, a decision-making model 403, a user behavior recognition model 404, and an image recognition model 405.
[0113] Among them, the image sensor 401 is used to obtain the facial information of the object to be recognized and send the facial information of the object to be recognized to the image recognition model 405. The facial information at least includes the motion information of the face or a part of the face. Specifically, the image sensor 401 captures the face of the object to be recognized and obtains the facial information of the object to be recognized, which at least includes the motion information of the face or a part of the face. Among them, the motion information of the face refers to all the motion information of the face of the object to be recognized according to the instructions, and the motion information refers to the information of the face making movements, for example, making facial expressions, shaking the head, etc. The motion information of a part of the face refers to the motion information of a part of the face of the object to be recognized according to the instructions, and the motion information refers to the information of the face making movements, for example, blinking, opening the mouth, etc.
[0114] The image recognition model 405 is used to receive the facial information of the object to be recognized, obtain a first recognition result including the identity information of the object to be recognized and whether it is a live body according to the facial information of the object to be recognized, and send the first recognition result including the identity information of the object to be recognized and whether it is a live body to the decision-making model;
[0115] After obtaining the facial information of the object to be recognized, a first recognition result including the identity information of the object to be recognized and whether it is a live body corresponding to the facial information of the user can be obtained through the image recognition model 405. Among them, the identity information of the object to be recognized is used to represent the identity of the object to be recognized, for example, it is XX user. It should be noted that this first recognition result corresponding to the facial information of the user can determine whether the object to be recognized is a live body under the corresponding detection conditions. For example, the detection condition is set to detect a picture. When the terminal instructs the object to be recognized to make corresponding actions such as shaking the head and smiling, the object to be detected on the picture cannot complete them, and then it can be determined whether the object to be recognized is a live body through this first recognition result.
[0116] The timing sensor 402 is used to obtain the monitoring data of the timing sensor 402 by monitoring the object to be recognized while acquiring the facial information of the object to be recognized, and input the monitoring data of the timing sensor 402 into the user behavior recognition model 404. Specifically, while the image sensor 401 acquires the facial information of the object to be recognized, the timing sensor 402 of the system also detects the face of the object to be recognized, and correspondingly obtains the monitoring data of the timing sensor 402. Among them, the monitoring data of the timing sensor 402 at least includes gyroscope monitoring data and linear accelerator monitoring data. The monitoring data of the timing sensor 402 is the data corresponding to the monitoring when the control terminal of the object to be recognized moves. For example, when the object to be recognized uses the terminal to take a face photo to obtain facial information, there is a phenomenon of hand shaking of the object to be recognized, so that the monitoring data of the timing sensor 402 is the data corresponding to the monitoring when the terminal moves. After obtaining the monitoring data, the monitoring data of the timing sensor 402 is input into the user behavior recognition model 404.
[0117] The user behavior recognition model 404 is used to receive the monitoring data of the timing sensor 402, obtain a second recognition result including the behavior category of the object to be recognized according to the monitoring data, and send the second recognition result to the decision model 403. It should be noted that the second recognition result obtained according to the user behavior recognition model is coordinated with the first recognition result, which can be a supplement to the first recognition result or assist the image recognition model to further determine whether the object to be recognized is a live body. For example, when verifying in the form of video recording, there are actions such as shaking the head and smiling in the video recording. At this time, the first recognition result cannot correctly determine whether the object to be recognized is a live body. However, the behavior classification information determined by the second recognition result has a specific relationship and specific law with the behaviors of the user shaking the head, smiling, etc., so that it can correctly determine whether the object to be recognized is a live body.
[0118] The decision model 403 is used to receive the first recognition result and the second recognition result, and obtain the identity information of the object to be recognized and a third recognition result for determining whether the object to be recognized is a live body according to the first recognition result and the second recognition result.
[0119] Corresponding to the training method of the user behavior recognition model provided in the first embodiment of the present application, the fourth embodiment of the present application correspondingly provides a training device for the user behavior recognition model. Since the device embodiment is basically similar to the first embodiment, the description is relatively simple, and for the related parts, please refer to the partial description of the first embodiment. The device embodiments described below are only illustrative.
[0120] Please refer to Figure 6, which is a schematic diagram of a training device for a user behavior recognition model provided in the fourth embodiment of the present application. The training device for the user behavior recognition model includes: an unlabeled data acquisition unit 501, configured to acquire data of a time-series sensor without labels as original data; a preprocessing unit 502, configured to preprocess the original data to obtain time-domain data, frequency-domain data, and statistic data; a pre-training unit 503, configured to input the time-domain data, frequency-domain data, and statistic data into a model respectively for pre-training to obtain a pre-trained model; a labeled data acquisition unit 504, configured to acquire time-domain data of a time-series sensor with labels, and obtain corresponding frequency-domain data and statistic data; a re-training unit 505, configured to input the time-domain data, frequency-domain data, and statistic data of the time-series sensor with labels into the pre-trained model respectively for re-training to obtain the user behavior recognition model.
[0121] Optionally, the pre-training unit 503 is configured to extract time-domain features, frequency-domain features, and statistic features corresponding to the time-domain data, frequency-domain data, and statistic data; aggregate the time-domain features, frequency-domain features, and statistic features as training data.
[0122] Optionally, the pre-training unit 503 is further configured to acquire similarity features and difference features between the time-domain features, frequency-domain features, and statistic features; aggregate the similarity features and difference features as training data.
[0123] Optionally, the pre-training unit 503 is configured to input the time-domain data, frequency-domain data, and statistic data into different processing units of the model in parallel.
[0124] Optionally, the model is a transformer model, and the processing unit is an encoder of the transformer.
[0125] Optionally, it further includes: a first normalization processing unit, configured to preprocess the time-domain data, its frequency-domain data, and statistic data before inputting the time-domain data, frequency-domain data, and statistic data into the model respectively.
[0126] Optionally, it further includes: a second normalization processing unit, configured to preprocess the time-domain data, frequency-domain data, and statistic data of the time-series sensor with labels before inputting the time-domain data, frequency-domain data, and statistic data of the time-series sensor with labels into the pre-trained model respectively.
[0127] Optionally, the data of the time-series sensor without labels and the time-domain data of the time-series sensor with labels respectively include timestamp information.
[0128] Corresponding to the method for face recognition provided in the second embodiment of the present application, the fifth embodiment of the present application correspondingly provides a device for face recognition. Since the device embodiment is basically similar to the second embodiment, the description is relatively simple. For related parts, please refer to the partial description of the second embodiment. The device embodiments described below are merely illustrative.
[0129] Please refer to Figure 7 , which is a schematic diagram of a device for face recognition provided in the fifth embodiment of the present application. The device for face recognition includes: a face information acquisition unit 601, configured to acquire face information of an object to be recognized through an image sensor, where the face information at least includes motion information of the face or a part of the face; a first recognition result acquisition unit 602, configured to input the face information into an image recognition model to obtain a first recognition result including the identity information of the object to be recognized and whether it is a live body; a monitoring data acquisition unit 603, configured to monitor the object to be recognized through a time series sensor while acquiring the face information of the object to be recognized, and obtain monitoring data of the time series sensor; a second recognition result acquisition unit 604, configured to input the monitoring data into a user behavior recognition model to obtain a second recognition result including the behavior category of the object to be recognized; and a recognition result determination unit 605, configured to obtain the identity information of the object to be recognized and a third recognition result for determining whether the object to be recognized is a live body according to the first recognition result and the second recognition result.
[0130] Corresponding to the training method of the user behavior recognition model in the first embodiment of the present application and the method for face recognition in the second embodiment of the present application, the sixth embodiment of the present application further provides an electronic device. As Figure 8 shown, Figure 8 is a schematic diagram of an electronic device provided in the sixth embodiment of the present application. The electronic device includes: a processor 701; a memory 702, configured to store a computer program, and the computer program is run by the processor to execute the training method of the user behavior recognition model in the first embodiment and the method for face recognition in the second embodiment.
[0131] Corresponding to the training method of the user behavior recognition model in the first embodiment of the present application and the method for face recognition in the second embodiment of the present application, the seventh embodiment of the present application further provides a computer storage medium. The computer storage medium stores a computer program, and the computer program is run by the processor to execute the training method of the user behavior recognition model in the first embodiment and the method for face recognition in the second embodiment.
[0132] Although the present application is disclosed above in preferred embodiments, it is not intended to limit the present application. Any person skilled in the art can make possible changes and modifications without departing from the spirit and scope of the present application. Therefore, the protection scope of the present application shall be subject to the scope defined by the claims of the present application.
[0133] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory. The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0134] Computer-readable media includes both permanent and non-permanent, removable and non-removable media implemented by any method or technology for information storage. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile discs (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.
[0135] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
Claims
1. A training method for a user behavior recognition model, characterized in that, comprising: Obtaining data of a time-series sensor without labels as original data; Preprocessing the original data to obtain time-domain data, frequency-domain data, and statistic data, wherein the statistic data extracts features with prior knowledge through the calculation of statistics; Respectively inputting the time-domain data, frequency-domain data, and statistic data into a model for pre-training to obtain a pre-trained model; Obtaining the time-domain data of a time-series sensor with labels, and obtaining corresponding frequency-domain data and statistic data; Respectively inputting the time-domain data, frequency-domain data, and statistic data of the time-series sensor with labels into the pre-trained model for re-training to obtain the user behavior recognition model; Wherein, the time-series sensor is used to obtain the facial information of an object to be recognized, and at the same time, by monitoring the object to be recognized, obtain the monitoring data of the time-series sensor, and input the monitoring data of the time-series sensor into the user behavior recognition model, and the monitoring data of the time-series sensor is the data corresponding to the movement of the control terminal of the object to be recognized.
2. The training method for a user behavior recognition model according to claim 1, characterized in that, The step of respectively inputting the time-domain data, frequency-domain data, and statistic data into a model for pre-training to obtain a pre-trained model includes: Extracting time-domain features, frequency-domain features, and statistic features corresponding to the time-domain data, frequency-domain data, and statistic data; Aggregating the time-domain features, frequency-domain features, and statistic features as training data.
3. The training method for a user behavior recognition model according to claim 2, characterized in that, After extracting the time-domain features, frequency-domain features, and statistic features corresponding to the time-domain data, frequency-domain data, and statistic data, it further includes: Obtaining the similar features and different features among the time-domain features, frequency-domain features, and statistic features; Aggregating the similar features and different features as training data.
4. The training method for a user behavior recognition model according to any one of claims 1 to 3, characterized in that, The step of respectively inputting the time-domain data, frequency-domain data, and statistic data into a model for pre-training to obtain a pre-trained model includes: Parallelly inputting the time-domain data, frequency-domain data, and statistic data into different processing units of the model.
5. The training method for a user behavior recognition model according to claim 4, characterized in that, The model is a transformer model, and the processing unit is the encoder of the transformer.
6. The training method for a user behavior recognition model according to claim 1, characterized in that, Before respectively inputting the time-domain data, frequency-domain data, and statistic data into the model, it further includes: preprocessing the time-domain data, its frequency-domain data, and statistic data; Correspondingly, before inputting the time-domain data, frequency-domain data, and statistical data of the tagged time-series sensor into the pre-trained model respectively, the method further includes: preprocessing the time-domain data, frequency-domain data, and statistical data of the tagged time-series sensor.
7. The training method of the user behavior recognition model according to claim 1, wherein, the data of the untagged time-series sensor and the time-domain data of the tagged time-series sensor respectively include timestamp information.
8. A method for face recognition, wherein, it includes: acquiring facial information of an object to be recognized through an image sensor, where the facial information at least includes motion information of the face or a part of the face. Among them, the motion information of the face refers to all the motion information of the face that the object to be recognized makes according to the instructions of the terminal, and the motion information of a part of the face refers to the motion information of a part of the face that the object to be recognized makes according to the instructions of the terminal; inputting the facial information into an image recognition model to obtain a first recognition result including the identity information of the object to be recognized and whether it is a live body; while acquiring the facial information of the object to be recognized, monitoring the object to be recognized through a time-series sensor to obtain monitoring data of the time-series sensor. Among them, the time-series sensor is used to obtain monitoring data of the time-series sensor by monitoring the object to be recognized while acquiring the facial information of the object to be recognized, and input the monitoring data of the time-series sensor into the user behavior recognition model. The monitoring data of the time-series sensor is the corresponding monitoring data when the object to be recognized controls the terminal to move; inputting the monitoring data into the user behavior recognition model generated by the training method of the user behavior recognition model according to any one of claims 1 to 7 above to obtain a second recognition result including the behavior category of the object to be recognized; obtaining a third recognition result including the identity information of the object to be recognized and determining whether the object to be recognized is a live body according to the first recognition result and the second recognition result.
9. A face recognition system, wherein, it includes: an image sensor, an image recognition model, a time-series sensor, a decision-making model, and a user behavior recognition model; the image sensor is used to acquire facial information of an object to be recognized and send the facial information of the object to be recognized to the image recognition model; the facial information at least includes motion information of the face or a part of the face. Among them, the motion information of the face refers to all the motion information of the face that the object to be recognized makes according to the instructions of the terminal, and the motion information of a part of the face refers to the motion information of a part of the face that the object to be recognized makes according to the instructions of the terminal; the image recognition model is used to receive the facial information of the object to be recognized, obtain a first recognition result including the identity information of the object to be recognized and whether it is a live body according to the facial information of the object to be recognized, and send the first recognition result including the identity information of the object to be recognized and whether it is a live body to the decision-making model; The timing sensor is used to obtain the monitoring data of the timing sensor by monitoring the object to be recognized while obtaining the facial information of the object to be recognized, and input the monitoring data of the timing sensor into the user behavior recognition model. The monitoring data of the timing sensor is the data corresponding to the monitoring when the object to be recognized controls the terminal to move. The user behavior recognition model is used to receive the monitoring data of the timing sensor, obtain a second recognition result including the behavior category of the object to be recognized according to the monitoring data, and send the second recognition result to the decision-making model. The decision-making model is used to receive the first recognition result and the second recognition result, and obtain the identity information of the object to be recognized and a third recognition result for determining whether the object to be recognized is a living body according to the first recognition result and the second recognition result.
10. A training device for a user behavior recognition model Characterized in that It includes: An unlabeled data acquisition unit, configured to acquire the data of the timing sensor without labels as the original data; A preprocessing unit, configured to preprocess the original data to obtain time-domain data, frequency-domain data, and statistic data. Among them, the statistic data extracts features with prior knowledge through the calculation of statistics; A pre-training unit, configured to input the time-domain data, frequency-domain data, and statistic data into the model respectively for pre-training to obtain a pre-trained model; A labeled data acquisition unit, configured to acquire the time-domain data of the timing sensor with labels, and obtain the corresponding frequency-domain data and statistic data; A re-training unit, configured to input the time-domain data, frequency-domain data, and statistic data of the timing sensor with labels into the pre-trained model respectively for re-training to obtain the user behavior recognition model; Wherein, the timing sensor is used to obtain the monitoring data of the timing sensor by monitoring the object to be recognized while obtaining the facial information of the object to be recognized, and input the monitoring data of the timing sensor into the user behavior recognition model. The monitoring data of the timing sensor is the data corresponding to the monitoring when the object to be recognized controls the terminal to move.
11. A device for facial recognition Characterized in that It includes: A facial information acquisition unit, configured to acquire the facial information of the object to be recognized through an image sensor. The facial information at least includes the motion information of the face or a part of the face. Among them, the motion information of the face refers to all the motion information of the face of the object to be recognized according to the instructions of the terminal, and the motion information of a part of the face refers to the motion information of a part of the face of the object to be recognized according to the instructions of the terminal; A first recognition result acquisition unit, configured to input the facial information into an image recognition model to obtain a first recognition result including the identity information of the object to be recognized and whether it is a living body. A monitoring data acquisition unit is configured to monitor the object to be recognized through a timing sensor while acquiring the facial information of the object to be recognized, so as to obtain the monitoring data of the timing sensor. The timing sensor is configured to monitor the object to be recognized while acquiring the facial information of the object to be recognized, so as to obtain the monitoring data of the timing sensor, and input the monitoring data of the timing sensor into the user behavior recognition model. The monitoring data of the timing sensor is the data corresponding to the monitoring when the control terminal of the object to be recognized moves; A second recognition result acquisition unit is configured to input the monitoring data into the user behavior recognition model generated by the training method of the user behavior recognition model according to any one of claims 1 to 7 above, so as to obtain a second recognition result including the behavior category of the object to be recognized; A recognition result determination unit is configured to obtain the identity information of the object to be recognized and a third recognition result for determining whether the object to be recognized is a living body according to the first recognition result and the second recognition result.
12. An electronic device, characterized in that, the electronic device includes: a processor; a memory for storing a computer program, and the computer program is run by the processor to execute the method according to any one of claims 1-7, 8.
13. A computer storage medium, characterized in that, the computer storage medium stores a computer program, and the computer program is run by the processor to execute the method according to any one of claims 1-7, 8.
Citation Information
Patent Citations
Tone mapping image quality evaluation method based on structural similarity difference degree
CN110996096A
Multi-person behavior recognition method based on Transform network
CN113033657A
Tooth brushing behavior monitoring method, device, equipment, medium and chip system
CN113257277A