Owner identification method and terminal equipment
By combining the multimodal fusion method of IMU, touch screen and respiratory audio data with face recognition, the problem of the owner's recognition rate dropping under light changes or environmental interference by terminal equipment is solved, achieving higher recognition accuracy and stability, and protecting the owner's privacy.
Patent Information
- Application Number
- CN202311870258.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-29
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2043-12-29
AI Technical Summary
During the use of terminal devices, the owner's private information is easily seen by others, resulting in privacy leakage. The existing facial and fingerprint recognition solutions decrease in recognition rate under light changes or environmental interference.
Combining the IMU data, touch screen data, breathing audio data and face recognition of the terminal device, a multi-modal fusion owner recognition method is adopted to comprehensively identify it through a cross-modal encoder and a decision fusion model.
It improves the accuracy and stability of the owner's identification, expands the recognition scenario, ensures that the owner can still be accurately identified under light changes or environmental interference, and protects private information security.
Smart Images

Figure CN120277648A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular, to a method for identifying the device owner and a terminal device. Background Art
[0002] Terminal devices, especially mobile phones, have gradually become the center of the personal digital world, carrying a large amount of personal information and privacy data. During the use of terminal devices, the private information of the device owner may be displayed. If the private information of the device owner is seen by someone other than the device owner, it will lead to the leakage of the device owner's privacy.
[0003] Therefore, during the use of terminal devices, it is necessary to identify the device owner to protect the security of the device owner's private information. Summary of the Invention
[0004] In view of this, the embodiments of this application provide a method for identifying the device owner and a terminal device, which can identify the device owner during the use of the terminal device to protect the security of the device owner's private information.
[0005] In a first aspect, the embodiments of this application provide a method for identifying the device owner, which is applied to a terminal device. The method for identifying the device owner may include: the terminal device can collect user data of the user to be tested in response to an unlocking operation of the user to be tested on the target content, where the target content is in a locked state, and the user data includes at least one of the following: a face image of the user to be tested, breathing audio data of the user to be tested, IMU data of the user to be tested, and touch screen data of the user to be tested; then, the terminal device can perform device owner identification based on the user data and a pre-recorded device owner template, where the device owner template includes at least one of the following: a face template, a breathing sound feature template, and a cross-modal feature template, and the cross-modal feature template is used to describe the comprehensive features of the device owner in multiple modalities; finally, the terminal device can unlock the target content when the device owner is recognized.
[0006] Among them, the user to be tested generally refers to the user to be identified. The face image of the user to be tested is the face image of the user to be tested. The breathing audio data of the user to be tested is the breathing audio data of the user to be tested. The IMU data of the user to be tested is the IMU data of the terminal device in response to the user to be tested. The touch screen data of the user to be tested is the touch screen data of the terminal device in response to the user to be tested.
[0007] The device owner template is a pre-recorded device owner template. The face template is a pre-recorded face image of the device owner. The breathing sound feature template is a pre-recorded breathing sound feature of the device owner. The cross-modal feature template is a pre-recorded comprehensive feature of the device owner in multiple modalities.
[0008] In an embodiment of the present application, when a user to be measured wants to unlock target content in a locked state, the terminal device can collect user data of the user to be measured, thereby performing owner identification on the user to be measured through the collected user data. And only when the owner is identified, the terminal device will unlock the target content, which can ensure the security of the owner's private information. In addition, the terminal device can use one or more of the face data, breathing audio data, IMU data, and touch screen data corresponding to the user to be measured to perform owner identification on the user to be measured, which can achieve owner identification with multiple modalities of data, and the advantages of different modalities of data complement each other, thereby improving the stability and accuracy of owner identification.
[0009] Optionally, the terminal device can preferentially use face data for owner identification. When the face recognition result is not ideal, it continues to use one or more of the breathing audio data, IMU data, and touch screen data for owner identification. In this way, it can be realized that even when the face recognition result is not ideal, owner identification can still be performed, which can improve the accuracy and stability of owner identification.
[0010] In addition, the IMU data and touch screen data of the terminal device reflect the usage habits of the user to be measured for the terminal device. Combining these two pieces of data to further determine whether it is the owner can make the recognition result more accurate and reliable, that is, it can further improve the accuracy rate of owner identification.
[0011] In the first possible implementation manner of the first aspect, the target content includes any one of the following: a locked notification message, a locked service card, a locked application, a locked function, and a locked interface.
[0012] In an embodiment of the present application, there can be one or more target contents, that is to say, there can be many application scenarios that can trigger the terminal device to perform owner identification. That is to say, the terminal device can lock one or more contents to ensure the security of the owner's private information.
[0013] As an optional embodiment of the present application, the terminal device can lock the content set by the lock setting operation in response to the lock setting operation of the owner.
[0014] Among them, the lock setting operation is used to lock the content set by the owner. For example, one or more applications set by the owner can be locked. For another example, the service card function set by the owner can be locked. For another example, the passwordless payment function set by the owner can be locked, and so on.
[0015] In a second possible implementation of the first aspect, the process of the terminal device inputting the face template of the device owner may include: First, in response to the user's face input operation, the terminal device displays a face input interface, where the face input interface includes a face collection frame; After that, the terminal device may input the collected face image as a face template in response to the collected face image.
[0016] In the embodiments of the present application, the device owner can input the face template through human-computer interaction, which is intuitive and convenient to operate, and helps to improve the user experience.
[0017] As an optional embodiment of the present application, the terminal device can input the device owner template of one device owner, or can input the device owner templates of multiple device owners. Among them, usually one device owner corresponds to one device owner template.
[0018] In a third possible implementation of the first aspect, the process of the terminal device inputting the cross-modal feature template of the device owner may include: First, when the multi-modal feature input condition is met, the terminal device can collect the breathing audio data of the device owner, the IMU data of the terminal device, and the touch screen data of the terminal device; Then, the terminal device can extract the breathing sound features from the breathing audio data; Finally, the terminal device can encode the IMU data, the touch screen data, and the breathing sound features through a cross-modal encoder to obtain cross-modal features, and input the cross-modal features as a cross-modal feature template, where the cross-modal encoder is used to encode data of multiple modalities.
[0019] Among them, the multi-modal feature input condition includes any one of the following: the face template of the device owner is successfully input, the device owner's face unlocking is successful and the cross-modal feature template of the device owner currently performing the unlocking has not been input into the terminal device.
[0020] In the embodiments of the present application, when the face template of the device owner is successfully input, the terminal device can continue to collect the breathing audio data of the device owner, so as to extract the breathing sound features from the collected breathing audio data, and encode the IMU data, the touch screen data, and the breathing sound features through a cross-modal encoder to obtain a cross-modal feature template.
[0021] Since the breathing audio data collected when the device owner is successfully input may be interfered by various factors, for example, the environment is noisy, resulting in no breathing sound features being extracted, which will also cause the cross-modal encoder to be unable to encode and obtain cross-modal features. Therefore, when the cross-modal feature template of the device owner has not been successfully input, when the device owner's face unlocking is successful, the breathing audio data of the device owner can be continuously collected, and the breathing sound features therein can be extracted, so as to encode and obtain cross-modal features in combination with the breathing sound features, and input them as the cross-modal feature template of the device owner. In this way, the successful input of the cross-modal feature template of the device owner can be guaranteed.
[0022] In a fourth possible implementation of the first aspect, the cross-modal encoder may include a cross-channel attention encoder. The cross-channel attention encoder includes parallel first slice encoding module, second slice encoding module, and third slice encoding module. The first slice encoding module includes a first Group layer and a first Linear layer connected in sequence; the second slice encoding module includes a second Group layer and a second Linear layer connected in sequence; the third slice encoding module includes a third Group layer and a third Linear layer connected in sequence.
[0023] Wherein, the input end of the first Linear layer is connected to the output end of the first Group layer, the input end of the second Linear layer is connected to the output end of the first Linear layer and the output end of the second Group layer, and the input end of the third Linear layer is connected to the output end of the second Linear layer and the output end of the third Group layer.
[0024] In the embodiments of the present application, the output data of the first Group layer is encoded the most times, and the output data of the third Group layer is encoded the least times.
[0025] That is to say, when performing multiple feature fusions, the cross-modal encoder can deploy good-quality features in the third Group layer and poor-quality features in the first Group layer. In this way, the number of Linear layers that the features deployed in the third Group layer need to pass through is the least, that is, the number of encoding times is the least. The number of Linear layers that the features deployed in the first Group layer need to pass through is the most, that is, the number of encoding times is also the most. In this way, the cross-channel attention encoder can combine the quality of each feature and encode each feature to different degrees to better effectively fuse good-quality features and poor-quality features. That is to say, the cross-modal encoder can achieve feature enhancement, thereby improving the accuracy of the output features.
[0026] Optionally, after encoding and fusing good-quality features and poor-quality features to different degrees, the fused features can be transposed, and then the transposed fused data is input into the Features layer, which helps to achieve further effective fusion between features.
[0027] In a fifth possible implementation of the first aspect, the cross-modal encoder further includes a cross-channel linear encoder, wherein the cross-channel linear encoder includes a Linear layer, an activation function, a Linear layer, a norm layer, a conv layer, and an activation function connected in sequence.
[0028] In the embodiments of the present application, the cross-channel linear encoder uses a Linear layer for encoding, aligns the feature dimensions of different modalities, and uses a conv layer to implement one-dimensional pointwise convolution for information fusion between channels. That is to say, the cross-channel linear encoder can increase the abstraction level of the output features.
[0029] In the sixth possible implementation manner of the first aspect, the terminal device can perform owner identification based on user data and a pre-recorded owner template, which may include: First, when there is a face in the face image to be measured, the terminal device can perform owner identification based on the face image to be measured and the face template to obtain a face recognition score; Then, when the face recognition score is greater than a pre-set face recognition lower threshold and less than a pre-set face recognition upper threshold, the terminal device can extract the measured breathing sound characteristics of the measured user from the measured breathing audio data; Finally, input the measured breathing sound characteristics, the measured IMU data, the measured touch screen data, and the face recognition score into a pre-trained multi-modal fusion owner identification model to obtain an owner indication information, which is used to indicate whether the measured user is the owner.
[0030] In the embodiments of the present application, only when the face recognition effect is poor, for example, strong light, backlight, low light and other situations lead to poor face recognition, will the terminal device further collect the measured breathing audio data, touch screen data, and IMU data of the measured user, so as to combine multi-modal data to jointly perform owner identification to improve the accuracy of owner identification. Since the accuracy of face recognition is high most of the time, that is, the situation of using multi-modal data for owner identification is relatively rare, it is possible to save the computing power of the mobile phone and reduce the power consumption of the mobile phone while ensuring the accuracy of owner identification, thereby improving the practicality of the mobile phone and helping to further enhance the user experience.
[0031] Optionally, the above face recognition upper threshold may be the FAR threshold of face recognition. The above face recognition lower threshold may be the FRR threshold of face recognition.
[0032] In practice, when the face recognition score is between the above face recognition lower threshold and the above face recognition upper threshold, it usually indicates that the face recognition result is poor. That is to say, the terminal device can further perform accurate owner identification by combining multi-modal data when the face recognition effect is poor, which can effectively reduce the situation that the user cooperates with the terminal device for repeated operations of owner identification, and can not only effectively protect the security of the owner's private information, but also improve the user experience.
[0033] Optionally, when there is no face in the face image to be measured, the terminal device can continue to use the breathing sound feature in the breathing audio data to be measured for the identification of the device owner. In this way, when the face recognition result is not ideal, the device owner can still be identified, which can improve the accuracy and stability of the device owner identification.
[0034] Optionally, when there is no face in the face image to be measured and the breathing audio data to be measured does not contain breathing sounds, that is, when the breathing sound feature cannot be extracted, the terminal device can output information prompting the user to be measured to operate again. This can ensure the accuracy rate of the device owner identification.
[0035] Optionally, when the face recognition score is greater than the lower face recognition threshold and less than the upper face recognition threshold, the terminal device can also use the breathing sound feature in the audio data to be measured for the identification of the device owner. For example, the terminal device can determine the similarity between the breathing sound feature to be measured and the breathing sound feature template. When the obtained similarity is greater than the preset second determination threshold, it is determined that the user to be measured is the device owner. Conversely, when the obtained similarity is less than or equal to the second determination threshold, it is determined that the user to be measured is not the device owner.
[0036] Optionally, when the face recognition score is greater than the lower face recognition threshold and less than the upper face recognition threshold, if the audio data to be measured does not contain breathing sounds, that is, the breathing sound feature is not extracted, the terminal device can combine the face recognition score to continue to determine whether the user to be measured is the device owner. For example, when the face recognition score is greater than the preset face recognition score threshold, the terminal device can determine that the user to be measured is the device owner. Conversely, when the face recognition score is less than or equal to the face recognition score threshold, the terminal device can determine that the user to be measured is not the device owner.
[0037] Optionally, when the face recognition score is greater than the lower face recognition threshold and less than the upper face recognition threshold, if the audio data to be measured does not contain breathing sounds, that is, the breathing sound feature is not extracted, the terminal device can also output information prompting the user to be measured to operate again. This can ensure the accuracy rate of the device owner identification.
[0038] In the seventh possible implementation manner of the first aspect, when the terminal device performs device owner identification based on the user data and the pre-recorded device owner template, it may further include: when the face recognition score is greater than the upper face recognition threshold, the terminal device can determine that the user to be measured is the device owner; when the face recognition score is less than the lower face recognition threshold, the terminal device can determine that the user to be measured is not the device owner.
[0039] In the embodiments of the present application, if the face recognition score is greater than the face recognition upper threshold or the face recognition score is less than the face recognition lower threshold, it indicates that the face recognition result is very clear and accurate. The terminal device can directly determine whether the user to be measured is the owner, which can improve the efficiency and accuracy of owner recognition.
[0040] In the eighth possible implementation manner of the first aspect, the data processing process of the multi-modal fusion owner recognition model may include the following steps: First, encode the to-be-measured IMU data, the to-be-measured touch screen data, and the to-be-measured breath sound feature through a cross-modal encoder to obtain the to-be-measured cross-modal feature; then, determine the cross-modal similarity between the to-be-measured cross-modal feature and the cross-modal feature template in the owner template, and determine the breath sound similarity between the to-be-measured breath sound feature and the breath sound feature template in the owner template; after that, determine the multi-modal recognition score according to the cross-modal similarity, the breath sound similarity, and the face recognition score, where the multi-modal recognition score is used to describe the comprehensive recognition result of the user to be measured in multiple modalities; finally, determine whether the user to be measured is the owner according to the multi-modal recognition score and a preset fusion threshold.
[0041] In the embodiments of the present application, the multi-modal fusion owner recognition model uses a cross-modal encoder to perform feature-level fusion on IMU data, touch screen data, and breath sound features, and then performs decision-level fusion on the cross-modal similarity (or cross-modal recognition score), the breath sound similarity (or breath recognition score), and the face recognition score to obtain the multi-modal recognition score. Among them, a stronger modality, such as the face modality, does not participate in the feature-level fusion, which can reduce the coupling between data and effectively utilize the complementary advantages between multiple modalities, improving the comprehensive recognition success rate of the model.
[0042] As an optional embodiment of the present application, the multi-modal fusion owner recognition model fuses the cross-modal similarity (or cross-modal recognition score), the breath sound similarity (or breath recognition score), and the face recognition score through a decision fusion model to obtain the multi-modal recognition score.
[0043] Among them, the fusion decision model may include a Stem layer, a 12-layer encoding layer, a Head layer, and a Sigmoid function, where the encoding layer includes a Linear layer, an activation function, and a Linear layer.
[0044] In the embodiments of the present application, the fusion decision model uses the Stem layer to expand the feature dimensions of the cross-modal recognition score, the breath recognition score, and the face recognition score, then encodes them through the 12-layer encoding layer of the feed-forward neural network with a residual structure, and finally outputs the multi-modal recognition score using the classification head (Head layer) and the Sigmoid function. The overall decision fusion model adopts a lightweight design, which can save storage resources.
[0045] Optionally, when a host template is input in the terminal device, the terminal device may adopt the feature templates in the host template, such as the breath sound feature template and the cross-modal feature template, to determine the multi-modal recognition score.
[0046] Optionally, when multiple host templates are input in the terminal device, the terminal device may select the host template with the highest corresponding face recognition score to determine the multi-modal recognition score. In this way, the computational effort consumed in the host recognition process can be reduced.
[0047] As an optional embodiment of the present application, when multiple host templates are input in the terminal device, the terminal device may determine a multi-modal recognition score for each host template. After that, the terminal device can combine the obtained multiple multi-modal recognition scores to determine whether the user to be measured is the host and which host it is specifically.
[0048] In the embodiment of the present application, using each host template to separately determine the corresponding multi-modal recognition score can ensure the accuracy of host recognition.
[0049] As an optional embodiment of the present application, when there are multiple feature templates in the host template, the terminal device may calculate the similarity between the mean value of the multiple feature templates in the host template and the to-be-measured feature of the same modality (such as the to-be-measured breath sound feature) to obtain the recognition score of the corresponding modality.
[0050] As an optional embodiment of the present application, when there are multiple feature templates in the host template, the terminal device may also calculate the similarity between each feature template in the host template and the to-be-measured feature of the same modality to obtain multiple similarities; after that, the terminal device calculates the recognition score of the corresponding modality through the obtained multiple similarities. For example, the maximum value among the multiple similarities may be used as the recognition score of the corresponding modality.
[0051] In the ninth possible implementation manner of the first aspect, when multiple host templates are input in the terminal device, there are multiple multi-modal recognition scores. At this time, when the terminal device determines whether the user to be measured is the host according to the multi-modal recognition score and a preset fusion threshold, it may include: selecting the maximum value among the multiple multi-modal recognition scores, and if the selected multi-modal recognition score is greater than the fusion threshold, determining that the user to be measured is the host.
[0052] In an embodiment of the present application, when multiple owner templates are recorded in the terminal device, the terminal device can obtain a multi-modal recognition score for each owner template. Generally, the larger the multi-modal recognition score, the more similar the user to be measured is to the owner indicated by the owner template corresponding to the multi-modal recognition score. That is to say, using the largest multi-modal recognition score for owner recognition can improve the accuracy and efficiency of owner recognition.
[0053] Optionally, when multiple owner templates are recorded in the terminal device, there are multiple multi-modal recognition scores. For each multi-modal recognition score, the terminal device can compare the multi-modal recognition score with a preset fusion threshold. If the multi-modal recognition score is greater than the fusion threshold, it is considered that the user to be measured is the owner, and specifically, the owner indicated by the owner template corresponding to the multi-modal recognition score.
[0054] In the tenth possible implementation manner of the first aspect, the process of training the multi-modal fusion owner recognition model may include the following steps: First, select m batches of training samples from the sample dataset used to train the multi-modal fusion owner recognition model. Each batch of training samples includes 1 positive sample and n negative samples, where n>1 and m>1; then, during the process of training the multi-modal fusion owner recognition model with each batch of training samples, the negative sample that is recognized as a positive sample and has the highest corresponding recognition probability is determined as a difficult sample; finally, determine the model training loss of the multi-modal fusion owner recognition model based on the positive sample and the difficult sample, and adjust the model parameters of the multi-modal fusion owner recognition model according to the determined model training loss.
[0055] In an embodiment of the present application, during the process of training the multi-modal fusion owner recognition model, for the same sample user, m batches of training samples can be collected. That is, the model is trained using m groups of training data for each sample user, rather than only using one group of data for training, which can improve the accuracy of the model. In addition, in each batch of training samples, difficult samples are mined as negative samples, and the dynamic difficult samples and positive samples are jointly used to calculate the model training loss, and the model parameters are updated based on the model training loss, which can effectively reduce the false recognition rate and improve the TAR index at a lower false recognition rate. That is, the trained multi-modal fusion owner recognition model can be made more accurate and the stability of owner recognition is also stronger.
[0056] In the eleventh possible implementation manner of the first aspect, the terminal device can also perform the following operations: when the template update condition is met, determine the mean value of the breathing sound feature template in the owner template and the current breathing sound feature as the new breathing sound feature template, and add the new breathing sound feature template to the owner template, where the template update condition includes: successful face unlocking and reaching the update cycle.
[0057] Among them, the current breath sound feature is the breath sound feature in the breath audio data collected when the face unlocking is successful.
[0058] In the embodiments of the present application, the terminal device uses the mean value of the current feature and the old feature template as the new feature template, which can accurately and stably update the owner template, and helps to improve the accuracy and stability of owner recognition.
[0059] In the twelfth possible implementation manner of the first aspect, the terminal device can also perform the following operations: when the template update condition is met, the terminal device can also determine the mean value of the cross-modal feature template in the owner template and the current cross-modal feature as the new cross-modal feature template, and add the new cross-modal feature template to the owner template.
[0060] Among them, the current cross-modal feature is the cross-modal feature corresponding to the multi-modal data collected when the face unlocking is successful.
[0061] In the embodiments of the present application, the terminal device uses the mean value of the current feature and the old feature template as the new feature template, which can accurately and stably update the owner template, and helps to improve the accuracy and stability of owner recognition.
[0062] In some optional implementation manners of the embodiments of the present application, before updating the feature template in the owner template, the terminal device can first perform a feature quality judgment on the current feature of the owner, and determine whether to update the feature template according to the feature quality judgment result. For example, the current feature of the owner can be deleted when the feature quality judgment result is not good. For another example, when the feature quality judgment result is good, the current feature can be used to update the owner template.
[0063] In a second aspect, the embodiments of the present application provide an owner recognition device applied to a terminal device. Among them, the device includes:
[0064] A data acquisition unit, configured to acquire user data of a user to be measured in response to an unlocking operation of the user to be measured on a target content, where the target content is in a locked state, and the user data includes at least one of the following: a face image to be measured, breath audio data to be measured, IMU data to be measured, and touch screen data to be measured;
[0065] An identification execution unit, configured to perform owner identification according to the user data and a pre-recorded owner template, where the owner template includes at least one of the following: a face template, a breath sound feature template, and a cross-modal feature template, and the cross-modal feature template is used to describe the comprehensive features of the owner in multiple modalities;
[0066] An unlocking execution unit for unlocking the target content when the device owner is recognized.
[0067] As an embodiment of the present application, the device owner recognition device can implement the method according to any one of the above first aspects.
[0068] In a third aspect, an embodiment of the present application provides a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the method according to any one of the above first aspects is implemented.
[0069] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the method according to any one of the above first aspects is implemented.
[0070] In a fifth aspect, an embodiment of the present application provides a chip system, which includes a processor. The processor is coupled to a memory, and the processor executes a computer program stored in the memory to implement the method according to any one of the above first aspects. The chip system can be a single chip or a chip module composed of multiple chips.
[0071] In a sixth aspect, an embodiment of the present application provides a computer program product, which, when running on a terminal device, causes the terminal device to execute the method according to any one of the above first aspects.
[0072] It can be understood that the beneficial effects of the above second to sixth aspects can be referred to the relevant descriptions in the above first aspect, and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0073] Figure 1A It is a schematic diagram of an information display effect in a scenario without device owner recognition provided by an embodiment of the present application;
[0074] Figure 1B It is a schematic diagram of an information display effect in a scenario with device owner recognition provided by an embodiment of the present application;
[0075] Figure 1C It is another schematic diagram of an information display effect in a scenario without device owner recognition provided by an embodiment of the present application;
[0076] Figure 1D It is another schematic diagram of an information display effect in a scenario with device owner recognition provided by an embodiment of the present application;
[0077] Figure 2 It is a schematic diagram of a scenario of a multi-modal fusion device owner recognition method provided by an embodiment of the present application;
[0078] Figure 3Schematic diagram of the process of inputting a face template provided by an embodiment of the present application;
[0079] Figure 4A Schematic diagram of the process of the audio processing module processing respiratory audio data provided by an embodiment of the present application;
[0080] Figure 4B Schematic diagram of the process of the audio processing module preprocessing respiratory audio data provided by an embodiment of the present application;
[0081] Figure 4C Schematic diagram of the respiratory sound recognition model provided by an embodiment of the present application;
[0082] Figure 4D Schematic diagram of the respiratory sound feature extraction model provided by an embodiment of the present application;
[0083] Figure 4E Schematic diagram of a structure of the multi-scale residual layer in the respiratory sound feature extraction model provided by an embodiment of the present application;
[0084] Figure 5A Schematic diagram of the structure of the cross-modal encoder provided by an embodiment of the present application;
[0085] Figure 5B Schematic diagram of the structure of the cross-channel linear encoder in the cross-modal encoder provided by an embodiment of the present application;
[0086] Figure 5C Schematic diagram of the structure of the cross-channel attention encoder in the cross-modal encoder provided by an embodiment of the present application;
[0087] Figure 6A Schematic diagram of the process of the mobile phone performing owner recognition based on multi-modal fusion provided by an embodiment of the present application;
[0088] Figure 6B Schematic diagram of the process of recognizing the mobile phone owner based on respiratory sound provided by an embodiment of the present application;
[0089] Figure 6C Schematic diagram of the process of the multi-modal fusion mobile phone owner recognition model performing owner recognition provided by an embodiment of the present application;
[0090] Figure 6D Schematic diagram of the structure of the fusion decision model provided by an embodiment of the present application;
[0091] Figure 6E Model training architecture for positive and negative sample pairing based on dynamic hard sample mining provided by an embodiment of the present application;
[0092] Figure 7 ROC curve graphs of mobile phone owner recognition corresponding to various modalities provided by an embodiment of the present application;
[0093] Figure 8 PR curves of the owner recognition corresponding to various modalities provided by the embodiments of the present application;
[0094] Figure 9 Schematic structural diagram of the owner recognition device provided by the embodiments of the present application;
[0095] Figure 10 Schematic structural diagram of the terminal device provided by the embodiments of the present application;
[0096] Figure 11 Software structural block diagram of the terminal device provided by the embodiments of the present application. Detailed implementation manners
[0097] In the following description, for the purpose of illustration rather than limitation, specific details such as specific system structures and technologies are put forward so as to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits and methods are omitted so as not to impede the description of the present application with unnecessary details.
[0098] Some concepts that the embodiments of the present application may involve are described as follows:
[0099] (1) Multiple: Unless otherwise specified, in the embodiments of the present application, multiple means two or more.
[0100] (2) Difficult sample: In the embodiments of the present application, a difficult sample refers to a negative sample that is misrecognized as a positive sample by the model and has the highest corresponding recognition probability during the model training process.
[0101] (3) False Acceptance Rate (FAR): In the embodiments of the present application, FAR refers to the proportion of samples that are wrongly classified into other categories in a classification problem. In the fields of machine learning and pattern recognition, we usually focus on two types of errors: False Positive and False Negative. False Positive refers to the situation where a negative sample is wrongly classified as a positive sample, while False Negative refers to the situation where a positive sample is wrongly classified as a negative sample. The false acceptance rate can be calculated by the following formula: false acceptance rate = (number of false positive samples + number of false negative samples) / total number of samples.
[0102] In practice, the lower the false acceptance rate, the better the performance of the classifier. In practical applications, we usually weigh the weights of false positives and false negatives according to specific requirements, select a suitable classifier or adjust the parameters of the model to achieve the best false acceptance rate.
[0103] (4) Receiver Operating Characteristic curve (ROC curve): In the embodiments of the present application, the ROC curve is a curve plotted with the True Positive Rate (TPR) as the vertical axis and the False Positive Rate (FPR) as the horizontal axis. It is a visualization tool for evaluating the performance of a binary classification model at different thresholds.
[0104] When plotting the ROC curve, it is necessary to sort the samples according to the output probability or decision value of the classification model. Then, starting from the lowest threshold, the threshold is gradually increased, and the samples above the current threshold are classified as positive examples, and the samples below the current threshold are classified as negative examples. At each threshold, the corresponding true positive rate and false positive rate are calculated and plotted as coordinates. The finally obtained curve is the ROC curve.
[0105] (5) Area Under the Curve (AUC): In the embodiments of the present application, the AUC generally refers to the area under the ROC curve, and the AUC is an index for evaluating the performance of a binary classification model. The value range of the AUC is between 0.5 and 1, and the larger the value, the better the model performance.
[0106] (6) Precision-Recall curve (PR curve): In the embodiments of the present application, the PR curve is another visualization tool for evaluating the performance of a binary classification model. Different from the ROC curve, the PR curve is plotted with Precision as the vertical axis and Recall as the horizontal axis.
[0107] When plotting the PR curve, it is necessary to sort the samples according to the scores output by the model. Then, starting from the lowest threshold, the threshold is gradually increased, and the samples above the current threshold are classified as positive examples, and the samples below the current threshold are classified as negative examples. At each threshold, the corresponding precision and recall are calculated and plotted as coordinates. The finally obtained curve is the PR curve. Among them, each point on the PR curve represents the precision corresponding to different recall rates. The closer the PR curve is to the upper right corner, the better the performance of the classifier in various situations.
[0108] (7) Area Under the PR Curve (AUC-PR, AP): In the embodiments of the present application, the AP is an index for comprehensively evaluating the performance of the model. In practice, generally, the larger the AP value, the better the classification performance of the model.
[0109] (8) Service Card: In the embodiments of this application, a service card is a service form that provides a more fine-grained atomic service capability than an application (APP), and directly presents the services or content that users care most about to users in the form of an interactive service card. Service cards can be embedded in various APPs or interactive scenarios to better meet user needs. A service card is generally a page in an application or the corresponding APP of that application. When a user clicks on a service card, they can be redirected to the corresponding application or a certain service page within the APP. Unless otherwise specified, the "application" referred to in the embodiments of this application can be either an application or the corresponding APP of that application.
[0110] The generation process of service cards is as follows: The provider of the application service uploads the service content that needs to be pushed to users on the server. The server generates corresponding service cards based on the service content uploaded by the provider and pushes these service cards to the user's terminal device for display, enabling the user's terminal device to present the services pushed by the provider of the application service in the form of service cards. Generally speaking, one service card corresponds to a part of the service functions of an application. The provider of the application can split the services of the application into multiple service modules, with each service module corresponding to one service card, and then push different service cards to users according to different service scenarios.
[0111] (9) TAR, TAR@FAR: In the embodiments of this application, the True Accept Rate (TAR) is used to indicate the proportion of correct judgments. TAR@FAR refers to the value of TAR under a specific FAR. For example, TAR@FAR = 1% means that when FAR = 1%, the corresponding value of TAR.
[0112] (10) Neural Network Layers, Activation Functions: In the embodiments of this application, the neural network layers involved include Linear layer, Group layer, norm layer, Inputs layer, Features layer, conv layer, Stem layer, Head layer, etc.
[0113] Among them, the Linear layer, usually referred to as the linear layer, can also be called the fully connected layer or the dense layer. The Linear layer is usually used to perform a linear transformation on the input data to obtain the output.
[0114] The Group layer is usually used to group the input data. Each group is calculated independently. For example, the grouped data can be sent to different convolutional kernels for convolution operations. This grouping can improve the computational efficiency and the expressive power of the model. Through the grouping operation, the Group layer can increase the non-linear characteristics of the model and improve the model's ability to learn complex patterns. In addition, the Group layer can also reduce the number of parameters, lower the computational cost, and improve the generalization ability of the model.
[0115] The norm layer, usually referred to as the regularization layer or normalization layer, is typically used to normalize the input data. The norm layer aims to adjust the distribution of the input data to reduce redundancy and correlation, thereby improving the robustness and generalization ability of the network to better adapt to the training or inference process of the network model.
[0116] The Inputs layer, usually called the input layer, is used to receive external input data and convert it into a tensor form that can be processed by the neural network. The input layer is usually the first layer of the network structure and defines the format, dimensions, and attributes of the input data, etc.
[0117] The Features layer, usually referred to as the feature layer, is an intermediate or hidden layer in the neural network and is used to extract discriminative features from the input data. They can capture the key information of the input data through convolution, pooling, embedding, etc., and provide richer and more meaningful feature representations for the subsequent network layers.
[0118] The conv layer, usually called the Convolutional layer, extracts local features of the input data by applying convolution operations and generates a feature map as the output. The conv layer is widely used in fields such as computer vision and speech processing.
[0119] The Stem layer usually extracts richer features by performing spatial transformation and local feature fusion on the feature map. In the embodiments of this application, the Stem layer can be used to expand the feature dimension.
[0120] The Head layer is usually used to describe the last or top layer of the model. The main function of the Head layer is to convert the previous feature representation into the final output result.
[0121] The activation functions involved include the Sigmoid function, softmax function, ReLU function, etc.
[0122] Among them, the Sigmoid function is a common non-linear activation function that maps the input real value to the output between 0 and 1. The Sigmoid function is usually used in the output layer of binary classification problems to convert the output of the model into a probability value.
[0123] The softmax function is a commonly used activation function that converts the input real vector into a probability distribution such that the sum of all output probabilities is equal to 1. The softmax function is usually used in the output layer of multi-class classification problems, and it can convert the output of the model into the probability of each class for the final classification decision.
[0124] The Rectified Linear Unit (ReLU), also known as the rectified linear unit, is a commonly used activation function.
[0125] (11) Inertial Measurement Unit (IMU): In the embodiments of this application, the IMU is a sensor used to measure and track the attitude of an object (such as direction, angle, and speed). The IMU contains multiple measurement units, including accelerometers, gyroscopes, and magnetometers, etc. Among them, the accelerometer can measure the linear acceleration of the object along the three axes of x, y, and z; the gyroscope can measure the angular velocity of the object around the three axes of x, y, and z; the magnetometer can measure the magnetic field strength and direction of the object in the earth's magnetic field.
[0126] In the embodiments of this application, the IMU data can include the acceleration, pitch angle, roll angle, etc. of the terminal device in each direction. The IMU data is the data collected by the IMU in the terminal device.
[0127] (12) Multimodal: In the embodiments of this application, multimodal usually refers to multiple types. Multimodal data usually refers to multiple different types of data. For example, face data, breathing audio data, touch screen data, and IMU data are 4 different types of data.
[0128] (13) Audio, time-frequency diagram, Mel spectrogram, log Mel spectrogram, Mel cepstrum, Mel cepstral coefficient: In the embodiments of this application, performing a short-time Fourier transform on the audio can convert the audio into a time-frequency diagram. Performing a Mel scale transform on the time-frequency diagram can obtain the corresponding Mel spectrogram. Performing a logarithmic transform on the Mel spectrogram can obtain the corresponding log Mel spectrogram, or Mel cepstrum. Performing a discrete cosine transform on the Mel cepstrum can obtain the corresponding Mel cepstral coefficient.
[0129] Terminal devices, especially mobile phones, have gradually become the center of the personal digital world, carrying a large amount of personal information and privacy data. During the use of terminal devices, the private information of the device owner may be displayed. If the private information of the device owner is seen by someone other than the device owner, it will undoubtedly lead to the leakage of the device owner's privacy.
[0130] Taking the terminal device as a mobile phone as an example, the following lists several usage scenarios that may disclose the private information of the device owner. It can be understood that the terminal device in the embodiments of the present application is not limited to only mobile phones, and can also be other devices, such as tablets, laptop computers, etc.
[0131] Usage scenario 1: The mobile phone supports the service card function. When the mobile phone displays the service card, if the service card contains private information such as the device owner's bank card number, password, and express delivery pickup code. At this time, if the user next to the mobile phone is not the device owner or there are other people besides the device owner, when the mobile phone displays the service card, the private information of the device owner displayed by the service card may be inadvertently seen by others, resulting in the leakage of the device owner's private information.
[0132] Usage scenario 2: The device owner shares the mobile phone content with others. For example, the device owner uses their own mobile phone to report or display materials to the company leader, or the device owner watches a video with a friend. Due to some applications on the mobile phone, such as WeChat, when receiving a notification message, it usually previews the notification message at the upper end of the mobile phone interface. In this usage scenario, if the mobile phone receives a notification message from a certain application and previews the notification message, the content of the notification message may be seen by others, resulting in the leakage of the device owner's private information, and the scene will be rather embarrassing.
[0133] Usage scenario 3: The device owner reads the mobile phone in a crowded place. For example, the device owner reads news or short videos in an elevator, subway, or bus. In this usage scenario, if the mobile phone receives a notification message from a certain application and previews the notification message, the content of the notification message may be inadvertently seen by others, resulting in the leakage of the device owner's private information.
[0134] Usage scenario 4: The device owner's mobile phone is borrowed by others. For example, the device owner's friend borrows the device owner's mobile phone to watch short videos, and the device owner is not present. In this usage scenario, if the mobile phone receives a notification message from a certain application and previews the notification message, the content of the notification message may be inadvertently seen by others, resulting in the leakage of the device owner's private information.
[0135] Combined with the above several usage scenarios, it can be known that during the use of terminal devices, it is very necessary to timely identify the device owner to protect the security of the device owner's private information.
[0136] The following further combines specific usage scenarios to illustrate the uses or possible effects of the device owner identification solution according to the embodiments of the present application by way of comparison.
[0137] Figure 1A Fig. shows a schematic diagram of an information display effect in the case of no device owner identification scenario. Figure 1A In, device owner A is looking at the mobile phone, and there are other people beside, such as B and C. The mobile phone directly displays the service card of the memo application. As Figure 1A shown, the service card has A's bank card number and password. At this time, B and C can peek at the content of the service card, that is, other people can peek at the private information of the device owner.
[0138] Figure 1B Fig. shows a schematic diagram of an information display effect in the case of device owner identification scenario. Figure 1B In, device owner A is looking at the mobile phone, and there are other people beside, such as B and C. When the mobile phone does not recognize the device owner, it locks the content of the service card, that is, hides the content of the service card. In this way, the security of the device owner's private information can be protected.
[0139] It can be understood that, as Figure 1B shown, if A wants to view the content of the service card, he can further click on the icon. The mobile phone can respond to the user's click on the icon and display the content of the service card after recognizing the device owner.
[0140] Figure 1C Fig. shows another schematic diagram of an information display effect in the case of no device owner identification scenario. Figure 1C In, device owner A is looking at the mobile phone, and there are other people beside, such as B and C. The mobile phone receives a notification message from the information application and directly previews the notification message. At this time, both A and B and C beside can see the content of the notification message, that is, other people can peek at the private information of the device owner.
[0141] Figure 1D Fig. shows another schematic diagram of an information display effect in the case of device owner identification scenario. Figure 1D In, device owner A is looking at the mobile phone, and there are other people beside, such as B and C. The mobile phone receives a notification message from the information application. When the device owner is not recognized, it automatically hides the content of the notification message. It can be understood that if A wants to view the content of the notification message, he can click on the notification message whose content is hidden. The mobile phone can respond to the user's click on the notification message and display the content of the notification message after recognizing the device owner.
[0142] Combined with the above Figures 1A - 1DBy comparing the information displays in the two usage scenarios shown, it can be clearly seen that during the use of the terminal device, timely identification of the device owner can effectively protect the security of the device owner's private information.
[0143] In fact, when the terminal device identifies the device owner's identity and treats the device owner and non-device owner differently, in addition to protecting the device owner's private information from being leaked, in some application scenarios, it can also provide a more convenient device usage experience for the device owner. For example, for the device owner, payment can be made without entering a password, while for non-device owners, a password must be entered to make a payment.
[0144] Currently, the more commonly used device owner identification schemes are face recognition or fingerprint recognition. However, in existing schemes that use the face for device owner identification, in scenarios such as low light, backlight, strong light, and deviation of distance and angle, the recognition rate will decrease due to poor face image quality. And in existing schemes that use fingerprints for device owner identification, in scenarios where there is dirt on the finger or on the fingerprint touch area of the terminal device, such as when there are sweat stains, water stains, etc. on the finger, or dust, foreign objects, etc. on the fingerprint touch area, the recognition rate will also decrease. That is to say, as the scenario changes, such as when the light becomes darker or stronger, the recognition rate of existing device owner identification schemes may decrease, and it may even be impossible to correctly identify the device owner.
[0145] In view of the above technical problems in the related art, in order to improve the accuracy and stability of the terminal device's identification of the device owner, in the embodiments of this application, the terminal device combines the user's face data, the IMU of the terminal device, the touch screen data of the terminal device, and the user's breathing audio data to comprehensively identify the device owner, thereby effectively improving the success rate of device owner identification. That is, the embodiments of this application can use multi-modal data to identify the device owner.
[0146] Figure 2 This is a schematic diagram of the scenario of the multi-modal fusion device owner identification method provided by the embodiments of this application. Figure 2 In it, the terminal device is a mobile phone. The mobile phone can combine the current user's IMU data, touch screen data, breathing audio data, and face recognition score to comprehensively determine whether the current user is the device owner. If the mobile phone identifies that the current user is the device owner, the privacy information is normally displayed. If the mobile phone identifies that the current user is not the device owner, the privacy information is hidden.
[0147] The embodiments of this application have at least the following beneficial effects:
[0148] 1. The terminal device combines the user's face data, the IMU data of the terminal device, the touch screen data of the terminal device, and the user's breathing audio data to comprehensively identify the device owner, which can still perform device owner identification when the scenario changes, and can improve the accuracy and stability of device owner identification.
[0149] 2. When the face recognition result is not ideal, at least one of the IMU data of the terminal device, the touch screen data of the terminal device, and the user's breathing audio data can be combined to further determine whether it is the device owner, which can expand the device owner recognition scenario and improve the accuracy and stability of device owner recognition.
[0150] 3. The IMU data and touch screen data of the terminal device reflect the user's usage habits of the terminal device. Combining these two data to further determine whether it is the device owner can make the recognition result more accurate and reliable, that is, the accuracy rate of device owner recognition can be improved.
[0151] The embodiments of the present application are applicable to scenarios involving the recognition of the device owner in the terminal device.
[0152] The device owner recognition method provided by the embodiments of the present application can be applied to terminal devices such as mobile phones, tablet computers, and wearable devices. At this time, the terminal device is the execution subject of the device owner recognition method provided by the embodiments of the present application, and the specific type of the terminal device is not specifically limited in the embodiments of the present application.
[0153] Taking the terminal device as a mobile phone as an example, the device owner recognition method provided by the present application will be specifically described from three aspects: device owner template entry, device owner recognition based on the device owner template, and device owner template update.
[0154] (1) Device owner template entry
[0155] The device owner template is usually pre-entered or stored information used to describe the characteristics of the device owner. The device owner template can include a face template, a breathing sound feature template, and a cross-modal feature template. Among them, the face template is pre-entered or stored information used to describe the face characteristics of the device owner. For example, it can be the device owner's face image. The breathing sound feature template is pre-entered or stored information used to describe the breathing sound characteristics of the device owner. The cross-modal feature template is pre-entered or stored information used to describe the comprehensive characteristics of the device owner in multiple modalities.
[0156] In practice, a mobile phone can enter the device owner template of one device owner, or can enter the device owner templates of multiple device owners. That is to say, a mobile phone can have one device owner template or multiple device owner templates.
[0157] In practice, one device owner template usually corresponds to one device owner.
[0158] Optionally, in one device owner template, there can be one or more face templates of the device owner, one or more breathing sound feature templates, and one or more cross-modal feature templates of the device owner.
[0159] In the embodiments of the present application, the process of entering the machine owner template will be described from the following aspects: 1. Entering the face template; 2. Entering the breath sound feature template; 3. Entering the cross-modal feature template.
[0160] 1. Entering the face template.
[0161] In the embodiments of the present application, the machine owner can enter the faces of multiple people one after another. For example, the machine owner can first enter their own face, and then enter the faces of their father and mother. For example, the machine owner, Xiao A, can enter their own face, and then enter the faces of their father, Da B, and mother, Da C. When Xiao A only enters their own face, there is one machine owner. At this time, the mobile phone can retain one machine owner identity identifier (IDentifier, ID), for example, A. If after Xiao A enters their own face, they enter the faces of their father, Da B, and mother, Da C one after another, then the mobile phone will retain 3 machine owner IDs, for example, A, B, C. It can be understood that the embodiments of the present application do not limit the representation method of the machine owner ID. For example, it can be Chinese characters, letters, numbers, or a combination of numbers and letters, etc.
[0162] It can be understood that in the embodiments of the present application, face entry is the same as entering the face template.
[0163] Please refer to Figure 3 , Figure 3 which shows a schematic diagram of the process of entering the face template. As shown in Figure 3 (a) in, the machine owner, Xiao A, can click on the biometric and password control 302 on the settings interface 301 of the settings application (an application dedicated to performing system settings functions). The mobile phone can respond to the user's click on the biometric and password control 302 and enter the biometric and password interface 303. After that, as shown in Figure 3 (b) in, the mobile phone can respond to Xiao A's further click on the "Face Recognition" control 304 and enter the input lock screen password interface 305. After that, as shown in Figure 3 (c) in, the mobile phone can respond to Xiao A's further input of the correct lock screen password and enter the face recognition interface 306. After that, as shown in Figure 3 (d) in, the mobile phone can respond to Xiao A's further click on the "Start Entry" control 307 and enter the face entry interface 308. As shown in Figure 3 (e) in, on the face entry interface 308, Xiao A aligns their face with the face collection frame, and the mobile phone captures Xiao A's face image, completing the entry of Xiao A's face template. As shown in Figure 3 (f) in, when the mobile phone completes the entry of Xiao A's face template, it can display "Entry Successful" on the face entry interface 308.
[0164] According to asFigure 3 In a similar operation mode, Little A can help his father and mother input their faces on the mobile phone, so as to complete the input of the face templates of his father and mother by the mobile phone.
[0165] In the embodiment of the present application, after multiple face templates are input, the mobile phone can identify the owners corresponding to each face template. For example, if the face templates of Little A, Big B, and Big C are stored on the mobile phone. When Little A holds the mobile phone for face unlocking, the mobile phone can identify that the current operator is the owner, and specifically, the owner is Little A. That is to say, the mobile phone can determine which specific owner is operating the mobile phone. Of course, any owner can perform face unlocking on the mobile phone.
[0166] In some application scenarios, the same user can be input once or multiple times. For example, the same user can perform multiple face inputs, and each face input can be assigned a corresponding owner ID by the mobile phone. That is, there can be multiple owner IDs for the same user. Or rather, one user can correspond to multiple owner templates. In the case where a user is assigned multiple owner IDs, a user can be regarded as multiple owners by the mobile phone.
[0167] 2. Input of breath sound feature template.
[0168] In the embodiment of the present application, when the input of the face template of the owner is successful, the mobile phone can continue to collect the breath audio data of the owner through the microphone, and judge whether the breath audio data contains a breath sound. If it contains a breath sound, further extract the breath sound feature therefrom, and determine the extracted breath sound feature as the owner's breath sound feature template. That is to say, the trigger condition for inputting the breath sound feature template (or called the breath feature input condition) can be the successful input of the face template of the owner.
[0169] It should be noted that, within a short period of time after the successful input of the owner template, for example, within 5 seconds, the owner who has just input the face template usually will not immediately put down the mobile phone or leave the mobile phone. That is, the breath audio data collected within the short period of time after the successful input of the face template usually has the breath sound of the owner. That is to say, when the input of the face template of the owner is successful, the mobile phone continues to collect the breath audio data within the short period of time after the successful input of the face template, which can ensure that the collected breath audio data contains the breath sound of the owner, thereby ensuring the accurate and effective input of the owner's breath sound feature template.
[0170] Among them, the above-mentioned shorter duration is usually a preset shorter duration value. In practice, the above-mentioned shorter duration can be any duration from 0.5 seconds to 5 seconds. Since the longer the audio data, the more accurate the provided audio features, in order to ensure the accuracy of the breathing sound feature template of the registered phone owner, when registering the breathing sound feature template of the phone owner, it is usually to collect the breathing audio data within 5 seconds after the successful registration of the face template, that is, to collect the breathing audio data for a relatively long period of time to ensure the accuracy of the breathing sound feature template of the phone owner.
[0171] It can be understood that if there is no breathing sound in the collected breathing audio data, the phone may not register the breathing sound feature template of the phone owner.
[0172] In some application scenarios, in the case of detecting that there is no breathing sound in the above-mentioned breathing audio data, the phone can continue to collect breathing audio data and determine whether the breathing audio data contains a breathing sound when the subsequent face unlock is successful. If it contains a breathing sound, further extract the breathing sound features from it, and determine the extracted breathing sound features as the breathing sound feature template of the phone owner. That is to say, the trigger condition for registering the breathing sound feature template (or called the breathing feature registration condition) can also be that the face unlock is successful and the phone has not registered the breathing sound feature template of the phone owner currently performing the unlock when the face unlock is successful.
[0173] Generally speaking, the trigger condition for registering the breathing sound feature template (or called the breathing feature registration condition) can be that the face template of the phone owner is successfully registered, or the face unlock is successful and the phone has not registered the breathing sound feature template of the phone owner currently performing the unlock when the face unlock is successful.
[0174] Regardless of the trigger condition for collecting the breathing audio data, the collection parameters for obtaining the breathing sound feature template can be fixed. For example, the collection parameters can be: single-channel collection, sampling rate of 16 kHz, and collection duration of 5 seconds.
[0175] It should be noted that since the breathing sound is usually relatively weak, if the environment where the phone owner performs face registration is complex, for example, the phone owner is in a noisy scene such as a concert venue or a vegetable market, in this case, it is difficult for the phone to identify the breathing sound from the further collected breathing audio data of the phone owner. Therefore, first judge whether there is a breathing sound in the breathing audio data, and only when it is determined that there is a breathing sound, perform the extraction of the breathing sound features, which can reduce the loss of unnecessary computing resources.
[0176] The following will further illustrate how to extract the breathing sound feature template corresponding to the phone owner from the breathing audio data of the phone owner.
[0177] In an embodiment of the present application, during the process of inputting the respiratory sound feature template, the mobile phone can call a pre-stored audio processing module. The audio processing module extracts the respiratory sound features from the collected respiratory audio data of the phone owner. After that, the mobile phone can use the extracted respiratory sound features as the respiratory sound feature template of the phone owner.
[0178] For further reference Figure 4A 。 Figure 4A FIG. is a schematic diagram of the process of the audio processing module provided by the embodiment of the present application for processing respiratory audio data. The audio processing module can extract the respiratory sound features from the respiratory audio data of the phone owner by executing S401-S405. Details are as follows:
[0179] S401, the audio processing module performs audio preprocessing on the respiratory audio data to obtain the mel cepstrum and mel cepstral coefficients of the respiratory audio data.
[0180] Figure 4B FIG. shows a schematic diagram of the process of the audio processing module performing audio preprocessing on the respiratory audio data. As Figure 4B shown, the process of the audio processing module performing audio preprocessing on the respiratory audio data may include the following steps S4011-S4014.
[0181] S4011, the audio processing module performs a short-time Fourier transform on the respiratory audio data to convert the respiratory audio data into a time-frequency diagram form.
[0182] In an embodiment of the present application, the collected respiratory audio data is usually a time-domain audio signal. Therefore, the short-time Fourier transform can be used to convert the respiratory audio data of the time-domain audio signal into a two-dimensional time-frequency signal form. After that, the two-dimensional time-frequency signal can be converted into a time-frequency diagram in combination with the amplitude of the two-dimensional time-frequency signal.
[0183] Among them, the short-time Fourier transform (STFT) is a classic signal analysis method that decomposes a signal into components in time and frequency. STFT performs local analysis on the signal with a moving window of a certain length, improving the resolution of time-domain and frequency-domain analysis. The main parameters of STFT include the window function, window length, and step size. Commonly used window functions include the Hanning Window, Hamming Window, Blackman Window, etc.
[0184] In some application scenarios, when performing short-time Fourier transform, the window function used can be a Hanning window, the window length can be 0.025 seconds, and the step size can be 0.01 seconds. The number of time domain points in the time-frequency signal can be represented by num_samples, where num_samples = sampling duration ÷ step size + 1. In this application scenario, the value of num_samples can be 501, where 501 = 5 ÷ 0.01 + 1.
[0185] S4012. The audio processing module performs a frequency-Mel scale transformation on the time-frequency diagram to convert the frequency of the time-frequency diagram into Mel scale, thereby obtaining a Mel spectrogram.
[0186] Here, the mapping formula between frequency and Mel scale is as follows:
[0187]
[0188] where f is the frequency, mel(f) is the Mel scale, and × is the multiplication sign.
[0189] In some application scenarios, during the process of converting the time-frequency diagram into a Mel spectrogram, the parameters involved can include: the minimum frequency is 125.0 Hz, the maximum frequency is 7500.0 Hz, and the number of Mel filter banks is 64. Here, the number of Mel filter banks is 64, which means that the number of Mel filter banks is relatively large. It should be noted that the more the number of Mel filter banks, the more accurate the obtained Mel spectrogram can be, which helps to refine the low-frequency information and coarsen the high-frequency information of the audio signal.
[0190] S4013. The audio processing module performs a logarithmic transformation on the Mel spectrogram to obtain a logarithmic Mel spectrogram (also known as Mel cepstrum).
[0191] Among them, the logarithmic Mel spectrogram (log mel-specgram) is a method for representing audio features. It converts the spectral information of the audio signal into the energy distribution under the Mel scale and takes the logarithm to enhance the distinguishability of the features.
[0192] Here, taking the logarithm of the Mel spectrogram in S4012 obtains a logarithmic Mel spectrogram. The size of the obtained logarithmic Mel spectrogram can be num_samples × n_mels. Among them, num_samples is the number of time domain points in the time-frequency signal, and n_mels is the number of Mel filter banks.
[0193] S4014. The audio processing module performs a discrete cosine transform on the Mel cepstrum to obtain Mel cepstral coefficients.
[0194] Among them, the Mel cepstral coefficients are characteristic coefficients obtained through a series of signal processing steps, which are used to describe the energy distribution of sound at different frequencies. The Mel cepstral coefficients are widely used in speech recognition and audio feature extraction. They can effectively capture the important features of audio signals while reducing the dimension of features, which helps to improve the accuracy and efficiency of pattern recognition.
[0195] Here, perform a discrete cosine transform (DCT) on the Mel cepstrum obtained in step S4013, and retain the first 40 MFCC coefficients, that is, n_mfcc = 40, to obtain the MFCC coefficients to be used. At this time, the size of the MFCC coefficients is num_samples × n_mfcc. Among them, num_samples is the number of time domain points in the time-frequency signal.
[0196] It should be noted that the larger the value of n_mfcc, the more information retained by the MFCC coefficients, which helps to make the extracted breath sound features more accurate.
[0197] It can be understood that the embodiments of the present application do not specifically limit the values of the various parameters involved in audio preprocessing. For example, in some scenarios, the number of Mel filter banks can also be other values except 64. For another example, in some other scenarios, the number of MFCC coefficients retained can also be other values except 40. The values of the above parameters can be adjusted according to actual needs.
[0198] S402, the audio processing module divides the Mel cepstrum into multiple Mel cepstrum segments.
[0199] Here, since the data volume of the complete Mel cepstrum is usually relatively large, dividing the Mel cepstrum into multiple Mel cepstrum segments can simplify feature representation, reduce computational costs, and can more effectively capture the local features of audio signals, thereby improving the accuracy of breath sound recognition.
[0200] In practical applications, it is found that if the segmentation granularity is too fine, on the one hand, the computational amount will increase, and on the other hand, the breath sound recognition accuracy will decrease. Therefore, in practice, the length of the Mel cepstrum segment is usually set to about 1 second. For example, it can be 0.96 seconds. The reason is that the duration of a person's exhalation or inhalation is about 1 second. Changing the Mel cepstrum segment to a segment with a length of about 1 second can enable each Mel cepstrum segment to more completely describe the breath characteristics, thereby further improving the accuracy of breath sound recognition.
[0201] In practice, the Mel cepstrum can be segmented with a window length of 0.96 seconds and a step size of 0.24 seconds to divide the Mel cepstrum into multiple Mel cepstrum segments.
[0202] In the process of segmenting the mel cepstrum, when the mel cepstrum cannot be exactly segmented into an integer number of segments, that is, when it cannot be segmented into an integer number of segments of 0.96 seconds, zeros can be padded to the mel cepstrum, and the padded mel cepstrum will be segmented into an integer number of mel cepstrum segments.
[0203] In addition, since the model input size of the breath sound recognition model is usually constant. For example, the model input size of the breath sound recognition model can be the number of mel cepstrum segments corresponding to 5 seconds of breath sound. However, in some application scenarios, since 5 seconds of audio data is not collected when collecting breath audio data. For example, maybe only 1 second of audio data is collected. To ensure that there is enough data volume to input into the breath sound recognition model, it is necessary to further pad zeros to the mel cepstrum.
[0204] S403, the audio processing module inputs the mel cepstrum segments into a pre-trained breath sound recognition model to obtain the breath sound confidence that the breath audio data is recognized as a breath sound.
[0205] Wherein, the above-mentioned breath sound confidence is the confidence or probability that the breath audio data is determined to be a breath sound. The input of the breath sound recognition model is multiple consecutive mel cepstrum segments, and the output at least includes the breath sound confidence.
[0206] Here, inputting the mel cepstrum segments of the breath audio data into the breath sound recognition model, the breath sound recognition model can output the confidence or probability that the breath audio data is recognized as a breath sound, and thus it can be further determined whether the breath audio data contains a breath sound.
[0207] It should be noted that since the mel cepstrum usually carries rich audio information, identifying whether the audio is a breath sound through the mel cepstrum segments of the audio can ensure the accuracy of breath sound recognition, that is, it can enable the breath sound recognition model to more accurately determine whether the audio is a breath sound.
[0208] S404, the audio processing module combines the breath sound confidence with a preset threshold a to further determine whether the breath audio data is a breath sound.
[0209] Wherein, the threshold a is a preset threshold for determining whether the audio is a breath sound. In some usage scenarios of the embodiments of the present application, the threshold a can also be referred to as the first determination threshold.
[0210] Here, if the breath sound confidence is greater than the threshold a, it is considered that the breath audio data is a breath sound. On the contrary, if the breath sound confidence is less than or equal to the threshold a, it is considered that the breath audio data does not contain a breath sound.
[0211] S405. When the respiratory audio data is determined to be breath sounds, the audio processing module inputs the mel cepstral coefficients in S401 into a pre-trained breath sound feature extraction model to extract the breath sound features of the respiratory audio data.
[0212] Here, when the mobile phone is inputting the breath sound feature template of the machine owner, the breath sound features extracted in S405 can be input as the breath sound feature template of the machine owner.
[0213] It should be noted that mel cepstral coefficients can effectively capture the important features of audio signals and have a small data dimension. When it has been determined that the respiratory audio data is breath sounds, inputting the mel cepstral coefficients of the respiratory audio data into the breath sound feature extraction model to extract the breath sound features of the respiratory audio data can achieve fast and accurate extraction of the breath sound features of the respiratory audio data.
[0214] Next, further combined with Figure 4C to illustrate the breath sound recognition model in S403 and the training process of the breath sound recognition model.
[0215] In the embodiment of the present application, in order to increase the generalization of the breath sound recognition model, in addition to the breath sound category, other common human-related sound events are added, a total of seven types of sounds, including quiet, speaking, breath sounds, coughing, sneezing, snoring, and environmental noise. In the scenario where the breath sound recognition model can recognize 7 types of sounds, the output of the breath sound recognition model can be the confidence levels of these 7 types of sounds. Specifically, for each type of sound, a corresponding confidence level can be output, and for 7 types of sounds, 7 confidence levels can be output. Generally, the sum of the confidence levels of the 7 types of sounds output by the breath sound recognition model is usually 1.
[0216] It can be understood that in some application scenarios, the output of the breath sound recognition model can also be sound category information. For example, it can output "breath", indicating that the final recognition result of the breath sound recognition model is breath.
[0217] Since it is found in actual applications that the generation principle of breath sounds is that during the inhalation and exhalation processes, the air in the environment is disturbed, thus generating sounds, similar to the sound of wind. The disturbance caused by exhalation is greater than that of inhalation, so generally the audio amplitude of exhalation is higher than that of inhalation. From a large time scale, breathing is a periodic and slow continuous action event, inhaling and exhaling alternately at a certain frequency, and there are regular pauses in between. This characteristic of breathing is significantly different from coughing, sneezing, etc. From a small time scale, each inhalation and exhalation causes air disturbance, and the breath sounds caused by this disturbance are different from the sounds of sneezing, coughing, etc.
[0218] Therefore, in order to improve the accuracy of breath sound recognition, the breath sound recognition model can include two parts. One is the frame-level sound representation (or single-segment sound representation); the other is the overall sound representation (also known as continuous-segment sound representation). The frame-level sound representation represents a single Mel cepstrum segment obtained from S402. When the segment length is 0.96 seconds, the frame-level sound representation can represent the refined sound features of 0.96 seconds of audio. Since a complete breath is longer than 0.96 seconds, in order to represent the sound features of long audio as a whole, each 0.96-second audio segment in S402 can be re-spliced together after frame-level representation to fuse and obtain the context sound features in the long audio, that is, to obtain the overall sound representation.
[0219] Figure 4C Schematic diagram of the breath sound recognition model provided by the embodiments of the present application. Figure 4C In [the figure], the breath sound recognition model can include a feature extraction layer, a feature splicing layer, and a classifier connected in sequence. Among them, the classifier includes a bidirectional long short-term memory network (BiLSTM) and a fully connected layer connected in sequence.
[0220] Combined with Figure 4C , the input of the breath sound recognition model is a continuous plurality of Mel cepstrum segments, and the output is the confidence of 7 types of sounds. After the continuous plurality of Mel cepstrum segments are input into the breath sound recognition model, for each Mel cepstrum segment, the breath sound recognition model can call a pre-trained feature extractor to extract the frame-level sound representation of the Mel cepstrum segment. In this way, each Mel cepstrum segment can obtain a frame-level sound representation. Then, the breath sound recognition model can sequentially splice the frame-level sound representations corresponding to each Mel cepstrum segment to obtain the overall sound representation (also known as continuous-segment sound representation). After that, the continuous-segment sound representation can be input into the classifier for sound classification, so as to obtain the confidence of being determined as various sounds, that is, the confidence corresponding to 7 types of sounds respectively.
[0221] The training process of the breath sound recognition model can include the following aspects.
[0222] (1) Obtain training samples
[0223] Here, in order to increase the generalization of the breath sound recognition model and improve the accuracy of the breath sound recognition model for recognizing breath sounds, a rich training sample is needed to train the breath sound recognition model, and the training sample specifically needs to cover the following 7 types of sounds: quiet, speaking, breath sound, cough, sneeze, snore, and environmental noise.
[0224] In the process of collecting training samples, for the same user, audio of various types of sounds of the user can be collected. Subsequently, the collected audio can be segmented into many audio segments of a fixed length, and the above fixed length can be 5 seconds. It can be understood that the length of the audio segment is related to the specific usage scenario of the breath sound recognition model. For example, when the breath sound recognition model is used to recognize the sound category of 5-second breath audio data, the length of the audio segment can be 5 seconds.
[0225] In the process of processing the collected audio into training samples, for each audio segment, the operations of S401 - S402 can be performed on the audio segment to obtain a continuous plurality of mel cepstrum segments corresponding to the audio segment. Subsequently, the audio segment, the continuous plurality of mel cepstrum segments corresponding to the audio segment, and the sound category corresponding to the audio segment can be used as training samples. Thus, a sample set for training the breath sound recognition model can be obtained. For the convenience of description, the sample set for training the breath sound recognition model can be referred to as an audio sample set.
[0226] (2) Training the feature extractor
[0227] Here, the initial feature extractor can be various models that can be used to extract features. For example, it can be a Convolutional Neural Network (CNN), a Recurrent Neural Network (RNN), etc.
[0228] Further combined with Figure 4C , the process of training the feature extractor is equivalent to training a single-frame breath sound recognition model with only single-frame input. Among them, the input of the single-frame breath sound recognition model is a single mel cepstrum segment, and the output is the sound category of the corresponding audio. The input of the feature extractor is a single mel cepstrum segment, and the output is the frame-level sound representation corresponding to the input mel cepstrum segment.
[0229] Figure 4C In, the single-frame breath sound recognition model includes a sequentially connected feature extractor and a fully connected layer. In the process of training the single-frame breath sound recognition model, for the audio samples in the audio sample set, a single mel cepstrum segment in the audio sample can be used as the input, and the sound category in the audio sample can be used as the output to complete the training of the single-frame breath sound recognition model, thereby completing the training of the feature extractor.
[0230] After the feature extractor is trained, the parameters of the feature extractor can be solidified or encapsulated so that the feature extractor can be called or used as a feature extraction tool.
[0231] (3) Training the breath sound recognition model
[0232] Here, the feature extractor in the initial breath sound recognition model is the feature extractor trained through the operation in the above step (2). That is to say, during the training of the breath sound recognition model, the feature extractor is only called or used as a feature extraction tool.
[0233] During the process of training the breath sound recognition model using the audio sample set, for each audio sample, multiple consecutive Mel cepstrum segments in the audio sample can be used as the input of the breath sound recognition model, and the corresponding sound category of the audio sample can be used as the output to train the initial breath sound recognition model to obtain the trained breath sound recognition model.
[0234] In practice, the trained breath sound recognition model can be called or used as a tool for recognizing breath sounds.
[0235] It can be understood that in the embodiments of the present application, the execution entity for training the feature extractor and the breath sound feature recognition model can be a mobile phone or other devices. For example, it can be a server. That is to say, the model can be trained on a mobile phone or trained on other devices and then the trained model can be moved into the mobile phone.
[0236] The embodiments of the present application do not make specific limitations on which device the model is trained on.
[0237] Next, further combined with Figure 4D and Figure 4E to illustrate the breath sound feature extraction model in S405 and the training process of the breath sound feature extraction model.
[0238] Figure 4D is a schematic diagram of the breath sound feature extraction model provided by the embodiments of the present application. Figure 4D In, the breath sound feature extraction model may include a convolutional layer, two multiscale residual layers (Multiscale Resnet Layer), and a bottleneck layer (Bottleneck Layer) connected in sequence. It can be understood that in the breath sound feature extraction model, the number of multiscale residual layers may be more than two.
[0239] Figure 4E is a schematic structural diagram of a multiscale residual layer in the breath sound feature extraction model provided by the embodiments of the present application. Figure 4E In, cross-fusion of convolutional features of multiple scales is adopted. For example, cross-fusion of 5×5 convolutional features and 3×3 convolutional features is adopted, which helps to extract richer and more comprehensive features.
[0240] Combined with Figure 4D and Figure 4E, the input of the breath sound feature extraction model is the Mel cepstral coefficients, and the output is the breath sound features corresponding to the Mel cepstral coefficients. The breath sound feature extraction model compresses the input Mel cepstral coefficients through a convolutional layer; then, through multiple multi-scale residual layers, it extracts the features expressed by the corresponding audio from the compressed data. Among them, the multi-scale residual module adds cross-fusion of convolutional features of different scales on the framework of the residual network, and can obtain richer and more comprehensive breath sound features; finally, through the bottleneck layer, the features of each module are fused to output the breath sound features.
[0241] See further Figure 4D , after the breath sound feature extraction model outputs the breath sound features, the obtained breath sound features can further pass through a fully connected layer and a softmax function layer (or called the softmax layer) to output the user ID confidence and the recognition result. For example, it can output the confidence of being recognized as user 1, the confidence of being recognized as user 2... the confidence of being recognized as user N, etc. At the same time, the user ID with the highest corresponding confidence can also be used as the recognition result.
[0242] It can be understood that Figure 4D In, adding a sequentially connected fully connected layer and a softmax function layer on the basis of the breath sound feature extraction model is actually a user recognition model. The input of the user recognition model is the Mel cepstral coefficients of the audio, and the output is the user ID confidence and / or the recognition result.
[0243] The training process of the breath sound feature extraction model can include the following aspects.
[0244] (1) Obtain training samples
[0245] In the process of obtaining training samples, breath audio generated by the same user breathing in different ways (such as normal breath sound, rapid breathing, and deep breathing) at different distances (such as the face is 30 cm, 50 cm, 100 cm, etc. away from the mobile phone) can be obtained. In this way, multiple breath audio can be obtained for each user.
[0246] After that, each obtained breath audio can be cut into multiple audio segments of a fixed length, and the above fixed length can be 5 seconds. It can be understood that the length of the audio segment is related to the specific usage scenario of the breath sound feature extraction model. For example, when the breath sound feature extraction model is used to extract the breath sound features of 5-second breath audio data, the length of the audio segment can be 5 seconds.
[0247] In the process of processing the obtained respiratory audio into training samples, for each audio segment, the operations of S401 above (that is, the operations of audio preprocessing) can be performed on the audio segment to obtain the Mel cepstral coefficients corresponding to the audio segment. Then, the audio segment, the Mel cepstral coefficients corresponding to the audio segment, and the user ID corresponding to the audio segment can be used as training samples. Thus, a sample set of the user recognition model including the respiratory sound feature extraction model can be obtained. For the convenience of description, the sample set for training the user recognition model can be called the respiratory sample set.
[0248] (2) Training the respiratory sound feature extraction model
[0249] Continuing to combine Figure 4E , training the respiratory sound feature extraction model is actually training the user recognition model.
[0250] In some application scenarios, the respiratory sound feature extraction model can be trained by the following training method 1.
[0251] Training method 1: In the process of training the user recognition model, for each respiratory sample in the respiratory sample set, the Mel cepstral coefficients in the respiratory sample can be used as the input, and the user ID in the respiratory sample can be used as the output to complete the training of the user recognition model, thereby completing the training of the respiratory sound feature extraction model.
[0252] Due to the differences in respiratory sounds between different individuals, which mainly come from physiological differences (such as the lungs, chest and abdominal muscles, and the upper respiratory tract) and habit differences (such as chest breathing or abdominal breathing, etc.), the above differences are reflected in respiratory sounds in the following three points: The first point is the duration of inhalation, exhalation, and pause; the second point is the change in air flow velocity during inhalation and exhalation; the third point is the change in air flow volume during inhalation and exhalation.
[0253] Also, since the difficulty of respiratory sound identity authentication lies in that even for the respiratory sounds of the same user, there are certain differences when the user's current state changes, the position of the mobile phone relative to the face changes, or the external environment changes.
[0254] Therefore, in the process of training the respiratory sound feature extraction model, data augmentation and contrastive learning of respiratory audio are required to improve the generalization ability and robustness of the respiratory sound feature extraction model.
[0255] In the embodiments of the present application, the following three data augmentation methods can be used to perform data augmentation on the audio in the training samples.
[0256] Data augmentation method 1: Add random noise with different signal-to-noise ratio levels at random positions in the audio to mimic breathing audio in various scenarios. The added noise can include one or more of the following: household noise, outdoor noise, office noise, transportation noise, animal noise, etc.
[0257] Data augmentation method 2: Randomly change the speed of the audio to mimic fast breathing and slow breathing.
[0258] Data augmentation method 3: Randomly change the volume range of the audio to mimic audio collected at close range and audio collected at long range.
[0259] In some other application scenarios, the breathing sound feature extraction model can also be trained through the following training method 2.
[0260] Training method 2: During the iterative training of the breathing sound feature extraction model, the breathing sound feature extraction model can be trained in batches. For example, there can be K batches of training samples, that is, K times of training. Among them, each batch of training samples can have M training samples, and these M training samples come from N users. Both M and N are integers greater than 1.
[0261] It should be noted that the M training samples in each batch come from N users, which means that in the case of multi-batch training, the audio of the same user may be selected multiple times. In this way, the accuracy and stability of the trained model can be further improved.
[0262] In practice, in order to further improve the generalization ability of the model, for each batch of M training samples, one of them can be labeled as a positive sample and the rest can be labeled as negative samples. That is, in this training method 2, the training samples used for training can have two labeling tags, one for labeling the user ID and the other for labeling positive and negative samples.
[0263] As an example, for each batch of M training samples, one can be randomly selected from these M training samples and labeled as a positive sample first, and then the remaining training samples in these M training samples can be labeled as negative samples.
[0264] In addition, for each training sample (or called breathing sample), data augmentation can be performed on this training sample by using data augmentation methods, so as to obtain an augmented sample (or called augmented breathing sample) for this training sample. For example, if the above 3 data augmentation methods are used to perform data augmentation on this training sample, then 3 augmented samples can be obtained for each training sample. The augmented samples can be used to train the breathing sound feature extraction model.
[0265] When the above three data augmentation methods are used to perform data augmentation on each training sample, M training samples in each batch can be expanded into 3M augmented samples.
[0266] It should be noted that training the breath sound feature extraction model with augmented samples can improve the generalization ability of the trained model. That is to say, it can enable the trained breath sound feature extraction model to accurately extract the breath sound features of users in various situations.
[0267] In the embodiment of the present application, training the breath sound feature extraction model using the above training method two can have the following beneficial effects:
[0268] 1. When training the breath sound feature extraction model, marking one of the M training samples in the same batch as the positive sample and the rest as negative samples can make the trained breath sound feature extraction model have stronger recognition ability for breath sounds. That is to say, it can improve the generalization ability and robustness of the model, thereby helping the trained breath sound feature extraction model to extract breath sound features more accurately and effectively.
[0269] 2. Performing data augmentation on each training sample and training the breath sound feature extraction model with the augmented samples obtained by data augmentation can improve the generalization ability and robustness of the trained breath sound feature extraction model. That is to say, it can enable the trained breath sound feature extraction model to accurately extract the breath sound features of users in various situations.
[0270] In some optional implementation manners of the embodiment of the present application, during the process of training the breath sound feature extraction model (belonging to the user recognition model) using the above training method two, the loss function used may include the following two parts of losses: one is the user recognition cross-entropy loss, and the other is the similarity loss of contrast learning. The user recognition cross-entropy loss is used to describe the accuracy of the breath sound feature extraction model in user recognition, and the contrast learning similarity loss is used to describe the similarity between the features extracted by the breath sound feature extraction model and the true features.
[0271] Exemplarily, the user recognition cross-entropy loss can be shown as in formula (1):
[0272]
[0273] where loss au is the user recognition cross-entropy loss, lable(y ji ) is the true label, predict(y ji) is the user identification prediction probability output by the softmax layer. The significance of this loss function is to evaluate the gap between the true label and the prediction result, loss au The smaller it is, the more accurate the prediction result.
[0274] Exemplarily, the contrastive learning similarity loss can be as shown in formula group (2):
[0275]
[0276]
[0277]
[0278] Among them, loss cl is the contrastive learning similarity loss, the S matrix is the cosine similarity matrix, and each element in the S matrix is used to indicate the cosine similarity between the breath sound features corresponding to two training samples. For example, s M,3M is used to indicate the cosine similarity between the breath sound features corresponding to the Mth training sample and the breath sound features corresponding to the 3Mth training sample. The values of the elements in the S matrix can be any value from 0 to 1.
[0279] The L matrix is the target similarity matrix, and each element in the L matrix is used to indicate whether two training samples come from the same user. If two training samples come from the same user, they are considered similar features, and the similarity value is 1; if two training samples come from different users, they are considered dissimilar features, and the similarity value is 0. That is, the values of the elements in the L matrix are 0 or 1.
[0280] The above contrastive learning similarity loss loss cl can be calculated based on the cosine similarity matrix S and the target similarity matrix L. When the contrastive learning similarity loss loss cl is smaller, the similar training samples are closer in the feature space, while the dissimilar training samples are farther away in the feature space.
[0281] In practice, the user identification cross-entropy loss loss au and the contrastive learning similarity loss loss cl can each correspond to a weight.
[0282] In practice, if the weight corresponding to the user identification cross-entropy loss loss au can be a1, and the weight corresponding to the contrastive learning similarity loss loss clThe corresponding weight can be a2. Exemplarily, the overall model training loss loss can be as shown in the following formula (3). Among them, a1 and a2 are preset weight values. In practice, the value ranges of a1 and a2 are 0-1, and the sum of a1 and a2 can be 1.
[0283] loss = a1×loss au +a2×loss cl (3)
[0284] In the embodiment of the present application, during the iterative training of the breath sound feature extraction model, the model parameters can be updated by the gradient descent method, so that the overall model training loss loss becomes smaller and smaller, realizing the parameter optimization of the model.
[0285] In the embodiment of the present application, during the model training process, the loss of the model is described from two dimensions of user recognition and contrast learning, so as to optimize and adjust the model, which can ensure that the trained user recognition model can accurately recognize the user and has strong recognition of the breath sound. That is, the breath sound features extracted by the trained breath sound feature extraction model can have strong recognition and can more comprehensively and accurately describe the personal breathing characteristics of the user.
[0286] It should be noted that the higher the recognition of the breath sound by the breath sound feature extraction model, the higher the accuracy or recognition of the extracted breath sound features. When performing the main user recognition based on the breath sound features, the recognition result will be more accurate.
[0287] 3. Cross-modal feature template entry.
[0288] In the embodiment of the present application, when the face template of the main user is successfully entered, the mobile phone can collect breath audio data through the microphone, collect touch screen data through the touch screen sensor, and collect IMU data through the IMU, and obtain the cross-modal feature template of the main user through the above three types of data.
[0289] In the embodiment of the present application, when the collection frequency of the breath audio data is 16khz and the collection duration is 5 seconds, the size of the breath audio data collected by the mobile phone is 80000 dimensions, where 80000 = 16k×5.
[0290] The size of the IMU data collected by the IMU can be 50×90, where 50 indicates that 50 times of IMU data are collected, 90 indicates the dimension of the IMU data, and the IMU data can include the acceleration, pitch angle, roll angle, etc. of the mobile phone in each direction. The collection frequency of the IMU data can be 50 times per second. That is, the mobile phone can obtain the above 50×90-dimensional IMU data by collecting the IMU data within 1 second after the successful entry of the main user's face template.
[0291] The size of the touch screen data collected by the touch screen sensor can be 200×17. Among them, 200 indicates the most recent 200 touch screen data, and 17 indicates the dimension of the touch screen data. Each touch screen data can include the number of swipes and clicks of the touch screen, etc. That is, after the face template of the phone owner is successfully entered, the phone can collect the most recent 200 touch screen data that occurred on the phone to obtain the above-mentioned 200×17-dimensional touch screen data.
[0292] It can be understood that in the embodiments of the present application, the sizes of the above three types of modal data can be adaptively adjusted in combination with the actual application scenario, and the embodiments of the present application do not specifically limit the sizes of various modal data.
[0293] Next, it will be further described how to fuse the breathing audio data of the phone owner, the touch screen data of the phone, and the IMU data of the phone to obtain the cross-modal feature template of the phone owner.
[0294] In the embodiments of the present application, due to the large amount of breathing audio data, in order to improve the data processing efficiency, the phone can call the pre-stored audio processing module, and the audio processing module processes the breathing audio data to extract the breathing sound features in the above-mentioned breathing audio data. After that, the phone can call the pre-stored cross-modal encoder, and the cross-modal encoder encodes the breathing sound features, touch screen data, and IMU data to obtain the corresponding cross-modal features.
[0295] It should be noted that during the process of first entering the cross-modal feature template of a certain phone owner, if there is no breathing sound in the collected breathing audio data or it is not recognized as a breathing sound, the breathing sound features may not be obtained. In such a case, the phone can temporarily not enter the breathing sound feature template and cross-modal feature template of this phone owner.
[0296] In some application scenarios, when it is detected that there is no breathing sound in the above-mentioned breathing audio data, the phone can continue to collect breathing audio data and determine whether the breathing audio data is a breathing sound when the subsequent face unlocking is successful. If it is a breathing sound, further extract the breathing sound features and enter the extracted breathing sound features as the breathing sound feature template of the phone owner. At the same time, the cross-modal features of the phone owner can be determined by combining the breathing sound feature template of this phone owner, the IMU data, and the touch screen data after the face unlocking is successful, and the cross-modal features can be entered as the cross-modal feature template of the phone owner.
[0297] In combination with the above description, in the embodiments of the present application, the triggering conditions for triggering the entry of cross-modal feature templates (or multi-modal feature entry conditions) include any one of the following: the face template of the device owner is successfully entered, the device owner's face unlocking is successful and the cross-modal feature template of the device owner currently performing the unlocking has not been entered in the terminal device.
[0298] The following will be combined Figure 5A 、 Figure 5B and Figure 5C to illustrate the structure of the cross-modal encoder. Figure 5A FIG. is a schematic structural diagram of the cross-modal encoder provided by the embodiments of the present application. Figure 5B FIG. is a schematic structural diagram of the cross-channel linear encoder in the cross-modal encoder provided by the embodiments of the present application. Figure 5C FIG. is a schematic structural diagram of the cross-channel attention encoder in the cross-modal encoder provided by the embodiments of the present application.
[0299] As Figure 5A shown, the input of the cross-modal encoder is IMU data, touch screen data, and breath sound features, and the output is cross-modal features. The cross-modal encoder can encode the IMU data and touch screen data through two cross-channel linear encoders from top to bottom. For ease of description, the cross-channel linear encoder for encoding the IMU data can be denoted as the first cross-channel linear encoder, and the cross-channel linear encoder for encoding the touch screen data can be denoted as the second cross-channel linear encoder.
[0300] After that, the cross-modal encoder can perform feature extraction on the encoded IMU data, encoded touch screen data, and breath sound features through three Stem layers from top to bottom. For ease of description, the Stem layer for performing feature extraction on the encoded IMU data can be denoted as the first Stem layer, the Stem layer for performing feature extraction on the encoded touch screen data can be denoted as the second Stem layer, and the Stem layer for performing feature extraction on the breath sound features can be denoted as the third Stem layer.
[0301] After that, the cross-modal encoder can fuse the IMU features extracted by the first Stem layer, the touch features extracted by the second Stem layer, and the breath sound features extracted by the third Stem layer to obtain fused features.
[0302] After that, the cross-modal encoder can process the fused features through the fourth Stem layer, the third cross-channel linear encoder, and the multi-target operation module in sequence, and output cross-modal features. As Figure 5A shown, the target operation module can include a downsampling module, a cross-channel linear encoder, and a cross-channel attention encoder connected in sequence.
[0303] It can be understood that the above first, second, third, and fourth are only used to distinguish one element from another in description.
[0304] Combined with Figure 5A , the two important components of the cross-modal encoder are: the cross-channel linear encoder and the cross-channel attention encoder.
[0305] Figure 5B In , the cross-channel linear encoder may include a Linear layer, an activation function, a Linear layer, a norm layer, a conv layer, and an activation function connected in sequence. The cross-channel linear encoder uses the Linear layer for encoding, aligns the feature dimensions of different modalities, and uses the conv layer to implement one-dimensional pointwise convolution for information fusion between channels. Figure 5B In , is the symbol indicating addition, is the symbol indicating the activation function.
[0306] In the embodiments of the present application, the role of the cross-channel linear encoder is mainly to increase the degree of abstraction of features. The input of the cross-channel linear encoder is feature data, and the output is the feature data after further abstraction.
[0307] Figure 5C In , combined with the structural part in the dashed box, the cross-channel attention encoder slices the data input to the input layer (also called the Inputs layer) through the Group layer. Specifically, the data of each modality corresponds to one channel, realizing the cutting of features from the channel dimension; then encoding through the Linear layer; then, merging the sliced features; finally, using the attention mechanism in the channel dimension to promote the fusion of multi-modal information between channels. Figure 5C In , is the symbol indicating addition, is the symbol indicating the activation function, is the symbol indicating multiplication, is the symbol indicating positional encoding.
[0308] Further combined with Figure 5C , Figure 5C Three Group layers from top to bottom are shown in the dashed box, as well as the Linear layers respectively connected to each Group layer. For ease of description, the topmost Group layer can be denoted as the first Group layer, the topmost Linear layer as the first Linear layer, the middle Group layer as the second Group layer, the middle Linear layer as the second Linear layer, the bottommost Group layer as the third Group layer, and the bottommost Linear layer as the third Linear layer.
[0309] Among them, the input end of the first Linear layer is connected to the output end of the first Group layer, the input end of the second Linear layer is connected to the output ends of the first Linear layer and the second Group layer, and the input end of the third Linear layer is connected to the output ends of the second Linear layer and the third Group layer.
[0310] From Figure 5C It can be found that the output data from the topmost first Group layer is encoded the most times, and the output data from the bottommost third Group layer is encoded the fewest times.
[0311] In the embodiments of the present application, features with good quality can be deployed in the lower Group layer, and features with poor quality can be deployed in the upper Group layer. That is, the features deployed in the topmost Group layer are relatively the worst, and the features deployed in the bottommost Group layer are relatively the best. In this way, the features deployed in the bottommost Group layer require the fewest number of Linear layers to pass through, that is, the fewest number of encoding times. The features deployed in the topmost Group layer require the most number of Linear layers to pass through, that is, the most number of encoding times. In this way, the cross-channel attention encoder can combine the quality of each feature and encode each feature to different degrees to better effectively fuse features with good quality and features with poor quality.
[0312] In the embodiments of the present application, after encoding and fusing features with good quality and features with poor quality to different degrees, the fused features can be transposed, and then the transposed fused data is input into the Features layer, which helps to further effectively fuse the features.
[0313] In the embodiments of the present application, the main function of the cross-channel attention encoder is to effectively fuse features with good quality and features with poor quality, strengthen the features, so as to improve the accuracy of the output features.
[0314] In addition, combined with Figure 5A 、 Figure 5B and Figure 5C , the cross-modal encoder adopts a lightweight design as a whole, and lightweight deployment can be achieved.
[0315] (2) Use the owner template for owner identification
[0316] In the embodiments of the present application, the triggering conditions for the mobile phone to enter the owner identification may include the following Conditions 1 - 6. It can be understood that only some triggering conditions are listed in the embodiments of the present application, and there may be other triggering conditions with the change of the actual application scenario.
[0317] Condition 1: The continuous authentication function of the mobile phone is enabled.
[0318] In the embodiments of the present application, when the continuous authentication function is enabled, the mobile phone can automatically lock the target content during use. For example, when the continuous authentication function of the mobile phone is enabled, the mobile phone can lock the service card with the owner's private information.
[0319] Condition 2: The mobile phone detects an operation to open the target application.
[0320] In the embodiments of the present application, when an operation to open the target application is detected, it indicates that the user needs to open the target application at this time. At this time, the mobile phone can be triggered to enter the owner identification. If it is identified as the owner, the target application is opened. If it is identified as a non-owner, the target application is not opened. In this way, the security of the owner's private information can be guaranteed.
[0321] Among them, the target application usually refers to an application that may involve the owner's privacy, such as WeChat, Alipay, photo album, etc. It can be understood that the target application can be set by the mobile phone itself. For example, the mobile phone can set the instant messaging application as the target application, or set the application with a usage frequency higher than a certain frequency threshold as the target application. Of course, the target application can also be an application set by the user according to his own needs. For example, the user can set the photo album application with a relatively high privacy level as the target application. It can be understood that the embodiments of the present application do not make specific limitations on the target application.
[0322] Condition 3: The mobile phone receives a notification message.
[0323] In the embodiments of the present application, when a notification message of a certain application is received, the mobile phone can be triggered to enter the owner identification. If it is identified as the owner, the notification message is previewed. If it is identified as a non-owner, the notification message is hidden and not previewed. In this way, the security of the owner's private information can be guaranteed.
[0324] Condition 4: The mobile phone needs to display a service card, and the service card contains private information, such as bank card number, password, etc.
[0325] In the embodiments of the present application, before the mobile phone displays the service card carrying private information, it can perform owner identification. If it is identified as the owner, the service card is directly displayed. If it is identified as a non-owner, the private information in the service card is hidden. In this way, the security of the owner's private information can be guaranteed.
[0326] Condition 5: The mobile phone detects that the target function is triggered.
[0327] Among them, the target function may be a password-free payment function. Of course, the target function may also be other functions involving the information security or financial security of the owner.
[0328] In an embodiment of the present application, when it is necessary to execute the target function, the mobile phone can be triggered to enter the owner identification. If the person is identified as the owner, password-free payment can be made. If the person is identified as not the owner, password-free payment cannot be made. For example, a password must be entered to complete the payment.
[0329] Condition 6: The mobile phone detects that the target interface is triggered.
[0330] The target interface may be an interface with the private information of the owner, for example, an interface displaying the content of a short message.
[0331] In an embodiment of the present application, when it is necessary to open the target interface, the mobile phone can be triggered to enter the owner identification. If the owner is identified, the target interface is opened. If the owner is not identified, the target interface is not opened, thereby ensuring the security of the owner's privacy information.
[0332] When the triggering conditions for owner identification are met, the mobile phone enters owner identification. The following will further explain how the mobile phone performs owner identification.
[0333] Figure 6A A schematic diagram of the process of mobile phone owner identification based on multimodal fusion is shown.
[0334] S601, the mobile phone determines whether a face is detected, and if no face is detected, S602 is executed. If a face is detected, S605 is executed.
[0335] Here, the mobile phone can analyze whether the detected image is a face image, for example, by analyzing the pixel value of the image to determine whether it is a face image. If it is a face image, the mobile phone can identify the face of the owner in combination with the pre-recorded face template of the owner and obtain a face recognition score. The face recognition score is used to describe the degree of similarity between the face to be tested and the face template.
[0336] S602, when no face is detected, the mobile phone determines whether breathing sound is detected. If breathing sound is detected, step S603 is executed. If no breathing sound is detected, it indicates that the detection result is unreliable, and step S604 is executed.
[0337] In an embodiment of the present application, when no face is detected, the mobile phone can continue to collect the breathing audio data of the user to be tested, and extract the breathing sound features to be tested in the breathing audio data to be tested, so as to determine whether the user to be tested is the owner of the machine by comparing the similarity between the breathing sound features to be tested and the breathing sound feature template in the owner template.
[0338] Among them, the mobile phone can call a pre-stored audio processing module (see Figure 4A ), and the audio processing module processes the to-be-detected respiratory audio data to extract the to-be-detected respiratory sound features in the above-mentioned to-be-detected respiratory audio data. Among them, the above-mentioned to-be-detected respiratory audio data is the respiratory audio data of the to-be-detected user. The to-be-detected respiratory sound features are the respiratory sound features in the to-be-detected respiratory audio data.
[0339] In the embodiments of the present application, the acquisition parameters of the to-be-detected respiratory audio data can be: single-channel acquisition, the sampling rate is 16 kHz, and the acquisition duration is t seconds. Among them, the value range of t is 0.5 second to 5 seconds.
[0340] In practice, the acquisition duration of the to-be-detected respiratory audio data is not fixed. The acquisition duration of the to-be-detected respiratory audio data is usually related to the usage scenario, and the acquisition duration is usually positively correlated with the recognition accuracy of the owner recognition in the usage scenario and negatively correlated with the recognition speed of the owner recognition in the usage scenario. For example, for a usage scenario with very high requirements for the recognition accuracy of the owner, the acquisition duration can be as long as possible; for a usage scenario with very high requirements for the recognition speed of the owner, the acquisition duration can be as short as possible.
[0341] For another example, for the usage scenario of a user logging in to an account, this usage scenario has very high requirements for the recognition accuracy of the owner. At this time, the to-be-detected respiratory audio data collected can be 5 seconds. For the usage scenario of viewing photos, in this usage scenario, if the user waits for a long time to open the photo, such as waiting for 5 seconds, obviously the user experience will be relatively poor. That is to say, this usage scenario has high requirements for the recognition speed of the owner. At this time, in order to improve the recognition speed of the owner, the to-be-detected respiratory audio data collected can be 0.5 second.
[0342] Since the respiratory sound recognition model in the audio processing module may recognize that the to-be-detected respiratory audio data collected does not contain respiratory sounds, in the case where the respiratory sound recognition model recognizes that the to-be-detected respiratory audio data does not contain respiratory sounds, it is considered that no respiratory sounds are detected. On the contrary, in the case where the respiratory sound recognition model recognizes that the to-be-detected respiratory audio data contains respiratory sounds, it is considered that respiratory sounds are detected.
[0343] S603, recognize the owner based on the to-be-detected respiratory sound features and output the owner recognition result.
[0344] In the embodiments of the present application, for the sake of differential description, the respiratory sound features of the to-be-detected user can be denoted as the to-be-detected respiratory sound features.
[0345] Here, in the case where respiratory sounds are detected in S602, it indicates that the respiratory sound recognition model in the audio processing module recognizes that the to-be-detected respiratory audio data contains respiratory sounds. Then, the audio processing module can finally output the to-be-detected respiratory sound features.
[0346] After that, the mobile phone can adopt the to-be-detected breathing sound feature corresponding to the user to be detected to identify whether the user to be detected is the owner of the mobile phone.
[0347] Figure 6B The process schematic diagram of identifying the owner of the mobile phone based on the breathing sound is shown. Combining Figure 6B , the mobile phone can compare the to-be-detected breathing sound feature corresponding to the user to be detected with the breathing sound feature template of the owner stored in advance to determine whether the user to be detected is the owner of the mobile phone.
[0348] In some optional implementation manners of the embodiments of the present application, the mobile phone can determine whether the user to be detected is the owner of the mobile phone in the following manner: when there is a breathing sound feature template in the owner template, determine the similarity between the to-be-detected breathing sound feature and the breathing sound feature template. When the obtained similarity is greater than a preset second determination threshold, determine that the user to be detected is the owner of the mobile phone. On the contrary, when the obtained similarity is less than or equal to the second determination threshold, determine that the user to be detected is not the owner of the mobile phone.
[0349] Wherein, the above-mentioned second determination threshold is usually a preset similarity value. For example, it can be 0.6, and the value range of the second determination threshold is usually 0-1.
[0350] In practice, calculating the similarity between the to-be-detected breathing sound feature and the breathing sound feature template is usually calculating the cosine similarity between the vector corresponding to the to-be-detected breathing sound feature and the vector corresponding to the breathing sound feature template.
[0351] In some optional implementation manners of the embodiments of the present application, the mobile phone can also determine whether the user to be detected is the owner of the mobile phone in the following manner: when there are multiple breathing sound feature templates in the owner template, respectively determine the similarity between each breathing sound feature template and the to-be-detected breathing sound feature. When the average value of the obtained multiple similarities is greater than a preset fourth determination threshold, determine that the user to be detected is the owner of the mobile phone. On the contrary, when the average value of the obtained multiple similarities is less than or equal to the fourth determination threshold, determine that the user to be detected is not the owner of the mobile phone.
[0352] Wherein, the above-mentioned fourth determination threshold is usually a preset similarity value. For example, it can be 0.3, and the value range of the fourth determination threshold is 0-1.
[0353] Further, for example, if there are 3 breath sound feature templates in the owner template, namely breath sound feature template 1, breath sound feature template 2, and breath sound feature template 3, and the similarity between the breath sound feature to be measured and breath sound feature template 1 is S1, the similarity between the breath sound feature to be measured and breath sound feature template 2 is S2, and the similarity between the breath sound feature to be measured and breath sound feature template 3 is S3, then when the average value of S1, S2, and S3 is greater than the fourth determination threshold, the user to be measured is considered the owner; otherwise, the user to be measured is considered a non-owner.
[0354] It should be noted that in the case where there are multiple breath sound feature templates in the owner template, by using the average value of the similarities between the breath sound feature to be measured and each breath sound feature template to determine whether the user to be measured is the owner, the accuracy of owner identification can be improved.
[0355] Optionally, in the case where the owner template includes multiple breath sound feature templates, the mobile phone can also respectively determine the similarity between each breath sound feature template and the breath sound feature to be measured. When the maximum value among the obtained multiple similarities is greater than the second determination threshold, it is determined that the user to be measured is the owner. Conversely, when the maximum value among the obtained multiple similarities is less than or equal to the second determination threshold, it is determined that the user to be measured is a non-owner.
[0356] In some optional implementation manners of the embodiments of the present application, when there are multiple owner templates stored in the mobile phone, the mobile phone can also determine whether the user to be measured is the owner in the following manner: for each owner template, in the case where there is one breath sound feature template in the owner template, determine the similarity between the breath sound feature to be measured and the breath sound feature template. When the obtained similarity is greater than the second determination threshold, determine that the user to be measured is the owner, and specifically, it is the owner indicated by the owner template; in the case where there are multiple breath sound feature templates in the owner template, respectively determine the similarity between each breath sound feature template in the owner template and the breath sound feature to be measured. When the average value of the obtained multiple similarities is greater than the fourth determination threshold, determine that the user to be measured is the owner, and specifically, it is the owner indicated by the owner template;
[0357] Optionally, when the user to be measured is not the owner indicated by any owner template, the mobile phone can determine that the user to be measured is a non-owner.
[0358] In the embodiments of the present application, when the user to be measured is identified as the owner, the mobile phone can continue to perform operations related to the owner's permissions, such as previewing notification messages, making password-free payments, etc. When the user to be measured is identified as a non-owner, the mobile phone hides the private information of the owner, such as hiding the content of communication messages, hiding the content of service cards, etc.
[0359] S604, prompt that the detection result is unknown.
[0360] In the embodiment of the present application, during the process of the mobile phone performing the owner identification, if neither a human face nor a breathing sound is detected, it can be considered that the detection result is unreliable. At this time, the mobile phone can output a prompt message such as "Detection result unknown", or "Retry once", or "Do it again".
[0361] After that, the user to be tested can, in combination with the prompt of the mobile phone, further cooperate with the mobile phone to perform the owner identification.
[0362] S605, the mobile phone inputs the face recognition score into the threshold discriminator to determine whether the user to be tested is the owner.
[0363] Here, if the face recognition score is greater than the pre-set face recognition upper threshold, such as 0.65, it is considered that the user to be tested is the owner; if the face recognition score is less than the pre-set face recognition lower threshold, such as 0.05, it is considered that the user to be tested is not the owner. If the face recognition score is between the above-mentioned face recognition lower threshold and the above-mentioned face recognition upper threshold, step S606 can be continued to be executed.
[0364] Among them, the above-mentioned face recognition upper threshold can be the FAR threshold of face recognition. The above-mentioned face recognition lower threshold can be the False Rejection Rate (FRR) threshold of face recognition. In practice, when the face recognition score is between the above-mentioned face recognition lower threshold and the above-mentioned face recognition upper threshold, it usually indicates that the face detection result is poor.
[0365] S606, in the case of a poor face detection result, the mobile phone determines whether a breathing sound is detected. If no breathing sound is detected, S607 is executed; conversely, if a breathing sound is detected, S608 is executed.
[0366] In the embodiment of the present application, in the case of a poor face detection result, the mobile phone can continue to collect the breathing audio data to be tested of the user to be tested, and call the pre-stored audio processing module (see Figure 4A ), and the audio processing module processes the breathing audio data to be tested to extract the breathing sound characteristics in the breathing audio data to be tested.
[0367] In practice, if the breathing sound recognition model recognizes that the breathing audio data to be tested does not contain a breathing sound, it is considered that no breathing sound is detected. Conversely, if the breathing sound recognition model recognizes that the breathing audio data to be tested is a breathing sound, it is considered that a breathing sound is detected.
[0368] S607, in the case of no breathing sound detected, the mobile phone directly outputs the face detection result.
[0369] Here, when the mobile phone combines the pre-recorded face template of the owner to detect and recognize the face, it can initially determine whether it is the owner by combining the face recognition score. In the case where no breathing sound is detected, the mobile phone can output this initial judgment. For example, if the face recognition score is greater than the pre-set face recognition score threshold, it is determined to be the owner; conversely, if the face recognition score is less than or equal to the face recognition score threshold, it is determined not to be the owner. Among them, the above-mentioned face recognition score threshold can be a pre-set threshold, for example, it can be 0.5.
[0370] Optionally, in some application scenarios where the accuracy of owner recognition is very high, when the face detection result is poor and no breathing sound is detected, the mobile phone can also output information to prompt re-detection, for example, it can output "Please try again".
[0371] S608, in the case where a breathing sound is detected, the mobile phone combines the multi-modal data of the user to be tested and calls the multi-modal fusion owner recognition model to perform owner recognition.
[0372] Among them, the multi-modal data can include face recognition score, the breathing sound feature to be tested, the touch screen data to be tested, and the IMU data to be tested.
[0373] In the embodiments of the present application, only when the face recognition effect is poor, for example, in cases such as strong light, backlight, and low light that cause poor face recognition, will the breathing audio data, touch screen data, and IMU data of the user to be tested be further collected, and then combined with the multi-modal data to jointly perform owner recognition to improve the accuracy of owner recognition. Since the accuracy of face recognition is high most of the time, that is, the situation of using multi-modal data for owner recognition is relatively rare, it is possible to save the computing power of the mobile phone and reduce the power consumption of the mobile phone while ensuring the accuracy of owner recognition, thereby improving the practicality of the mobile phone and helping to further enhance the user experience.
[0374] For the sake of distinction and description, the touch screen data collected during the owner recognition process can be denoted as the touch screen data to be tested, and the IMU data collected can be denoted as the IMU data to be tested.
[0375] The operation of collecting the touch screen data to be tested during the owner recognition process can be basically the same as the operation of collecting the touch screen data during the aforementioned owner template entry process, and will not be elaborated here. In addition, the operation of collecting the IMU data to be tested during the owner recognition process can also be basically the same as the operation of collecting the IMU data during the aforementioned owner template entry process.
[0376] It can be understood that in some alternative implementation manners, in S608, when breath sound is detected, the mobile phone can also directly use the breath sound feature of the user to be measured for the identification of the phone owner. At this time, S608 can also be replaced with the following steps: when breath sound is detected, the mobile phone identifies the phone owner based on the breath sound feature to be measured and outputs the phone owner identification result. This operation is basically the same as the operation of the aforementioned S603 and will not be elaborated here.
[0377] Figure 6C Fig. shows a schematic process diagram of identifying the phone owner by using a multi-modal fusion phone owner identification model.
[0378] S6081, calculate the cross-modal feature to be measured.
[0379] Wherein, the cross-modal feature to be measured is the cross-modal feature corresponding to the user to be measured.
[0380] Here, the mobile phone can use a cross-modal encoder (see Figure 5A ) to encode the touch screen data to be measured, the IMU data to be measured and the breath sound feature to be measured, so as to obtain the cross-modal feature to be measured.
[0381] S6082, calculate the cross-modal similarity between the cross-modal feature to be measured and the cross-modal feature template.
[0382] Wherein, the cross-modal similarity is the similarity between the cross-modal feature to be measured and the cross-modal feature template. In the embodiments of the present application, the cross-modal similarity can also be referred to as the cross-modal recognition score.
[0383] Here, if the cross-modal feature template S m = E(I m , T m , R m ), and the cross-modal feature to be measured S t = E(I t , T t , R t ), wherein, E(,,) represents a cross-modal encoder, I m is the IMU data of the phone owner, T m is the touch screen data of the phone owner, R m is the breath sound feature template of the phone owner, I t is the IMU data to be measured of the user to be measured, T t is the touch screen data to be measured of the user to be measured, and R t is the breath sound feature to be measured of the user to be measured.
[0384] Exemplarily, the cross-modal similarity can be calculated by the following formula (4):
[0385]
[0386] Among them, D c is the cross-modal similarity. C(,) represents the cosine similarity comparator, and || || is the L2 norm, which is used to calculate the magnitude of the vector norm.
[0387] S6083, calculate the breath sound similarity between the breath sound feature to be measured and the breath sound feature template of the machine owner.
[0388] Among them, the breath sound similarity is the similarity between the breath sound feature to be measured and the breath sound feature template of the machine owner. In the embodiments of the present application, the breath sound similarity can also be referred to as the breath recognition score.
[0389] Exemplarily, the breath sound similarity can be calculated by the following formula (5):
[0390]
[0391] Among them, D r is the breath sound similarity.
[0392] S6084, input the cross-modal similarity, the breath sound similarity, and the face recognition score into the fusion decision model to obtain the multi-modal recognition score.
[0393] Among them, the input of the fusion decision model is the cross-modal similarity (also called the cross-modal recognition score), the breath sound similarity (also called the breath recognition score), and the face recognition score, and the output is the multi-modal recognition score. The multi-modal recognition score is used to describe the comprehensive recognition result of the user to be measured in multiple modalities.
[0394] In the embodiments of the present application, the multi-modal recognition score can be expressed as: D = M(D c , D r , D f ), where D is the multi-modal recognition score, D c is the cross-modal similarity, D r is the breath sound similarity, D f is the face recognition score, and M(,,) represents the fusion decision model.
[0395] S6085, combine the multi-modal recognition score to determine whether the user to be measured is the machine owner.
[0396] In the embodiments of the present application, if the multi-modal recognition score is greater than or equal to the preset fusion threshold, it is recognized as the machine owner; if the multi-modal recognition score is less than the fusion threshold, it is recognized as a non-machine owner.
[0397] The above-mentioned fusion threshold is usually a preset threshold. For example, it can be 0.6.
[0398] Combine Figure 6CIt can be found that for the main recognition model of the multi-modal fusion machine, a cross-modal encoder is used to perform feature-level fusion on IMU data, touch screen data, and breath sound features, and then a decision fusion model is used to perform decision-level fusion on the cross-modal recognition score, breath recognition score, and face recognition score to obtain the multi-modal recognition score. Among them, stronger modalities, such as the face modality, do not participate in the feature-level fusion, which can reduce the coupling between data and effectively utilize the complementary advantages between multi-modalities, improving the comprehensive recognition success rate of the model.
[0399] Considering that there may be one or multiple owner templates of the owner entered in the mobile phone. That is to say, the mobile phone can have one owner template or multiple owner templates.
[0400] In some alternative implementation manners of the embodiments of the present application, when multiple owner templates are pre-entered in the mobile phone, in combination with the above Figure 6C , during the execution of S6081 - S6084 by the mobile phone, one owner template can be selected first, and the multi-modal recognition score can be calculated using the selected owner template.
[0401] In some application scenarios, the mobile phone can select the owner template corresponding to the highest face recognition score. Further, for example, if there are three owner templates, namely owner template 1, owner template 2, and owner template 3, when the mobile phone executes S601 and identifies the user to be tested as the owner corresponding to owner template 1, then owner template 1 can be directly selected.
[0402] In this application scenario, during the execution of S6081 - S6084 by the mobile phone, a multi-modal recognition score for owner template 1 can be calculated. Then, the mobile phone can further determine whether it is the owner through this multi-modal recognition score. If it is the owner, it means that the user to be tested is the owner indicated by owner template 1. It is not difficult to find that in this application scenario, only the multi-modal recognition score needs to be calculated for one owner template, and the calculation amount is relatively small.
[0403] In other alternative implementation manners of the embodiments of the present application, when multiple owner templates are pre-entered in the mobile phone, the mobile phone can calculate a multi-modal recognition score for each owner template. Then, the mobile phone can combine the obtained multiple multi-modal recognition scores to determine whether the user to be tested is the owner and which specific owner it is.
[0404] For further example, if there are three owner templates, namely owner template 1, owner template 2, and owner template 3, the mobile phone can perform the operations of S6081 - S6084 using owner template 1 to obtain a multi-modal recognition score, such as 99 points. Perform the operations of S6081 - S6084 using owner template 2 to obtain another multi-modal recognition score, such as 78 points, and perform the operations of S6081 - S6084 using owner template 3 to obtain yet another multi-modal recognition score, such as 60 points. At this time, it can be determined that the user to be measured is the owner, and specifically the owner indicated by owner template 1.
[0405] In this implementation, since the above-mentioned step S608 is only executed to further perform owner recognition using the multi-modal fusion model when the face detection effect is not good, and when the face detection effect in step S601 is not good, selecting the owner template based on the face detection result may not be accurate enough. Therefore, in this implementation, when the face detection accuracy in S601 is not high, calculating the multi-modal recognition score using each owner template can ensure the accuracy rate of owner recognition.
[0406] Considering that there may be multiple feature templates in an owner template. For example, there are multiple face templates; for another example, there are multiple breathing sound feature templates; for yet another example, there are multiple cross-modal feature templates.
[0407] In some alternative implementation manners of the embodiments of the present application, when there are multiple feature templates in the owner template, during the process of the mobile phone executing S6081 - S6084, the mobile phone can calculate the similarity between the mean value of the multiple feature templates in the owner template and the feature to be measured of the same modality (such as the breathing sound feature to be measured) to obtain the recognition score of the corresponding modality.
[0408] For further example, if there are 3 breathing sound feature templates in the owner template, the mobile phone can calculate the similarity between the mean value of the 3 breathing sound feature templates and the breathing sound feature to be measured to obtain the recognition score corresponding to the breathing modality, or it can be called the breathing recognition score.
[0409] Optionally, the mobile phone can also calculate the similarity between each feature template in the owner template and the feature to be measured of the same modality, so as to obtain multiple similarities; then, calculate the recognition score of the corresponding modality through the multiple similarities obtained. For example, the maximum value among the multiple similarities can be used as the recognition score of the corresponding modality.
[0410] For further example, if there are 3 breathing sound feature templates in the owner template, the mobile phone can calculate the similarity between each breathing sound feature template and the breathing sound feature to be measured, obtaining 3 similarities. Then, the maximum value among the 3 similarities can be used as the recognition score corresponding to the breathing mode, or referred to as the breathing recognition score.
[0411] In some alternative implementation manners of the embodiments of the present application, when multiple owner templates are pre-recorded in the mobile phone, combined with the above Figure 6C , during the execution of S6085 by the mobile phone, since each owner template may correspond to a multi-modal recognition score, at this time, the mobile phone can determine whether the user to be measured is the owner in the following manner: for the multi-modal recognition scores respectively corresponding to each owner template, if the multi-modal recognition score is greater than the fusion threshold, the mobile phone can determine that the user to be measured is the owner. At this time, the user to be measured can be recognized as: the owner indicated by the owner template corresponding to the multi-modal recognition score.
[0412] In some application scenarios, if there are multiple situations where the multi-modal recognition scores are simultaneously greater than the fusion threshold, the mobile phone can determine that the user to be measured is the owner. At this time, the user to be measured can be recognized as: the owner indicated by the owner template corresponding to the maximum value among the multiple multi-modal recognition scores.
[0413] The following combines Figure 6D to further illustrate the fusion decision model in S6084. Figure 6D This is a schematic structural diagram of the fusion decision model provided by the embodiments of the present application. Combining Figure 6C , the input of the fusion decision model can be cross-modal recognition scores, breathing recognition scores, and face recognition scores, and the output is multi-modal recognition scores.
[0414] Combining Figure 6D , the fusion decision model includes a Stem layer, 12 encoding layers, a Head layer, and a Sigmoid function, where the encoding layer includes a Linear layer, an activation function, and a Linear layer. Figure 6D In , the fusion decision model uses the Stem layer to expand the feature dimensions of the cross-modal recognition scores, breathing recognition scores, and face recognition scores, and then encodes them through the 12 encoding layers of the feed-forward neural network with a residual structure. Finally, the classification head (Head layer) and the Sigmoid function are used to output the multi-modal recognition scores. The overall decision fusion model adopts a lightweight design, which can save storage resources.
[0415] The following describes the training process of the multi-modal fusion owner recognition model. From Figure 6CIt can be seen that the main recognition model of the multimodal fusion machine mainly includes a cross-modal encoder and a fusion decision-making model. That is to say, training the main recognition model of the multimodal fusion machine is also training the cross-modal encoder and the fusion decision-making model.
[0416] During the training process of the main recognition model of the multimodal fusion machine, difficult samples can be dynamically mined and used as negative samples, so as to further train the main recognition model of the multimodal fusion machine in the way of pairing positive and negative samples.
[0417] Figure 6E The model training architecture of pairing positive and negative samples based on dynamic difficult sample mining is shown. As Figure 6E shown, the process of training the main recognition model of the multimodal fusion machine can include the following steps:
[0418] In the first step, a batch of training samples is obtained from all the sample data sets corresponding to the sample users.
[0419] Combined with Figure 6E , the "m×" in the upper right corner of the dashed box indicates sampling m batches of training samples. Each batch of training samples includes 1 positive sample and n negative samples.
[0420] In the embodiments of the present application, each sample user can correspond to a sample data set. Each sample data in the sample data set can include sample IMU data, sample touch screen data, sample breath sound features, and sample face recognition scores.
[0421] Here, a positive sample i and n negative samples j∈(1,m) can be obtained from all the sample data sets corresponding to the sample user u . Among them, the sample label corresponding to the positive sample is and the sample label corresponding to the negative sample is . Among them, k is the total number of sample users, and m is the total number of training batches.
[0422] Exemplarily, the main recognition model of the multimodal fusion machine can be represented as a non-linear mapping . Where x is the input of the model, and θ represents the parameters of the model, then the output score of the main recognition model of the multimodal fusion machine can be expressed by the following formula (6):
[0423] f θ (x) = σ(φ θ (x)) (6)
[0424] Among them, f θ (x) is the multimodal recognition score, and σ is the Sigmoid function.
[0425] In the second step, input a batch of acquired training samples into the main recognition model of the multimodal fusion machine to obtain the recognition probability corresponding to each sample, and use the negative sample that is misjudged as a positive sample by the model and has the highest corresponding recognition probability as a difficult sample (or called a dynamic difficult sample).
[0426] Exemplarily, the difficult sample can be represented by the following formula (7):
[0427]
[0428] In the third step, use the difficult sample and the positive sample to jointly calculate the loss of this training.
[0429] Exemplarily, when using the difficult sample and the positive sample to jointly calculate the loss of this training, the loss calculation formula (8) can be as follows:
[0430]
[0431] where L total,ij is the loss of this model training, and L(,) is the cross-entropy loss function.
[0432] In the fourth step, use the model training loss to update the model parameters.
[0433] Exemplarily, the model parameters can be updated based on the following formula (9).
[0434]
[0435] where θ * is the updated model parameter.
[0436] Combining the above training process, it can be found that during the training of the main recognition model of the multimodal fusion machine, for the same sample user, m batches of training samples can be collected. That is, m groups of training data of each sample user are used for model training, rather than only using one group of data for training, which can improve the accuracy of the model.
[0437] In addition, in each batch of training samples, mining difficult samples as negative samples, jointly calculating the model training loss with dynamic difficult samples and positive samples, and updating the model parameters based on the model training loss can effectively reduce the misrecognition rate and improve the TAR index at a lower misrecognition rate. That is, it can make the main recognition model of the multimodal fusion machine obtained by training more accurate and the stability of the owner recognition stronger.
[0438] Next, further combine Table 1, Figure 7 and Figure 8 to illustrate the effect of the owner recognition method based on multimodal fusion provided by the embodiments of the present application.
[0439] Combined with Table 1, it can be found that compared with the best-performing face modality in the unimodal case, when the FAR (i.e., the false recognition rate) is 1%, the TAR (recognition success rate) of the multi-modal fusion-based owner recognition method has increased by 4.17%. When the FAR is 1.5%, the TAR of the multi-modal fusion-based owner recognition method has increased by 4.54%. When the FAR is 5%, the TAR of the multi-modal fusion-based owner recognition method has increased by 5.83%. That is to say, the multi-modal fusion-based owner recognition method has improved the TAR index at a lower false recognition rate.
[0440] Table 1 Performance Evaluation Results of the Owner Recognition Method Based on Multi-Modal Fusion
[0441]
[0442] Among them, the values in Table 1 are percentage values. For example, the AUC under the respiration unimodal is 89.36, which means 89.36%, or 0.8936.
[0443] Figure 7 This is the ROC curve graph for owner recognition corresponding to various modalities provided by the embodiments of the present application. Figure 8 This is the PR curve graph for owner recognition corresponding to various modalities provided by the embodiments of the present application.
[0444] Combined with Table 1, Figure 7 and Figure 8 it can be found that compared with the best-performing face modality in the unimodal case, the AUC of the multi-modal fusion-based owner recognition method has increased by 3.42%, and the AP has increased by 2.20%.
[0445] Based on the comprehensive Table 1, Figure 7 and Figure 8 it can be found that the multi-modal fusion-based owner recognition method can effectively improve the comprehensive recognition success rate by utilizing the complementary advantages between multi-modalities.
[0446] (3) Update of the owner template.
[0447] In practical applications, since the characteristics of the owner may change over time, such as changes in facial features and breathing sounds, etc., it is necessary to update the owner template.
[0448] In the embodiments of the present application, the timing for updating the owner template recorded in the mobile phone can be: update when a face is entered, and update regularly when the face is unlocked.
[0449] The update period for updating the owner template can be one day or one week. The embodiments of the present application do not make specific limitations on the update period.
[0450] It should be noted that the presence of face entry indicates the entry of a new owner template, and at this time, the owner template in the mobile phone must be updated. When face unlocking occurs and the update cycle is reached, the entered owner template can be updated.
[0451] Since the process of entering the owner template has been described in detail in the owner template entry in the foregoing part (I). Therefore, the following mainly describes how to update the entered owner template, and mainly describes how to update the cross-modal feature template and the breath sound feature template.
[0452] In practice, for any owner, the face template in the owner template can be updated synchronously or asynchronously with the cross-modal feature template and the breath sound feature template. For example, when the update condition (or called the template update condition) is met, the face template can be updated first, and then the cross-modal feature template and the breath sound feature template can be updated. Another example is that when the face template is updated, the update of the cross-modal feature template and the breath sound feature template may be postponed because no breath sound is detected in the collected breath audio data.
[0453] In the embodiment of the present application, during the process of updating the existing owner template, if the update cycle is reached and face unlocking occurs, the update of the owner template of the corresponding owner can be triggered. For example, if the owner templates of three owners A, B, and C have been entered in the mobile phone, after the update cycle is reached and face unlocking occurs, for example, when the owner A is recognized during face unlocking, the update of the owner template of the owner A can be triggered. At this time, the mobile phone can collect the IMU data, touch screen data, and breath audio data of the corresponding owner when face unlocking is successful. Then, the pre-stored audio processing module can be called (see Figure 4A ), and the breath sound feature in the breath audio data can be extracted by the audio processing module processing the breath audio data. The breath sound feature is the current breath sound feature of the owner. The current breath sound feature is the breath sound feature in the breath audio data collected when face unlocking is successful. Then, the mobile phone can call the pre-stored cross-modal encoder (see Figure 5A ), and the cross-modal encoder fuses the breath sound feature, touch screen data, and IMU data to obtain the current cross-modal feature of the owner. The current cross-modal feature is the cross-modal feature corresponding to the multi-modal data collected when face unlocking is successful.
[0454] In some alternative implementation manners of the embodiments of the present application, in order to improve the accuracy of updating the owner template, the update of the breath sound feature template may be performed when the breath sound confidence corresponding to the current breath sound feature is greater than the third determination threshold c, where c > a, and a is the lower threshold for determining whether the breath sound audio data includes breath sounds. In some scenarios of the present application, a is also referred to as the first determination threshold.
[0455] That is to say, in some application scenarios, the update condition (or referred to as the template update condition) for updating the already entered owner template includes, in addition to face unlocking and reaching the update period, that the breath sound confidence corresponding to the current breath sound feature is greater than c.
[0456] It can be understood that the mobile phone usually updates the owner template of the owner recognized when face unlocking is successful. For example, when face unlocking is successful and the recognized owner is A, then, when the template update condition is met, the mobile phone can update the owner template of owner A.
[0457] (1) Update of the owner template in the single feature template scenario
[0458] A single feature template usually means that the owner template includes one feature template. For example, the owner template includes one breath sound feature template, or in other words, one owner corresponds to one breath sound feature template.
[0459] In some alternative implementation manners of the embodiments of the present application, when the above-mentioned update conditions (or referred to as template update conditions) for updating the owner template are met, the mobile phone can directly replace the original feature template with the current feature. The current feature may include the current breath sound feature and the current cross-modal feature.
[0460] For example, the original breath sound feature template can be replaced with the current breath sound feature, and the original cross-modal feature template can be replaced with the current cross-modal feature. That is to say, the current breath sound feature is used as the breath sound feature template, and the current cross-modal feature is used as the cross-modal feature template.
[0461] In some alternative implementation manners, when the template update condition is met, the mobile phone can also determine the mean value of the feature template in the owner template and the current feature of the same modality as the new feature template of the corresponding modality, and replace the old feature template of the corresponding modality with the new feature template.
[0462] For example, the mobile phone can determine the mean of the breathing sound feature template in the owner template and the current breathing sound feature as the new breathing sound feature template, and replace the old breathing sound feature template with the new one. At the same time, the mobile phone can determine the mean of the cross-modal feature template in the owner template and the current cross-modal feature as the new cross-modal feature template, and replace the old cross-modal feature template with the new one.
[0463] It should be noted that taking the mean of the current feature and the old feature template as the new feature template can accurately and stably update the owner template, which helps to improve the accuracy and stability of owner recognition.
[0464] (2) Update of the owner template in the scenario of multiple feature templates
[0465] Multiple feature templates usually mean that the owner template includes multiple feature templates. For example, the owner template includes multiple breathing sound feature templates, or in other words, one owner corresponds to multiple breathing sound feature templates.
[0466] In some optional implementation manners of the embodiments of the present application, when the template update condition is met, the mobile phone can determine the mean of each feature template in the owner template and the current feature of the same modality as the new feature template of the corresponding modality, and add the new feature template to the owner template.
[0467] For example, the mobile phone can determine the mean of the breathing sound feature template in the owner template and the current breathing sound feature as the new breathing sound feature template, and add the new breathing sound feature template to the owner template. At the same time, the mobile phone can determine the mean of the cross-modal feature template in the owner template and the current cross-modal feature as the new cross-modal feature template, and add the new cross-modal feature template to the owner template.
[0468] Further, for example, if the owner template records p - 1 cross-modal feature templates, which are S m,1 , S m,2 ,..., S m,p-1 , and the current cross-modal feature is S m,p , then the new cross-modal feature template S m can be calculated by the following formula (10).
[0469]
[0470] Similarly, if the owner template records q - 1 breathing sound feature templates, which are R m,1 , R m,2 ,..., R m,q-1 , and the current breathing sound feature is R m,q , then the new breathing sound feature template Rm It can be calculated through the following formula (11).
[0471]
[0472] It should be noted that using the mean value of the current feature and the old feature template as the new feature template can accurately and stably update the owner template, which helps to improve the accuracy and stability of owner recognition.
[0473] In some optional implementation manners of the embodiments of the present application, the mobile phone can also first perform feature quality judgment on the current feature of the owner, and determine whether to update the feature template according to the feature quality judgment result. In the case where the feature quality judgment result is not good, the current feature of the owner can be deleted. And in the case where the feature quality judgment result is good, the current feature can be used to update the owner template.
[0474] As an example, if the feature quality judgment result of the current cross-modal feature of the owner is not good, the mobile phone can directly delete the current cross-modal feature.
[0475] As another example, if the feature quality judgment is performed on the current breathing sound feature of the owner and it is determined that the feature quality is very good, at this time, the current breathing sound feature can be used to update the owner template. For example, the current breathing sound feature can replace the original breathing sound feature template, or the current breathing sound feature can be directly added to the owner template.
[0476] In some optional implementation manners of the embodiments of the present application, when the template update condition is satisfied, the mobile phone can respectively determine the similarity between each feature template in the owner template and the current feature of the same modality, and then select the similarity with the largest value from multiple similarities. For the convenience of description, it can be called the maximum similarity.
[0477] Update method 1: If the maximum similarity is greater than the fifth determination threshold e, the current feature replaces the feature template corresponding to the maximum similarity. Among them, the fifth determination threshold is a preset threshold, and the value range of the fifth determination threshold is 0-1. In the scenario of updating using update method 1, the current feature replaces the feature template corresponding to the maximum similarity.
[0478] Update method 2: If the maximum similarity is less than or equal to e, and the number of feature templates of the corresponding modality stored in the current mobile phone is less than the preset template quantity threshold, the feature template corresponding to the maximum similarity is retained, and the current feature is added to the owner template. In the scenario of updating using update method 2, the number of feature templates in the owner template increases by one.
[0479] Update method 3: If the maximum similarity is less than or equal to e, and the number of feature templates stored in the current mobile phone is equal to the template quantity threshold, then delete the feature template with the longest storage time in the owner template, that is, the one farthest from the current time, and add the current feature to the owner template. In the scenario of updating using Update method 3, the current feature replaces the feature template with the longest storage time.
[0480] The following further gives examples in combination with the breathing modality.
[0481] If the similarities between the current breathing sound feature R0 and the existing breathing sound feature templates R1, R2, and R3 in the owner template are S1, S2, and S3 respectively, where the maximum similarity is S3, and S3 corresponds to the breathing sound feature template R3.
[0482] If S3 is greater than the threshold e, the mobile phone can replace the breathing sound feature template R3 with the current breathing sound feature R0.
[0483] If S3 is less than or equal to the threshold e, and the number of breathing sound feature templates in the current mobile phone is less than the template quantity threshold, the mobile phone can add the current breathing sound feature R0 to the owner template. At this time, there are 4 breathing sound feature templates in the owner template.
[0484] If S3 is less than or equal to the threshold e, and the number of breathing sound feature templates in the current mobile phone is equal to the template quantity threshold, the mobile phone can first delete the breathing sound feature template with the longest storage time. For example, it can delete the breathing sound feature template R1, and then add the current breathing sound feature R0 to the owner template.
[0485] In the embodiments of the present application, by combining the number of feature templates stored in the mobile phone and the storage time of the feature templates, the owner template is updated, which can realize flexible and effective updating of each owner template.
[0486] Corresponding to the owner recognition method in the above embodiments, Figure 9 FIG. shows a schematic structural diagram of an owner recognition device in a host provided by an embodiment of the present application. For ease of illustration, only the parts related to the embodiments of the present application are shown.
[0487] Refer to Figure 9 , the owner recognition device includes:
[0488] A data acquisition unit 901, configured to collect user data of a user to be measured in response to an unlocking operation of the user to be measured on a target content, where the target content is in a locked state, and the user data includes at least one of the following: a face image to be measured, breathing audio data to be measured, IMU data to be measured, and touch screen data to be measured;
[0489] An identification execution unit 902, configured to perform owner identification according to user data and a pre-recorded owner template, where the owner template includes at least one of the following: a face template, a breathing sound feature template, and a cross-modal feature template, and the cross-modal feature template is used to describe the comprehensive features of the owner in multiple modalities;
[0490] An unlocking execution unit 903, configured to unlock the target content when the owner is identified.
[0491] As an embodiment of the present application, the owner identification device can implement the embodiments shown in Figures 1A - 8 the embodiments and other related method embodiments.
[0492] In the owner identification device provided by the embodiments of the present application, the process of each module implementing its respective function can be specifically referred to the description of the embodiments shown in the foregoing Figures 1A - 8 the embodiments and other related method embodiments, and will not be elaborated here.
[0493] It should be noted that the content such as owner identification and execution process between the above-mentioned devices / units, due to being based on the same concept as the method embodiments of the present application, for their specific functions and the technical effects brought, reference can be specifically made to the method embodiment part, and will not be elaborated here.
[0494] It should be understood that the sequence numbers of the steps in the above embodiments do not mean the order of execution is prior or subsequent. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0495] It should be understood that when used in the specification and claims of the present application, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or their combinations.
[0496] It should also be understood that the term "and / or" used in the specification and claims of the present application refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0497] As used in the specification and claims of the present application, the term "if" can be interpreted as "when", "once", "in response to determining", or "in response to detecting" according to the context. Similarly, the phrase "if determined" or "if detected [the described condition or event]" can be interpreted as meaning "once determined", "in response to determining", "once detected [the described condition or event]", or "in response to detecting [the described condition or event]" according to the context.
[0498] In addition, in the description of the specification and the appended claims of this application, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and should not be construed as indicating or implying relative importance. It should also be understood that although the terms "first", "second", etc. are used in the text in some embodiments of this application to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, the first status code can be named the second status code, and similarly, the second status code can be named the first status code without departing from the scope of the various described embodiments. The first status code and the second status code are both status codes, but they are not the same status code.
[0499] The reference to "one embodiment" or "some embodiments" etc. described in the specification of this application means that a specific feature, structure, or characteristic described in connection with that embodiment is included in one or more embodiments of this application. Thus, the statements "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0500] Figure 10 It is a schematic structural diagram of a terminal device provided by an embodiment of this application.
[0501] The terminal device may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a sensor module 180, a key 190, a camera 193, a touch screen 194, and a SIM card interface 195, etc.
[0502] Among them, the audio module 170 may include a speaker, a microphone, etc.
[0503] Among them, the sensor module 180 may include a pressure sensor, an IMU, a distance sensor, a proximity light sensor, a fingerprint sensor, a touch sensor, an ambient light sensor (of course, the terminal device may also include other sensors, such as a temperature sensor, a bone conduction sensor, etc., which are not shown in the figure).
[0504] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors. Among them, the controller may be the nerve center and command center of the terminal device. The controller may generate operation control signals according to the instruction operation code and timing signal to complete the control of fetching and executing instructions.
[0505] A memory may also be provided in the processor 110 for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory may save the instructions or data that the processor 110 has just used or recycled. If the processor 110 needs to use the instruction or data again, it can be directly called from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.
[0506] The processor 110 may run the device owner identification method provided in the embodiments of the present application to facilitate the quick and accurate creation of schedules and enhance the user experience. The processor 110 may include different devices. For example, when the CPU and GPU are integrated, the CPU and GPU may cooperate to execute the device owner identification method provided in the embodiments of the present application. For example, some algorithms in the device owner identification method are executed by the CPU and another part of the algorithms are executed by the GPU to obtain a faster processing efficiency.
[0507] It should be understood that the illustrated terminal device is only an example, and the terminal device may have more or fewer components than those shown in the figure, may combine two or more components, or may have different component configurations. The various components shown in the figure may be implemented in hardware, software, or a combination of hardware and software including one or more signal processing and / or application specific integrated circuits.
[0508] The software system of the terminal device may adopt a layered architecture, an event-driven architecture, a microkernel architecture, a microservices architecture, or a cloud architecture. In the embodiments of the present application, the Android system with a layered architecture is taken as an example to exemplarily illustrate the software structure of the terminal device.
[0509] Figure 11It is a software structure block diagram of the terminal device according to an embodiment of the present application.
[0510] The layered architecture divides the software into several layers, and each layer has a clear role and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers, from top to bottom, namely the application layer, the application framework layer, the Android runtime and system libraries, and the kernel layer.
[0511] The application layer may include a series of application packages.
[0512] Such as Figure 11 As shown, the application packages may include multiple applications such as a gallery, a calendar, a map, WeChat, QQ, text messages, etc. It may also include multiple models such as a breath sound recognition model, a breath sound feature extraction model, a multi-modal fusion machine main recognition model, a cross-modal encoder, and a fusion decision model.
[0513] The application framework layer provides application programming interfaces (APIs) and programming frameworks for the applications in the application layer. The application framework layer includes some predefined functions.
[0514] In some implementation manners, multiple models such as a breath sound recognition model, a breath sound feature extraction model, a multi-modal fusion machine main recognition model, a cross-modal encoder, and a fusion decision model may also be deployed in the application framework layer.
[0515] Such as Figure 11 As shown, the application framework layer may include a window manager, a content provider, a view system, a phone manager, a resource manager, a notification manager, etc.
[0516] The window manager is used to manage window programs. The window manager can obtain the display screen size, determine whether there is a status bar, lock the screen, capture the screen, etc.
[0517] The content provider is used to store and obtain data, and make this data accessible to applications. The data may include videos, images, audio, dialed and received calls, browsing history and bookmarks, phone books, etc.
[0518] The view system includes visible controls, such as controls for displaying text, controls for displaying pictures, etc. The view system can be used to build applications. The display interface can be composed of one or more views. For example, a display interface including a text message notification icon may include a view for displaying text and a view for displaying pictures.
[0519] The phone manager is used to provide the communication function of the terminal device. For example, the management of call status (including answering, hanging up, etc.).
[0520] The resource manager provides various resources for applications, such as localized strings, icons, pictures, layout files, video files, etc.
[0521] The notification manager enables applications to display notification information in the status bar. It can be used to convey informative messages, which can automatically disappear after a short stay without user interaction. For example, the notification manager is used to inform that the download is completed, message reminders, etc. The notification manager can also be a notification that appears in the system top status bar in the form of a chart or scroll bar text, such as the notification of a background running application, or a notification that appears in the form of a dialogue window on the screen. For example, it prompts text information in the status bar, emits a prompt tone, the terminal device vibrates, the indicator light flashes, etc.
[0522] Android Runtime includes a core library and a virtual machine. Android Runtime is responsible for the scheduling and management of the Android system.
[0523] The core library consists of two parts: one part is the functional functions that need to be called by the Java language, and the other part is the core library of Android.
[0524] The application layer and the application framework layer run in the virtual machine. The virtual machine executes the Java files of the application layer and the application framework layer as binary files. The virtual machine is used to perform functions such as the management of object life cycles, stack management, thread management, security and exception management, and garbage collection.
[0525] The system library can include multiple functional modules. For example: surface manager, Media Libraries, 3D graphics processing library (such as: OpenGL ES), 2D graphics engine (such as: SGL), etc.
[0526] The surface manager is used to manage the display subsystem and provides the fusion of 2D and 3D layers for multiple applications.
[0527] The media library supports the playback and recording of various common audio and video formats, as well as static image files, etc. The media library can support multiple audio and video coding formats, such as: MPEG4, H.164, MP3, AAC, AMR, JPG, PNG, etc.
[0528] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, synthesis, and layer processing, etc.
[0529] The 2D graphics engine is a drawing engine for 2D drawing.
[0530] The kernel layer is the layer between hardware and software. The kernel layer includes at least a display driver, a camera driver, an audio driver, and a sensor driver.
[0531] In addition, those skilled in the art can clearly understand that, for the convenience and simplicity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules as needed, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. In each embodiment of the present application, each functional unit can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0532] The embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments can be implemented.
[0533] The embodiment of the present application provides a computer program product. When the computer program product runs on a terminal device, the terminal device can execute the steps in the above method embodiments.
[0534] The embodiment of the present application also provides a chip system. The chip system includes a processor, the processor is coupled to a memory, and the processor executes a computer program stored in the memory to implement the steps in the above method embodiments.
[0535] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above method embodiments of the present application, it can also be completed by a computer program instructing relevant hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps in the above method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable storage medium can include: any entity or device that can carry computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc.
[0536] In the above embodiments, the descriptions of the various embodiments have their own focuses. For parts not described in detail or recorded in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0537] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0538] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0539] The above embodiments are only used to illustrate the technical solutions of this application, rather than to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included in the protection scope of this application.
Claims
1. A method for identifying a machine owner, characterized in that, Applied to a terminal device, the method includes: In response to an unlocking operation of a target content by a user to be measured, collect user data of the user to be measured. Wherein, the target content is in a locked state, and the user data includes at least one of the following: a face image to be measured, respiratory audio data to be measured, IMU data to be measured, and touch screen data to be measured; Perform owner identification based on the user data and a pre-recorded owner template. The owner template includes at least one of the following: a face template, a respiratory sound feature template, a cross-modal feature template, and the cross-modal feature template is used to describe the comprehensive features of the owner in multiple modalities; Unlock the target content when the owner is identified.
2. The method for identifying the machine owner according to claim 1, wherein The target content includes any one of the following: a locked notification message, a locked service card, a locked application, a locked function, and a locked interface.
3. The method for identifying the machine owner according to claim 1, characterized in that, The process of entering the face template includes: In response to a face entry operation of a user, display a face entry interface, and the face entry interface includes a face collection frame; In response to collecting a face image, enter the collected face image as the face template.
4. The method for identifying the machine owner according to claim 1, wherein The process of entering the cross-modal feature template includes: When the multi-modal feature entry condition is met, collect the respiratory audio data of the owner, the IMU data of the terminal device, and the touch screen data of the terminal device. Wherein, the multi-modal feature entry condition includes any one of the following: the face template entry of the owner is successful, the face unlock of the owner is successful and the cross-modal feature template of the owner currently performing the unlock has not been entered in the terminal device; Extract the respiratory sound feature from the respiratory audio data; Encode the IMU data, the touch screen data, and the respiratory sound feature through a cross-modal encoder to obtain cross-modal features, and enter the cross-modal features as the cross-modal feature template. Wherein, the cross-modal encoder is used to encode data of multiple modalities.
5. The method for identifying the machine owner according to claim 4, wherein, The cross-modal encoder includes a cross-channel attention encoder; The cross-channel attention encoder includes a first slice encoding module, a second slice encoding module, and a third slice encoding module in parallel; The first slice encoding module includes a first Group layer and a first Linear layer connected in sequence; the second slice encoding module includes a second Group layer and a second Linear layer connected in sequence; the third slice encoding module includes a third Group layer and a third Linear layer connected in sequence; Wherein, the input end of the first Linear layer is connected to the output end of the first Group layer, the input end of the second Linear layer is connected to the output end of the first Linear layer and the output end of the second Group layer, and the input end of the third Linear layer is connected to the output end of the second Linear layer and the output end of the third Group layer.
6. The method for identifying the machine owner according to claim 4, wherein The cross-modal encoder includes a cross-channel linear encoder, wherein the cross-channel linear encoder includes a Linear layer, an activation function, a Linear layer, a norm layer, a conv layer, and an activation function connected in sequence.
7. The method for identifying the machine owner according to claim 1, characterized in that, Performing owner identification based on the user data and a pre - entered owner template includes: In the case where a face exists in the face image to be measured, performing owner identification based on the face image to be measured and the face template to obtain a face recognition score; In the case where the face recognition score is greater than a pre - set lower face recognition threshold and less than a pre - set upper face recognition threshold, extracting the measured breathing sound features of the measured user from the measured breathing audio data; Inputting the measured breathing sound features, the measured IMU data, the measured touch - screen data, and the face recognition score into a pre - trained multi - modal fusion owner identification model to obtain an owner indication information, where the owner indication information is used to indicate whether the measured user is the owner.
8. The method for identifying the machine owner according to claim 7, characterized in that, Performing owner identification based on the user data and a pre - entered owner template further includes: In the case where the face recognition score is greater than the upper face recognition threshold, determining that the measured user is the owner; In the case where the face recognition score is less than the lower face recognition threshold, determining that the measured user is not the owner.
9. The method for identifying the machine owner according to claim 7, wherein The data - processing process of the multi - modal fusion owner identification model includes: Encoding the measured IMU data, the measured touch - screen data, and the measured breathing sound features through a cross - modal encoder to obtain measured cross - modal features; Determining the cross - modal similarity between the measured cross - modal features and the cross - modal feature template in the owner template, and determining the breathing sound similarity between the measured breathing sound features and the breathing sound feature template in the owner template; Determining a multi - modal recognition score according to the cross - modal similarity, the breathing sound similarity, and the face recognition score, where the multi - modal recognition score is used to describe the comprehensive recognition result of the measured user in multiple modalities; Determining whether the measured user is the owner according to the multi - modal recognition score and a pre - set fusion threshold.
10. The method for identifying the machine owner according to claim 9, characterized in that, In the case where multiple owner templates are entered in the terminal device, there are multiple multi - modal recognition scores; Determining whether the measured user is the owner according to the multi - modal recognition score and a pre - set fusion threshold includes: Selecting the maximum value among multiple multi - modal recognition scores. If the selected multi - modal recognition score is greater than the fusion threshold, determining that the measured user is the owner.
11. The method for identifying the machine owner according to claim 7, wherein The training process of the multi - modal fusion owner identification model includes: Selecting m batches of training samples from the sample dataset used to train the multi - modal fusion owner identification model. Each batch of training samples includes 1 positive sample and n negative samples, where n > 1 and m > 1; During the process of training the multi - modal fusion owner identification model with each batch of training samples, determining the hard sample as the negative sample that is recognized as a positive sample and has the highest corresponding recognition probability; Determining the model training loss of the multi - modal fusion owner identification model according to the positive sample and the hard sample, and adjusting the model parameters of the multi - modal fusion owner identification model according to the determined model training loss.
12. The method for identifying a machine owner according to any one of claims 1-11, characterized in that, The method further includes: When the template update condition is satisfied, the breathing sound feature template in the owner template and the mean value of the current breathing sound feature are determined as the new breathing sound feature template, and the new breathing sound feature template is added to the owner template, where the template update condition includes: face unlocking is successful and the update period is reached; The cross-modal feature template in the owner template and the mean value of the current cross-modal feature are determined as the new cross-modal feature template, and the new cross-modal feature template is added to the owner template; Wherein, the current breathing sound feature is the breathing sound feature in the breathing audio data collected when face unlocking is successful.
13. The method for identifying the machine owner according to any one of claims 1-12, characterized in that, The method further includes: When the template update condition is satisfied, the cross-modal feature template in the owner template and the mean value of the current cross-modal feature are determined as the new cross-modal feature template, and the new cross-modal feature template is added to the owner template; Wherein, the current cross-modal feature is the cross-modal feature corresponding to the multi-modal data collected when face unlocking is successful.
14. A terminal device, characterized in that, The terminal device includes a memory and a processor, and a computer program capable of running on the processor is stored on the memory. When the processor executes the computer program, the owner recognition method according to any one of claims 1 to 13 is implemented.
Citation Information
Patent Citations
User management method and system for mobile terminal
CN106127006A
Identity authentication method and device
CN111199032A
Cross-modal face recognition method and device, equipment and storage medium
CN112016523A
Living body identification method and system
CN116503962A
User identity based on human breath analytics
US20220092162A1