Owner identification method and terminal equipment
By combining multimodal recognition methods with face, IMU, touch screen and respiratory audio data, the problem of unstable recognition of the owner in the terminal device is solved, the recognition accuracy and stability are improved, and the security of the owner's private information is protected.
Patent Information
- Application Number
- CN202311871268.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-29
- Publication Date
- 2025-07-08
AI Technical Summary
During the use of terminal devices, the private information of the owner is easily seen by others, resulting in privacy leakage. The existing facial and fingerprint recognition solutions have unstable recognition rates in different scenarios.
Combining the user's face data, terminal device IMU, touch screen data and user breathing audio data, the owner is comprehensively identified and multi-modal data is used for identity verification.
It improves the accuracy and stability of the owner's identification, ensuring that the owner can still be effectively identified when the scene changes, and protects the security of private information.
Smart Images

Figure CN120277649A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular, to a method for identifying the device owner and a terminal device. Background Art
[0002] Terminal devices, especially mobile phones, have gradually become the center of the personal digital world, carrying a large amount of personal information and privacy data. During the use of terminal devices, the private information of the device owner may be displayed. If the private information of the device owner is seen by someone other than the device owner, it will lead to the leakage of the device owner's privacy.
[0003] Therefore, during the use of terminal devices, it is necessary to identify the device owner to protect the security of the device owner's private information. Summary of the Invention
[0004] In view of this, the embodiments of this application provide a method for identifying the device owner and a terminal device, which can identify the device owner during the use of the terminal device to protect the security of the device owner's private information.
[0005] In a first aspect, the embodiments of this application provide a method for identifying the device owner, which is applied to a terminal device. The method for identifying the device owner may include: the terminal device collects user data of a user to be tested in response to an unlocking operation of the user to be tested on the target content, where the target content is in a locked state, and the user data includes at least one of the following: a face image of the user to be tested and respiratory audio data of the user to be tested; then, the terminal device can perform device owner identification based on the user data and a pre-recorded device owner template, where the device owner template may include at least one of the following: a face template and a respiratory sound feature template; finally, the terminal device can unlock the target content when the device owner is recognized.
[0006] Herein, the user to be tested generally refers to the user to be identified. The face image of the user to be tested is the face image of the user to be tested. The respiratory audio data of the user to be tested is the respiratory audio data of the user to be tested. The device owner template is a pre-recorded device owner template. The face template is a pre-recorded face image of the device owner. The respiratory sound feature template is a pre-recorded respiratory sound feature of the device owner.
[0007] In an embodiment of the present application, when a user to be detected wants to unlock target content in a locked state, the terminal device can collect user data of the user to be detected, thereby performing owner identification on the user to be detected through the collected user data, and only when the owner is identified, the terminal device unlocks the target content, which can ensure the security of the owner's private information. In addition, the terminal device can use at least one of the face data and breathing audio data of the user to be detected to perform owner identification on the user to be detected. For example, face data can be preferentially used for owner identification. When the face recognition result is not ideal, breathing audio data is continuously used for owner identification. In this way, owner identification can still be performed when the face recognition result is not ideal, which can improve the accuracy and stability of owner identification. In addition, based on both face data and breathing audio data at the same time, these two modalities of data are used for owner identification, which can achieve the complementary advantages of different modalities of data, thereby improving the stability and accuracy of owner identification.
[0008] In the first possible implementation manner of the first aspect, the target content may include any one of the following: a locked notification message, a locked service card, a locked application, a locked function, and a locked interface.
[0009] In an embodiment of the present application, there may be one or more target contents, that is to say, there are many application scenarios that can trigger the terminal device to perform owner identification. That is to say, the terminal device can lock one or more contents to ensure the security of the owner's private information.
[0010] As an optional embodiment of the present application, the terminal device can lock the content set by the lock setting operation in response to the lock setting operation of the owner.
[0011] Among them, the lock setting operation is used to lock the content set by the owner. For example, one or more applications set by the owner can be locked. For another example, the service card function set by the owner can be locked. For another example, the passwordless payment function set by the owner can be locked, and so on.
[0012] As an optional embodiment of the present application, the process of the terminal device entering the face template of the owner may include: First, the terminal device displays a face entry interface in response to the user's face entry operation, where the face entry interface includes a face capture frame; then, the terminal device can enter the captured face image as a face template in response to the captured face image.
[0013] In an embodiment of the present application, the owner can enter the face template through human-computer interaction, which is intuitive and convenient to operate, and helps to improve the user experience.
[0014] As an optional embodiment of the present application, the terminal device can input the template of the device owner of one device owner, or can input the templates of multiple device owners. Usually, one device owner corresponds to one device owner template.
[0015] In the second possible implementation manner of the first aspect, the process of the terminal device inputting the breathing sound feature template of the device owner may include: First, when the breathing feature input condition is met, the terminal device can collect the breathing audio data of the device owner. Then, the terminal device can extract the breathing sound features from the breathing audio data and input the extracted breathing sound features as the breathing sound feature template.
[0016] The breathing feature input condition includes that the face template of the device owner is successfully input, or the device owner's face unlocking is successful and the breathing sound feature template of the device owner currently performing the unlocking has not been input into the terminal device.
[0017] In the embodiment of the present application, when the face template of the device owner is successfully input, the terminal device can continue to collect the breathing audio data of the device owner, so as to extract the breathing sound features from the collected breathing audio data as the breathing sound feature template. Since the breathing audio data collected when the device owner is successfully input may be interfered by various factors, for example, the environment is noisy, resulting in no breathing sound features being extracted. Therefore, when the breathing sound feature template of the device owner has not been successfully input, when the device owner's face unlocking is successful, the breathing audio data of the device owner can be continuously collected, and the breathing sound features therein can be extracted and input as the breathing sound feature template of the device owner. In this way, the successful input of the breathing sound feature template of the device owner can be ensured.
[0018] In the third possible implementation manner of the first aspect, the terminal device can extract the breathing sound features from the breathing audio data by performing the following operations: First, extract the mel cepstrum and mel cepstral coefficients from the breathing audio data; then, divide the mel cepstrum into multiple continuous mel cepstrum segments; after that, input the multiple mel cepstrum segments into a pre-trained breathing sound recognition model to obtain the breathing sound confidence, where the breathing sound confidence is used to indicate the probability that the breathing audio data is classified as breathing sound; finally, when the breathing sound confidence is greater than a pre-set first determination threshold, input the mel cepstral coefficients into a pre-trained breathing sound feature extraction model to obtain the breathing sound features.
[0019] In the embodiments of the present application, since the complete mel cepstrum usually has a relatively large amount of data, splitting the mel cepstrum into multiple mel cepstrum segments can simplify the feature representation, reduce the computational cost, and can more effectively capture the local features of the audio signal, thereby improving the accuracy of breath sound recognition. Additionally, since the mel cepstrum usually carries rich audio information, identifying whether an audio is a breath sound through the mel cepstrum segments of the audio can ensure the accuracy rate of breath sound recognition, that is, it can enable the breath sound recognition model to more accurately determine whether an audio is a breath sound. Moreover, since the mel cepstrum coefficients can effectively capture the important features of the audio signal and the data dimension is small, in the case where the respiratory audio data has been determined to be a breath sound, inputting the mel cepstrum coefficients of the respiratory audio data into the breath sound feature extraction model to extract the breath sound features of the respiratory audio data can achieve fast and accurate extraction of the breath sound features of the respiratory audio data.
[0020] As an optional embodiment of the present application, the length of the mel cepstrum segment is usually about 1 second.
[0021] In the embodiments of the present application, if the segmentation granularity is too fine, on the one hand, the computational amount will increase, and on the other hand, the breath sound recognition accuracy will decrease. Therefore, in practice, the length of the mel cepstrum segment is usually set to about 1 second. For example, it can be 0.96 seconds. The reason is that the duration of a person's exhalation or inhalation is approximately 1 second. Switching the mel cepstrum segment to a segment with a length of about 1 second can enable each mel cepstrum segment to more completely describe the breath characteristics, thereby further improving the accuracy of breath sound recognition.
[0022] In the fourth possible implementation manner of the first aspect, the data processing process of the breath sound recognition model may include the following steps: First, through a pre-trained feature extractor, extract the single-segment sound representations corresponding to each mel cepstrum segment respectively; then, in the time order of each mel cepstrum segment, splice the single-segment sound representations corresponding to each mel cepstrum segment respectively to obtain a continuous-segment sound representation; finally, according to the continuous-segment sound representation, determine the confidence levels corresponding to various sound types, where the sound types include: quiet, speaking, breath sound, cough, sneeze, snore, and environmental noise, and the confidence level corresponding to the sound type is used to indicate the probability that the respiratory audio data is determined to be the corresponding sound type.
[0023] In the embodiment of the present application, the output of the breath sound recognition model is the confidence levels of 7 types of sounds. That is to say, during the training process of the breath sound recognition model, 7 types of sounds are used to train the model. In this way, it helps to increase the generalization of the breath sound recognition model, that is, it can enable the breath sound recognition model to accurately distinguish breath sounds. In addition, the breath sound recognition model extracts refined single-segment sound representations corresponding to each Mel cepstrum segment through a feature extractor, and then splices the single-segment sound representations to obtain the overall sound representation corresponding to multiple Mel cepstrum segments, which can achieve the extraction of accurate features in the breath audio data. That is to say, it helps to further improve the accuracy of the breath sound recognition model in distinguishing breath sounds.
[0024] In the fifth possible implementation manner of the first aspect, the breath sound feature extraction model may include a convolutional layer, a multi-scale residual layer, and a bottleneck layer connected in sequence, where there are at least two multi-scale residual layers.
[0025] In the embodiment of the present application, the breath sound feature extraction model can compress the input Mel cepstrum coefficients through a convolutional layer; then, extract the features expressed by the corresponding audio from the compressed data through a multi-layer multi-scale residual layer. Among them, the multi-scale residual module adds cross-fusion of convolutional features of different scales on the framework of the residual network, and can obtain richer and more comprehensive breath sound features; finally, fuse the features of each module through the bottleneck layer and output the breath sound features. This helps to ensure the accuracy of the extracted breath sound features.
[0026] In the sixth possible implementation manner of the first aspect, the training process of the breath sound feature extraction model may include the following steps: First, select K batches of breath samples from the breath sample set used to train the breath sound recognition model, where K>1; then, perform data augmentation processing on the audio in each breath sample based on the features of the breath sound to obtain at least one augmented breath sample corresponding to each breath sample, and the augmented breath samples are used to train the breath sound feature extraction model; finally, during the process of training the breath sound feature extraction model based on the augmented breath samples corresponding to each batch of breath samples, determine the model training loss, and adjust the model parameters of the breath sound feature extraction model based on the obtained model training loss.
[0027] Among them, the model training loss includes a user recognition cross-entropy loss and a contrast learning similarity loss. The user recognition cross-entropy loss is used to describe the accuracy of the breath sound feature extraction model in user recognition, and the contrast learning similarity loss is used to describe the similarity between the features extracted by the breath sound feature extraction model and the true features.
[0028] In the embodiments of the present application, data augmentation is performed on each training sample, and the augmented samples obtained by data augmentation are used to train the breath sound feature extraction model, which can improve the generalization ability and robustness of the trained breath sound feature extraction model. That is, the trained breath sound feature extraction model can accurately extract the breath sound features of the user in various situations. In addition, during the model training process, the loss of the model is described from two dimensions of user recognition and contrast learning, so as to optimize and adjust the model, which can make the breath sound features extracted by the trained breath sound feature extraction model have strong recognition ability and can more comprehensively and accurately describe the personal breath characteristics of the user. When using the breath sound features extracted by the breath sound feature extraction model for the host identification, it helps to improve the accuracy of the host identification.
[0029] Optionally, the user recognition cross-entropy loss and the contrast learning similarity loss can correspond to the first weight coefficient and the second weight coefficient respectively. The model training loss can be the sum of the first loss weight value and the second loss weight value. Among them, the first loss weight value is the product of the user recognition cross-entropy loss and the first weight coefficient, and the second loss weight value is the product of the contrast learning similarity loss and the second weight coefficient.
[0030] Optionally, there can be M training samples in each batch of training samples, and the M training samples come from N users. Both M and N are integers greater than 1. The fact that the M training samples in each batch come from N users means that in the case of multi-batch training, the audio of the same user may be selected multiple times. In this way, the accuracy and stability of the trained model can be further improved.
[0031] Optionally, for each batch of M training samples, one of them can be marked as the positive sample and the rest are marked as the negative samples. In this way, the trained breath sound feature extraction model can have stronger recognition ability for the breath sound. That is, it can improve the generalization ability and robustness of the model, thus helping the trained breath sound feature extraction model to extract the breath sound features more accurately and effectively.
[0032] It can be understood that in the embodiments of the present application, the execution entity for training each model, such as the feature extractor, the breath sound feature recognition model, and the breath sound feature extraction model, can be the terminal device or other devices, such as a server. That is to say, the model can be directly trained in the terminal device, or the model can be trained in other devices and then moved into the terminal device for use.
[0033] In the seventh possible implementation manner of the first aspect, the breath sound feature extraction model can perform data augmentation processing on the breath samples through one or more of the following methods: In the first method, random noise with a random signal-to-noise ratio level is added to a random position in the audio of the breath sample to obtain an enhanced breath sample. The random noise includes at least one of the following: household noise, outdoor noise, office noise, transportation noise, and animal noise. This data augmentation method can mimic breath audio in various scenarios.
[0034] In the second method, the speed of the audio in the breath sample is adjusted. This data augmentation method can mimic rapid breathing and slow breathing.
[0035] In the third method, the volume range of the audio in the breath sample is adjusted. This data augmentation method can mimic audio collected at close range and audio collected at a distance.
[0036] In the embodiments of this application, the breath sound feature extraction model can imitate various situations that may exist when the same user breathes. Therefore, the enhanced samples obtained by imitation can be used to train the breath sound feature extraction model, which can improve the generalization ability and robustness of the trained breath sound feature extraction model. That is to say, the trained breath sound feature extraction model can accurately extract the breath sound features of the user in various situations.
[0037] In the eighth possible implementation manner of the first aspect, the terminal device can collect the user data of the user to be measured through the following steps: First, according to the usage scenario corresponding to the unlocking operation, determine the audio collection duration of the breath audio data to be measured, where the audio collection duration is less than or equal to the target duration, and the target duration is the duration of the breath audio data collected when the breath sound feature template is entered; then, collect the breath audio data to be measured of the user to be measured according to the audio collection duration.
[0038] In the embodiments of this application, the terminal device combines different usage scenarios to determine the collection duration of the breath audio data to be measured, which can achieve flexible collection of the breath audio data to be measured and help improve the user experience.
[0039] Optionally, the collection duration of the breath audio data to be measured is usually related to the usage scenario, and the collection duration is usually positively correlated with the recognition accuracy of the usage scenario for the device owner identification and negatively correlated with the recognition speed of the usage scenario for the device owner identification. For example, for a usage scenario with very high requirements for the recognition accuracy of the device owner, the collection duration can be as long as possible, and for a usage scenario with very high requirements for the recognition speed of the device owner, the collection duration can be as short as possible.
[0040] Optionally, the mapping relationship between the usage scenario and the collection duration can be pre-stored in the terminal device. In this way, the terminal device can directly determine the collection duration of the breath audio data to be measured in combination with the usage scenario.
[0041] In the ninth possible implementation manner of the first aspect, the terminal device performs the owner identification according to the user data and the pre-recorded owner template, which can be implemented through the following steps: First, when there is no face in the to-be-tested face image, the terminal device can extract the to-be-tested breath sound features from the to-be-tested breath audio data; then, the terminal device can determine whether the to-be-tested user is the owner according to the to-be-tested breath sound features and the breath sound feature template in the owner template.
[0042] In the embodiments of the present application, the terminal device can continue to use the breath audio data for owner identification when no face is recognized. In this way, it can be realized that the owner identification can still be performed when the face recognition result is not ideal, which can improve the accuracy and stability of the owner identification.
[0043] Optionally, when there is a face in the to-be-tested face image, the terminal device can use the to-be-tested face image and the face template to determine whether the to-be-tested user is the owner through face recognition.
[0044] In the tenth possible implementation manner of the first aspect, the terminal device determines whether the to-be-tested user is the owner according to the to-be-tested breath sound features and the breath sound feature template in the owner template, which can be implemented through the following steps: First, when there is one breath sound feature template in the owner template, the terminal device can determine the similarity between the to-be-tested breath sound features and the breath sound feature template, and when the obtained similarity is greater than the pre-set second determination threshold, determine that the to-be-tested user is the owner; then, when there are multiple breath sound feature templates in the owner template, the terminal device can respectively determine the similarity between each breath sound feature template and the to-be-tested breath sound features, and when the average value of the obtained multiple similarities is greater than the pre-set fourth determination threshold, determine that the to-be-tested user is the owner.
[0045] In the embodiments of the present application, when there are multiple breath sound feature templates in the owner template, by using the average value of the similarities between the to-be-tested breath sound features and each breath sound feature template to determine whether the to-be-tested user is the owner, the accuracy of the owner identification can be improved.
[0046] Optionally, when there are multiple breath sound feature templates in the owner template, the terminal device can also respectively determine the similarity between each breath sound feature template and the to-be-tested breath sound features, and when the maximum value of the obtained multiple similarities is greater than the second determination threshold, determine that the to-be-tested user is the owner. On the contrary, when the maximum value of the obtained multiple similarities is less than or equal to the second determination threshold, determine that the to-be-tested user is not the owner.
[0047] In some alternative implementation manners of the embodiments of the present application, when multiple owner templates are recorded in the terminal device, the terminal device can also determine whether the user to be tested is the owner in the following manner: For each owner template, when the owner template includes a breathing sound feature template, determine the similarity between the breathing sound feature to be tested and the breathing sound feature template. When the obtained similarity is greater than the second determination threshold, determine that the user to be tested is the owner, and specifically, the owner indicated by this owner template; when the owner template includes multiple breathing sound feature templates, respectively determine the similarity between each breathing sound feature template in the owner template and the breathing sound feature to be tested. When the average value of the obtained multiple similarities is greater than the fourth determination threshold, determine that the user to be tested is the owner, and specifically, the owner indicated by this owner template; Optionally, when the user to be tested is not the owner indicated by any owner template, the terminal device can determine that the user to be tested is not the owner.
[0048] In the eleventh possible implementation manner of the first aspect, when the owner template includes a breathing sound feature template and meets the template update condition, the terminal device can also switch the breathing sound feature template in the owner template to the current breathing sound feature, where the current breathing sound feature is the breathing sound feature in the breathing audio data collected when the face unlock is successful.
[0049] Among them, the template update condition includes: the face unlock is successful, the update period is reached, and the breathing sound confidence corresponding to the current breathing sound feature is greater than a preset third determination threshold, where the third determination threshold is greater than the first determination threshold.
[0050] In the embodiments of the present application, since over time, the features of the owner may change, such as face changes, breathing sound changes, etc., therefore, timely updating of the owner template can ensure the accuracy of owner recognition and improve the user experience.
[0051] It can be understood that the terminal device usually updates the owner template of the owner recognized when the face unlock is successful. For example, when the face unlock is successful and the recognized owner is A, then, when the template update condition is met, the terminal device can update the owner template of owner A.
[0052] In the twelfth possible implementation manner of the first aspect, when the owner template includes multiple breathing sound feature templates and meets the template update condition, the terminal device can first determine the similarity between the current breathing sound feature and each breathing sound feature template respectively; then, determine the maximum similarity with the largest corresponding value and the target breathing sound feature template corresponding to the maximum similarity; finally, update the owner template according to the maximum similarity and the target breathing sound feature template.
[0053] In an embodiment of the present application, the terminal device can update the owner template by combining the maximum similarity and the corresponding respiratory sound feature template of the maximum similarity. For example, the current respiratory sound feature can directly replace the respiratory sound feature template corresponding to the maximum similarity, so as to effectively update the owner template.
[0054] In the thirteenth possible implementation manner of the first aspect, the terminal device updates the owner template according to the maximum similarity and the target respiratory sound feature template, which can be: when the maximum similarity is greater than a preset fifth determination threshold, the terminal device can replace the target respiratory sound feature template in the owner template with the current respiratory sound feature. When the maximum similarity is less than or equal to the fifth determination threshold, and the number of stored respiratory sound feature templates in the terminal device is less than a preset template number threshold, the terminal device can add the current respiratory sound feature to the owner template. When the maximum similarity is less than or equal to the fifth determination threshold, and the number of stored respiratory sound feature templates in the terminal device is equal to the template number threshold, the terminal device can delete the respiratory sound feature template with the longest storage time in the owner template and add the current respiratory sound feature to the owner template.
[0055] In an embodiment of the present application, the terminal device can update the owner template by combining the number of stored respiratory sound feature templates and the storage time of the respiratory sound feature templates, so as to realize flexible and effective update of the owner template.
[0056] In a second aspect, an embodiment of the present application provides an owner recognition device applied to a terminal device. The device includes: A data acquisition unit for collecting user data of a user to be measured in response to an unlocking operation of the user to be measured on a target content, where the target content is in a locked state, and the user data includes at least one of the following: a face image to be measured and respiratory audio data to be measured; An identification execution unit for performing owner identification according to the user data and a pre-recorded owner template, where the owner template includes at least one of the following: a face template and a respiratory sound feature template; An unlocking execution unit for unlocking the target content when the owner is recognized.
[0057] As an embodiment of the present application, the owner recognition device can implement the method of any item in the above first aspect.
[0058] In a third aspect, an embodiment of the present application provides a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the method of any item in the above first aspect is implemented.
[0059] Fourthly, an embodiment of the present application provides a computer-readable storage medium storing a computer program, which when executed by a processor, implements the method according to any one of the above first aspects.
[0060] Fifthly, an embodiment of the present application provides a chip system, which includes a processor coupled to a memory. The processor executes a computer program stored in the memory to implement the method according to any one of the above first aspects. The chip system may be a single chip or a chip module composed of multiple chips.
[0061] Sixthly, an embodiment of the present application provides a computer program product, which when running on a terminal device, causes the terminal device to execute the method according to any one of the above first aspects.
[0062] It can be understood that the beneficial effects of the above second to sixth aspects can be referred to the relevant descriptions in the above first aspect and will not be elaborated here. Description of Drawings
[0063] Figure 1A It is a schematic diagram of an information display effect in an inorganic host recognition scenario provided by an embodiment of the present application; Figure 1B It is a schematic diagram of an information display effect in an organic host recognition scenario provided by an embodiment of the present application; Figure 1C It is another schematic diagram of an information display effect in an inorganic host recognition scenario provided by an embodiment of the present application; Figure 1D It is another schematic diagram of an information display effect in an organic host recognition scenario provided by an embodiment of the present application; Figure 2 It is a schematic diagram of a scenario of a multi-modal fusion host recognition method provided by an embodiment of the present application; Figure 3 It is a schematic diagram of the process of inputting a face template provided by an embodiment of the present application; Figure 4A It is a schematic diagram of the process of the audio processing module processing respiratory audio data provided by an embodiment of the present application; Figure 4B It is a schematic diagram of the process of the audio processing module preprocessing respiratory audio data provided by an embodiment of the present application; Figure 4C It is a schematic diagram of a respiratory sound recognition model provided by an embodiment of the present application; Figure 4D It is a schematic diagram of a respiratory sound feature extraction model provided by an embodiment of the present application; Figure 4ESchematic diagram of a structure of a multi-scale residual layer in a breath sound feature extraction model provided by an embodiment of the present application; Figure 5A Schematic diagram of a structure of a cross-modal encoder provided by an embodiment of the present application; Figure 5B Schematic diagram of a structure of a cross-channel linear encoder in a cross-modal encoder provided by an embodiment of the present application; Figure 5C Schematic diagram of a structure of a cross-channel attention encoder in a cross-modal encoder provided by an embodiment of the present application; Figure 6A Schematic diagram of a process of mobile phone owner identification based on multi-modal fusion provided by an embodiment of the present application; Figure 6B Schematic diagram of a process of identifying the phone owner based on breath sound provided by an embodiment of the present application; Figure 6C Schematic diagram of a process of identifying the phone owner by a multi-modal fusion phone owner identification model provided by an embodiment of the present application; Figure 6D Schematic diagram of a structure of a fusion decision model provided by an embodiment of the present application; Figure 6E Model training architecture of positive and negative sample pairing based on dynamic hard sample mining provided by an embodiment of the present application; Figure 7 ROC curve graph of phone owner identification corresponding to various modalities provided by an embodiment of the present application; Figure 8 PR curve graph of phone owner identification corresponding to various modalities provided by an embodiment of the present application; Figure 9 Schematic diagram of a structure of a phone owner identification device provided by an embodiment of the present application; Figure 10 Schematic diagram of a structure of a terminal device provided by an embodiment of the present application; Figure 11 Software structure block diagram of a terminal device provided by an embodiment of the present application. Detailed implementation manners
[0064] In the following description, specific details such as specific system structures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.
[0065] Some concepts that may be involved in the embodiments of the present application are described as follows: (1)Multiple: Unless otherwise specified, in the embodiments of this application, "multiple" means two or more.
[0066] (2)Difficult sample: In the embodiments of this application, a difficult sample refers to a negative sample that is misrecognized as a positive sample by the model and has the highest corresponding recognition probability during the model training process.
[0067] (3)False Acceptance Rate (FAR): In the embodiments of this application, FAR refers to the proportion of samples that are incorrectly classified as other categories in a classification problem. In the fields of machine learning and pattern recognition, we usually focus on two types of errors: False Positive and False Negative. False Positive refers to the situation where a negative sample is incorrectly classified as a positive sample, while False Negative refers to the situation where a positive sample is incorrectly classified as a negative sample. The false acceptance rate can be calculated by the following formula: False acceptance rate = (number of false positive samples + number of false negative samples) / total number of samples.
[0068] In practice, the lower the false acceptance rate, the better the performance of the classifier. In practical applications, we usually weigh the weights of false positives and false negatives according to specific requirements, select a suitable classifier or adjust the parameters of the model to achieve the best false acceptance rate.
[0069] (4)Receiver Operating Characteristic (ROC) curve: In the embodiments of this application, the ROC curve is a curve plotted with the True Positive Rate (TPR) as the vertical axis and the False Positive Rate (FPR) as the horizontal axis. It is a visualization tool for evaluating the performance of a binary classification model at different thresholds.
[0070] When plotting the ROC curve, it is necessary to sort the samples according to the output probability or decision value of the classification model. Then, starting from the lowest threshold, gradually increase the threshold, classify the samples above the current threshold as positive examples, and classify the samples below the current threshold as negative examples. At each threshold, calculate the corresponding true positive rate and false positive rate, and plot them as coordinates. The finally obtained curve is the ROC curve.
[0071] (5)Area Under the Curve (AUC): In the embodiments of this application, AUC usually refers to the area under the ROC curve, and AUC is an index used to evaluate the performance of a binary classification model. The value range of AUC is between 0.5 and 1, and the larger the value, the better the model performance.
[0072] (6) Precision-Recall curve (PR curve): In the embodiments of the present application, the PR curve is another visualization tool for evaluating the performance of a binary classification model. Different from the ROC curve, the PR curve is plotted with precision on the vertical axis and recall on the horizontal axis.
[0073] When plotting the PR curve, the samples need to be sorted according to the scores output by the model. Then, starting from the lowest threshold, the threshold is gradually increased, and the samples above the current threshold are classified as positive examples, and the samples below the current threshold are classified as negative examples. At each threshold, the corresponding precision and recall are calculated and plotted as coordinates. The finally obtained curve is the PR curve. Among them, each point on the PR curve represents the precision corresponding to different recall rates. The closer the PR curve is to the upper right corner, the better the performance of the classifier in various situations.
[0074] (7) Area under the PR curve (AUC-PR, AP): In the embodiments of the present application, AP is an index for comprehensively evaluating the performance of the model. In practice, generally, the larger the AP value, the better the classification performance of the model.
[0075] (8) Service card: In the embodiments of the present application, the service card is a service form that provides a more fine-grained atomic service capability than an application (APP), and directly presents the services or content that the user cares most about to the user in the form of an interactive service card. The service card can be embedded in various APPs or interactive scenarios to better meet the user's needs. The service card is generally a page in the application or the corresponding APP of the application. When the user clicks on the service card, they can be redirected to the corresponding application or a certain service page within the APP. Unless otherwise specified, the "application" referred to in the embodiments of the present application can be either the application or the corresponding APP of the application.
[0076] The generation process of the service card is as follows: The provider of the application service uploads the service content that needs to be pushed to the user on the server. The server generates the corresponding service card according to the service content uploaded by the provider and pushes the service card to the user's terminal device for display, so that the user's terminal device presents the service pushed by the provider of the application service in the form of a service card. Generally speaking, one service card corresponds to a part of the service functions of the application. The provider of the application program can split the services of the application into multiple service modules, each service module corresponding to one service card, and then push different service cards to the user according to different service scenarios.
[0077] (9) TAR, TAR@FAR: In the embodiments of this application, the True Accept Rate (TAR) is used to indicate the proportion of correct judgments. TAR@FAR refers to the value of TAR under a specific FAR. For example, TAR@FAR = 1% means that when FAR = 1%, the corresponding value of TAR.
[0078] (10) Neural network layers, activation functions: In the embodiments of this application, the neural network layers involved include Linear layer, Group layer, norm layer, Inputs layer, Features layer, conv layer, Stem layer, Head layer, etc.
[0079] Among them, the Linear layer is usually referred to as the linear layer and can also be called the fully connected layer or the dense layer. The Linear layer is usually used to perform a linear transformation on the input data to obtain the output.
[0080] The Group layer is usually used to group the input data, and each group is calculated independently. For example, the grouped data can be sent to different convolutional kernels for convolution operations. This grouping can improve the calculation efficiency and the expressive ability of the model. Through the grouping operation, the Group layer can increase the non-linear characteristics of the model and improve the model's learning ability for complex patterns. In addition, the Group layer can also reduce the number of parameters, lower the calculation cost, and improve the generalization ability of the model.
[0081] The norm layer is usually referred to as the regularization layer or the normalization layer. The norm layer is usually used to perform normalization processing on the input data. The norm layer aims to adjust the distribution of the input data to reduce redundancy and correlation, thereby improving the robustness and generalization ability of the network, so as to better adapt to the training or inference process of the network model.
[0082] The Inputs layer is usually referred to as the input layer. The Inputs layer is used to receive external input data and convert it into a tensor form that can be processed by the neural network. The input layer is usually the first layer of the network structure and defines the format, dimension, attributes, etc. of the input data.
[0083] The Features layer is usually referred to as the feature layer. The Features layer is an intermediate layer or a hidden layer in the neural network and is used to extract discriminative features from the input data. They can capture the key information of the input data through convolution, pooling, embedding, etc., and provide a richer and more meaningful feature representation for the subsequent network layers.
[0084] The conv layer, usually known as the Convolutional layer, extracts local features of the input data by applying convolutional operations and generates a feature map as the output. The conv layer is widely used in fields such as computer vision and speech processing.
[0085] The Stem layer usually extracts richer features by performing spatial transformation and local feature fusion on the feature map. In the embodiments of this application, the Stem layer can be used to expand the feature dimension.
[0086] The Head layer is usually used to describe the last or top layer of the model. The main function of the Head layer is to transform the previous feature representation into the final output result.
[0087] The activation functions involved include the Sigmoid function, the softmax function, the ReLU function, etc.
[0088] Among them, the Sigmoid function is a common non-linear activation function that maps the input real value to the output between 0 and 1. The Sigmoid function is usually used in the output layer of binary classification problems to transform the output of the model into a probability value.
[0089] The softmax function is a commonly used activation function that transforms the input real vector into a probability distribution such that the sum of all output probabilities is equal to 1. The softmax function is usually used in the output layer of multi-class classification problems. It can transform the output of the model into the probability of each class for the final classification decision.
[0090] The Rectified Linear Unit (ReLU), also known as the rectified linear unit, is a commonly used activation function.
[0091] (11) Inertial Measurement Unit (IMU): In the embodiments of this application, the IMU is a sensor used to measure and track the posture of an object (such as direction, angle, and speed). The IMU contains multiple measurement units, including accelerometers, gyroscopes, and magnetometers, etc. Among them, the accelerometer can measure the linear acceleration of the object along the three axes of x, y, and z; the gyroscope can measure the angular velocity of the object around the three axes of x, y, and z; the magnetometer can measure the magnetic field strength and direction of the object in the earth's magnetic field.
[0092] In the embodiments of this application, the IMU data can include the acceleration, pitch angle, roll angle, etc. of the terminal device in each direction. The IMU data is the data collected by the IMU in the terminal device.
[0093] (12) Multimodality: In the embodiments of the present application, multimodality generally refers to multiple types. Multimodal data usually refers to data of multiple different types. For example, face data, breathing audio data, touch screen data, and IMU data are four different types of data.
[0094] (13) Audio, time-frequency diagram, Mel spectrogram, logarithmic Mel spectrogram, Mel cepstrum, Mel cepstral coefficient: In the embodiments of the present application, performing a short-time Fourier transform on audio can convert the audio into a time-frequency diagram. Performing a Mel scale transform on the time-frequency diagram can obtain the corresponding Mel spectrogram. Performing a logarithmic transform on the Mel spectrogram can obtain the corresponding logarithmic Mel spectrogram, or Mel cepstrum for short. Performing a discrete cosine transform on the Mel cepstrum can obtain the corresponding Mel cepstral coefficient.
[0095] Terminal devices, especially mobile phones, have gradually become the center of the personal digital world, carrying a large amount of personal information and privacy data. During the use of terminal devices, the private information of the device owner may be displayed. If the private information of the device owner is seen by someone other than the device owner, it will undoubtedly lead to the leakage of the device owner's privacy.
[0096] Taking the terminal device as a mobile phone as an example, the following lists several usage scenarios in which the private information of the device owner may be leaked. It can be understood that the terminal device in the embodiments of the present application is not limited to being a mobile phone, and can also be other devices, such as tablets, laptop computers, etc.
[0097] Usage scenario 1: The mobile phone supports the service card function. When the mobile phone displays a service card, if the service card contains private information such as the device owner's bank card number, password, and express delivery pick-up code. At this time, if the user next to the mobile phone is not the device owner or there are other people besides the device owner, when the mobile phone displays the service card, the private information of the device owner displayed on the service card may be seen by others intentionally or unintentionally, resulting in the leakage of the device owner's private information.
[0098] Usage scenario 2: The device owner shares the content of the mobile phone with others. For example, the device owner uses their own mobile phone to report or display materials to the company leader, or the device owner watches a video with a friend. Due to some applications on the mobile phone, such as WeChat, when receiving a notification message, it usually previews the notification message at the upper end of the mobile phone interface. In this usage scenario, if the mobile phone receives a notification message from an application and previews the notification message, the content of the notification message may be seen by others, resulting in the leakage of the device owner's private information and an embarrassing situation.
[0099] Usage scenario three: The phone owner is looking at the phone in a crowded place. For example, the phone owner is reading news or short videos in an elevator, subway, or bus. In this usage scenario, if the phone receives a notification message from a certain application and previews the notification message, the content of the notification message may be seen intentionally or unintentionally by others, resulting in the leakage of the phone owner's private information.
[0100] Usage scenario four: The phone of the phone owner is borrowed by others. For example, the phone owner's friend borrows the phone owner's phone to watch short videos, and the phone owner is not present. In this usage scenario, if the phone receives a notification message from a certain application and previews the notification message, the content of the notification message may be seen intentionally or unintentionally by others, resulting in the leakage of the phone owner's private information.
[0101] Combining the above usage scenarios, it can be known that during the use of the terminal device, it is very necessary to timely identify the phone owner to protect the security of the phone owner's private information.
[0102] The following further combines specific usage scenarios to illustrate the use or possible effects of the phone owner identification scheme of the embodiments of the present application in a comparative manner.
[0103] Figure 1A Fig. shows a schematic diagram of an information display effect in the case of no phone owner identification scenario. Figure 1A In, phone owner A is looking at the phone, and there are other people beside, such as B and C. The phone directly displays the service card of the memo application. As Figure 1A shown, the service card has A's bank card number and password. At this time, B and C can see the content of the service card, that is, others can see the phone owner's private information.
[0104] Figure 1B Fig. shows a schematic diagram of an information display effect in the case of phone owner identification scenario. Figure 1B In, phone owner A is looking at the phone, and there are other people beside, such as B and C. When the phone does not recognize the phone owner, it locks the content of the service card, that is, hides the content of the service card. In this way, the security of the phone owner's private information can be protected.
[0105] It can be understood that, as Figure 1B shown, if A wants to view the content of the service card, he can further click the " " icon. The phone can respond to the user's click on the " " icon and display the content of the service card after recognizing the phone owner.
[0106] Figure 1C Fig. shows another schematic diagram of an information display effect in the case of no phone owner identification scenario. Figure 1CIn this case, the owner, Little A, is looking at the mobile phone, and there are other people beside, such as Little B and Little C. The mobile phone receives a notification message from the messaging application and directly previews the notification message. At this time, both Little A and Little B and Little C beside can see the content of the notification message. That is to say, others can peek at the owner's private information.
[0107] Figure 1D It shows another schematic diagram of the information display effect in the scenario of owner recognition. Figure 1D In this case, the owner, Little A, is looking at the mobile phone, and there are other people beside, such as Little B and Little C. The mobile phone receives a notification message from the messaging application. In the case where the owner is not recognized, the content of the notification message is automatically hidden. It can be understood that if Little A wants to view the content of the notification message, he can click on the notification message whose content is hidden. The mobile phone can respond to the user's click on the notification message and display the content of the notification message after recognizing the owner.
[0108] Combined with the above Figures 1A - 1D By comparing the information display in the two usage scenarios shown, it can be clearly seen that during the use of the terminal device, timely owner recognition can effectively protect the security of the owner's private information.
[0109] In fact, the terminal device identifies the owner's identity and treats the owner and non-owner differently. In addition to protecting the owner's private information from being leaked, in some application scenarios, it can also provide a more convenient device usage experience for the owner. For example, for the owner, payment can be made without a password, while for non-owners, a password must be entered to make a payment.
[0110] Currently, the more commonly used owner recognition solutions are face recognition or fingerprint recognition. However, in existing solutions for using the face to recognize the owner, in scenarios such as low light, backlight, strong light, and deviation of the distance and angle, the recognition rate will decrease due to poor face image quality. And in existing solutions for using fingerprints to recognize the owner, in scenarios where there is dirt on the finger or on the fingerprint touch area of the terminal device, such as sweat or water stains on the finger, or dust or foreign objects on the fingerprint touch area, the recognition rate will also decrease. That is to say, as the scenario changes, such as the light getting darker or stronger, the recognition rate of existing owner recognition solutions may decrease, and even the owner may not be correctly recognized.
[0111] In view of the above technical problems in the related art, in order to improve the accuracy and stability of the terminal device's recognition of the owner, in the embodiments of the present application, the terminal device combines the user's face data, the IMU of the terminal device, the touch screen data of the terminal device, and the user's breathing audio data to comprehensively recognize the owner, thereby effectively improving the success rate of owner recognition. That is to say, the embodiments of the present application can use multi-modal data to recognize the owner.
[0112] Figure 2 This is a schematic diagram of the scenario of the multi-modal fusion-based device owner recognition method provided by the embodiments of the present application. Figure 2 In this scenario, the terminal device is a mobile phone. The mobile phone can comprehensively judge whether the current user is the device owner by combining the current user's IMU data, touch screen data, breathing audio data, and face recognition score. If the mobile phone recognizes that the current user is the device owner, the privacy information is normally displayed. If the mobile phone recognizes that the current user is not the device owner, the privacy information is hidden.
[0113] The embodiments of the present application at least have the following beneficial effects: 1. The terminal device combines the user's face data, the terminal device's IMU data, the terminal device's touch screen data, and the user's breathing audio data to comprehensively identify the device owner, so that the device owner can still be identified when the scenario changes, which can improve the accuracy and stability of device owner recognition.
[0114] 2. When the face recognition result is not ideal, at least one of the terminal device's IMU data, the terminal device's touch screen data, and the user's breathing audio data can be combined to further determine whether it is the device owner, which can expand the device owner recognition scenario and improve the accuracy and stability of device owner recognition.
[0115] 3. The IMU data and touch screen data of the terminal device reflect the user's usage habits of the terminal device. Combining these two pieces of data to further determine whether it is the device owner can make the recognition result more accurate and reliable, that is, it can improve the accuracy rate of device owner recognition.
[0116] The embodiments of the present application are applicable to scenarios where the device owner needs to be recognized in the terminal device.
[0117] The device owner recognition method provided by the embodiments of the present application can be applied to terminal devices such as mobile phones, tablet computers, and wearable devices. At this time, the terminal device is the execution subject of the device owner recognition method provided by the embodiments of the present application. The specific type of the terminal device is not specifically limited in the embodiments of the present application.
[0118] Taking the terminal device as a mobile phone as an example, the device owner recognition method provided by the present application will be specifically described from three aspects: device owner template entry, device owner recognition based on the device owner template, and device owner template update.
[0119] (1) Device owner template entry The owner template is usually pre - entered or stored information used to describe the characteristics of the owner. The owner template can include a face template, a breathing sound feature template, and a cross - modal feature template. Among them, the face template is pre - entered or stored information used to describe the facial features of the owner. For example, it can be an image of the owner's face. The breathing sound feature template is pre - entered or stored information used to describe the breathing sound features of the owner. The cross - modal feature template is pre - entered or stored information used to describe the comprehensive features of the owner in multiple modalities.
[0120] In practice, a mobile phone can have one owner's template entered, or multiple owners' templates entered. That is to say, a mobile phone can have one owner template or multiple owner templates.
[0121] In practice, one owner template usually corresponds to one owner.
[0122] Optionally, in one owner template, there can be one or more face templates of the owner, one or more breathing sound feature templates, and one or more cross - modal feature templates.
[0123] In the embodiments of this application, the owner template entry process will be described from the following aspects: 1. Face template entry; 2. Breathing sound feature template entry; 3. Cross - modal feature template entry.
[0124] 1. Face template entry.
[0125] In the embodiments of this application, the owner can successively enter the faces of multiple people. For example, the owner can first enter his / her own face, and then enter the faces of his / her father and mother. For example, owner A can enter his / her own face, and then enter the faces of his father B and mother C. When A only enters his / her own face, there is one owner. At this time, the mobile phone can retain an owner identity identifier (IDentifier, ID), such as A. If after A enters his / her own face, he / she then enters the faces of his father B and mother C successively, at this time, the mobile phone will retain 3 owner IDs, such as A, B, C. It can be understood that the embodiments of this application do not limit the representation method of the owner ID. For example, it can be Chinese characters, letters, numbers, or a combination of numbers and letters, etc.
[0126] It can be understood that in the embodiments of this application, face entry is the same as face template entry.
[0127] Please refer to Figure 3 , Figure 3 which shows a schematic diagram of the process of entering the face template. As Figure 3In (a) as shown, the phone owner, Little A, can click on the biometric and password control 302 on the settings interface 301 of the settings app (an app dedicated to performing system settings functions). The phone can respond to the user's click on the biometric and password control 302 and enter the biometric and password interface 303. After that, as Figure 3 shown in (b), the phone can respond to Little A's further click on the "Face Recognition" control 304 and enter the input lock screen password interface 305. After that, as Figure 3 shown in (c), the phone can respond to Little A's further input of the correct lock screen password and enter the face recognition interface 306. After that, as Figure 3 shown in (d), the phone can respond to Little A's further click on the "Start Enrollment" control 307 and enter the face enrollment interface 308. As Figure 3 shown in (e), on the face enrollment interface 308, Little A aligns the face with the face capture frame, and the phone captures Little A's face image, completing the entry of Little A's face template. As Figure 3 shown in (f), when the phone completes the entry of Little A's face template, it can display "Enrollment Successful" on the face enrollment interface 308.
[0128] According to a similar operation method as Figure 3 shown, Little A can help mom and dad enter their faces on the phone, thus completing the entry of their face templates by the phone.
[0129] In the embodiments of this application, after multiple face templates are entered, the phone can identify the phone owners corresponding to each face template. For example, if the phone stores the face templates of Little A, Big B, and Big C. When Little A holds the phone for face unlocking, the phone can identify that the current operator is the phone owner, and specifically, it is the phone owner Little A. That is to say, the phone can determine which specific phone owner is operating the phone. Of course, any phone owner can perform face unlocking on the phone.
[0130] In some application scenarios, the same user can be enrolled once or multiple times. For example, the same user can perform multiple face enrollments, and each face enrollment can be assigned a corresponding phone owner ID by the phone. That is to say, there can be multiple phone owner IDs for the same user. Or rather, one user can correspond to multiple phone owner templates. In the case where a user is assigned multiple phone owner IDs, one user can be regarded as multiple phone owners by the phone.
[0131] 2. Respiratory sound feature template entry.
[0132] In the embodiments of the present application, when the face template of the phone owner is successfully entered, the phone can continue to collect the breathing audio data of the phone owner through the microphone, and determine whether the breathing audio data contains breathing sounds. If it contains breathing sounds, the breathing sound features are further extracted therefrom, and the extracted breathing sound features are determined as the breathing sound feature template of the phone owner. That is to say, the trigger condition for entering the breathing sound feature template (or called the breathing feature entry condition) can be the successful entry of the face template of the phone owner.
[0133] It should be noted that within a short period of time after the successful entry of the phone owner's template, for example, within 5 seconds, the phone owner who has just entered the face template usually does not immediately put down the phone or leave the phone. That is to say, the breathing audio data collected within the short period of time after the successful entry of the face template usually has the breathing sounds of the phone owner. That is to say, when the face template of the phone owner is successfully entered, the phone continues to collect the breathing audio data within the short period of time after the successful entry of the face template, which can ensure that the collected breathing audio data contains the breathing sounds of the phone owner, thereby ensuring the accurate and effective entry of the breathing sound feature template of the phone owner.
[0134] Among them, the above-mentioned short period of time is usually a preset short time value. In practice, the above-mentioned short period of time can be any time from 0.5 seconds to 5 seconds. Since the longer the audio data, the more accurate the provided audio features, in order to ensure the accuracy of the breathing sound feature template of the phone owner entered, when entering the breathing sound feature template of the phone owner, the breathing audio data within 5 seconds after the successful entry of the face template is usually collected. That is to say, the breathing audio data for a relatively long period of time is collected to ensure the accuracy of the breathing sound feature template of the phone owner.
[0135] It can be understood that if the above-mentioned collected breathing audio data does not have breathing sounds, the phone can not enter the breathing sound feature template of the phone owner.
[0136] In some application scenarios, when it is detected that the above-mentioned breathing audio data does not have breathing sounds, the phone can continue to collect the breathing audio data and determine whether the breathing audio data contains breathing sounds when the subsequent face unlocking is successful. If it contains breathing sounds, the breathing sound features are further extracted therefrom, and the extracted breathing sound features are determined as the breathing sound feature template of the phone owner. That is to say, the trigger condition for entering the breathing sound feature template (or called the breathing feature entry condition) can also be the successful face unlocking, and the phone has not entered the breathing sound feature template of the phone owner who currently performs the unlocking when the face unlocking is successful.
[0137] Generally speaking, the trigger condition for breathing sound feature template entry (or the breathing feature entry condition) can be that the face template of the phone owner is successfully entered, or the face unlock is successful, and when the face unlock is successful, the breathing sound feature template of the phone owner who is currently performing the unlock has not been entered into the phone.
[0138] Regardless of the trigger condition for collecting breathing audio data, the collection parameters for obtaining the breathing sound feature template can be fixed. For example, the collection parameters can be: single-channel collection, sampling rate of 16 kHz, and collection duration of 5 seconds.
[0139] It should be noted that since breathing sounds are usually relatively weak, if the environment where the phone owner performs face entry is complex, for example, the phone owner is in a noisy scene such as a concert venue or a vegetable market, in this case, it is difficult for the phone to recognize the breathing sound from the further collected breathing audio data of the phone owner. Therefore, first judge whether there is a breathing sound in the breathing audio data, and only when it is determined that there is a breathing sound, perform the extraction of the breathing sound feature, which can reduce the loss of unnecessary computing resources.
[0140] The following will further illustrate how to extract the breathing sound feature template corresponding to the phone owner from the breathing audio data of the phone owner.
[0141] In the embodiment of the present application, during the process of breathing sound feature template entry, the phone can call the pre-stored audio processing module. The audio processing module extracts the breathing sound feature from the collected breathing audio data of the phone owner. After that, the phone can use the extracted breathing sound feature as the breathing sound feature template of this phone owner.
[0142] For further reference Figure 4A 。 Figure 4A FIG. is a schematic diagram of the process of the audio processing module provided in the embodiment of the present application for processing breathing audio data. The audio processing module can extract the breathing sound feature from the breathing audio data of the phone owner by executing S401-S405. Details are as follows: S401, the audio processing module performs audio preprocessing on the breathing audio data to obtain the mel cepstrum and mel cepstral coefficients of the breathing audio data.
[0143] Figure 4B FIG. shows a schematic diagram of the process of the audio processing module performing audio preprocessing on the breathing audio data. As Figure 4B shown, the process of the audio processing module performing audio preprocessing on the breathing audio data can include the following steps S4011-step S4014.
[0144] S4011, the audio processing module performs a short-time Fourier transform on the breathing audio data to convert the breathing audio data into a time-frequency diagram form.
[0145] In the embodiments of the present application, the collected respiratory audio data is usually a time-domain audio signal. Therefore, the short-time Fourier transform can be used to convert the respiratory audio data of the time-domain audio signal into a two-dimensional time-frequency signal form. Then, the two-dimensional time-frequency signal can be converted into a time-frequency diagram by combining the amplitude of the two-dimensional time-frequency signal.
[0146] Among them, the short-time Fourier transform (STFT) is a classic signal analysis method that decomposes a signal into components in time and frequency. STFT performs local analysis on the signal with a moving window of a certain length, improving the resolution of time-domain and frequency-domain analysis. The main parameters of STFT include the window function, window length, and step size. Commonly used window functions include the Hanning window, Hamming window, Blackman window, etc.
[0147] In some application scenarios, when performing the short-time Fourier transform, the window function used can be the Hanning window, the window length can be 0.025 seconds, and the step size can be 0.01 seconds. The number of time-domain points in the time-frequency signal can be represented by num_samples, where num_samples = sampling duration ÷ step size + 1. In this application scenario, the value of num_samples can be 501, where 501 = 5 ÷ 0.01 + 1.
[0148] S4012, the audio processing module performs a frequency-Mel scale transformation on the time-frequency diagram to convert the frequency of the time-frequency diagram into the Mel scale, thereby obtaining the Mel spectrogram.
[0149] Here, the mapping formula between frequency and Mel scale is as follows:
[0150] Among them, f is the frequency, mel(f) is the Mel scale, and × is the multiplication sign.
[0151] In some application scenarios, in the process of converting the time-frequency diagram into the Mel spectrogram, the parameters involved can include: the minimum frequency is 125.0 Hz, the maximum frequency is 7500.0 Hz, and the number of Mel filter banks is 64. Here, the number of Mel filter banks is 64, that is to say, the number of Mel filter banks is relatively large. It should be noted that the more the number of Mel filter banks, the more accurate the obtained Mel spectrogram can be, which helps to refine the low-frequency information and coarsen the high-frequency information of the audio signal.
[0152] S4013, the audio processing module performs a logarithmic transformation on the Mel spectrogram to obtain the logarithmic Mel spectrogram (also known as the Mel cepstrum).
[0153] Among them, the log mel - specgram is a method for representing audio features. It converts the spectral information of an audio signal into an energy distribution under the mel scale and takes the logarithm to enhance the discriminability of the features.
[0154] Here, the logarithm of the mel - specgram in S4012 is taken to obtain the log mel - specgram. The size of the obtained log mel - specgram can be num_samples×n_mels. Among them, num_samples is the number of time - domain points in the time - frequency signal, and n_mels is the number of mel filter banks.
[0155] S4014, the audio processing module performs a discrete cosine transform on the mel - cepstrum to obtain mel - cepstral coefficients.
[0156] Among them, the mel - cepstral coefficients are characteristic coefficients obtained through a series of signal processing steps and are used to describe the energy distribution of sound at different frequencies. Mel - cepstral coefficients are widely used in speech recognition and audio feature extraction. They can effectively capture the important features of an audio signal while reducing the dimension of the features, which helps to improve the accuracy and efficiency of pattern recognition.
[0157] Here, a discrete cosine transform (Discrete Cosine Transform, DCT) is performed on the mel - cepstrum obtained in step S4013, and the first 40 MFCC coefficients are retained, that is, n_mfcc = 40, to obtain the MFCC coefficients required for use. At this time, the size of the MFCC coefficients is num_samples×n_mfcc. Among them, num_samples is the number of time - domain points in the time - frequency signal.
[0158] It should be noted that the larger the value of n_mfcc, the more information the MFCC coefficients retain, which helps to make the extracted breath sound features more accurate.
[0159] It can be understood that the embodiments of the present application do not specifically limit the values of the parameters involved in audio pre - processing. For example, in some scenarios, the number of mel filter banks can also take other values other than 64. For another example, in some other scenarios, the number of MFCC coefficients retained can also be other values other than 40. The values of the above - mentioned parameters can be adjusted according to actual needs.
[0160] S402, the audio processing module divides the mel - cepstrum into multiple mel - cepstrum segments.
[0161] Here, due to the complete mel cepstrum, the amount of data is usually relatively large. Splitting the mel cepstrum into multiple mel cepstrum segments can simplify the feature representation, reduce the computational cost, and can more effectively capture the local features of the audio signal, thereby improving the accuracy of respiratory sound recognition.
[0162] In practical applications, it is found that if the segmentation granularity is too fine, on the one hand, it will increase the amount of calculation, and on the other hand, it will also reduce the accuracy of respiratory sound recognition. Therefore, in practice, the length of the mel cepstrum segment is usually set to about 1 second. For example, it can be 0.96 seconds. The reason is that the duration of a person's exhalation or inhalation is about 1 second. Switching the mel cepstrum segment to a segment with a length of about 1 second can enable each mel cepstrum segment to more completely describe the respiratory characteristics, thereby further improving the accuracy of respiratory sound recognition.
[0163] In practice, the mel cepstrum can be segmented with a window length of 0.96 seconds and a step size of 0.24 seconds to achieve splitting the mel cepstrum into multiple mel cepstrum segments.
[0164] During the process of splitting the mel cepstrum, when the mel cepstrum cannot be exactly split into an integer number of segments, that is, when it cannot be split into an integer number of 0.96 - second segments, zeros can be padded to the mel cepstrum, and the padded mel cepstrum will be split into an integer number of mel cepstrum segments.
[0165] In addition, since the model input size of the respiratory sound recognition model is usually constant. For example, the model input size of the respiratory sound recognition model can be the number of mel cepstrum segments corresponding to 5 - second respiratory audio data. However, in some application scenarios, since 5 - second audio data is not collected when collecting respiratory audio data. For example, maybe only 1 - second audio data is collected. To ensure that there is enough data volume to input into the respiratory sound recognition model, it is necessary to further pad zeros to the mel cepstrum.
[0166] S403, the audio processing module inputs the mel cepstrum segment into a pre - trained respiratory sound recognition model to obtain the confidence level of the respiratory audio data being recognized as a respiratory sound.
[0167] Among them, the above - mentioned confidence level of the respiratory sound is the confidence level or probability that the respiratory audio data is judged to be a respiratory sound. The input of the respiratory sound recognition model is a continuous series of mel cepstrum segments, and the output includes at least the confidence level of the respiratory sound.
[0168] Here, inputting the mel cepstrum segment of the respiratory audio data into the respiratory sound recognition model, the respiratory sound recognition model can output the confidence level or probability that the respiratory audio data is recognized as a respiratory sound, and thus it can be further judged whether the respiratory audio data contains a respiratory sound.
[0169] It should be noted that since the Mel cepstrum usually carries rich audio information, identifying whether the audio is a breath sound through the Mel cepstrum segment of the audio can ensure the accuracy of breath sound recognition, that is, it can enable the breath sound recognition model to more accurately determine whether the audio is a breath sound.
[0170] S404. The audio processing module combines the breath sound confidence with a preset threshold a to further determine whether the breath audio data is a breath sound.
[0171] Among them, the threshold a is a preset threshold for determining whether the audio is a breath sound. In some usage scenarios of the embodiments of the present application, the threshold a can also be referred to as the first determination threshold.
[0172] Here, if the breath sound confidence is greater than the threshold a, it is considered that the breath audio data is a breath sound. Conversely, if the breath sound confidence is less than or equal to the threshold a, it is considered that the breath audio data does not contain a breath sound.
[0173] S405. In the case where the breath audio data is determined to be a breath sound, the audio processing module inputs the Mel cepstrum coefficients in S401 into a pre-trained breath sound feature extraction model to extract the breath sound features of the breath audio data.
[0174] Here, during the process of the mobile phone inputting the breath sound feature template of the phone owner, the breath sound features extracted in S405 can be input as the breath sound feature template of the phone owner.
[0175] It should be noted that the Mel cepstrum coefficients can effectively capture the important features of the audio signal and have a small data dimension. In the case where it has been determined that the breath audio data is a breath sound, inputting the Mel cepstrum coefficients of the breath audio data into the breath sound feature extraction model to extract the breath sound features of the breath audio data can achieve fast and accurate extraction of the breath sound features of the breath audio data.
[0176] The following further combines Figure 4C to illustrate the breath sound recognition model in S403 and the training process of the breath sound recognition model.
[0177] In the embodiments of the present application, in order to increase the generalization of the breath sound recognition model, in addition to the breath sound category, other common human-related sound events are added, a total of seven types of sounds, including quiet, speaking, breath sound, coughing, sneezing, snoring, and environmental noise. In the scenario where the breath sound recognition model can recognize 7 types of sounds, the output of the breath sound recognition model can be the confidence of these 7 types of sounds. Specifically, for each type of sound, a corresponding confidence can be output, and for 7 types of sounds, 7 confidences can be output. Generally, the sum of the confidences of the 7 types of sounds output by the breath sound recognition model is usually 1.
[0178] Understandably, in some application scenarios, the output of the breath sound recognition model can also be sound category information. For example, it can output "breath", indicating that the final recognition result of the breath sound recognition model is breath.
[0179] Since it is found in practical applications that the generation principle of breath sounds is that the air in the environment is disturbed during the inhalation and exhalation processes, thus generating sounds, similar to the sound of wind. The disturbance caused by exhalation is greater than that caused by inhalation. Therefore, generally, the audio amplitude of exhalation is higher than that of inhalation. From a large time scale, breathing is a periodic, slow, continuous action event, inhaling and exhaling alternately at a certain frequency, and there are regular pauses in between. This characteristic of breathing is significantly different from coughing, sneezing, etc. From a small time scale, each inhalation and exhalation causes air disturbance, and the breath sounds caused by this disturbance are different from the sounds of sneezing, coughing, etc.
[0180] Therefore, in order to improve the accuracy of breath sound recognition, the breath sound recognition model can include two parts. One is the frame-level sound representation (or called single-segment sound representation); the other is the overall sound representation (also called continuous-segment sound representation). The frame-level sound representation represents a single Mel cepstrum segment obtained from S402. When the segment length is 0.96 seconds, the frame-level sound representation can represent the refined sound features of 0.96 seconds of audio. Since a complete breath is longer than 0.96 seconds, in order to represent the sound features of long audio as a whole, each 0.96-second audio segment in S402 can be re-spliced together after frame-level representation to fuse and obtain the context sound features in the long audio, that is, to obtain the overall sound representation.
[0181] Figure 4C It is a schematic diagram of the breath sound recognition model provided by the embodiments of the present application. Figure 4C In it, the breath sound recognition model can include a feature extraction layer, a feature splicing layer, and a classifier connected in sequence. Among them, the classifier includes a bidirectional long short-term memory network (BiLSTM) and a fully connected layer connected in sequence.
[0182] Combined with Figure 4C, the input of the breath sound recognition model is multiple consecutive Mel cepstrum segments, and the output is the confidence levels of 7 types of sounds. After multiple consecutive Mel cepstrum segments are input into the breath sound recognition model, for each Mel cepstrum segment, the breath sound recognition model can call a pre-trained feature extractor to extract the frame-level sound representation of this Mel cepstrum segment. In this way, each Mel cepstrum segment can obtain a frame-level sound representation. After that, the breath sound recognition model can sequentially splice the frame-level sound representations corresponding to each Mel cepstrum segment to obtain an overall sound representation (also known as a continuous segment sound representation). Then, the continuous segment sound representation can be input into a classifier for sound classification, so as to obtain the confidence levels determined for each type of sound, that is, the confidence levels corresponding to 7 types of sounds respectively.
[0183] The training process of the breath sound recognition model can include the following aspects.
[0184] (1) Obtain training samples Here, in order to increase the generalization of the breath sound recognition model and improve the accuracy of the breath sound recognition model in recognizing breath sounds, it is necessary to train the breath sound recognition model with rich training samples, and the training samples specifically need to cover the following 7 types of sounds: quiet, speaking, breath sound, cough, sneeze, snore, and environmental noise.
[0185] In the process of collecting training samples, for the same user, audio of various types of sounds of this user can be collected. After that, the collected audio can be cut into many audio segments of a fixed length, and the above fixed length can be 5 seconds. It can be understood that the length of the audio segment is related to the specific usage scenario of the breath sound recognition model. For example, when the breath sound recognition model is used to identify the sound category of 5-second breath audio data, the length of the audio segment can be 5 seconds.
[0186] In the process of processing the collected audio into training samples, for each audio segment, the operations of S401 - S402 can be performed on this audio segment to obtain multiple consecutive Mel cepstrum segments corresponding to this audio segment. After that, the audio segment, the multiple consecutive Mel cepstrum segments corresponding to the audio segment, and the sound category corresponding to the audio segment can be used as training samples. In this way, a sample set for training the breath sound recognition model can be obtained. For the convenience of description, the sample set for training the breath sound recognition model can be called an audio sample set.
[0187] (2) Train the feature extractor Here, the initial feature extractor can be various models that can be used to extract features. For example, it can be a Convolutional Neural Network (CNN), a Recurrent Neural Network (RNN), etc.
[0188] Further combined with Figure 4C , the process of training the feature extractor is equivalent to training a single-frame breath sound recognition model with only single-frame input. Among them, the input of the single-frame breath sound recognition model is a single mel cepstrum segment, and the output is the sound category of the corresponding audio. The input of the feature extractor is a single mel cepstrum segment, and the output is the frame-level sound representation corresponding to the input mel cepstrum segment.
[0189] Figure 4C In, the single-frame breath sound recognition model includes a sequentially connected feature extractor and a fully connected layer. During the process of training the single-frame breath sound recognition model, for the audio samples in the audio sample set, the single mel cepstrum segment in the audio sample can be used as the input, and the sound category in the audio sample can be used as the output to complete the training of the single-frame breath sound recognition model, thereby completing the training of the feature extractor.
[0190] After the feature extractor is trained, the parameters of the feature extractor can be solidified or encapsulated so that the feature extractor can be called or used as a feature extraction tool.
[0191] (3) Training the breath sound recognition model Here, the feature extractor in the initial breath sound recognition model is the feature extractor trained through the operation in the above step (2). That is to say, during the process of training the breath sound recognition model, the feature extractor is only called or used as a feature extraction tool.
[0192] During the process of training the breath sound recognition model with the audio sample set, for each audio sample, multiple consecutive mel cepstrum segments in the audio sample can be used as the input of the breath sound recognition model, and the sound category corresponding to the audio sample can be used as the output to train the initial breath sound recognition model to obtain the trained breath sound recognition model.
[0193] In practice, the trained breath sound recognition model can be called or used as a tool for recognizing breath sounds.
[0194] It can be understood that in the embodiments of the present application, the execution entity for training the feature extractor and the breath sound feature recognition model can be a mobile phone or other devices. For example, it can be a server. That is to say, the model can be trained on a mobile phone or on other devices and the trained model can be transferred to the mobile phone.
[0195] The embodiments of the present application do not make specific limitations on the device on which the model is trained.
[0196] The following further combines Figure 4D and Figure 4E to illustrate the breath sound feature extraction model in S405 and the training process of the breath sound feature extraction model.
[0197] Figure 4D is a schematic diagram of the breath sound feature extraction model provided by the embodiments of the present application. Figure 4D In, the breath sound feature extraction model may include a convolutional layer, two multi-scale residual layers (Multiscale Resnet Layer), and a bottleneck layer (Bottleneck Layer) connected in sequence. It can be understood that in the breath sound feature extraction model, the number of multi-scale residual layers may be more than two.
[0198] Figure 4E is a schematic structural diagram of a multi-scale residual layer in the breath sound feature extraction model provided by the embodiments of the present application. Figure 4E In, cross-fusion of convolutional features of multiple scales is adopted. For example, cross-fusion of 5×5 convolutional features and 3×3 convolutional features is adopted, which helps to extract richer and more comprehensive features.
[0199] Combined with Figure 4D and Figure 4E The input of the breath sound feature extraction model is the mel cepstrum coefficients, and the output is the breath sound features corresponding to the mel cepstrum coefficients. The breath sound feature extraction model compresses the input mel cepstrum coefficients through a convolutional layer; then, through multiple multi-scale residual layers, features expressed by the corresponding audio are extracted from the compressed data. Among them, the multi-scale residual module adds cross-fusion of convolutional features of different scales on the framework of the residual network, and can obtain richer and more comprehensive breath sound features; finally, the features of each module are fused through a bottleneck layer to output the breath sound features.
[0200] Further refer to Figure 4DAfter the breath sound feature extraction model outputs the breath sound features, the obtained breath sound features can further pass through a fully connected layer and a softmax function layer (or a softmax layer) to output the user ID confidence and the recognition result. For example, it can output the confidence of being recognized as user 1, the confidence of being recognized as user 2... the confidence of being recognized as user N, etc. At the same time, the user ID with the highest corresponding confidence can also be used as the recognition result.
[0201] It can be understood that Figure 4D In, adding a sequentially connected fully connected layer and a softmax function layer on the basis of the breath sound feature extraction model is actually a user recognition model. The input of the user recognition model is the mel cepstral coefficients of the audio, and the output is the user ID confidence and / or the recognition result.
[0202] The training process of the breath sound feature extraction model can include the following aspects.
[0203] (1) Obtain training samples In the process of obtaining training samples, breath audio generated by the same user breathing in different ways (such as normal breath sound, rapid breathing, and deep breathing, etc.) at different distances (such as the face is 30 cm, 50 cm, 100 cm, etc. away from the mobile phone) can be obtained. In this way, multiple breath audio can be obtained for each user.
[0204] After that, each obtained breath audio can be cut into multiple audio segments of a fixed length, and the above fixed length can be 5 seconds. It can be understood that the length of the audio segment is related to the specific usage scenario of the breath sound feature extraction model. For example, when the breath sound feature extraction model is used to extract the breath sound features of 5-second breath audio data, the length of the audio segment can be 5 seconds.
[0205] In the process of processing the obtained breath audio into training samples, for each audio segment, the operation of S401 (that is, the audio preprocessing operation) can be performed on the audio segment to obtain the mel cepstral coefficients corresponding to the audio segment. After that, the audio segment, the mel cepstral coefficients corresponding to the audio segment, and the user ID corresponding to the audio segment can be used as training samples. Thus, a sample set of the user recognition model including the breath sound feature extraction model can be obtained. For the convenience of description, the sample set for training the user recognition model can be called the breath sample set.
[0206] (2) Train the breath sound feature extraction model Continue to combine Figure 4E to train the breath sound feature extraction model, which is actually training the user recognition model.
[0207] In some application scenarios, the respiratory sound feature extraction model can be trained through the following training methods.
[0208] Training method 1: During the training of the user recognition model, for each respiratory sample in the respiratory sample set, the mel cepstral coefficients in the respiratory sample can be used as the input, and the user ID in the respiratory sample can be used as the output to complete the training of the user recognition model, thereby completing the training of the respiratory sound feature extraction model.
[0209] Since the differences in respiratory sounds between different individuals mainly come from physiological differences (such as the lungs, chest and abdominal muscles, and upper respiratory tract) and habitual differences (such as chest breathing or abdominal breathing, etc.), the manifestations of the above differences in respiratory sounds are as follows: First, the duration of inhalation, exhalation, and pause; Second, the change in air flow velocity during inhalation and exhalation; Third, the change in air flow volume during inhalation and exhalation.
[0210] Also, since the difficulty of respiratory sound authentication lies in that even for the same user, there are certain differences in respiratory sounds when the user's current state changes, the position of the mobile phone relative to the face changes, or the external environment changes.
[0211] Therefore, during the training of the respiratory sound feature extraction model, data augmentation and contrastive learning need to be performed on the respiratory audio to improve the generalization ability and robustness of the respiratory sound feature extraction model.
[0212] In the embodiments of the present application, the following three data augmentation methods can be used to augment the audio in the training samples.
[0213] Data augmentation method 1: Add random noise with different signal-to-noise ratio levels at random positions in the audio to simulate respiratory audio in various scenarios. Among them, the added noise can include one or more of the following: household noise, outdoor noise, office noise, transportation noise, and animal noise, etc.
[0214] Data augmentation method 2: Randomly change the speed of the audio to simulate rapid breathing and slow breathing.
[0215] Data augmentation method 3: Randomly change the volume range of the audio to simulate audio collected at close range and audio collected at long range.
[0216] In some other application scenarios, the respiratory sound feature extraction model can also be trained through the following training method 2.
[0217] In Training Method 2, during the iterative training of the breath sound feature extraction model, the breath sound feature extraction model can be trained in batches. For example, there can be K batches of training samples, that is, K times of training are performed. Among them, there can be M training samples in each batch of training samples, and these M training samples come from N users. Both M and N are integers greater than 1.
[0218] It should be noted that the M training samples in each batch come from N users, which means that in the case of multi-batch training, the audio of the same user may be selected multiple times. In this way, the accuracy and stability of the trained model can be further improved.
[0219] In practice, in order to further improve the generalization ability of the model, for the M training samples in each batch, one of them can be labeled as a positive sample and the rest can be labeled as negative samples. That is, in this Training Method 2, the training samples used for training can have two labeling tags, one for labeling the user ID and the other for labeling positive and negative samples.
[0220] As an example, for the M training samples in each batch, one can be randomly selected from these M training samples and labeled as a positive sample. After that, the remaining training samples among these M training samples are labeled as negative samples.
[0221] In addition, for each training sample (or called breath sample), data augmentation can be used to perform data augmentation processing on this training sample, so as to obtain an augmented sample (or called augmented breath sample) for this training sample. For example, the above 3 data augmentation methods can be used to perform data augmentation processing on this training sample, then 3 augmented samples can be obtained for each training sample. The augmented samples can be used to train the breath sound feature extraction model.
[0222] In the case of performing data augmentation processing on each training sample using the above 3 data augmentation methods, the M training samples in each batch can be expanded to 3M augmented samples.
[0223] It should be noted that using the augmented samples to train the breath sound feature extraction model can improve the generalization ability of the trained model, that is, it can enable the trained breath sound feature extraction model to accurately extract the breath sound features of users in various situations.
[0224] In the embodiments of the present application, training the breath sound feature extraction model using the above Training Method 2 can have the following beneficial effects: 1. When training the breath sound feature extraction model, marking one of the M training samples in the same batch as the positive sample and the rest as negative samples can make the trained breath sound feature extraction model have stronger recognition ability for breath sounds. That is to say, it can improve the generalization ability and robustness of the model, thus helping the trained breath sound feature extraction model to extract breath sound features more accurately and effectively.
[0225] 2. Data augmentation is performed on each training sample, and the augmented samples obtained by data augmentation are used to train the breath sound feature extraction model, which can improve the generalization ability and robustness of the trained breath sound feature extraction model. That is to say, it can make the trained breath sound feature extraction model accurately extract the breath sound features of users in various situations.
[0226] In some optional implementation manners of the embodiments of the present application, during the process of training the breath sound feature extraction model (which belongs to the user recognition model) using the above training method 2, the loss function used may include the losses of the following two parts: one is the user recognition cross-entropy loss, and the other is the similarity loss of contrast learning. The user recognition cross-entropy loss is used to describe the accuracy of the breath sound feature extraction model in user recognition, and the contrast learning similarity loss is used to describe the similarity between the features extracted by the breath sound feature extraction model and the true features.
[0227] Exemplarily, the user recognition cross-entropy loss can be shown as formula (1): (1) Where, is the user recognition cross-entropy loss, is the true label, is the user recognition prediction probability output by the softmax layer. The meaning of this loss function is to evaluate the gap between the true label and the prediction result. The smaller it is, the more accurate the prediction result.
[0228] Exemplarily, the contrast learning similarity loss can be shown as formula group (2): (2) Where, is the contrast learning similarity loss, the S matrix is the cosine similarity matrix, and each element in the S matrix is used to indicate the cosine similarity between the breath sound features corresponding to two training samples. For example, is used to indicate the cosine similarity between the breath sound feature corresponding to the Mth training sample and the breath sound feature corresponding to the 3Mth training sample. The values of the elements in the S matrix can be any value from 0 to 1.
[0229] The L matrix is the target similarity matrix. Each element in the L matrix is used to indicate whether two training samples come from the same user. If two training samples come from the same user, they are considered similar features, and the similarity value is 1; if two training samples come from different users, they are considered dissimilar features, and the similarity value is 0. That is, the values of the elements in the L matrix are either 0 or 1.
[0230] The above contrastive learning similarity loss can be calculated based on the cosine similarity matrix S and the target similarity matrix L. When the contrastive learning similarity loss is smaller, the similar training samples are closer in the feature space, while the dissimilar training samples are farther apart in the feature space.
[0231] In practice, the user recognition cross-entropy loss and the contrastive learning similarity loss can each correspond to a weight.
[0232] In practice, if the weight corresponding to the user recognition cross-entropy loss can be , and the weight corresponding to the contrastive learning similarity loss can be . Exemplarily, the overall model training loss can be as shown in formula (3) below. Among them, and are preset weight values. In practice, and range from 0 to 1, and and sum to 1.
[0233] (3) In the embodiments of the present application, during the iterative training of the breath sound feature extraction model, the model parameters can be updated by the gradient descent method, so that the overall model training loss becomes smaller and smaller, realizing the parameter optimization of the model.
[0234] In the embodiments of the present application, during the model training process, the loss of the model is described from two dimensions of user recognition and contrastive learning, so as to optimize and adjust the model, which can ensure that the trained user recognition model can accurately identify users and has strong recognition of breath sounds. That is, it can make the breath sound features extracted by the trained breath sound feature extraction model have strong recognition and can more comprehensively and accurately describe the individual breath characteristics of users.
[0235] It should be noted that the higher the recognition rate of the breath sound feature extraction model for breath sounds, the higher the accuracy or recognition rate of the extracted breath sound features. When performing the owner recognition based on the breath sound features, the recognition result will be more accurate.
[0236] 3. Cross-modal feature template entry.
[0237] In the embodiment of the present application, when the face template of the owner is successfully entered, the mobile phone can collect breath audio data through the microphone, collect touch screen data through the touch screen sensor, and collect IMU data through the IMU, and obtain the cross-modal feature template of the owner through the above three types of data.
[0238] In the embodiment of the present application, when the collection frequency of the breath audio data is 16 kHz and the collection duration is 5 seconds, the size of the breath audio data collected by the mobile phone is 80,000 dimensions, where 80,000 = 16k × 5.
[0239] The size of the IMU data collected by the IMU can be 50 × 90, where 50 indicates that 50 times of IMU data are collected, and 90 indicates the dimension of the IMU data. The IMU data can include the acceleration, pitch angle, roll angle, etc. of the mobile phone in each direction. The collection frequency of the IMU data can be 50 times per second, that is, the mobile phone can obtain the above 50 × 90-dimensional IMU data by collecting the IMU data within 1 second after the face template of the owner is successfully entered.
[0240] The size of the touch screen data collected by the touch screen sensor can be 200 × 17, where 200 indicates the latest 200 touch screen data, and 17 indicates the dimension of the touch screen data. Each touch screen data can include the number of sliding times and the number of click times of the touch screen. That is, after the face template of the owner is successfully entered, the mobile phone can collect the latest 200 touch screen data of the mobile phone to obtain the above 200 × 17-dimensional touch screen data.
[0241] It can be understood that in the embodiment of the present application, the sizes of the above three types of modal data can be adaptively adjusted according to the actual application scenario, and the embodiment of the present application does not make specific limitations on the sizes of various modal data.
[0242] Next, it will be further described how to fuse the breath audio data of the owner, the touch screen data of the mobile phone, and the IMU data of the mobile phone to obtain the cross-modal feature template of the owner.
[0243] In the embodiments of the present application, due to the large amount of breathing audio data, in order to improve the data processing efficiency, the mobile phone can call a pre-stored audio processing module, and the audio processing module processes the breathing audio data to extract the breathing sound features in the above breathing audio data. Then, the mobile phone can call a pre-stored cross-modal encoder, and the cross-modal encoder encodes the breathing sound features, touch screen data, and IMU data to obtain corresponding cross-modal features.
[0244] It should be noted that during the process of initially entering the cross-modal feature template of a certain phone owner, if there is no breathing sound in the collected breathing audio data or it is not recognized as a breathing sound, the breathing sound features may not be obtained. In such a case, the mobile phone can temporarily not enter the breathing sound feature template and cross-modal feature template of this phone owner.
[0245] In some application scenarios, when it is detected that there is no breathing sound in the above breathing audio data, the mobile phone can continue to collect breathing audio data and determine whether the breathing audio data is a breathing sound when the subsequent face unlock is successful. If it is a breathing sound, the breathing sound features are further extracted, and the extracted breathing sound features are entered as the breathing sound feature template of the phone owner. At the same time, the cross-modal features of the phone owner can be determined by combining the breathing sound feature template of this phone owner, the IMU data and touch screen data after the face unlock is successful, and the cross-modal features are entered as the cross-modal feature template of the phone owner.
[0246] Combined with the above description, in the embodiments of the present application, the triggering conditions (or multi-modal feature entry conditions) for triggering the entry of the cross-modal feature template include any one of the following: the face template of the phone owner is successfully entered, the phone owner's face unlock is successful and the cross-modal feature template of the phone owner currently performing the unlock has not been entered in the terminal device.
[0247] The following will be combined with Figure 5A 、 Figure 5B and Figure 5C to illustrate the structure of the cross-modal encoder. Figure 5A FIG. is a schematic structural diagram of a cross-modal encoder provided by an embodiment of the present application. Figure 5B FIG. is a schematic structural diagram of a cross-channel linear encoder in the cross-modal encoder provided by an embodiment of the present application. Figure 5C FIG. is a schematic structural diagram of a cross-channel attention encoder in the cross-modal encoder provided by an embodiment of the present application.
[0248] Such as Figure 5AAs shown, the input of the cross-modal encoder is IMU data, touch screen data, and breath sound features, and the output is cross-modal features. The cross-modal encoder can encode the IMU data and touch screen data through two cross-channel linear encoders from top to bottom. For ease of description, the cross-channel linear encoder for encoding the IMU data can be denoted as the first cross-channel linear encoder, and the cross-channel linear encoder for encoding the touch screen data can be denoted as the second cross-channel linear encoder.
[0249] After that, the cross-modal encoder can perform feature extraction on the encoded IMU data, encoded touch screen data, and breath sound features through three Stem layers from top to bottom. For ease of description, the Stem layer for performing feature extraction on the encoded IMU data can be denoted as the first Stem layer, the Stem layer for performing feature extraction on the encoded touch screen data can be denoted as the second Stem layer, and the Stem layer for performing feature extraction on the breath sound features can be denoted as the third Stem layer.
[0250] After that, the cross-modal encoder can fuse the IMU features extracted by the first Stem layer, the touch features extracted by the second Stem layer, and the breath sound features extracted by the third Stem layer to obtain fused features.
[0251] After that, the cross-modal encoder can process the fused features through the fourth Stem layer, the third cross-channel linear encoder, and the multi-target operation module in sequence to output cross-modal features. As Figure 5A shown, the target operation module can include a downsampling module, a cross-channel linear encoder, and a cross-channel attention encoder connected in sequence.
[0252] It can be understood that the above first, second, third, and fourth are only used to distinguish one element from another in description.
[0253] Combined Figure 5A , two important components of the cross-modal encoder are: the cross-channel linear encoder and the cross-channel attention encoder.
[0254] Figure 5B In, the cross-channel linear encoder can include a Linear layer, an activation function, a Linear layer, a norm layer, a conv layer, and an activation function connected in sequence. The cross-channel linear encoder uses the Linear layer for encoding and aligns the feature dimensions of different modalities, and uses the conv layer to implement one-dimensional pointwise convolution for information fusion between channels. Figure 5B In, " " is the symbol indicating addition, and " " is the symbol indicating the activation function.
[0255] In the embodiments of the present application, the role of the cross-channel linear encoder is mainly to increase the level of feature abstraction. The input of the cross-channel linear encoder is feature data, and the output is further abstracted feature data.
[0256] Figure 5C In [description], in combination with the structural part in the dashed box, the cross-channel attention encoder slices the data input to the input layer (also called the Inputs layer) through the Group layer. Specifically, each type of modality data corresponds to one channel, realizing the slicing of features from the channel dimension; then it is encoded through the Linear layer; after that, the sliced features are merged; finally, the attention mechanism is used in the channel dimension to promote the fusion of multimodal information among channels. Figure 5C In [description], " " is the symbol indicating addition, " " is the symbol indicating the activation function, " " is the symbol indicating multiplication, " " is the symbol indicating positional encoding (PositionalEncoding).
[0257] Further in combination with Figure 5C , Figure 5C Three Group layers from top to bottom are shown in the dashed box, as well as the Linear layers respectively connected to each Group layer. For the convenience of description, the topmost Group layer can be denoted as the first Group layer, the topmost Linear layer can be denoted as the first Linear layer, the middle Group layer can be denoted as the second Group layer, the middle Linear layer can be denoted as the second Linear layer, the bottommost Group layer can be denoted as the third Group layer, and the bottommost Linear layer can be denoted as the third Linear layer.
[0258] Among them, the input end of the first Linear layer is connected to the output end of the first Group layer, the input end of the second Linear layer is connected to the output end of the first Linear layer and the output end of the second Group layer, and the input end of the third Linear layer is connected to the output end of the second Linear layer and the output end of the third Group layer.
[0259] From Figure 5C it can be found that the output data from the topmost first Group layer is encoded the most times, and the output data from the bottommost third Group layer is encoded the fewest times.
[0260] In the embodiments of the present application, features with good quality can be deployed in the lower Group layer, and features with poor quality can be deployed in the upper Group layer. That is, the features deployed in the topmost Group layer are relatively the worst, and the features deployed in the bottommost Group layer are relatively the best. In this way, the features deployed in the bottommost Group layer require the fewest number of Linear layers, that is, the fewest number of encoding times. The features deployed in the topmost Group layer require the most number of Linear layers, that is, the most number of encoding times. In this way, the cross-channel attention encoder can combine the quality of each feature and perform different degrees of encoding on each feature respectively, so as to better effectively fuse features with good quality and features with poor quality.
[0261] In the embodiments of the present application, after encoding and fusing features with good quality and features with poor quality to different degrees, the fused features can be transposed, and then the transposed fused data can be input into the Features layer, which helps to further effectively fuse the features.
[0262] In the embodiments of the present application, the main function of the cross-channel attention encoder is to effectively fuse features with good quality and features with poor quality, strengthen the features, so as to improve the accuracy of the output features.
[0263] In addition, combined with Figure 5A 、 Figure 5B and Figure 5C , the cross-modal encoder adopts a lightweight design as a whole, which can achieve lightweight deployment.
[0264] (2) Using the owner template for owner identification In the embodiments of the present application, the triggering conditions for the mobile phone to enter the owner identification can include the following Conditions 1 - 6. It can be understood that only some triggering conditions are listed in the embodiments of the present application, and there can be other triggering conditions with the change of the actual application scenario.
[0265] Condition 1, the mobile phone has enabled the continuous authentication function.
[0266] In the embodiments of the present application, when the continuous authentication function is enabled, the mobile phone can automatically lock the target content during use. For example, when the continuous authentication function of the mobile phone is enabled, the mobile phone can lock the service card with the owner's private information.
[0267] Condition 2, the mobile phone detects an operation to open the target application.
[0268] In an embodiment of the present application, when an operation to open a target application is detected, it indicates that the user needs to open the target application at this time. At this time, the mobile phone can be triggered to enter the owner identification. If it is identified as the owner, the target application is opened. If it is identified as a non-owner, the target application is not opened. In this way, the security of the owner's private information can be ensured.
[0269] Among them, the target application usually refers to an application that may involve the owner's privacy. For example, WeChat, Alipay, photo album, etc. It can be understood that the target application can be set by the mobile phone itself. For example, the mobile phone can set an instant messaging application as the target application, or set an application with a usage frequency higher than a certain frequency threshold as the target application. Of course, the target application can also be an application set by the user according to his own needs. For example, the user can set the photo album application with a relatively high privacy level as the target application. It can be understood that the embodiments of the present application do not make specific limitations on the target application.
[0270] Condition 3: The mobile phone receives a notification message.
[0271] In an embodiment of the present application, when a notification message of a certain application is received, the mobile phone can be triggered to enter the owner identification. If it is identified as the owner, the notification message is previewed. If it is identified as a non-owner, the notification message is hidden and not previewed. In this way, the security of the owner's private information can be ensured.
[0272] Condition 4: The mobile phone needs to display a service card, and the service card contains private information, such as bank card numbers, passwords, etc.
[0273] In an embodiment of the present application, before the mobile phone displays a service card carrying private information, it can perform owner identification. If it is identified as the owner, the service card is directly displayed. If it is identified as a non-owner, the private information in the service card is hidden. In this way, the security of the owner's private information can be ensured.
[0274] Condition 5: The mobile phone detects that a target function is triggered.
[0275] Among them, the target function can be the passwordless payment function. Of course, the target function can also be other functions related to the security of the owner's information or funds.
[0276] In an embodiment of the present application, when the target function needs to be executed, the mobile phone can be triggered to enter the owner identification. If it is identified as the owner, passwordless payment is performed. If it is identified as a non-owner, passwordless payment cannot be performed. For example, a password must be entered to complete the payment.
[0277] Condition 6: The mobile phone detects that a target interface is triggered.
[0278] Among them, the target interface can be an interface with the owner's privacy information. For example, it can be the content display interface of a certain short message.
[0279] In the embodiments of the present application, when it is necessary to open the target interface, the mobile phone can be triggered to enter the owner identification. If the owner is recognized, the target interface is opened. If the owner is not recognized, the target interface is not opened, which can ensure the security of the owner's privacy information.
[0280] When the triggering condition for owner identification is currently met, the mobile phone enters the owner identification. How the mobile phone performs owner identification will be further described below.
[0281] Figure 6A The schematic flow diagram of the mobile phone's owner identification based on multi-modal fusion is shown.
[0282] S601, the mobile phone determines whether a human face is detected. If no human face is detected, S602 is executed. If a human face is detected, S605 is executed.
[0283] Here, the mobile phone can analyze whether the detected image is a human face image. For example, by analyzing the pixel values of the image, it can be determined whether it is a human face image. If it is a human face image, the mobile phone can combine the pre-recorded owner's face template to identify the owner's face and obtain a face recognition score. The face recognition score is used to describe the similarity between the face to be measured and the face template.
[0284] S602, when no human face is detected, the mobile phone determines whether a breathing sound is detected. If a breathing sound is detected, step S603 is executed. If no breathing sound is detected, it indicates that the detection result is unreliable. At this time, step S604 is executed.
[0285] In the embodiments of the present application, when no human face is detected, the mobile phone can continue to collect the to-be-measured breathing audio data of the to-be-measured user and extract the to-be-measured breathing sound features in the to-be-measured breathing audio data, so as to determine whether the to-be-measured user is the owner by comparing the similarity between the to-be-measured breathing sound features and the breathing sound feature template in the owner template.
[0286] Among them, the mobile phone can call the pre-stored audio processing module (see Figure 4A ), and the audio processing module processes the to-be-measured breathing audio data to extract the to-be-measured breathing sound features in the to-be-measured breathing audio data. Among them, the above-mentioned to-be-measured breathing audio data is the breathing audio data of the to-be-measured user. The to-be-measured breathing sound features are the breathing sound features in the to-be-measured breathing audio data.
[0287] In the embodiments of the present application, the acquisition parameters of the to-be-detected respiratory audio data may be: single-channel acquisition, sampling rate of 16 kHz, and acquisition duration of t seconds. Among them, the value range of t is from 0.5 second to 5 seconds.
[0288] In practice, the acquisition duration of the to-be-detected respiratory audio data is not fixed. The acquisition duration of the to-be-detected respiratory audio data is usually related to the usage scenario, and the acquisition duration is usually positively correlated with the recognition accuracy of the owner recognition in the usage scenario, and inversely correlated with the recognition speed of the owner recognition in the usage scenario. For example, for a usage scenario with very high requirements for the recognition accuracy of the owner, the acquisition duration can be as long as possible; for a usage scenario with very high requirements for the recognition speed of the owner, the acquisition duration can be as short as possible.
[0289] Further taking examples, for the usage scenario of a user logging in to an account, this usage scenario has very high requirements for the recognition accuracy of the owner. At this time, the to-be-detected respiratory audio data collected can be 5 seconds. For the usage scenario of viewing photos, in this usage scenario, if the user waits for a long time to open a photo, such as waiting for 5 seconds, obviously the user experience will be relatively poor, that is, this usage scenario has high requirements for the recognition speed of the owner. At this time, in order to improve the recognition speed of the owner, the to-be-detected respiratory audio data collected can be 0.5 second.
[0290] Since the respiratory sound recognition model in the audio processing module may recognize that the to-be-detected respiratory audio data collected does not contain respiratory sounds, in the case where the respiratory sound recognition model recognizes that the to-be-detected respiratory audio data does not contain respiratory sounds, it is considered that no respiratory sounds are detected. On the contrary, in the case where the respiratory sound recognition model recognizes that the to-be-detected respiratory audio data contains respiratory sounds, it is considered that respiratory sounds are detected.
[0291] S603, recognize the owner based on the to-be-detected respiratory sound features, and output the owner recognition result.
[0292] In the embodiments of the present application, for the sake of differential description, the respiratory sound features of the to-be-detected user can be denoted as the to-be-detected respiratory sound features.
[0293] Here, in the case where respiratory sounds are detected in S602, it means that the respiratory sound recognition model in the audio processing module recognizes that the to-be-detected respiratory audio data contains respiratory sounds. Then the audio processing module can finally output the to-be-detected respiratory sound features.
[0294] After that, the mobile phone can use the to-be-detected respiratory sound features corresponding to the to-be-detected user to recognize whether the to-be-detected user is the owner.
[0295] Figure 6B Shows a schematic process diagram of recognizing the owner based on respiratory sounds. Combining Figure 6B, the mobile phone can compare the breathing sound feature to be measured corresponding to the user to be measured with the breathing sound feature template of the owner stored in advance to determine whether the user to be measured is the owner.
[0296] In some optional implementation manners of the embodiments of the present application, the mobile phone can determine whether the user to be measured is the owner in the following manner: when the owner template includes a breathing sound feature template, determine the similarity between the breathing sound feature to be measured and the breathing sound feature template. When the obtained similarity is greater than a preset second determination threshold, determine that the user to be measured is the owner. Conversely, when the obtained similarity is less than or equal to the second determination threshold, determine that the user to be measured is not the owner.
[0297] Wherein, the above-mentioned second determination threshold is usually a preset similarity value. For example, it can be 0.6, and the value range of the second determination threshold is usually 0-1.
[0298] In practice, calculating the similarity between the breathing sound feature to be measured and the breathing sound feature template is usually to calculate the cosine similarity between the vector corresponding to the breathing sound feature to be measured and the vector corresponding to the breathing sound feature template.
[0299] In some optional implementation manners of the embodiments of the present application, the mobile phone can also determine whether the user to be measured is the owner in the following manner: when the owner template includes multiple breathing sound feature templates, respectively determine the similarity between each breathing sound feature template and the breathing sound feature to be measured. When the average value of the obtained multiple similarities is greater than a preset fourth determination threshold, determine that the user to be measured is the owner. Conversely, when the average value of the obtained multiple similarities is less than or equal to the fourth determination threshold, determine that the user to be measured is not the owner.
[0300] Wherein, the above-mentioned fourth determination threshold is usually a preset similarity value. For example, it can be 0.3, and the value range of the fourth determination threshold is 0-1.
[0301] Further, for example, if there are 3 breathing sound feature templates in the owner template, and they are breathing sound feature template 1, breathing sound feature template 2, and breathing sound feature template 3 respectively. If the similarity between the breathing sound feature to be measured and breathing sound feature template 1 is S1, the similarity between the breathing sound feature to be measured and breathing sound feature template 2 is S2, and the similarity between the breathing sound feature to be measured and breathing sound feature template 3 is S3, then when the average value of S1, S2, and S3 is greater than the fourth determination threshold, it is considered that the user to be measured is the owner; otherwise, it is considered that the user to be measured is not the owner.
[0302] It should be noted that when there are multiple breathing sound feature templates in the owner template, by using the average value of the similarities between the breathing sound feature to be measured and each breathing sound feature template to determine whether the user to be measured is the owner, the accuracy of owner identification can be improved.
[0303] Optionally, when multiple breath sound feature templates are included in the owner template, the mobile phone can also determine the similarity between each breath sound feature template and the breath sound feature to be measured respectively. When the maximum value among the obtained multiple similarities is greater than the second determination threshold, it is determined that the user to be measured is the owner. On the contrary, when the maximum value among the obtained multiple similarities is less than or equal to the second determination threshold, it is determined that the user to be measured is not the owner.
[0304] In some optional implementation manners of the embodiments of the present application, when multiple owner templates are recorded in the mobile phone, the mobile phone can also determine whether the user to be measured is the owner in the following manner: for each owner template, when a breath sound feature template is included in the owner template, determine the similarity between the breath sound feature to be measured and the breath sound feature template. When the obtained similarity is greater than the second determination threshold, it is determined that the user to be measured is the owner, and specifically, the owner indicated by the owner template; when multiple breath sound feature templates are included in the owner template, determine the similarity between each breath sound feature template in the owner template and the breath sound feature to be measured respectively. When the average value of the obtained multiple similarities is greater than the fourth determination threshold, it is determined that the user to be measured is the owner, and specifically, the owner indicated by the owner template; Optionally, when the user to be measured is not the owner indicated by any owner template, the mobile phone can determine that the user to be measured is not the owner.
[0305] In the embodiments of the present application, when the user to be measured is recognized as the owner, the mobile phone can continue to perform operations related to the owner's permissions, such as previewing notification messages, making password-free payments, and so on. When the user to be measured is recognized as not the owner, the mobile phone hides the private information of the owner, such as hiding the content of communication messages, hiding the content of service cards, and so on.
[0306] S604, prompt that the detection result is unknown.
[0307] In the embodiments of the present application, during the process of owner identification by the mobile phone, if neither a face nor a breath sound is detected, it can be considered that the detection result is unreliable. At this time, the mobile phone can output a prompt message such as "the detection result is unknown", "try again", or "do it again".
[0308] After that, the user to be measured can further cooperate with the mobile phone for owner identification in combination with the prompt of the mobile phone.
[0309] S605, the mobile phone inputs the face recognition score into the threshold discriminator to determine whether the user to be measured is the owner.
[0310] Here, if the face recognition score is greater than a pre-set upper threshold for face recognition, such as 0.65, it is considered that the user to be tested is the owner of the device. If the face recognition score is less than a pre-set lower threshold for face recognition, such as 0.05, it is considered that the user to be tested is not the owner of the device. If the face recognition score is between the above-mentioned lower threshold and upper threshold for face recognition, step S606 can be continued.
[0311] Among them, the above-mentioned upper threshold for face recognition can be the FAR threshold of face recognition. The above-mentioned lower threshold for face recognition can be the False Rejection Rate (FRR) threshold of face recognition. In practice, when the face recognition score is between the above-mentioned lower threshold and upper threshold for face recognition, it usually indicates that the face detection result is not good.
[0312] S606. When the face detection result is not good, the mobile phone determines whether breathing sound is detected. If no breathing sound is detected, S607 is executed. On the contrary, if breathing sound is detected, S608 is executed.
[0313] In the embodiment of the present application, when the face detection result is not good, the mobile phone can continue to collect the breathing audio data to be tested of the user to be tested, and call the pre-stored audio processing module (see Figure 4A ), and the audio processing module processes the breathing audio data to be tested to extract the breathing sound features in the breathing audio data to be tested.
[0314] In practice, if the breathing sound recognition model recognizes that the breathing audio data to be tested does not contain breathing sound, it is considered that no breathing sound is detected. On the contrary, if the breathing sound recognition model recognizes that the breathing audio data to be tested is breathing sound, it is considered that breathing sound is detected.
[0315] S607. When no breathing sound is detected, the mobile phone directly outputs the face detection result.
[0316] Here, when the mobile phone combines the pre-recorded face template of the device owner to detect and recognize the face, it can initially judge whether it is the device owner in combination with the face recognition score. When no breathing sound is detected, the mobile phone can output this initial judgment. For example, if the face recognition score is greater than a pre-set face recognition score threshold, it is determined to be the device owner. On the contrary, if the face recognition score is less than or equal to the face recognition score threshold, it is determined not to be the device owner. Among them, the above-mentioned face recognition score threshold can be a pre-set threshold, for example, it can be 0.5.
[0317] Optionally, in some application scenarios where the accuracy of the owner identification is very high, when the face detection result is poor and no breathing sound is detected, the mobile phone can also output information to prompt re-detection. For example, it can output "Please try again" and so on.
[0318] S608. When the breathing sound is detected, the mobile phone combines the multi-modal data of the user to be measured and calls the multi-modal fusion owner identification model to perform owner identification.
[0319] Among them, the multi-modal data may include the face recognition score, the breathing sound feature to be measured, the touch screen data to be measured, and the IMU data to be measured.
[0320] In the embodiments of the present application, only when the face recognition effect is poor, for example, when there are strong light, backlight, low light and other situations resulting in poor face recognition, will the breathing audio data to be measured, the touch screen data, and the IMU data of the user to be measured be further collected, and then combined with the multi-modal data to jointly perform owner identification to improve the accuracy of owner identification. Since the accuracy of face recognition is already high most of the time, that is, the situation of using multi-modal data for owner identification is relatively rare, it can save the computing power of the mobile phone and reduce the power consumption of the mobile phone while ensuring the accuracy of owner identification, thereby improving the practicality of the mobile phone and helping to further enhance the user experience.
[0321] For the purpose of distinguishing and describing, the touch screen data collected during the owner identification process can be recorded as the touch screen data to be measured, and the IMU data collected can be recorded as the IMU data to be measured.
[0322] The operation of collecting the touch screen data to be measured during the owner identification process can be basically the same as the operation of collecting the touch screen data during the aforementioned owner template entry process, and will not be elaborated here. In addition, the operation of collecting the IMU data to be measured during the owner identification process can also be basically the same as the operation of collecting the IMU data during the aforementioned owner template entry process.
[0323] It can be understood that in some optional implementation manners, in S608, when the breathing sound is detected, the mobile phone can also directly use the breathing sound feature to be measured of the user to be measured to perform owner identification. At this time, S608 can also be replaced with the following steps: when the breathing sound is detected, the mobile phone identifies the owner based on the breathing sound feature to be measured and outputs the owner identification result. This operation is basically the same as the operation of the aforementioned S603 and will not be elaborated here.
[0324] Figure 6C The process schematic diagram of the multi-modal fusion owner identification model for performing owner identification is shown.
[0325] S6081. Calculate the cross-modal feature to be measured.
[0326] Among them, the cross-modal feature to be measured is the cross-modal feature corresponding to the user to be measured.
[0327] Here, the mobile phone can use a cross-modal encoder (see Figure 5A ) to encode the touch screen data to be measured, the IMU data to be measured, and the breath sound feature to be measured, so as to obtain the cross-modal feature to be measured.
[0328] S6082, calculate the cross-modal similarity between the cross-modal feature to be measured and the cross-modal feature template.
[0329] Among them, the cross-modal similarity is the similarity between the cross-modal feature to be measured and the cross-modal feature template. In the embodiments of the present application, the cross-modal similarity can also be referred to as the cross-modal recognition score.
[0330] Here, if the cross-modal feature template , and the cross-modal feature to be measured , where represents the cross-modal encoder, is the IMU data of the owner of the machine, is the touch screen data of the owner of the machine, is the breath sound feature template of the owner of the machine, is the IMU data to be measured of the user to be measured, is the touch screen data to be measured of the user to be measured, is the breath sound feature to be measured of the user to be measured.
[0331] Exemplarily, the cross-modal similarity can be calculated by the following formula (4): (4) Where is the cross-modal similarity, represents the cosine similarity comparator, is the L2 norm, which is used to calculate the magnitude of the vector modulus.
[0332] S6083, calculate the breath sound similarity between the breath sound feature to be measured and the breath sound feature template of the owner of the machine.
[0333] Among them, the breath sound similarity is the similarity between the breath sound feature to be measured and the breath sound feature template of the owner of the machine. In the embodiments of the present application, the breath sound similarity can also be referred to as the breath recognition score.
[0334] Exemplarily, the breath sound similarity can be calculated by the following formula (5): (5) Where is the breath sound similarity.
[0335] S6084. Input the cross-modal similarity, breath sound similarity, and face recognition score into a fusion decision model to obtain a multi-modal recognition score.
[0336] Among them, the input of the fusion decision model is the cross-modal similarity (also known as the cross-modal recognition score), the breath sound similarity (also known as the breath recognition score), and the face recognition score, and the output is the multi-modal recognition score. The multi-modal recognition score is used to describe the comprehensive recognition result of the user to be measured in multiple modalities.
[0337] In the embodiments of the present application, the multi-modal recognition score can be expressed as follows: , where is the multi-modal recognition score, is the cross-modal similarity, is the breath sound similarity, is the face recognition score, represents the fusion decision model.
[0338] S6085. Combine the multi-modal recognition score to determine whether the user to be measured is the owner of the device.
[0339] In the embodiments of the present application, if the multi-modal recognition score is greater than or equal to a preset fusion threshold, it is recognized as the owner of the device; if the multi-modal recognition score is less than the fusion threshold, it is recognized as not the owner of the device.
[0340] The above fusion threshold is usually a preset threshold. For example, it can be 0.6.
[0341] Combine Figure 6C It can be found that the multi-modal fusion device owner recognition model uses a cross-modal encoder to perform feature-level fusion on IMU data, touch screen data, and breath sound features, and then uses a decision fusion model to perform decision-level fusion on the cross-modal recognition score, breath recognition score, and face recognition score to obtain a multi-modal recognition score. Among them, stronger modalities, such as the face modality, do not participate in feature-level fusion, which can reduce the coupling between data and effectively utilize the complementary advantages between multi-modalities, improving the comprehensive recognition success rate of the model.
[0342] Considering that there may be one device owner template or multiple device owner templates stored in the mobile phone. That is to say, there can be one device owner template or multiple device owner templates in the mobile phone.
[0343] In some optional implementation manners of the embodiments of the present application, when multiple device owner templates are pre-recorded in the mobile phone, in combination with the above Figure 6C , during the execution of S6081 - S6084 by the mobile phone, one device owner template can be selected first, and the multi-modal recognition score can be calculated using the selected device owner template.
[0344] In some application scenarios, the mobile phone can select the owner template with the highest corresponding face recognition score. Further, for example, if there are three owner templates, namely owner template 1, owner template 2, and owner template 3, when the mobile phone executes S601 and identifies the user to be tested as the owner corresponding to owner template 1, it can directly select owner template 1.
[0345] In this application scenario, during the process of the mobile phone executing S6081 - S6084, a multimodal recognition score for owner template 1 can be calculated. After that, the mobile phone can further determine whether it is the owner based on this multimodal recognition score. If it is the owner, it means that the user to be tested is the owner indicated by owner template 1. It is not difficult to find that in this application scenario, only one multimodal recognition score needs to be calculated for one owner template, and the calculation amount is relatively small.
[0346] In some other alternative implementation manners of the embodiments of the present application, when multiple owner templates are pre - recorded in the mobile phone, the mobile phone can calculate a multimodal recognition score for each owner template. After that, the mobile phone can combine the obtained multiple multimodal recognition scores to determine whether the user to be tested is the owner and which specific owner it is.
[0347] Further, for example, if there are three owner templates, namely owner template 1, owner template 2, and owner template 3, the mobile phone can perform the operations of S6081 - S6084 using owner template 1 to obtain a multimodal recognition score, such as 99 points. Perform the operations of S6081 - S6084 using owner template 2 to obtain another multimodal recognition score, such as 78 points, and perform the operations of S6081 - S6084 using owner template 3 to obtain yet another multimodal recognition score, such as 60 points. At this time, it can be determined that the user to be tested is the owner, and specifically, it is the owner indicated by owner template 1.
[0348] In this implementation manner, since the above - mentioned step S608 using the multimodal fusion model for further owner recognition is only performed when the face detection effect is not good, and when the face detection effect in step S601 is not good, selecting the owner template based on the face detection result may not be accurate enough. Therefore, in this implementation manner, when the face detection accuracy in S601 is not high, calculating the multimodal recognition score for each owner template separately can ensure the accuracy rate of owner recognition.
[0349] Considering that there may be multiple feature templates in one owner template. For example, there are multiple face templates; for another example, there are multiple breath sound feature templates; for yet another example, there are multiple cross - modal feature templates.
[0350] In some alternative implementation manners of the embodiments of the present application, when there are multiple feature templates in the owner template, during the execution of S6081 - S6084 by the mobile phone, the average value of the multiple feature templates in the owner template can be calculated for similarity with the to - be - measured feature of the same modality (for example, the to - be - measured breath sound feature), so as to obtain the recognition score corresponding to the modality.
[0351] For further illustration, if there are 3 breath sound feature templates in the owner template, the mobile phone can calculate the average value of the 3 breath sound feature templates for similarity with the to - be - measured breath sound feature, so as to obtain the recognition score corresponding to the breath modality, or called the breath recognition score.
[0352] Optionally, the mobile phone can also calculate the similarity between each feature template in the owner template and the to - be - measured feature of the same modality, so as to obtain multiple similarities; then, the recognition score corresponding to the modality is calculated through the obtained multiple similarities. For example, the maximum value among the multiple similarities can be used as the recognition score corresponding to the modality.
[0353] For further illustration, if there are 3 breath sound feature templates in the owner template, the mobile phone can calculate the similarity between each breath sound feature template and the to - be - measured breath sound feature, so as to obtain 3 similarities. Then, the maximum value among the 3 similarities can be used as the recognition score corresponding to the breath modality, or called the breath recognition score.
[0354] In some alternative implementation manners of the embodiments of the present application, when multiple owner templates are pre - recorded in the mobile phone, combined with the above Figure 6C , during the execution of S6085 by the mobile phone, since each owner template may correspond to a multi - modality recognition score, at this time, the mobile phone can determine whether the to - be - measured user is the owner in the following way: for the multi - modality recognition scores respectively corresponding to each owner template, if the multi - modality recognition score is greater than the fusion threshold, the mobile phone can determine that the to - be - measured user is the owner. At this time, the to - be - measured user can be identified as: the owner indicated by the owner template corresponding to the multi - modality recognition score.
[0355] In some application scenarios, if there are multiple multi - modality recognition scores that are greater than the fusion threshold at the same time, the mobile phone can determine that the to - be - measured user is the owner. At this time, the to - be - measured user can be identified as: the owner indicated by the owner template corresponding to the maximum value among the multiple multi - modality recognition scores.
[0356] The following combines Figure 6D to further illustrate the fusion decision model in S6084. Figure 6D It is a schematic structural diagram of the fusion decision model provided by the embodiments of the present application. Combining Figure 6C, the inputs of the fusion decision model can be cross-modal recognition scores, respiration recognition scores, and face recognition scores, and the output is the multi-modal recognition score.
[0357] Combine Figure 6D , the fusion decision model includes a Stem layer, 12 encoding layers, a Head layer, and a Sigmoid function. Among them, the encoding layer includes a Linear layer, an activation function, and a Linear layer. Figure 6D In, the fusion decision model uses the Stem layer to expand the feature dimensions of the cross-modal recognition score, respiration recognition score, and face recognition score, and then encodes them through 12 encoding layers of a feed-forward neural network with a residual structure. Finally, a classification head (Head layer) and a Sigmoid function are used to output the multi-modal recognition score. The overall decision fusion model adopts a lightweight design, which can save storage resources.
[0358] The training process of the multi-modal fusion machine main recognition model is described below. From Figure 6C It can be seen that the multi-modal fusion machine main recognition model mainly includes a cross-modal encoder and a fusion decision model. That is to say, training the multi-modal fusion machine main recognition model is also training the cross-modal encoder and the fusion decision model.
[0359] During the training of the multi-modal fusion machine main recognition model, hard samples can be dynamically mined and used as negative samples, so as to further train the multi-modal fusion machine main recognition model in the way of positive and negative sample pairing.
[0360] Figure 6E shows the model training architecture of positive and negative sample pairing based on dynamic hard sample mining. As Figure 6E shown, the training process of the multi-modal fusion machine main recognition model can include the following steps: The first step is to obtain a batch of training samples from all the sample data sets corresponding to the sample users.
[0361] Combine Figure 6E , the "m×" in the upper right corner of the dotted box indicates sampling m batches of training samples. Each batch of training samples includes 1 positive sample and n negative samples.
[0362] In the embodiments of the present application, each sample user can correspond to a sample data set. Each sample data in the sample data set can include sample IMU data, sample touch screen data, sample respiration sound features, and sample face recognition scores.
[0363] Here, a positive sample and n negative samples can be obtained from all the sample data sets corresponding to the sample user . Among them, the sample label corresponding to the positive sample is , the sample label corresponding to the negative sample is . Among them, is the total number of sample users, is the total number of training batches.
[0364] Exemplarily, the multi-modal fusion machine main recognition model can be represented as a non-linear mapping . Among them, is the input of the model. Let represent the parameters of the model. Then the output score of the multi-modal fusion machine main recognition model can be expressed by the following formula (6): (6) Among them, is the multi-modal recognition score, is the Sigmoid function.
[0365] In the second step, input a batch of acquired training samples into the multi-modal fusion machine main recognition model to obtain the recognition probability corresponding to each sample, and use the negative sample that is misjudged as a positive sample by the model and has the highest corresponding recognition probability as the hard sample (or called the dynamic hard sample).
[0366] Exemplarily, the hard sample can be expressed by the following formula (7): (7) In the third step, use the hard sample and the positive sample to jointly calculate the loss of this training.
[0367] Exemplarily, when using the hard sample and the positive sample to jointly calculate the loss of this training, the loss calculation formula (8) can be as follows: (8) Among them, is the loss of this model training, is the cross-entropy loss function.
[0368] In the fourth step, use the model training loss to update the model parameters.
[0369] Exemplarily, the model parameters can be updated based on the following formula (9).
[0370] (9) Among them, are the updated model parameters.
[0371] Combining the above training process, it can be found that during the training of the multi-modal fusion machine owner recognition model, for the same sample user, m training samples can be collected. That is, the model is trained using m groups of training data for each sample user, rather than only using a single set of data for training, which can improve the accuracy of the model.
[0372] In addition, in each batch of training samples, difficult samples are mined as negative samples, and the dynamic difficult samples and positive samples are jointly used to calculate the model training loss, and the model parameters are updated based on the model training loss, which can effectively reduce the misrecognition rate and improve the TAR index at a lower misrecognition rate. That is, the trained multi-modal fusion machine owner recognition model can be made more accurate and the stability of machine owner recognition is also stronger.
[0373] Next, further combined with Table 1, Figure 7 and Figure 8 illustrate the effect of the machine owner recognition method based on multi-modal fusion provided in the embodiments of the present application.
[0374] Table 1 Performance evaluation results of the machine owner recognition method based on multi-modal fusion
[0375] Among them, the values in Table 1 are percentage values. For example, the AUC under the single-modal respiration is 89.36, which means 89.36%, or 0.8936.
[0376] Combined with Table 1, it can be found that compared with the face modality with the best performance in the single-modal, when the FAR (that is, the misrecognition rate) is 1%, the TAR (recognition success rate) of the multi-modal fusion machine owner recognition method is increased by 4.17%. When the FAR is 1.5%, the TAR of the multi-modal fusion machine owner recognition method is increased by 4.54%. When the FAR is 5%, the TAR of the multi-modal fusion machine owner recognition method is increased by 5.83%. That is to say, the multi-modal fusion machine owner recognition method improves the TAR index at a lower misrecognition rate.
[0377] Figure 7 is the ROC curve graph corresponding to the machine owner recognition of each modality provided in the embodiments of the present application, Figure 8 is the PR curve graph corresponding to the machine owner recognition of each modality provided in the embodiments of the present application.
[0378] Combined with Table 1, Figure 7 and Figure 8 it can be found that compared with the face modality with the best performance in the single-modal, the AUC of the multi-modal fusion machine owner recognition method is increased by 3.42%, and the AP is increased by 2.20%.
[0379] Comprehensively considering Table 1,Figure 7 and Figure 8 It can be found that the multi-modal fusion mobile owner recognition method can effectively improve the comprehensive recognition success rate by utilizing the complementary advantages between multi-modalities.
[0380] (3) Update of Mobile Owner Template In practical applications, as time goes by, the features of the mobile owner may change. For example, facial changes, respiratory sound changes, etc. Therefore, it is necessary to update the mobile owner template.
[0381] In the embodiments of the present application, the timing for updating the mobile owner template entered in the mobile phone can be: update when face is entered, and update regularly when face is unlocked.
[0382] The update period for updating the mobile owner template can be one day or one week. The embodiments of the present application do not make specific limitations on the update period.
[0383] It should be noted that when there is face entry, it means that a new mobile owner template is entered, and at this time, the mobile owner template in the mobile phone must be updated. When there is face unlocking and the update period is reached, the entered mobile owner template can be updated.
[0384] Since the process of entering the mobile owner template has been described in detail in the mobile owner template entry in the foregoing part (1). Therefore, the following mainly describes how to update the entered mobile owner template, and mainly describes how to update the cross-modal feature template and the respiratory sound feature template.
[0385] In practice, for any mobile owner, the face template in the mobile owner template can be updated synchronously or asynchronously with the cross-modal feature template and the respiratory sound feature template. For example, when the update condition (or called template update condition) is met, the face template can be updated first, and then the cross-modal feature template and the respiratory sound feature template can be updated. For another example, it may also be that when updating the face template, since no respiratory sound is detected in the collected respiratory audio data, the update of the cross-modal feature template and the respiratory sound feature template is postponed.
[0386] In the embodiments of the present application, during the process of updating the existing mobile owner template, if the update period is reached and there is face unlocking, it can trigger the update of the mobile owner template of the corresponding mobile owner. For example, if the mobile owner templates of three mobile owners A, B, and C have been entered in the mobile phone, after the update period is reached and there is face unlocking, for example, when the mobile owner A is recognized during face unlocking, it can trigger the update of the mobile owner template of the mobile owner A. At this time, the mobile phone can collect the IMU data, touch screen data, and respiratory audio data of the corresponding mobile owner when face unlocking is successful, and then, the pre-stored audio processing module can be called (see Figure 4A), the audio processing module processes the breathing audio data to extract the breathing sound features in the above breathing audio data, and the breathing sound features are the current breathing sound features of the machine owner. The current breathing sound features are the breathing sound features in the breathing audio data collected when the face unlock is successful. After that, the mobile phone can call a pre-stored cross-modal encoder (see Figure 5A ), and the cross-modal encoder fuses the breathing sound features, touch screen data, and IMU data to obtain the current cross-modal features of the machine owner. The current cross-modal features are the cross-modal features corresponding to the multi-modal data collected when the face unlock is successful.
[0387] In some optional implementation manners of the embodiments of the present application, in order to improve the accuracy of updating the machine owner template, the breathing sound feature template can be updated when the breathing sound confidence corresponding to the current breathing sound features is greater than the third determination threshold c. Wherein, c > a, and a is the lower threshold for determining whether the breathing audio data includes breathing sounds. In some scenarios of the present application, a is also referred to as the first determination threshold.
[0388] That is to say, in some application scenarios, the update condition (or referred to as the template update condition) for updating the entered machine owner template includes not only face unlock and reaching the update cycle, but also the breathing sound confidence corresponding to the current breathing sound features being greater than c.
[0389] It can be understood that the mobile phone usually updates the machine owner template of the machine owner recognized when the face unlock is successful. For example, when the face unlock is successful and the recognized machine owner is A, then, when the template update condition is met, the mobile phone can update the machine owner template of the machine owner A.
[0390] (1) Update of the machine owner template in the single feature template scenario.
[0391] A single feature template usually means that the machine owner template includes one feature template. For example, the machine owner template includes one breathing sound feature template, or in other words, one machine owner corresponds to one breathing sound feature template.
[0392] In some optional implementation manners of the embodiments of the present application, when the above update conditions (or referred to as template update conditions) for updating the machine owner template are met, the mobile phone can directly replace the original feature template with the current feature. Wherein, the current feature can include the current breathing sound features and the current cross-modal features.
[0393] For example, the current breathing sound features can be used to replace the original breathing sound feature template, and the current cross-modal features can be used to replace the original cross-modal feature template. That is, the current breathing sound features are used as the breathing sound feature template, and the current cross-modal features are used as the cross-modal feature template.
[0394] In some alternative implementations, when the template update condition is met, the mobile phone can also determine the average value of the feature template in the owner template and the current feature of the same modality as the new feature template of the corresponding modality, and replace the old feature template of the corresponding modality with the new feature template.
[0395] For example, the mobile phone can determine the average value of the breath sound feature template in the owner template and the current breath sound feature as the new breath sound feature template, and replace the old breath sound feature template with the new breath sound feature template. At the same time, the mobile phone can determine the average value of the cross-modal feature template in the owner template and the current cross-modal feature as the new cross-modal feature template, and replace the old cross-modal feature template with the new cross-modal feature template.
[0396] It should be noted that taking the average value of the current feature and the old feature template as the new feature template can accurately and stably update the owner template, which helps to improve the accuracy and stability of owner recognition.
[0397] (2) Update of the owner template in the scenario of multiple feature templates.
[0398] Multiple feature templates generally refer to that the owner template includes multiple feature templates. For example, the owner template includes multiple breath sound feature templates, or in other words, one owner corresponds to multiple breath sound feature templates.
[0399] In some alternative implementations of the embodiments of the present application, when the template update condition is met, the mobile phone can determine the average value of each feature template in the owner template and the current feature of the same modality as the new feature template of the corresponding modality, and add the new feature template to the owner template.
[0400] For example, the mobile phone can determine the average value of the breath sound feature template in the owner template and the current breath sound feature as the new breath sound feature template, and add the new breath sound feature template to the owner template. At the same time, the mobile phone can determine the average value of the cross-modal feature template in the owner template and the current cross-modal feature as the new cross-modal feature template, and add the new cross-modal feature template to the owner template.
[0401] Further, for example, if there are p - 1 cross-modal feature templates recorded in the owner template, which are respectively , and the current cross-modal feature is , then the new cross-modal feature template can be calculated by the following formula (10).
[0402] (10) Similarly, if there are q - 1 breath sound feature templates recorded in the owner template, which are respectively , the current breath sound feature is , then the new breath sound feature template can be calculated by the following formula (11).
[0403] (11) It should be noted that taking the mean of the current feature and the old feature template as the new feature template can achieve accurate and stable updating of the owner template, which helps to improve the accuracy and stability of owner recognition.
[0404] In some optional implementation manners of the embodiments of the present application, the mobile phone can also first perform feature quality judgment on the current feature of the owner, and determine whether to update the feature template according to the feature quality judgment result. In the case where the feature quality judgment result is not good, the current feature of the owner can be deleted. And in the case where the feature quality judgment result is good, the current feature can be used to update the owner template.
[0405] As an example, if the feature quality judgment result of the current cross-modal feature of the owner is not good, the mobile phone can directly delete the current cross-modal feature.
[0406] As another example, if it is determined that the feature quality of the current breath sound feature of the owner is very good after the feature quality judgment, at this time, the current breath sound feature can be used to update the owner template. For example, the current breath sound feature can replace the original breath sound feature template, or the current breath sound feature can be directly added to the owner template.
[0407] In some optional implementation manners of the embodiments of the present application, when the template update condition is met, the mobile phone can respectively determine the similarity between each feature template in the owner template and the current feature of the same modality, and then select the similarity with the largest value from multiple similarities. For the convenience of description, it can be called the maximum similarity.
[0408] Update method 1: If the maximum similarity is greater than the fifth determination threshold e, then replace the feature template corresponding to the maximum similarity with the current feature. Among them, the fifth determination threshold is a preset threshold, and the value range of the fifth determination threshold is 0-1. In the scenario of updating by using update method 1, the current feature replaces the feature template corresponding to the maximum similarity.
[0409] Update method 2: If the maximum similarity is less than or equal to e, and the number of feature templates of the corresponding modality stored in the current mobile phone is less than the preset template quantity threshold, then retain the feature template corresponding to the maximum similarity, and add the current feature to the owner template. In the scenario of updating by using update method 2, the number of feature templates in the owner template increases by one.
[0410] Update method 3: If the maximum similarity is less than or equal to e, and the number of feature templates stored in the current mobile phone is equal to the template quantity threshold, then delete the feature template with the longest preservation time in the owner template, that is, the one with the farthest distance from the current time, and add the current feature to the owner template. In the scenario of updating using Update method 3, the current feature replaces the feature template with the longest preservation time.
[0411] The following is a further example in combination with the breathing modality.
[0412] If the similarity between the current breathing sound feature R0 and the existing breathing sound feature templates R1, R2, and R3 in the owner template is S1, S2, and S3 respectively, where the maximum similarity is S3, and S3 corresponds to the breathing sound feature template R3.
[0413] If S3 is greater than the threshold e, the mobile phone can replace the breathing sound feature template R3 with the current breathing sound feature R0.
[0414] If S3 is less than or equal to the threshold e, and the number of breathing sound feature templates in the current mobile phone is less than the template quantity threshold, the mobile phone can add the current breathing sound feature R0 to the owner template. At this time, there are 4 breathing sound feature templates in the owner template.
[0415] If S3 is less than or equal to the threshold e, and the number of breathing sound feature templates in the current mobile phone is equal to the template quantity threshold, the mobile phone can first delete the breathing sound feature template with the longest preservation time. For example, it can delete the breathing sound feature template R1, and then add the current breathing sound feature R0 to the owner template.
[0416] In the embodiments of the present application, by combining the number of feature templates stored in the mobile phone and the storage time of the feature templates, the owner template is updated, which can realize flexible and effective updating of each owner template.
[0417] Corresponding to the owner recognition method in the above embodiments, Figure 9 The structural schematic diagram of the owner recognition device in the host provided by the embodiments of the present application is shown. For the sake of convenience of description, only the parts related to the embodiments of the present application are shown.
[0418] Refer to Figure 9 , the owner recognition device includes: A data acquisition unit 901, configured to collect user data of a user to be measured in response to an unlocking operation of the user to be measured on the target content, where the target content is in a locked state, and the user data includes at least one of the following: a face image to be measured, breathing audio data to be measured, IMU data to be measured, and touch screen data to be measured; An identification execution unit 902, configured to perform owner identification according to user data and a pre-recorded owner template, where the owner template includes at least one of the following: a face template, a breathing sound feature template, and a cross-modal feature template, and the cross-modal feature template is used to describe the comprehensive features of the owner in multiple modalities; An unlocking execution unit 903, configured to unlock the target content when the owner is identified.
[0419] As an embodiment of the present application, the owner identification device can implement as Figures 1A - 8 shown in the embodiments and other related method embodiments.
[0420] In the owner identification device provided in the embodiments of the present application, the processes of each module implementing its respective functions can be specifically referred to the Figures 1A - 8 described embodiments shown above and other related method embodiments, and will not be elaborated here.
[0421] It should be noted that the content such as owner identification and execution process among the above-mentioned devices / units, due to being based on the same concept as the method embodiments of the present application, for their specific functions and the technical effects brought, reference can be specifically made to the method embodiment part, and will not be elaborated here.
[0422] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution is prior or posterior. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0423] It should be understood that when used in the specification and the appended claims of the present application, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0424] It should also be understood that the term "and / or" used in the specification and the appended claims of the present application refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0425] As used in the specification and the appended claims of the present application, the term "if" can be interpreted as "when...", "once", "in response to determining", or "in response to detecting" according to the context. Similarly, the phrase "if determined" or "if detecting [the described condition or event]" can be interpreted as meaning "once determined", "in response to determining", "once detecting [the described condition or event]", or "in response to detecting [the described condition or event]" according to the context.
[0426] In addition, in the description of the specification and the appended claims of the present application, the terms "first", "second", "third", etc. are only used for differential description and should not be construed as indicating or implying relative importance. It should also be understood that although the terms "first", "second", etc. are used in the text to describe various elements in some embodiments of the present application, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, the first status code can be named the second status code, and similarly, the second status code can be named the first status code, without departing from the scope of the various described embodiments. The first status code and the second status code are both status codes, but they are not the same status code.
[0427] The reference to "one embodiment" or "some embodiments" etc. described in the specification of the present application means that a specific feature, structure or characteristic described in connection with the embodiment is included in one or more embodiments of the present application. Thus, the statements "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments" etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0428] Figure 10 It is a schematic structural diagram of a terminal device provided by an embodiment of the present application.
[0429] The terminal device may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a sensor module 180, a key 190, a camera 193, a touch screen 194, and a SIM card interface 195, etc.
[0430] Among them, the audio module 170 may include a speaker and a microphone, etc.
[0431] Among them, the sensor module 180 may include a pressure sensor, an IMU, a distance sensor, a proximity light sensor, a fingerprint sensor, a touch sensor, an ambient light sensor (of course, the terminal device may also include other sensors, such as a temperature sensor, a bone conduction sensor, etc., which are not shown in the figure).
[0432] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors. Among them, the controller may be the nerve center and command center of the terminal device. The controller may generate operation control signals according to the instruction operation code and timing signals to complete the control of fetching and executing instructions.
[0433] A memory may also be provided in the processor 110 for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory may save the instructions or data that the processor 110 has just used or recycled. If the processor 110 needs to use the instruction or data again, it can be directly called from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.
[0434] The processor 110 may run the owner identification method provided in the embodiments of the present application to facilitate the quick and accurate creation of schedules and enhance the user experience. The processor 110 may include different devices. For example, when the CPU and GPU are integrated, the CPU and GPU may cooperate to execute the owner identification method provided in the embodiments of the present application. For example, some algorithms in the owner identification method are executed by the CPU and another part of the algorithms are executed by the GPU to obtain a faster processing efficiency.
[0435] It should be understood that the illustrated terminal device is only an example, and the terminal device may have more or fewer components than those shown in the figure, may combine two or more components, or may have different component configurations. The various components shown in the figure may be implemented in hardware, software, or a combination of hardware and software including one or more signal processing and / or application specific integrated circuits.
[0436] The software system of the terminal device may adopt a layered architecture, an event-driven architecture, a microkernel architecture, a microservices architecture, or a cloud architecture. In the embodiments of the present application, the Android system with a layered architecture is taken as an example to exemplarily illustrate the software structure of the terminal device.
[0437] Figure 11It is a software structure block diagram of the terminal device according to an embodiment of the present application.
[0438] The layered architecture divides the software into several layers, and each layer has a clear role and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers, from top to bottom, namely the application layer, the application framework layer, the Android runtime and the system libraries, and the kernel layer.
[0439] The application layer may include a series of application packages.
[0440] Such as Figure 11 shown, the application packages may include multiple applications such as a gallery, a calendar, a map, WeChat, QQ, text messages, etc. It may also include multiple models such as a breath sound recognition model, a breath sound feature extraction model, a multi-modal fusion machine main recognition model, a cross-modal encoder, and a fusion decision model.
[0441] The application framework layer provides application programming interfaces (APIs) and programming frameworks for the applications in the application layer. The application framework layer includes some predefined functions.
[0442] In some implementation manners, multiple models such as a breath sound recognition model, a breath sound feature extraction model, a multi-modal fusion machine main recognition model, a cross-modal encoder, and a fusion decision model may also be deployed in the application framework layer.
[0443] Such as Figure 11 shown, the application framework layer may include a window manager, a content provider, a view system, a phone manager, a resource manager, a notification manager, etc.
[0444] The window manager is used to manage window programs. The window manager can obtain the display screen size, determine whether there is a status bar, lock the screen, capture the screen, etc.
[0445] The content provider is used to store and obtain data, and make this data accessible to applications. The data may include videos, images, audio, dialed and answered calls, browsing history and bookmarks, phone books, etc.
[0446] The view system includes visible controls, such as controls for displaying text, controls for displaying pictures, etc. The view system can be used to build applications. The display interface can be composed of one or more views. For example, a display interface including a text message notification icon may include a view for displaying text and a view for displaying pictures.
[0447] The phone manager is used to provide the communication functions of the terminal device. For example, the management of call status (including answering, hanging up, etc.).
[0448] The resource manager provides various resources for applications, such as localized strings, icons, pictures, layout files, video files, etc.
[0449] The notification manager enables applications to display notification information in the status bar. It can be used to convey notification-type messages, which can automatically disappear after a short stay without user interaction. For example, the notification manager is used to inform that the download is completed, message reminders, etc. The notification manager can also be a notification that appears in the system top status bar in the form of a chart or scroll bar text, such as the notification of a background-running application, or a notification that appears in the form of a dialog window on the screen. For example, it prompts text information in the status bar, emits a prompt tone, the terminal device vibrates, the indicator light flashes, etc.
[0450] Android Runtime includes a core library and a virtual machine. Android Runtime is responsible for the scheduling and management of the Android system.
[0451] The core library consists of two parts: one part is the functional functions that need to be called by the Java language, and the other part is the core library of Android.
[0452] The application layer and the application framework layer run in the virtual machine. The virtual machine executes the Java files of the application layer and the application framework layer as binary files. The virtual machine is used to perform functions such as the management of object life cycles, stack management, thread management, security and exception management, and garbage collection.
[0453] The system library can include multiple functional modules. For example: surface manager, Media Libraries, 3D graphics processing library (such as: OpenGL ES), 2D graphics engine (such as: SGL), etc.
[0454] The surface manager is used to manage the display subsystem and provides the fusion of 2D and 3D layers for multiple applications.
[0455] The media library supports the playback and recording of multiple common audio and video formats, as well as static image files, etc. The media library can support multiple audio and video coding formats, such as: MPEG4, H.164, MP3, AAC, AMR, JPG, PNG, etc.
[0456] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, synthesis, and layer processing, etc.
[0457] The 2D graphics engine is a drawing engine for 2D drawing.
[0458] The kernel layer is the layer between hardware and software. The kernel layer includes at least a display driver, a camera driver, an audio driver, and a sensor driver.
[0459] In addition, those skilled in the art can clearly understand that, for the convenience and simplicity of description, only the above-mentioned division of each functional unit and module is used as an example for illustration. In actual applications, the above functions can be allocated to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. In each embodiment of the present application, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0460] The embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the above-mentioned various method embodiments can be implemented.
[0461] The embodiment of the present application provides a computer program product. When the computer program product runs on a terminal device, the terminal device can be made to execute the steps in the above-mentioned various method embodiments.
[0462] The embodiment of the present application also provides a chip system. The chip system includes a processor, the processor is coupled to a memory, and the processor executes a computer program stored in the memory to implement the steps in the above-mentioned various method embodiments.
[0463] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above-mentioned embodiment methods of the present application, it can also be completed by a computer program instructing relevant hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps in the above-mentioned various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable storage medium can include: any entity or device that can carry the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc.
[0464] In the above embodiments, the descriptions of the respective embodiments have their own focuses. For parts not detailed or recorded in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0465] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.
[0466] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0467] The above embodiments are only used to illustrate the technical solutions of this application, rather than to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A method for identifying a machine owner, characterized in that, Applied to a terminal device, the method includes: In response to an unlocking operation of a target content by a user to be measured, user data of the user to be measured is collected, where the target content is in a locked state, and the user data includes at least one of the following: a face image to be measured and breathing audio data to be measured; Based on the user data and a pre-recorded owner template, owner identification is performed, and the owner template includes at least one of the following: a face template and a breathing sound feature template; When the owner is recognized, the target content is unlocked.
2. The method for identifying the machine owner according to claim 1, wherein, The target content includes any one of the following: a locked notification message, a locked service card, a locked application, a locked function, and a locked interface.
3. The method for identifying the machine owner according to claim 1, wherein The process of inputting the breathing sound feature template includes: When the breathing feature input condition is met, breathing audio data of the owner is collected, where the breathing feature input condition includes any one of the following: the face template of the owner is successfully input, the face unlock of the owner is successful and the breathing sound feature template of the owner currently performing the unlock has not been input into the terminal device; Extract breathing sound features from the breathing audio data, and input the extracted breathing sound features as the breathing sound feature template.
4. The method for identifying the machine owner according to claim 3, wherein The extracting of the breathing sound features from the breathing audio data includes: Extract Mel cepstrum and Mel cepstral coefficients from the breathing audio data; Segment the Mel cepstrum into a plurality of consecutive Mel cepstrum segments; Input the plurality of Mel cepstrum segments into a pre-trained breathing sound recognition model to obtain a breathing sound confidence level, where the breathing sound confidence level is used to indicate the probability that the breathing audio data is classified as a breathing sound; When the breathing sound confidence level is greater than a preset first determination threshold, input the Mel cepstral coefficients into a pre-trained breathing sound feature extraction model to obtain the breathing sound features.
5. The method for identifying the machine owner according to claim 4, wherein The data processing process of the breathing sound recognition model includes: Extract single-segment sound representations respectively corresponding to each Mel cepstrum segment through a pre-trained feature extractor; Splice the single-segment sound representations respectively corresponding to each Mel cepstrum segment in the time order of each Mel cepstrum segment to obtain a continuous-segment sound representation; Based on the continuous-segment sound representation, determine the confidence levels respectively corresponding to various sound types, where the sound types include: quiet, speaking, breathing sound, coughing, sneezing, snoring, and environmental noise, and the confidence level corresponding to a sound type is used to indicate the probability that the breathing audio data is determined as the corresponding sound type.
6. The method for identifying the machine owner according to claim 4, wherein, The breathing sound feature extraction model includes a convolutional layer, a multi-scale residual layer, and a bottleneck layer connected in sequence, where there are at least two multi-scale residual layers.
7. The method for identifying the machine owner according to claim 4, wherein The training process of the breathing sound feature extraction model includes: Select K batches of breathing samples from the breathing sample set used for training the breathing sound recognition model, where K>1; Based on the features of the breathing sound, perform data augmentation processing on the audio in each of the breathing samples to obtain at least one enhanced breathing sample respectively corresponding to each of the breathing samples, and the enhanced breathing samples are used to train the breathing sound feature extraction model; During the process of training the breath sound feature extraction model based on the enhanced breath samples corresponding to each batch of breath samples, determine the model training loss, and based on the obtained model training loss, adjust the model parameters of the breath sound feature extraction model; Among them, the model training loss includes a user recognition cross-entropy loss and a contrast learning similarity loss. The user recognition cross-entropy loss is used to describe the accuracy of the breath sound feature extraction model for user recognition, and the contrast learning similarity loss is used to describe the similarity between the features extracted by the breath sound feature extraction model and the true features.
8. The method for identifying the machine owner according to claim 7, wherein, The data augmentation process includes at least one of the following: Adding random noise with a random signal-to-noise ratio level at a random position in the audio of the breath sample to obtain the enhanced breath sample. The random noise includes at least one of the following: household noise, outdoor noise, office noise, transportation noise, and animal noise; Adjusting the speed of the audio in the breath sample; Adjusting the volume range of the audio in the breath sample.
9. The method for identifying the machine owner according to claim 1, characterized in that The collection of the user data of the user to be measured includes: According to the usage scenario corresponding to the unlocking operation, determine the audio collection duration of the breath audio data to be measured. The audio collection duration is less than or equal to the target duration, and the target duration is the duration of the breath audio data collected when entering the breath sound feature template; Collect the breath audio data to be measured in the user data according to the audio collection duration.
10. The method for identifying the machine owner according to claim 1, wherein, The machine owner identification based on the user data and the pre-entered machine owner template includes: In the case where there is no face in the face image to be measured, extract the breath sound feature to be measured from the breath audio data to be measured; According to the breath sound feature to be measured and the breath sound feature template in the machine owner template, determine whether the user to be measured is the machine owner.
11. The method for identifying the machine owner according to claim 10, wherein The determination of whether the user to be measured is the machine owner according to the breath sound feature to be measured and the breath sound feature template in the machine owner template includes: In the case where there is one breath sound feature template in the machine owner template, determine the similarity between the breath sound feature to be measured and the breath sound feature template. When the obtained similarity is greater than a pre-set second determination threshold, determine that the user to be measured is the machine owner; In the case where there are multiple breath sound feature templates in the machine owner template, determine the similarity between each breath sound feature template and the breath sound feature to be measured respectively. When the average value of the obtained multiple similarities is greater than a pre-set fourth determination threshold, determine that the user to be measured is the machine owner.
12. The method for identifying the machine owner according to any one of claims 1-11, characterized in that, The method further includes: In the case where there is one breath sound feature template in the machine owner template and the template update condition is satisfied, switch the breath sound feature template in the machine owner template to the current breath sound feature, where the current breath sound feature is the breath sound feature in the breath audio data collected when the face unlock is successful; Among them, the template update condition includes: the face unlock is successful, the update period is reached, and the breath sound confidence corresponding to the current breath sound feature is greater than a pre-set third determination threshold, where the third determination threshold is greater than the first determination threshold.
13. The method for identifying the machine owner according to any one of claims 1-11, characterized in that, The method further includes: When there are multiple respiratory sound feature templates in the owner template and the template update condition is satisfied, determining the similarity between the current respiratory sound feature and each respiratory sound feature template; Determining the maximum similarity with the largest corresponding value and the target respiratory sound feature template corresponding to the maximum similarity; Updating the owner template according to the maximum similarity and the target respiratory sound feature template.
14. The method for identifying the machine owner according to claim 13, wherein The updating the owner template according to the maximum similarity and the target respiratory sound feature template includes: When the maximum similarity is greater than a preset fifth determination threshold, replacing the target respiratory sound feature template in the owner template with the current respiratory sound feature; When the maximum similarity is less than or equal to the fifth determination threshold and the number of respiratory sound feature templates stored in the terminal device is less than a preset template number threshold, adding the current respiratory sound feature to the owner template; When the maximum similarity is less than or equal to the fifth determination threshold and the number of respiratory sound feature templates stored in the terminal device is equal to the template number threshold, deleting the respiratory sound feature template with the longest storage duration in the owner template and adding the current respiratory sound feature to the owner template.
15. A terminal device, characterized in that, The terminal device includes a memory and a processor. A computer program that can run on the processor is stored on the memory. When the processor executes the computer program, the owner identification method according to any one of claims 1 to 14 is implemented.
Citation Information
Patent Citations
User management method and system for mobile terminal
CN106127006A
Breath anomaly classification model training method, breath anomaly detection method and electronic equipment
CN114333901A
User authentication method and related equipment
CN115982686A
Breathing sound signal identification method and system
CN116469416A
User identity authentication method and system based on breath sound
CN120030518A