Machine owner identification method and terminal device

By collecting multi-modal data for owner identification, the problem of owner privacy leakage in terminal devices has been solved, and accurate identification and low-power identification under different lighting conditions have been achieved, improving user experience and privacy protection.

CN120277648BActive Publication Date: 2026-06-02HONOR DEVICE CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HONOR DEVICE CO LTD
Filing Date
2023-12-29
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

During the use of terminal devices, the owner's private information may be seen by others, leading to privacy leaks. Existing technologies are insufficient to effectively protect the owner's privacy.

Method used

By collecting multimodal data from the user under test, such as facial images, respiratory audio data, IMU data, and touch screen data, and combining it with a pre-recorded owner template, the target content is unlocked only when the owner is identified. The complementary advantages of multimodal data are used to improve the stability and accuracy of the identification.

Benefits of technology

It enables accurate identification of the device owner even under different lighting conditions, reducing computational load and power consumption, improving user experience, and protecting the privacy and security of the device owner's information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277648B_ABST
    Figure CN120277648B_ABST
Patent Text Reader

Abstract

The embodiment of the application discloses a kind of master identification method and terminal equipment, it is applicable to computer technology field, the method comprises: terminal equipment responds to the unlocking operation of target content to the user data of the user to be measured of the user to be measured, wherein, target content is in the state of being locked, user data includes at least one of the following: the face image of the user to be measured, the respiratory audio data of the user to be measured, the IMU data of the user to be measured and the touch screen data of the user to be measured;According to user data and pre-recorded master template, master identification is carried out, master template includes at least one of the following: face template, respiratory sound feature template, cross-modal feature template, cross-modal feature template is used to describe the comprehensive characteristics of master in multiple modes;In the case where master is identified, target content is unlocked.The embodiment of the application can be used in the process of terminal equipment, and master identification is carried out based on respiratory sound feature, to protect the safety of master private information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a method for identifying the owner of a device and a terminal device. Background Technology

[0002] Terminal devices, especially mobile phones, are gradually becoming the center of the personal digital world, carrying a large amount of personal information and privacy data. During use, these devices may display the owner's private information, which could lead to privacy breaches if seen by someone other than the owner.

[0003] Therefore, during the use of terminal devices, it is necessary to identify the device owner in order to protect the security of the owner's private information. Summary of the Invention

[0004] In view of this, embodiments of this application provide a device owner identification method and a terminal device, which can identify the device owner during the use of the terminal device to protect the security of the device owner's private information.

[0005] In a first aspect, embodiments of this application provide a device owner identification method applied to a terminal device. The device owner identification method may include: the terminal device can, in response to a user's unlocking operation on target content, collect user data of the user, wherein the target content is in a locked state. The user data includes at least one of the following: a face image to be tested, respiratory audio data to be tested, IMU data to be tested, and touchscreen data to be tested; then, the terminal device can perform device owner identification based on the user data and a pre-recorded device owner template, the device owner template including at least one of the following: a face template, a respiratory sound feature template, and a cross-modal feature template, wherein the cross-modal feature template is used to describe the comprehensive features of the device owner in multiple modalities; finally, the terminal device can unlock the target content if the device owner is identified.

[0006] In this context, "user to be tested" typically refers to the user to be identified. "Face image to be tested" refers to the face image of the user to be tested. "Respiratory audio data to be tested" refers to the respiratory audio data of the user to be tested. "IMU data to be tested" refers to the IMU data received by the terminal device in response to the user to be tested. "Touchscreen data to be tested" refers to the touchscreen data received by the terminal device in response to the user to be tested.

[0007] The device owner template is a pre-recorded device owner template. The face template is a pre-recorded face image of the device owner. The breathing sound feature template is a pre-recorded breathing sound feature of the device owner. The cross-modal feature template is a pre-recorded composite feature of the device owner across multiple modalities.

[0008] In this embodiment, when a user wants to unlock locked content, the terminal device can collect the user's data to identify the user as the owner. Only after identifying the owner will the terminal device unlock the target content, ensuring the security of the owner's private information. Furthermore, the terminal device can use one or more of the user's facial data, breathing audio data, IMU data, and touchscreen data for owner identification. This allows for multi-modal data analysis, with different modalities complementing each other to improve the stability and accuracy of owner identification.

[0009] Optionally, the terminal device can prioritize using facial data for owner identification. If the facial recognition result is not ideal, it can continue to use one or more of the following: breathing audio data, IMU data, and touch screen data for owner identification. In this way, owner identification can still be performed even if the facial recognition result is not ideal, which can improve the accuracy and stability of owner identification.

[0010] In addition, the IMU data and touch screen data of the terminal device reflect the user's usage habits of the terminal device. Combining these two data to further determine whether the user is the owner can make the identification results more accurate and reliable, that is, it can further improve the accuracy of owner identification.

[0011] In a first possible implementation of the first aspect, the target content includes any of the following: a locked notification message, a locked service card, a locked application, a locked function, and a locked interface.

[0012] In this application embodiment, the target content can be one or more, meaning there can be many application scenarios that can trigger the terminal device to identify the owner. In other words, the terminal device can lock one or more types of content to ensure the security of the owner's private information.

[0013] As an optional embodiment of this application, the terminal device may lock the content set by the locking setting operation in response to the owner's locking setting operation.

[0014] The lock setting is used to lock the settings configured by the device owner. For example, it can lock one or more applications configured by the owner. It can also lock service card functions configured by the owner. Furthermore, it can lock password-free payment functions configured by the owner, and so on.

[0015] In a second possible implementation of the first aspect, the process of the terminal device recording the user's face template may include: first, the terminal device responds to the user's face recording operation by displaying a face recording interface, wherein the face recording interface includes a face capture frame; then, the terminal device may respond to the captured face image by recording the captured face image as a face template.

[0016] In this embodiment, the user can input a facial template through human-computer interaction, which is intuitive and convenient, and helps to improve the user experience.

[0017] As an optional embodiment of this application, the terminal device can input a user's user template for one user or multiple user templates. Typically, one user corresponds to one user template.

[0018] In the third possible implementation of the first aspect, the process of the terminal device recording the cross-modal feature template of the user may include: First, under the condition of meeting the multimodal feature recording conditions, the terminal device may collect the user's respiratory audio data, the terminal device's IMU data, and the terminal device's touch screen data; then, the terminal device may extract respiratory sound features from the respiratory audio data; finally, the terminal device may encode the IMU data, touch screen data, and respiratory sound features through a cross-modal encoder to obtain cross-modal features, and record the cross-modal features as a cross-modal feature template, wherein the cross-modal encoder is used to encode data of multiple modalities.

[0019] The multimodal feature input conditions include any of the following: the owner's face template is successfully entered, the owner's face unlock is successful, and the terminal device has not yet entered the cross-modal feature template of the owner currently performing the unlock.

[0020] In this embodiment of the application, the terminal device can continue to collect the breathing audio data of the owner after the owner's face template has been successfully registered, thereby extracting breathing sound features from the collected breathing audio data, and encoding the IMU data, touch screen data and breathing sound features through a cross-modal encoder to obtain a cross-modal feature template.

[0021] Because the breathing audio data collected when the user's face unlock is successfully activated may be affected by various factors, such as environmental noise, preventing the extraction of breathing sound features, the cross-modal encoder will be unable to encode cross-modal features. Therefore, even before the user's cross-modal feature template is successfully registered, breathing audio data can be collected again after the user's face unlock is successful. Breathing sound features can then be extracted and combined with these features to encode cross-modal features, which are then used to register the user's cross-modal feature template. This ensures the successful registration of the user's cross-modal feature template.

[0022] In a fourth possible implementation of the first aspect, the cross-modal encoder may include a cross-channel attention encoder. The cross-channel attention encoder includes a first slice encoding module, a second slice encoding module, and a third slice encoding module in parallel. The first slice encoding module includes a first group layer and a first linear layer connected in sequence; the second slice encoding module includes a second group layer and a second linear layer connected in sequence; and the third slice encoding module includes a third group layer and a third linear layer connected in sequence.

[0023] The input of the first Linear layer is connected to the output of the first Group layer, the input of the second Linear layer is connected to the output of the first Linear layer and the output of the second Group layer, and the input of the third Linear layer is connected to the output of the second Linear layer and the output of the third Group layer.

[0024] In this embodiment, the output data of the first Group layer is encoded the most times, and the output data of the third Group layer is encoded the least times.

[0025] In other words, when fusing multiple features, a cross-modal encoder can deploy high-quality features in the third group layer and low-quality features in the first group layer. This minimizes the number of linear layers (encoding times) required for features deployed in the third group layer, while minimizing the number of linear layers required for features deployed in the first group layer. Conversely, it maximizes the number of linear layers required for features deployed in the first group layer. This allows a cross-channel attention encoder to combine the quality of features and encode them at different levels, achieving better fusion of high-quality and low-quality features. In other words, a cross-modal encoder can enhance features, thereby improving the accuracy of the output features.

[0026] Optionally, after encoding and fusing high-quality and low-quality features to different degrees, the fused features can be transposed. The transposed fused data can then be input into the Features layer, which helps to achieve further effective fusion between features.

[0027] In a fifth possible implementation of the first aspect, the cross-modal encoder further includes a cross-channel linear encoder, wherein the cross-channel linear encoder comprises a sequentially connected Linear layer, an activation function, a Linear layer, a norm layer, a conv layer, and an activation function.

[0028] In the embodiments of this application, the cross-channel linear encoder uses a linear layer for encoding and aligns the feature dimensions of different modalities, and uses a conv layer to implement one-dimensional pointwise convolution to fuse information between channels. That is to say, the cross-channel linear encoder can increase the level of abstraction of the output features.

[0029] In the sixth possible implementation of the first aspect, the terminal device performs owner identification based on user data and a pre-entered owner template. This may include: First, if a face is present in the face image to be tested, the terminal device can perform owner identification based on the face image to be tested and the face template to obtain a face recognition score; then, if the face recognition score is greater than a pre-set lower threshold and less than a pre-set upper threshold, the terminal device can extract the test user's breath sound features from the test breath audio data; finally, the test breath sound features, test IMU data, test touchscreen data, and face recognition score are input into a pre-trained multimodal fusion owner identification model to obtain owner indication information, which is used to indicate whether the test user is the owner.

[0030] In this embodiment, only when facial recognition is ineffective, such as in situations with strong light, backlight, or low light, will the system further collect the user's breathing audio data, touchscreen data, and IMU data. This multimodal data is then combined to perform device owner identification, thereby improving the accuracy of owner recognition. Since facial recognition accuracy is already high enough most of the time, the use of multimodal data for owner identification is relatively rare. This approach allows for saving computational resources and reducing power consumption while maintaining accurate owner identification, thus improving the phone's usability and enhancing the user experience.

[0031] Optionally, the aforementioned upper limit threshold for face recognition can be the FAR (Facial Recognition Accuracy) threshold for face recognition. The aforementioned lower limit threshold for face recognition can be the FRR (Facial Recognition Response) threshold for face recognition.

[0032] In practice, a facial recognition score falling between the aforementioned lower and upper thresholds typically indicates a poor facial recognition result. This means that even with suboptimal facial recognition, the terminal device can combine multimodal data for more accurate owner identification. This significantly reduces the need for users to repeatedly cooperate with the device for owner identification, effectively protecting the owner's privacy while improving the user experience.

[0033] Optionally, the terminal device can continue to use the breathing sound features in the breathing audio data to identify the owner even if there is no face in the face image to be tested. In this way, the owner can still be identified even if the face recognition result is not ideal, which can improve the accuracy and stability of the owner identification.

[0034] Optionally, the terminal device may output a message prompting the user to repeat the operation if the face image to be tested does not contain a face and the breathing audio data to be tested does not contain breathing sounds, that is, if breathing sound features cannot be extracted. This is to ensure the accuracy of the device owner's recognition.

[0035] Optionally, if the face recognition score is greater than the lower threshold but less than the upper threshold, the terminal device can also use the breathing sound features in the audio data to be tested for owner identification. For example, the terminal device can determine the similarity between the breathing sound feature to be tested and the breathing sound feature template. If the obtained similarity is greater than a pre-set second determination threshold, the user to be tested is determined to be the owner. Conversely, if the obtained similarity is less than or equal to the second determination threshold, the user to be tested is determined to be a non-owner.

[0036] Optionally, if the face recognition score is greater than the lower threshold but less than the upper threshold, and the audio data being tested does not contain breathing sounds (i.e., no breathing sound features are extracted), the terminal device can combine the face recognition score to further determine whether the user being tested is the owner of the device. For example, if the face recognition score is greater than a pre-set threshold, the terminal device can determine that the user being tested is the owner. Conversely, if the face recognition score is less than or equal to the threshold, the terminal device can determine that the user being tested is not the owner.

[0037] Optionally, if the face recognition score is greater than the lower threshold but less than the upper threshold, and the audio data being tested does not contain breathing sounds (i.e., no breathing sound features were extracted), the terminal device can also output a message prompting the user to repeat the operation. This ensures the accuracy of the user's recognition.

[0038] In the seventh possible implementation of the first aspect, the terminal device performs owner identification based on user data and a pre-entered owner template, and may further include: if the face recognition score is greater than the upper limit threshold of face recognition, the terminal device can determine that the user to be tested is the owner; if the face recognition score is less than the lower limit threshold of face recognition, the terminal device can determine that the user to be tested is not the owner.

[0039] In this embodiment of the application, if the face recognition score is greater than the upper limit threshold or less than the lower limit threshold, it indicates that the face recognition result is clear and accurate. The terminal device can directly determine whether the user being tested is the owner of the device, which can improve the efficiency and accuracy of owner recognition.

[0040] In the eighth possible implementation of the first aspect, the data processing of the multimodal fusion host identification model may include the following steps: First, the IMU data to be tested, the touch screen data to be tested, and the breath sound features to be tested are encoded by a cross-modal encoder to obtain the cross-modal features to be tested; then, the cross-modal similarity between the cross-modal features to be tested and the cross-modal feature template in the host template is determined, as well as the breath sound similarity between the breath sound features to be tested and the breath sound feature template in the host template is determined; then, the multimodal identification score is determined based on the cross-modal similarity, the breath sound similarity, and the face recognition score, wherein the multimodal identification score is used to describe the comprehensive identification result of the user under multiple modalities; finally, the user under test is determined as the host based on the multimodal identification score and a pre-set fusion threshold.

[0041] In this embodiment, the multimodal fusion master recognition model employs a cross-modal encoder to perform feature-level fusion of IMU data, touchscreen data, and breath sound features. Then, it performs decision-level fusion of cross-modal similarity (or cross-modal recognition score), breath sound similarity (or breath recognition score), and face recognition score to obtain the multimodal recognition score. Stronger modalities, such as the face modality, do not participate in feature-level fusion. This reduces data coupling and effectively utilizes the complementary advantages of multiple modalities, improving the model's overall recognition success rate.

[0042] As an optional embodiment of this application, the multimodal fusion master recognition model uses a decision fusion model to fuse cross-modal similarity (or cross-modal recognition score), breath sound similarity (or breath recognition score), and face recognition score to obtain a multimodal recognition score.

[0043] The fusion decision model can include a Stem layer, a 12-layer encoding layer, a Head layer, and a Sigmoid function. The encoding layer includes a Linear layer, an activation function, and a Linear layer.

[0044] In this embodiment, the fusion decision model uses a Stem layer to expand the feature dimensions of cross-modal recognition scores, breathing recognition scores, and face recognition scores. These are then encoded using a 12-layer feedforward neural network with residual structures. Finally, a classification head layer and a sigmoid function are used to output the multimodal recognition scores. This decision fusion model employs a lightweight design, saving storage resources.

[0045] Optionally, if a master template is entered in the terminal device, the terminal device can use the feature templates in the master template, such as the breathing sound feature template and the cross-modal feature template, to determine the multimodal recognition score.

[0046] Optionally, if multiple owner templates are entered into the terminal device, the terminal device can select the owner template with the highest corresponding face recognition score to determine the multimodal recognition score. This can reduce the computational load consumed in the owner recognition process.

[0047] As an optional embodiment of this application, when multiple owner templates are entered in the terminal device, the terminal device can determine a multimodal recognition score for each owner template. Then, the terminal device can combine the obtained multiple multimodal recognition scores to determine whether the user being tested is the owner, and to determine which specific owner it is.

[0048] In this embodiment of the application, each host template is used to determine the corresponding multimodal recognition score, which can ensure the accuracy of host recognition.

[0049] As an optional embodiment of this application, when the master template has multiple feature templates, the terminal device can calculate the similarity between the average value of the multiple feature templates in the master template and the feature to be tested in the same modality (e.g., the feature of the breath sound to be tested) to obtain the recognition score of the corresponding modality.

[0050] As an optional embodiment of this application, when the host template has multiple feature templates, the terminal device can also perform similarity calculations between each feature template in the host template and the test feature of the same modality to obtain multiple similarities; then, the terminal device uses the obtained multiple similarities to calculate the recognition score of the corresponding modality. For example, the maximum value among the multiple similarities can be used as the recognition score of the corresponding modality.

[0051] In the ninth possible implementation of the first aspect, when multiple device owner templates are entered in the terminal device, there are multiple multimodal recognition scores. In this case, the terminal device determines whether the user to be tested is the device owner based on the multimodal recognition scores and a pre-set fusion threshold. This can include: selecting the largest of the multiple multimodal recognition scores; if the selected multimodal recognition score is greater than the fusion threshold, then the user to be tested is determined to be the device owner.

[0052] In this embodiment, when multiple owner templates are entered into the terminal device, the terminal device can obtain a multimodal recognition score for each owner template. Generally, the higher the multimodal recognition score, the more similar the user being tested is to the owner indicated by the owner template corresponding to that multimodal recognition score. In other words, using the highest multimodal recognition score for owner identification can improve the accuracy and efficiency of owner identification.

[0053] Optionally, if multiple owner templates are entered in the terminal device, there are multiple multimodal recognition scores. For each multimodal recognition score, the terminal device can compare the multimodal recognition score with a pre-set fusion threshold. If the multimodal recognition score is greater than the fusion threshold, the user to be tested is considered to be the owner, specifically the owner indicated by the owner template corresponding to the multimodal recognition score.

[0054] In the tenth possible implementation of the first aspect, the process of training the multimodal fusion machine master recognition model may include the following steps: First, select m batches of training samples from the sample dataset used to train the multimodal fusion machine master recognition model. Each batch of training samples includes one positive sample and n negative samples, where n>1 and m>1. Then, during the training of the multimodal fusion machine master recognition model using each batch of training samples, the negative samples that are identified as positive samples and have the highest corresponding recognition probability are identified as hard samples. Finally, determine the model training loss of the multimodal fusion machine master recognition model based on the positive samples and hard samples, and adjust the model parameters of the multimodal fusion machine master recognition model based on the determined model training loss.

[0055] In this embodiment, during the training of the multimodal fusion machine owner recognition model, m training samples can be collected for the same sample user. That is, m sets of training data for each sample user are used for model training, instead of just one set of data, which can improve the accuracy of the model. In addition, in each batch of training samples, difficult samples are identified as negative samples, and the model training loss is calculated together with the dynamic difficult samples and positive samples. The model parameters are then updated based on the model training loss, which can effectively reduce the false recognition rate and improve the TAR index at a lower false recognition rate. In other words, the trained multimodal fusion machine owner recognition model can be more accurate and more stable in machine owner recognition.

[0056] In the eleventh possible implementation of the first aspect, the terminal device may further perform the following operations: when the template update conditions are met, the average value of the breathing sound feature template in the main template and the current breathing sound feature is determined as the new breathing sound feature template, and the new breathing sound feature template is added to the main template, wherein the template update conditions include: successful face unlock and reaching the update cycle.

[0057] Among them, the current breathing sound feature is the breathing sound feature in the breathing audio data collected when face unlock is successful.

[0058] In this embodiment, the terminal device uses the average of the current feature and the old feature template as the new feature template, which can accurately and stably update the host template, thus helping to improve the accuracy and stability of host identification.

[0059] In the twelfth possible implementation of the first aspect, the terminal device may also perform the following operations: if the template update conditions are met, the terminal device may further determine the average value of the cross-modal feature template in the host template and the current cross-modal feature as the new cross-modal feature template, and add the new cross-modal feature template to the host template.

[0060] Among them, the current cross-modal features are the cross-modal features corresponding to the multimodal data collected when face unlock is successful.

[0061] In this embodiment, the terminal device uses the average of the current feature and the old feature template as the new feature template, which can accurately and stably update the host template, thus helping to improve the accuracy and stability of host identification.

[0062] In some optional implementations of this application's embodiments, before updating the feature template in the owner template, the terminal device can first perform a feature quality assessment on the owner's current features, and determine whether to update the feature template based on the feature quality assessment result. For example, if the feature quality assessment result is poor, the owner's current features can be deleted. Alternatively, if the feature quality assessment result is good, the current features can be used to update the owner template.

[0063] Secondly, embodiments of this application provide a device for identifying the owner of a device, applied to a terminal device, wherein the device includes:

[0064] The data acquisition unit is used to collect user data of the user under test in response to the unlocking operation of the target content by the user under test. The target content is in a locked state. The user data includes at least one of the following: the face image of the user under test, the breathing audio data of the user under test, the IMU data of the user under test, and the touch screen data of the user under test.

[0065] The recognition execution unit is used to perform owner recognition based on user data and pre-entered owner templates. The owner template includes at least one of the following: face template, breathing sound feature template, and cross-modal feature template. The cross-modal feature template is used to describe the comprehensive features of the owner in multiple modalities.

[0066] The unlock execution unit is used to unlock target content when the owner is identified.

[0067] As an embodiment of this application, the owner identification device can implement the method as described in any of the first aspects above.

[0068] Thirdly, embodiments of this application provide a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method as described in any of the first aspects above.

[0069] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method described in any of the first aspects above.

[0070] Fifthly, embodiments of this application provide a chip system including a processor coupled to a memory. The processor executes a computer program stored in the memory to implement the method described in any of the first aspects above. The chip system may be a single chip or a chip module composed of multiple chips.

[0071] Sixthly, embodiments of this application provide a computer program product that, when run on a terminal device, causes the terminal device to execute the method described in any of the first aspects above.

[0072] It is understood that the beneficial effects of the second to sixth aspects mentioned above can be found in the relevant descriptions in the first aspect mentioned above, and will not be repeated here. Attached Figure Description

[0073] Figure 1A This is a schematic diagram illustrating an information display effect in an inorganic master identification scenario provided in an embodiment of this application.

[0074] Figure 1B A schematic diagram illustrating an information display effect in an organic main recognition scenario provided in an embodiment of this application;

[0075] Figure 1C This is a schematic diagram illustrating another information display effect in an inorganic master identification scenario provided by an embodiment of this application;

[0076] Figure 1D This is a schematic diagram illustrating another information display effect in an organic main recognition scenario provided by an embodiment of this application;

[0077] Figure 2 A schematic diagram illustrating a scenario for the multimodal fusion-based owner identification method provided in this application embodiment;

[0078] Figure 3This is a schematic diagram illustrating the process of inputting a face template as provided in an embodiment of this application;

[0079] Figure 4A This is a schematic diagram illustrating the process by which the audio processing module provided in this embodiment processes respiratory audio data.

[0080] Figure 4B A schematic diagram illustrating the process by which the audio processing module provided in this embodiment performs audio preprocessing on respiratory audio data;

[0081] Figure 4C A schematic diagram of the breath sound recognition model provided in the embodiments of this application;

[0082] Figure 4D A schematic diagram of the respiratory sound feature extraction model provided in the embodiments of this application;

[0083] Figure 4E A schematic diagram of a multi-scale residual layer in the respiratory sound feature extraction model provided in this application embodiment;

[0084] Figure 5A This is a schematic diagram of the structure of a cross-modal encoder provided in an embodiment of this application;

[0085] Figure 5B This is a schematic diagram of the structure of a cross-channel linear encoder in a cross-modal encoder provided in an embodiment of this application;

[0086] Figure 5C This is a schematic diagram of the structure of the cross-channel attention encoder in the cross-modal encoder provided in the embodiments of this application;

[0087] Figure 6A A schematic diagram illustrating the process of mobile phone owner identification based on multimodal fusion provided in this application embodiment;

[0088] Figure 6B This is a schematic diagram of the process of the host device based on breath sound recognition provided in the embodiments of this application;

[0089] Figure 6C A schematic diagram illustrating the process of host identification using the multimodal fusion host identification model provided in this application embodiment;

[0090] Figure 6D This is a schematic diagram of the structure of the fusion decision model provided in the embodiments of this application;

[0091] Figure 6E The model training architecture for positive and negative sample pairing based on dynamic hard sample mining provided in the embodiments of this application;

[0092] Figure 7 ROC curves for the various modalities provided in the embodiments of this application, corresponding to the host identification.

[0093] Figure 8 The PR curves for the various modalities provided in the embodiments of this application are for the identification of the host.

[0094] Figure 9 This is a schematic diagram of the owner identification device provided in the embodiments of this application;

[0095] Figure 10 This is a schematic diagram of the structure of the terminal device provided in the embodiments of this application;

[0096] Figure 11 A software structure block diagram of a terminal device provided in an embodiment of this application. Detailed Implementation

[0097] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0098] The following explains some concepts that may be involved in the embodiments of this application:

[0099] (1) Multiple: Unless otherwise specified, in the embodiments of this application, multiple refers to two or more.

[0100] (2) Difficult sample: In this embodiment of the application, a difficult sample refers to a negative sample that is misidentified as a positive sample by the model during the model training process and has the highest corresponding recognition probability.

[0101] (3) False Acceptance Rate (FAR): In this embodiment, FAR refers to the proportion of samples incorrectly classified into other categories in a classification problem. In the fields of machine learning and pattern recognition, we typically focus on two types of errors: false positives and false negatives. A false positive is when a negative sample is incorrectly classified as a positive sample, while a false negative is when a positive sample is incorrectly classified as a negative sample. The false acceptance rate can be calculated using the following formula: False Acceptance Rate = (Number of False Positive Samples + Number of False Negative Samples) / Total Number of Samples.

[0102] In practice, the lower the false recognition rate, the better the classifier's performance. In practical applications, we usually weigh the weights of false positives and false negatives according to specific needs, and choose an appropriate classifier or adjust the model parameters to achieve the optimal false recognition rate.

[0103] (4) Receiver Operating Characteristic (ROC) curve: In this embodiment, the ROC curve is plotted with the True Positive Rate (TPR) on the vertical axis and the False Positive Rate (FPR) on the horizontal axis. It is a visualization tool for evaluating the performance of a binary classification model at different thresholds.

[0104] When plotting an ROC curve, the samples need to be sorted according to the output probability or decision value of the classification model. Then, starting from the lowest threshold, the threshold is gradually increased, classifying samples above the current threshold as positive examples and samples below the current threshold as negative examples. At each threshold, the corresponding true positive rate and false positive rate are calculated and plotted as coordinates. The final curve obtained is the ROC curve.

[0105] (5) Area Under the Curve (AUC): In this embodiment, AUC usually refers to the area under the ROC curve. AUC is an indicator used to evaluate the performance of a binary classification model. The value of AUC ranges from 0.5 to 1, and the larger the value, the better the model performance.

[0106] (6) Precision-Recall Curve (PR): In this embodiment, the PR curve is another visualization tool for evaluating the performance of a binary classification model. Unlike the ROC curve, the PR curve is plotted with precision on the vertical axis and recall on the horizontal axis.

[0107] When plotting the precision (PR) curve, samples need to be sorted based on the model's output scores. Then, starting from the lowest threshold, the threshold is progressively increased, classifying samples above the current threshold as positive examples and those below as negative examples. At each threshold, the corresponding precision and recall are calculated and plotted on a coordinate system. The resulting curve is the PR curve. Each point on the PR curve represents the precision at different recall rates. The closer the PR curve is to the upper right corner, the better the classifier's performance across various scenarios.

[0108] (7) Area under the PR curve (AUC-PR, AP): In this embodiment, AP is an indicator used to comprehensively evaluate the performance of the model. In practice, a larger AP value usually indicates better classification performance of the model.

[0109] (8) Service Card: In this embodiment, a service card is a service format that provides a more granular, atomic service capability than an application (APP). It directly displays the services or content that users care about most in an interactive service card format. Service cards can be embedded in various APPs or interactive scenarios to better meet user needs. A service card is generally a page within an application or its corresponding APP. Clicking a service card will take the user to the corresponding application or a specific service page within the APP. Unless otherwise specified, the term "application" in this embodiment can refer to an application or its corresponding APP.

[0110] The service card generation process is as follows: Application service providers upload the service content they want to push to users to their servers. The server generates corresponding service cards based on the uploaded service content and pushes these service cards to the user's terminal device for display. This allows the user's terminal device to present the services pushed by the application service provider to the user in the form of service cards. Generally, one service card corresponds to a portion of the application's service functions. Application providers can break down the application's services into multiple service modules, each corresponding to a service card, and then push different service cards to users based on different service scenarios.

[0111] (9) TAR, TAR@FAR: In this embodiment of the application, the True Accept Rate (TAR) is used to indicate the proportion of correct judgments. TAR@FAR refers to the value of TAR under a specific FAR. For example, TAR@FAR = 1% means the value of TAR when FAR = 1%.

[0112] (10) Neural network layers and activation functions: The neural network layers involved in the embodiments of this application include Linear layer, Group layer, norm layer, Inputs layer, Features layer, conv layer, Stem layer, Head layer, etc.

[0113] The Linear layer, also known as a fully connected layer or a dense layer, is typically used to perform a linear transformation on the input data to obtain the output.

[0114] Group layers are typically used to group input data, with each group being computed independently. For example, grouped data can be fed into different convolutional kernels for convolution operations. This grouping improves computational efficiency and the model's expressive power. Through grouping, group layers can increase the model's non-linear characteristics and enhance its ability to learn complex patterns. Furthermore, group layers can reduce the number of parameters, lower computational costs, and improve the model's generalization ability.

[0115] The norm layer, often called a regularization layer or normalization layer, is typically used to normalize the input data. The norm layer aims to adjust the distribution of the input data to reduce redundancy and correlation, thereby improving the robustness and generalization ability of the network to better adapt to the training or inference process of the network model.

[0116] The Inputs layer, often referred to as the input layer, receives external input data and transforms it into a tensor form that the neural network can process. The Inputs layer is typically the first layer of the network structure and defines the format, dimensions, and properties of the input data.

[0117] Features layers, often referred to as feature layers, are intermediate or hidden layers in neural networks used to extract discriminative features from input data. They can capture key information from the input data through techniques such as convolution, pooling, and embedding, providing richer and more meaningful feature representations for subsequent network layers.

[0118] A convolutional layer (conv) extracts local features from the input data by applying convolutional operations and generates feature maps as output. Convolutional layers are widely used in fields such as computer vision and speech processing.

[0119] Stem layers typically extract richer features by performing spatial transformations and local feature fusion on the feature map. In this embodiment, the Stem layer can be used to expand the feature dimension.

[0120] The Head layer is typically used to describe the last or top layer of a model. The main function of the Head layer is to transform the preceding feature representations into the final output.

[0121] The activation functions involved include the Sigmoid function, the softmax function, and the ReLU function.

[0122] The sigmoid function is a common non-linear activation function that maps real-valued inputs to a value between 0 and 1 for its output. The sigmoid function is typically used in the output layer of binary classification problems to transform the model's output into probability values.

[0123] The softmax function is a commonly used activation function that transforms an input real-valued vector into a probability distribution such that the sum of the probabilities of all outputs equals 1. The softmax function is typically used in the output layer of multi-class classification problems, where it transforms the model's output into the probability of each class for the final classification decision.

[0124] The Rectified Linear Unit (ReLU), also known as the Modified Linear Unit, is a commonly used activation function.

[0125] (11) Inertial Measurement Unit (IMU): In this embodiment, an IMU is a sensor used to measure and track the attitude (such as orientation, angle, and velocity) of an object. An IMU contains various measurement units, including an accelerometer, a gyroscope, and a magnetometer. The accelerometer measures the linear acceleration of an object along the x, y, and z axes; the gyroscope measures the angular velocity of an object around the x, y, and z axes; and the magnetometer measures the magnetic field strength and direction of the object in the Earth's magnetic field.

[0126] In this embodiment, the IMU data may include the terminal device's acceleration, pitch angle, roll angle, etc. in various directions. The IMU data is the data collected by the IMU in the terminal device.

[0127] (12) Multimodal: In the embodiments of this application, multimodal usually refers to multiple types. Multimodal data usually refers to multiple different types of data. For example, face data, breathing audio data, touch screen data, and IMU data are four different types of data.

[0128] (13) Audio, time-frequency graph, Mel spectrum, logarithmic Mel spectrum, Mel cepstral, Mel cepstral coefficients: In this embodiment, a short-time Fourier transform is performed on the audio to convert it into a time-frequency graph. A Mel scale transform is performed on the time-frequency graph to obtain the corresponding Mel spectrum. A logarithmic transform is performed on the Mel spectrum to obtain the corresponding logarithmic Mel spectrum, also known as the Mel cepstral. A discrete cosine transform is performed on the Mel cepstral to obtain the corresponding Mel cepstral coefficients.

[0129] Terminal devices, especially mobile phones, are gradually becoming the center of personal digital worlds, carrying a large amount of personal information and privacy data. During use, these devices may display the owner's private information, which, if seen by someone other than the owner, undoubtedly leads to a privacy breach.

[0130] The following uses a mobile phone as an example to list several usage scenarios that may leak the owner's private information. It is understood that the terminal device in this application embodiment is not limited to a mobile phone, but can also be other devices, such as tablets, laptops, etc.

[0131] Use Case 1: The phone supports a service card function. When the phone displays a service card, if the card contains the owner's private information such as bank card number, password, or package pickup code, then if the user next to the phone is not the owner or there is not only the owner, the owner's private information displayed on the service card may be intentionally or unintentionally seen by others, leading to the leakage of the owner's private information.

[0132] Use Case Two: The phone owner shares content with others, such as when using their phone to report to a company leader or present materials, or when watching videos with friends. Because some apps on the phone, such as WeChat, often preview notification messages at the top of the screen, if the phone receives a notification message from an app and previews it, the content of the notification message may be seen by others, potentially leading to the leakage of the phone owner's private information and an embarrassing situation.

[0133] Use Case 3: The user is using their phone in crowded places, such as in an elevator, subway, or bus, reading news or watching short videos. In this scenario, if the phone receives a notification from an app and previews it, the content of the notification may be intentionally or unintentionally seen by others, leading to the leakage of the user's private information.

[0134] Use Case 4: The phone owner's phone is borrowed by someone else, for example, a friend borrows the owner's phone to watch short videos when the owner is not present. In this scenario, if the phone receives a notification message from an app and previews it, the content of the notification message may be intentionally or unintentionally seen by others, leading to the leakage of the owner's private information.

[0135] Based on the above usage scenarios, it is clear that timely identification of the device owner during the use of the terminal device is essential to protect the security of the owner's private information.

[0136] The following further illustrates the purpose or potential effects of the owner identification scheme in this application embodiment by comparing specific usage scenarios.

[0137] Figure 1A This diagram illustrates an information display effect in an inorganic master identification scenario. Figure 1A In the image, the phone owner, A, is looking at their phone. Other people are nearby, such as B and C. The phone directly displays the service card from the Notes app. (Example:) Figure 1A As shown, the service card contains A's bank card number and password. At this time, B and C can see the contents of the service card, that is, other people can see the owner's private information.

[0138] Figure 1B This diagram illustrates an information display effect in an organic master recognition scenario. Figure 1B In this scenario, the phone owner, A, is looking at their phone while other people, such as B and C, are nearby. Without recognizing the phone owner, the phone locks and hides the contents of the service card. This protects the security of the phone owner's private information.

[0139] Understandable, such as Figure 1B As shown, if A wants to view the contents of the service card, they can click further. Icons that the phone can respond to when the user clicks. The icon displays the service card information after the owner is identified.

[0140] Figure 1C This diagram illustrates another information display effect for scenarios without inorganic master identification. Figure 1C In the scenario, the phone owner, A, is looking at their phone while other people, such as B and C, are nearby. The phone receives a notification from a messaging app and previews the message. At this moment, A and the other users (B and C) can all see the content of the notification, meaning they can access the phone owner's private information.

[0141] Figure 1D This diagram illustrates another information display effect for organic master recognition scenarios. Figure 1D In the scenario, the phone owner, A, is looking at their phone. Other people are nearby, such as B and C. The phone receives a notification from a messaging app, but without recognizing the owner, it automatically hides the notification content. Understandably, if A wants to view the notification, they can tap the hidden notification. The phone should respond to this tap, and after recognizing the owner, display the notification content.

[0142] In combination with the above Figures 1A-1DThe comparison of information display in the two usage scenarios clearly shows that timely identification of the device owner during the use of the terminal device can effectively protect the security of the owner's private information.

[0143] In fact, by identifying the owner of the device and treating the owner and non-owner differently, the terminal device can not only protect the owner's private information from being leaked, but also provide the owner with a more convenient device experience in some application scenarios. For example, the owner can make payments without a password, while non-owners must enter a password to make payments.

[0144] Currently, the most commonly used methods for owner identification are facial recognition or fingerprint recognition. However, existing facial recognition solutions suffer from reduced recognition rates in scenarios such as low light, backlight, strong light, and varying distances and angles due to poor facial image quality. Similarly, existing fingerprint recognition solutions experience reduced recognition rates when fingers or the fingerprint sensor itself is dirty (e.g., sweat, water, dust, or foreign objects). In other words, as the environment changes—for example, with dimmer or brighter lighting—the recognition rate of existing owner identification solutions may decrease, or they may even fail to correctly identify the owner.

[0145] In view of the aforementioned technical problems in related technologies, in order to improve the accuracy and stability of the terminal device's owner identification, in this embodiment of the application, the terminal device combines the user's facial data, the terminal device's IMU, the terminal device's touchscreen data, and the user's breathing audio data to comprehensively identify the owner, thereby effectively improving the success rate of owner identification. That is, this embodiment of the application can use multimodal data to identify the owner.

[0146] Figure 2 This is a schematic diagram of a scenario for the multimodal fusion-based owner identification method provided in an embodiment of this application. Figure 2 In this system, the terminal device is a mobile phone. The phone can combine the current user's IMU data, touchscreen data, breathing audio data, and facial recognition score to comprehensively determine whether the current user is the owner. If the phone identifies the current user as the owner, it will display the privacy information normally. If the phone identifies the current user as not the owner, it will hide the privacy information.

[0147] The embodiments of this application have at least the following beneficial effects:

[0148] 1. The terminal device combines user facial data, terminal device IMU data, terminal device touch screen data, and user breathing audio data to comprehensively identify the owner. This allows for owner identification even when the scene changes, improving the accuracy and stability of owner identification.

[0149] 2. If the facial recognition result is not ideal, at least one of the following can be combined: IMU data from the terminal device, touch screen data from the terminal device, and user breathing audio data, to further determine whether the user is the owner of the device. This can expand the owner recognition scenario and improve the accuracy and stability of owner recognition.

[0150] 3. The IMU data and touchscreen data of the terminal device reflect the user's usage habits of the terminal device. Combining these two data to further determine whether the user is the owner can make the identification results more accurate and reliable, that is, it can improve the accuracy of owner identification.

[0151] The embodiments of this application can be applied to scenarios involving the identification of the device owner in terminal devices.

[0152] The device owner identification method provided in this application embodiment can be applied to terminal devices such as mobile phones, tablets, and wearable devices. In this case, the terminal device is the executing subject of the device owner identification method provided in this application embodiment. This application embodiment does not impose specific restrictions on the specific type of terminal device.

[0153] The following uses a mobile phone as an example to illustrate the owner identification method provided in this application from three aspects: owner template input, owner identification based on the owner template, and owner template update.

[0154] (I) Inputting the client template

[0155] Owner templates are typically pre-recorded or stored information describing the characteristics of the owner. Owner templates can include face templates, breathing sound feature templates, and cross-modal feature templates. Face templates are pre-recorded or stored information describing the facial features of the owner, such as an image of the owner's face. Breathing sound feature templates are pre-recorded or stored information describing the breathing sound characteristics of the owner. Cross-modal feature templates are pre-recorded or stored information describing the combined characteristics of the owner across multiple modalities.

[0156] In practice, a mobile phone can have one owner's template or multiple owner's templates. That is to say, a mobile phone can have one owner's template or multiple owner's templates.

[0157] In practice, one host template usually corresponds to one host.

[0158] Optionally, a host template may contain one or more face templates, one or more breathing sound feature templates, and one or more cross-modal feature templates of the host.

[0159] In this application embodiment, the process of entering the host template will be described from the following aspects: 1. Face template entry; 2. Breath sound feature template entry; 3. Cross-modal feature template entry.

[0160] 1. Face template input.

[0161] In this embodiment, the phone owner can register the faces of multiple people sequentially. For example, they can register their own face first, and then register the faces of their father and mother. For instance, owner A can register their own face, and then register the faces of their father B and mother C. When A only registers their own face, there is only one owner ID, and the phone can retain one owner identifier (ID), such as A. If A registers their own face and then registers the faces of their father B and mother C, the phone will retain three owner IDs, such as A, B, and C. It is understood that this embodiment does not limit the representation of the owner ID; for example, it can be Chinese characters, letters, numbers, or a combination of numbers and letters, etc.

[0162] As is understandable, in this embodiment of the application, face registration refers to face template registration.

[0163] Please see Figure 3 , Figure 3 The diagram illustrates the process of inputting a facial template. Figure 3 In (a), the phone owner, A, can click on the biometrics and password control 302 in the settings interface 301 of the settings application (an application specifically used for performing system settings functions). The phone can respond to the user clicking the biometrics and password control 302 and enter the biometrics and password interface 303. Afterwards, as... Figure 3 As shown in (b), the phone can respond to A's further click on the "face recognition" control 304 to enter the lock screen password input interface 305. Then, as... Figure 3 As shown in (c), the phone can respond to A by further entering the correct lock screen password and enter the face recognition interface 306. Then, as... Figure 3 As shown in (d), the mobile phone can respond to A's further click on the "Start Enrollment" control 307 to enter the face enrollment interface 308. For example... Figure 3 As shown in (e), on the face registration interface 308, Xiao A aligns his face with the face capture frame, and the mobile phone captures Xiao A's face image, completing the face template registration for Xiao A. Figure 3 As shown in (f), when the mobile phone completes the face template input for Xiao A, it can display "Input successful" on the face input interface 308.

[0164] According to such Figure 3 In a similar way, Xiao A can help her parents register their faces on her phone, thus completing the registration of her parents' facial templates on the phone.

[0165] In this embodiment, after multiple face templates are entered, the mobile phone can identify the owner corresponding to each face template. For example, if the phone stores face templates for user A, user B, and user C, when user A uses the phone to unlock it with their face, the phone can identify that the person operating the phone is the owner, specifically user A. In other words, the phone can determine which owner is operating the phone. Of course, any owner can unlock the phone with their face.

[0166] In some application scenarios, a single user can register their face once or multiple times. For example, a single user can register their face multiple times, and each face registration can be assigned a corresponding owner ID by the phone. That is, a single user can have multiple owner IDs. Alternatively, a user can correspond to multiple owner templates. When a user is assigned multiple owner IDs, the phone can treat that user as having multiple owners.

[0167] 2. Input of breath sound feature template.

[0168] In this embodiment, when the user's face template is successfully registered, the mobile phone can continue to collect the user's breathing audio data through the microphone, and determine whether the breathing audio data contains breathing sounds. If it contains breathing sounds, it further extracts breathing sound features from the data and identifies the extracted breathing sound features as the user's breathing sound feature template. In other words, the trigger condition for registering the breathing sound feature template (or the breathing feature registration condition) is that the user's face template is successfully registered.

[0169] It should be noted that within a short period of time after successful face template registration, such as within 5 seconds, the user who just registered their face template usually does not immediately put down or leave the phone. In other words, the breathing audio data collected within this short period of successful face template registration usually contains the user's breathing sounds. Therefore, if the phone continues to collect breathing audio data within this short period of successful face template registration while the user's face template is being registered, it can ensure that the collected breathing audio data contains the user's breathing sounds, thereby ensuring the accurate and effective registration of the user's breathing sound feature template.

[0170] The aforementioned "shorter duration" is typically a pre-set short duration value. In practice, this "shorter duration" can be any duration between 0.5 seconds and 5 seconds. Since longer audio data provides more accurate audio characteristics, to ensure the accuracy of the recorded breathing sound feature template for the device owner, breathing audio data is usually collected within 5 seconds of successful facial template recording. In other words, a relatively long period of breathing audio data is collected to ensure the accuracy of the device owner's breathing sound feature template.

[0171] Understandably, if the collected respiratory audio data does not contain respiratory sounds, the mobile phone may not record the user's respiratory sound feature template.

[0172] In some application scenarios, if the phone detects that the breathing audio data does not contain breathing sounds, it can continue to collect breathing audio data and determine whether the breathing audio data contains breathing sounds after a successful face unlock. If breathing sounds are present, breathing sound features are extracted and identified as the owner's breathing sound feature template. In other words, the trigger condition for recording the breathing sound feature template (or breathing feature recording condition) can also be a successful face unlock, provided that the phone has not yet recorded the breathing sound feature template of the currently unlocking owner at the time of the successful face unlock.

[0173] In general, the triggering conditions for recording the breathing sound feature template (or breathing feature recording conditions) can be the successful recording of the owner's face template or the successful unlocking of the face, provided that the phone has not yet recorded the breathing sound feature template of the owner who is currently unlocking the phone when the face unlock is successful.

[0174] Regardless of the triggering conditions under which respiratory audio data is acquired, the acquisition parameters used to obtain the respiratory sound feature template can be fixed. For example, the acquisition parameters can be: single-channel acquisition, sampling rate of 16kHz, and acquisition duration of 5 seconds.

[0175] It should be noted that since breathing sounds are usually quite faint, if the environment in which the user is registering their face is complex, such as a noisy scene like a concert or a market, the phone may struggle to recognize breathing sounds from the further collected breathing audio data. Therefore, it is crucial to first determine whether breathing sounds are present in the breathing audio data. Only when breathing sounds are confirmed to be present should breathing sound feature extraction be performed, thus reducing unnecessary consumption of computational resources.

[0176] The following will further explain how to extract the corresponding breathing sound feature template from the owner's breathing audio data.

[0177] In this embodiment of the application, during the process of recording the breathing sound feature template, the mobile phone can call the pre-stored audio processing module, which extracts breathing sound features from the collected breathing audio data of the owner. Then, the mobile phone can use the extracted breathing sound features as the breathing sound feature template of the owner.

[0178] Further reference Figure 4A . Figure 4A This is a schematic diagram illustrating the process by which the audio processing module provided in this embodiment processes respiratory audio data. The audio processing module can extract respiratory sound features from the user's respiratory audio data by executing steps S401-S405. Details are as follows:

[0179] S401, the audio processing module performs audio preprocessing on the breathing audio data to obtain the Mel-Cepstral and Mel-Cepstral coefficients of the breathing audio data.

[0180] Figure 4B This diagram illustrates the audio preprocessing process of the audio processing module on breathing audio data. Figure 4B As shown, the audio preprocessing process performed by the audio processing module on the breathing audio data may include the following steps S4011-S4014.

[0181] S4011, the audio processing module performs a short-time Fourier transform on the breathing audio data to convert the breathing audio data into a time-frequency diagram.

[0182] In this embodiment, the collected respiratory audio data is typically a time-domain audio signal. Therefore, a short-time Fourier transform can be used to convert the time-domain audio signal into a two-dimensional time-frequency signal. Then, the amplitude of the two-dimensional time-frequency signal can be combined to convert it into a time-frequency graph.

[0183] The Short-Time Fourier Transform (STFT) is a classic signal analysis method that decomposes a signal into time and frequency components. STFT performs localized analysis of the signal using a moving window of a certain length, improving the resolution of both time and frequency domain analysis. The main parameters of STFT include the window function, window length, and step size. Commonly used window functions include the Hanning window, Hamming window, and Blackman window.

[0184] In some applications, the window function used for short-time Fourier transforms can be the Hanning window, with a window length of 0.025 seconds and a step size of 0.01 seconds. The number of time-domain points in the time-frequency signal can be represented by num_samples, where num_samples = sampling duration ÷ step size + 1. In this application, num_samples can take the value 501, where 501 = 5 ÷ 0.01 + 1.

[0185] S4012, the audio processing module performs frequency-Mel scale transformation on the time-frequency graph to convert the frequency of the time-frequency graph into Mel scale, thereby obtaining the Mel spectrum.

[0186] Here, the mapping formula between frequency and Mel scale is as follows:

[0187]

[0188] Where f is the frequency, mel(f) is the Mel scale, and × is the multiplication sign.

[0189] In some applications, the parameters involved in converting a time-frequency graph into a Mel spectrum may include: a minimum frequency of 125.0 Hz, a maximum frequency of 7500.0 Hz, and a number of Mel filter banks of 64. Here, 64 Mel filter banks indicates a relatively large number. It should be noted that a larger number of Mel filter banks results in a more accurate Mel spectrum, helping to refine low-frequency information and coarsen high-frequency information in audio signals.

[0190] S4013, the audio processing module performs a logarithmic transformation on the Mel spectrum to obtain the logarithmic Mel spectrum (also known as the Mel cepstral spectrum).

[0191] Log mel-specgram is an audio feature representation method that converts the spectral information of an audio signal into an energy distribution on a mel scale and takes a logarithm to enhance the discriminative power of the features.

[0192] Here, the logarithm of the Mel spectrum in S4012 is taken to obtain the logarithmic Mel spectrum. The size of the obtained logarithmic Mel spectrum can be num_samples × n_mels, where num_samples is the number of time-domain points in the time-frequency signal, and n_mels is the number of Mel filter banks.

[0193] S4014, the audio processing module performs discrete cosine transform on the Mel cepstral spectrum to obtain the Mel cepstral coefficients.

[0194] Mel-frequency cepstral coefficients (MFCCs) are feature coefficients obtained through a series of signal processing steps, used to describe the energy distribution of sound at different frequencies. MFCCs are widely used in speech recognition and audio feature extraction, effectively capturing important features of audio signals while reducing feature dimensionality, thus improving the accuracy and efficiency of pattern recognition.

[0195] Here, the Mel-frequency cepstral spectrum obtained in step S4013 is subjected to a Discrete Cosine Transform (DCT), and the first 40 MFCC coefficients are retained, i.e., n_mfcc = 40, to obtain the required MFCC coefficients. At this point, the size of the MFCC coefficients is num_samples × n_mfcc, where num_samples is the number of time-domain points in the time-frequency signal.

[0196] It should be noted that the larger the value of n_mfcc, the more information the MFCC coefficient retains, which helps to achieve more accurate extraction of breath sound features.

[0197] It is understood that the embodiments of this application do not specifically limit the values ​​of the parameters involved in audio preprocessing. For example, in some scenarios, the number of Mel filter banks can be other than 64. In other scenarios, the number of MFCC coefficients retained can be other than 40. The values ​​of the above parameters can be adjusted according to actual needs.

[0198] S402, the audio processing module divides the Mel cepstrum into multiple Mel cepstrum segments.

[0199] Here, since the complete Mel cepstrum usually involves a large amount of data, dividing the Mel cepstrum into multiple Mel cepstrum segments can simplify feature representation, reduce computational costs, and more effectively capture local features of the audio signal, thereby improving the accuracy of breath sound recognition.

[0200] In practical applications, it has been found that if the segmentation granularity is too fine, it will increase the computational load and reduce the accuracy of breath sound recognition. Therefore, in practice, the length of the Mel-Cepstral segment is usually set to about 1 second, for example, 0.96 seconds. The reason is that the duration of a person's exhalation or inhalation is approximately 1 second. Switching the Mel-Cepstral segment to a length of about 1 second allows each Mel-Cepstral segment to describe the breathing characteristics more completely, thereby further improving the accuracy of breath sound recognition.

[0201] In practice, the Mel cepstrum can be segmented with a window length of 0.96 seconds and a step size of 0.24 seconds to divide the Mel cepstrum into multiple Mel cepstrum segments.

[0202] When the Mel cepstrum cannot be divided into an integer number of segments, that is, when it cannot be divided into an integer number of 0.96-second segments, zeros can be added to the Mel cepstrum. The zero-added Mel cepstrum will then be divided into an integer number of Mel cepstrum segments.

[0203] Furthermore, since the input size of a breath sound recognition model is usually constant—for example, the input size could be the number of Mel-frequency cepstrum segments corresponding to 5 seconds of breath sounds—in some applications, 5 seconds of audio data may not have been collected during breath sound data acquisition; for example, only 1 second of audio data might have been collected. In order to ensure sufficient data input to the breath sound recognition model, it is necessary to further pad the Mel-frequency cepstrum with zeros.

[0204] S403, the audio processing module inputs the Mel cepstral segment into the pre-trained breath sound recognition model to obtain the breath sound confidence level that the breath audio data is recognized as a breath sound.

[0205] The aforementioned breath sound confidence score refers to the confidence or probability that a breath audio data point is identified as a breath sound. The input to the breath sound recognition model consists of multiple consecutive Mel-frequency cepstral segments, and the output includes at least the breath sound confidence score.

[0206] Here, the Mel-Cepstral segment of the respiratory audio data is input into the respiratory sound recognition model. The respiratory sound recognition model can output the confidence or probability that the respiratory audio data is recognized as a respiratory sound, thereby further determining whether the respiratory audio data contains a respiratory sound.

[0207] It should be noted that since Mel cepstrum usually carries a wealth of audio information, identifying whether an audio is a breath sound by using Mel cepstrum fragments can ensure the accuracy of breath sound recognition. In other words, it can enable the breath sound recognition model to more accurately determine whether an audio is a breath sound.

[0208] S404, the audio processing module combines the confidence level of the breath sound with the preset threshold a to further determine whether the breath audio data is a breath sound.

[0209] Here, threshold 'a' is a pre-set threshold used to determine whether the audio is a breathing sound. In some usage scenarios of this application embodiment, threshold 'a' can also be referred to as the first determination threshold.

[0210] Here, if the confidence level of breath sounds is greater than the threshold 'a', the breath audio data is considered to contain breath sounds. Conversely, if the confidence level of breath sounds is less than or equal to the threshold 'a', the breath audio data is considered not to contain breath sounds.

[0211] S405, when the breathing audio data is determined to be breathing sounds, the audio processing module inputs the Mel-frequency cepstral coefficients from S401 into the pre-trained breathing sound feature extraction model to extract the breathing sound features of the breathing audio data.

[0212] Here, during the process of recording the owner's breathing sound feature template, the mobile phone can record the breathing sound features extracted from S405 as the owner's breathing sound feature template.

[0213] It should be noted that Mel-Cepstral Coefficients can effectively capture important features of audio signals and have a small data dimension. When the respiratory audio data has been determined to be respiratory sounds, the Mel-Cepstral Coefficients of the respiratory audio data can be input into the respiratory sound feature extraction model to extract the respiratory sound features of the respiratory audio data, which can achieve fast and accurate extraction of the respiratory sound features of the respiratory audio data.

[0214] The following further combines Figure 4C This will explain the breath sound recognition model in S403 and the training process of the breath sound recognition model.

[0215] In this embodiment, to increase the generalization of the breath sound recognition model, in addition to breath sound categories, other common human-related sound events are added, resulting in seven sound categories: quiet, speaking, breath sounds, coughing, sneezing, snoring, and environmental noise. When the breath sound recognition model can identify all seven sound categories, its output can be the confidence score for each of these seven sound categories. Specifically, one confidence score is output for each sound category; for seven sound categories, seven confidence scores are output. Generally, the sum of the confidence scores for the seven sound categories output by the breath sound recognition model is usually 1.

[0216] Understandably, in some application scenarios, the output of a breath sound recognition model can also be sound category information. For example, it can output "breath," indicating that the final recognition result of the breath sound recognition model is breathing.

[0217] In practical applications, it has been found that breath sounds are produced by the disturbance of air in the environment during inhalation and exhalation, resulting in sound similar to the sound of wind. The disturbance caused by exhalation is greater than that caused by inhalation; therefore, the frequency amplitude of exhalation is generally higher than that of inhalation. On a large time scale, breathing is a periodic, slow, continuous action, alternating between inhalation and exhalation at a certain frequency, with regular pauses. This characteristic of breathing significantly distinguishes it from coughing, sneezing, and other similar sounds. On a small time scale, each inhalation and exhalation causes air disturbance, and the resulting breath sounds differ from those of sneezing, coughing, and other similar sounds.

[0218] Therefore, to improve the accuracy of breath sound recognition, the breath sound recognition model can include two parts: frame-level sound representation (or single-segment sound representation) and overall sound representation (also known as continuous segment sound representation). Frame-level sound representation represents a single Mel-frequency cepstral segment obtained from S402. When the segment length is 0.96 seconds, frame-level sound representation can represent the refined sound features of 0.96 seconds of audio. Since a complete breath is longer than 0.96 seconds, in order to represent the sound features of long audio as a whole, each 0.96-second audio segment in S402 can be reassembled after frame-level representation and fused to obtain the contextual sound features in the long audio, that is, to obtain the overall sound representation.

[0219] Figure 4C This is a schematic diagram of the breathing sound recognition model provided in the embodiments of this application. Figure 4C In this model, the breath sound recognition model can include a sequentially connected feature extraction layer, a feature concatenation layer, and a classifier. The classifier consists of a sequentially connected bidirectional long short-term memory (BiLSTM) network and a fully connected layer.

[0220] Combination Figure 4C The breathing sound recognition model takes multiple consecutive Mel-frequency cepstral segments as input and outputs confidence scores for seven sound categories. After these segments are input, the model uses a pre-trained feature extractor to extract a frame-level sound representation for each segment. Each Mel-frequency cepstral segment thus yields a frame-level sound representation. The model then concatenates these frame-level sound representations sequentially to obtain the overall sound representation (also known as the continuous segment sound representation). Finally, this continuous segment sound representation is input into a classifier for sound classification, resulting in confidence scores for each of the seven sound categories.

[0221] The training process of a breath sound recognition model can include the following aspects.

[0222] (1) Obtain training samples

[0223] To increase the generalization of the breathing sound recognition model and improve its accuracy in recognizing breathing sounds, it is necessary to train the model with a rich set of training samples. Specifically, the training samples need to cover the following seven types of sounds: quiet, speaking, breathing sounds, coughing, sneezing, snoring, and environmental noise.

[0224] During the training sample collection process, audio recordings of various types of sounds from the same user can be collected. The collected audio can then be divided into several fixed-length audio segments, such as 5 seconds long. Understandably, the length of each audio segment is related to the specific application scenario of the breathing sound recognition model. For example, if the breathing sound recognition model is used to identify the sound category of a 5-second breathing audio data set, the audio segment length can be 5 seconds.

[0225] In processing the collected audio into training samples, for each audio segment, the operations S401-S402 described above can be performed to obtain multiple consecutive Mel-frequency cepstral segments corresponding to that audio segment. Then, the audio segment, the multiple consecutive Mel-frequency cepstral segments corresponding to the audio segment, and the sound category corresponding to the audio segment can be used as training samples. This yields a sample set for training the breath sound recognition model. For ease of description, the sample set for training the breath sound recognition model can be referred to as the audio sample set.

[0226] (2) Training the feature extractor

[0227] Here, the initial feature extractor can be any model that can be used to extract features, such as a convolutional neural network (CNN), a recurrent neural network (RNN), etc.

[0228] Further integration Figure 4C The process of training the feature extractor is equivalent to training a single-frame breath sound recognition model with only a single frame of input. The single-frame breath sound recognition model takes a single Mel-Cepstral fragment as input and outputs the corresponding audio sound category. Similarly, the feature extractor takes a single Mel-Cepstral fragment as input and outputs a frame-level sound representation corresponding to the input Mel-Cepstral fragment.

[0229] Figure 4C In this model, the single-frame breath sound recognition model consists of a sequentially connected feature extractor and a fully connected layer. During training, for each audio sample in the audio sample set, a single Mel-Cepstral segment from the audio sample can be used as input, and the sound category of that audio sample as output, thus completing the training of the single-frame breath sound recognition model and consequently the training of the feature extractor.

[0230] After the feature extractor is trained, its parameters can be fixed or encapsulated so that the feature extractor can be called or used as a feature extraction tool.

[0231] (3) Training the breath sound recognition model

[0232] Here, the feature extractor in the initial breath sound recognition model is the feature extractor trained through the operations in step (2) above. That is to say, during the training of the breath sound recognition model, the feature extractor is only called or used as a feature extraction tool.

[0233] In the process of training the breathing sound recognition model using an audio sample set, for each audio sample, multiple consecutive Mel-frequency cepstral segments in the audio sample can be used as input to the breathing sound recognition model, and the sound category corresponding to the audio sample can be used as output to train the initial breathing sound recognition model, so as to obtain the trained breathing sound recognition model.

[0234] In practice, a well-trained breath sound recognition model can be used as a tool for recognizing breath sounds.

[0235] Understandably, in this embodiment, the entity that trains the feature extractor and the breath sound feature recognition model can be a mobile phone or other devices, such as a server. That is, the model can be trained on a mobile phone, or it can be trained on another device and then transferred to the mobile phone.

[0236] This application does not specifically limit the type of device on which the model is trained.

[0237] The following further combines Figure 4D and Figure 4E This will illustrate the respiratory sound feature extraction model in S405 and the training process of the respiratory sound feature extraction model.

[0238] Figure 4D This is a schematic diagram of the respiratory sound feature extraction model provided in the embodiments of this application. Figure 4D In a respiratory sound feature extraction model, there may be sequentially connected convolutional layers, two multiscale ResNet layers, and a bottleneck layer. Understandably, the number of multiscale residual layers in a respiratory sound feature extraction model can be more than two.

[0239] Figure 4E This is a schematic diagram of a multi-scale residual layer in the respiratory sound feature extraction model provided in this application embodiment. Figure 4E In this study, convolutional features of various scales were cross-fused. For example, 5×5 convolutional features and 3×3 convolutional features were cross-fused, which helps to extract richer and more comprehensive features.

[0240] Combination Figure 4D and Figure 4EThe respiratory sound feature extraction model takes Mel-frequency cepstral coefficients as input and outputs the corresponding respiratory sound features. The model compresses the input Mel-frequency cepstral coefficients using convolutional layers; then, it extracts the corresponding audio features from the compressed data using multiple multi-scale residual layers. The multi-scale residual module, within the residual network framework, adds cross-fusion of convolutional features at different scales, resulting in richer and more comprehensive respiratory sound features. Finally, a bottleneck layer fuses the features from each module to output the respiratory sound features.

[0241] Further reading Figure 4D After the respiratory sound feature extraction model outputs respiratory sound features, the obtained respiratory sound features can be further processed through a fully connected layer and a softmax function layer (or softmax layer) to output the user ID confidence and recognition result. For example, it can output the confidence of being identified as user 1, the confidence of being identified as user 2, ... the confidence of being identified as user N, etc. At the same time, the user ID with the highest confidence can be used as the recognition result.

[0242] Understandable. Figure 4D In this model, adding sequentially connected fully connected layers and a softmax function layer to the breath sound feature extraction model essentially creates a user identification model. The input to the user identification model is the Mel-Cepstral coefficients of the audio, and the output is the user ID confidence score and / or the identification result.

[0243] The training process of a respiratory sound feature extraction model can include the following aspects.

[0244] (1) Obtain training samples

[0245] During the acquisition of training samples, breathing audio can be obtained from the same user at different distances (e.g., the face is 30cm, 50cm, 100cm away from the phone, etc.) with different breathing patterns (e.g., normal breathing sounds, rapid breathing, and deep breathing, etc.). In this way, multiple breathing audio recordings can be obtained for each user.

[0246] Then, each obtained breathing audio recording can be divided into multiple audio segments of fixed length, such as 5 seconds. Understandably, the length of the audio segment is related to the specific application scenario of the breathing sound feature extraction model. For example, if the breathing sound feature extraction model is used to extract breathing sound features from 5 seconds of breathing audio data, the length of the audio segment can be 5 seconds.

[0247] In processing the obtained breathing audio as training samples, for each audio segment, the operation described in S401 (i.e., the audio preprocessing operation) can be performed to obtain the Mel-frequency cepstral coefficients corresponding to that audio segment. Then, the audio segment, its corresponding Mel-frequency cepstral coefficients, and the user ID corresponding to the audio segment can be used as training samples. This yields a sample set for a user recognition model that includes the breathing sound feature extraction model. For ease of description, the sample set used to train the user recognition model can be referred to as the breathing sample set.

[0248] (2) Training the breath sound feature extraction model

[0249] Continue to combine Figure 4E Training the respiratory sound feature extraction model is actually training the user recognition model.

[0250] In some application scenarios, a breathing sound feature extraction model can be trained using the following training method.

[0251] In the first training method, during the training of the user recognition model, for each breath sample in the breath sample set, the Mel-frequency cepstral coefficients in the breath sample can be used as input, and the user ID in the breath sample can be used as output to complete the training of the user recognition model, thereby completing the training of the breath sound feature extraction model.

[0252] The differences in breath sounds between individuals mainly stem from physiological differences (such as lungs, chest and abdominal muscles, and upper respiratory tract) and habitual differences (such as thoracic breathing or abdominal breathing). These differences are manifested in the following three aspects of breath sounds: First, the duration of inhalation, exhalation, and pauses; second, the changes in airflow velocity during inhalation and exhalation; and third, the changes in airflow volume during inhalation and exhalation.

[0253] Furthermore, the difficulty in identifying a user through breath sounds lies in the fact that even the same user's breath sounds can differ depending on the user's current state, the position of the phone relative to the face, or the external environment.

[0254] Therefore, during the training of the respiratory sound feature extraction model, it is necessary to perform data augmentation and contrastive learning on the respiratory audio to improve the generalization ability and robustness of the respiratory sound feature extraction model.

[0255] In the embodiments of this application, the following three data augmentation methods can be used to augment the audio in the training samples.

[0256] Data augmentation method 1 involves adding random noise of different signal-to-noise ratio levels at random locations in the audio to mimic breathing audio in various scenarios. The added noise may include one or more of the following: household noise, outdoor noise, office noise, vehicle noise, and animal noise, etc.

[0257] Data augmentation method 2: Randomly change the speed of the audio to mimic rapid and slow breathing.

[0258] Data augmentation method 3 involves randomly changing the volume range of the audio to mimic audio captured at close range and audio captured at a distance.

[0259] In other application scenarios, the respiratory sound feature extraction model can also be trained using the following training method two.

[0260] Training Method Two: During the iterative training of the breath sound feature extraction model, the model can be trained in batches. For example, there can be K batches of training samples, meaning K training iterations. Each batch of training samples can contain M training samples, and these M training samples come from N users. Here, M and N are both integers greater than 1.

[0261] It should be noted that each batch of M training samples comes from N users, which means that in the case of multiple batches of training, the audio of the same user may be selected multiple times. This can further improve the accuracy and stability of the trained model.

[0262] In practice, to further improve the model's generalization ability, for each batch of M training samples, one can be labeled as a positive sample and the rest as negative samples. That is, in this training method two, the training samples used for training can have two labels: one for labeling the user ID and the other for labeling the positive and negative samples.

[0263] As an example, for each batch of M training samples, one sample can be randomly selected from these M training samples and labeled as a positive sample. Then, the remaining training samples from these M training samples can be labeled as negative samples.

[0264] In addition, for each training sample (or breathing sample), data augmentation techniques can be used to enhance the training sample, resulting in an augmented sample (or enhanced breathing sample). For example, by using the three data augmentation techniques mentioned above, three augmented samples can be obtained for each training sample. These augmented samples can then be used to train the breathing sound feature extraction model.

[0265] When the above three data augmentation methods are used to augment each training sample, each batch of M training samples can be expanded into 3M augmented samples.

[0266] It should be noted that training the respiratory sound feature extraction model with enhanced samples can improve the generalization ability of the trained model. That is, it can enable the trained respiratory sound feature extraction model to accurately extract the respiratory sound features of users in various situations.

[0267] In this embodiment of the application, training the breath sound feature extraction model using the above-described training method two can have the following beneficial effects:

[0268] 1. When training the respiratory sound feature extraction model, label one of the M training samples in the same batch as a positive sample and the rest as negative samples. This can make the trained respiratory sound feature extraction model more recognizable for respiratory sounds. In other words, it can improve the model's generalization ability and robustness, thereby helping the trained respiratory sound feature extraction model to extract respiratory sound features more accurately and effectively.

[0269] 2. Perform data augmentation on each training sample. Use the augmented samples obtained from data augmentation to train the respiratory sound feature extraction model. This can improve the generalization ability and robustness of the trained respiratory sound feature extraction model. In other words, it can enable the trained respiratory sound feature extraction model to accurately extract the respiratory sound features of users in various situations.

[0270] In some optional implementations of this application, during the training of the breath sound feature extraction model (belonging to the user recognition model) using the above-described training method two, the loss function used may include the following two parts: one is the user recognition cross-entropy loss, and the other is the contrastive learning similarity loss. The user recognition cross-entropy loss is used to describe the accuracy of the breath sound feature extraction model in user recognition, and the contrastive learning similarity loss is used to describe the degree of similarity between the features extracted by the breath sound feature extraction model and the real features.

[0271] For example, the user identification cross-entropy loss can be expressed as shown in Equation (1):

[0272]

[0273] Where, loss au To identify the cross-entropy loss for the user, label(y) ji ) is the real label, predict(y) jiThe loss function () represents the user identification prediction probability output by the softmax layer. This loss function evaluates the difference between the true label and the predicted result. au The smaller the value, the more accurate the prediction.

[0274] For example, the contrastive learning similarity loss can be shown in the formula group (2):

[0275]

[0276]

[0277]

[0278] Where, loss cl To contrast the learning similarity loss, the S matrix is ​​a cosine similarity matrix. Each element in the S matrix indicates the cosine similarity between the breath sound features corresponding to two training samples. For example, s M,3M This is used to indicate the cosine similarity between the breath sound features corresponding to the Mth training sample and the breath sound features corresponding to the 3Mth training sample. The values ​​of each element in the S matrix can be any value between 0 and 1.

[0279] The L matrix is ​​the target similarity matrix. Each element in the L matrix indicates whether two training samples come from the same user. If two training samples come from the same user, they are considered to have similar features, and the similarity value is 1; if two training samples come from different users, they are considered to have dissimilar features, and the similarity value is 0. That is, each element in the L matrix takes the value of 0 or 1.

[0280] The above comparison learning similarity loss cl The similarity loss can be calculated based on the cosine similarity matrix S and the target similarity matrix L. When comparing the learning similarity loss... cl The smaller the value, the closer similar training samples are in the feature space, and the farther dissimilar training samples are in the feature space.

[0281] In practice, user identification cross-entropy loss au And contrastive learning similarity loss cl Each can be assigned a weight.

[0282] In practice, if the user identifies the cross-entropy loss... au The corresponding weight can be a1, and the comparison learning similarity loss is used. clThe corresponding weight can be a2. For example, the overall model training loss can be shown in the following formula (3). Where a1 and a2 are pre-set weight values. In practice, the values ​​of a1 and a2 are in the range of 0-1, and the sum of a1 and a2 can be 1.

[0283] loss = a1 × loss au +a2×loss cl (3)

[0284] In the embodiments of this application, during the iterative training of the respiratory sound feature extraction model, the model parameters can be updated using the gradient descent method, thereby reducing the overall model training loss and optimizing the model parameters.

[0285] In this embodiment of the application, during the model training process, the model loss is described from two dimensions: user identification and contrastive learning, so as to optimize and adjust the model. This can ensure that the trained user identification model can accurately identify users and has a strong ability to distinguish breathing sounds. In other words, the breathing sound features extracted by the trained breathing sound feature extraction model have a strong ability to distinguish breathing sounds and can comprehensively and accurately describe the user's personal breathing characteristics.

[0286] It should be noted that the higher the recognition rate of the breath sound feature extraction model, the higher the accuracy or recognition rate of the extracted breath sound features, and the more accurate the recognition results will be when the device owner is identified based on the breath sound features.

[0287] 3. Input of cross-modal feature templates.

[0288] In this embodiment of the application, when the face template of the device owner is successfully entered, the mobile phone can collect breathing audio data through the microphone, collect touch screen data through the touch screen sensor, and collect IMU data through the IMU, and obtain the cross-modal feature template of the device owner through the above three types of data.

[0289] In this embodiment of the application, when the acquisition frequency of the respiratory audio data is 16kHz and the acquisition duration is 5 seconds, the size of the respiratory audio data acquired by the mobile phone is 80,000 dimensions, where 80,000 = 16k × 5.

[0290] The size of the IMU data collected by the IMU can be 50×90, where 50 indicates 50 IMU data collections and 90 indicates the dimensions of the IMU data. The IMU data can include the phone's acceleration, pitch angle, roll angle, etc., in various directions. The IMU data collection frequency can be 50 times / second. That is, the phone can obtain the aforementioned 50×90 dimensional IMU data by collecting IMU data within one second after the user's face template is successfully registered.

[0291] The touchscreen data collected by the touchscreen sensor can be 200×17 in size, where 200 indicates the 200 most recent touchscreen data entries and 17 indicates the dimension of the touchscreen data. Each touchscreen data entry can include the number of touchscreen swipes, clicks, etc. That is, after the user's face template is successfully registered, the phone can collect the 200 most recent touchscreen data entries to obtain the aforementioned 200×17 dimensional touchscreen data.

[0292] It is understood that in the embodiments of this application, the size of the data of the above three modalities can be adaptively adjusted according to the actual application scenario. The embodiments of this application do not specifically limit the size of the data of various modalities.

[0293] The following section will further explain how to combine the user's breathing audio data, the phone's touchscreen data, and the phone's IMU data to obtain the user's cross-modal feature template.

[0294] In this embodiment, due to the large volume of respiratory audio data, to improve data processing efficiency, the mobile phone can call a pre-stored audio processing module to process the respiratory audio data and extract the respiratory sound features. Then, the mobile phone can call a pre-stored cross-modal encoder to encode the respiratory sound features, touchscreen data, and IMU data into corresponding cross-modal features.

[0295] It should be noted that during the initial recording of a user's cross-modal feature template, if the collected respiratory audio data does not contain respiratory sounds, or is not recognized as respiratory sounds, then respiratory sound features may not be obtained. In this case, the phone may temporarily not record the user's respiratory sound feature template and cross-modal feature template.

[0296] In some application scenarios, if the phone detects that the breathing audio data does not contain breathing sounds, it can continue to collect breathing audio data after a successful face unlock. It can also determine whether the breathing audio data is a breathing sound. If it is, it can further extract breathing sound features and record these features as the owner's breathing sound feature template. Simultaneously, by combining this owner's breathing sound feature template with the IMU data and touchscreen data after successful face unlock, the owner's cross-modal features can be determined and recorded as the owner's cross-modal feature template.

[0297] Based on the above description, in this embodiment of the application, the triggering condition (or multimodal feature entry condition) for triggering the entry of cross-modal feature templates includes any one of the following: the owner's face template is successfully entered, the owner's face unlock is successful, and the cross-modal feature template of the owner currently performing the unlock has not yet been entered in the terminal device.

[0298] The following combination Figure 5A , Figure 5B and Figure 5C This will illustrate the structure of a cross-modal encoder. Figure 5A This is a schematic diagram of the structure of the cross-modal encoder provided in the embodiments of this application. Figure 5B This is a schematic diagram of the structure of a cross-channel linear encoder in a cross-modal encoder provided in an embodiment of this application. Figure 5C This is a schematic diagram of the structure of a cross-channel attention encoder in a cross-modal encoder provided in an embodiment of this application.

[0299] like Figure 5A As shown, the inputs of the cross-modal encoder are IMU data, touchscreen data, and breath sound features, and the output is cross-modal features. The cross-modal encoder can encode the IMU data and touchscreen data separately using two cross-channel linear encoders from top to bottom. For ease of description, the cross-channel linear encoder that encodes the IMU data can be referred to as the first cross-channel linear encoder, and the cross-channel linear encoder that encodes the touchscreen data can be referred to as the second cross-channel linear encoder.

[0300] Subsequently, the cross-modal encoder can extract features from the encoded IMU data, encoded touchscreen data, and breath sound features through three STEM layers from top to bottom. For ease of description, the STEM layer for feature extraction of the encoded IMU data can be referred to as the first STEM layer, the STEM layer for feature extraction of the encoded touchscreen data as the second STEM layer, and the STEM layer for feature extraction of breath sound features as the third STEM layer.

[0301] Subsequently, the cross-modal encoder can fuse the IMU features extracted from the first Stem layer, the touch features extracted from the second Stem layer, and the breathing sound features extracted from the third Stem layer to obtain fused features.

[0302] Subsequently, the cross-modal encoder processes the fused features sequentially through the fourth Stem layer, the third cross-channel linear encoder, and the multi-target operation module, outputting cross-modal features. For example... Figure 5A As shown, the target operation module may include a downsampling module, a cross-channel linear encoder, and a cross-channel attention encoder connected in sequence.

[0303] Understandably, the first, second, third, and fourth points mentioned above are only used to distinguish one element from another in terms of description.

[0304] Combination Figure 5A The two important components of a cross-modal encoder are the cross-channel linear encoder and the cross-channel attention encoder.

[0305] Figure 5B In a cross-channel linear encoder, a sequentially connected linear layer, activation function, linear layer, normal layer, and conv layer can be used for encoding. The cross-channel linear encoder uses linear layers for encoding and aligns the feature dimensions of different modalities, and uses conv layers to implement one-dimensional pointwise convolution for information fusion between channels. Figure 5B middle, The sign indicating addition. The sign indicates the excitation function.

[0306] In this embodiment, the main function of the cross-channel linear encoder is to increase the level of abstraction of the features. The input of the cross-channel linear encoder is feature data, and the output is further abstracted feature data.

[0307] Figure 5C In the diagram, combined with the structural part within the dashed box, the cross-channel attention encoder slices the data input from the input layer (also known as the Inputs layer) through the Group layer, specifically by assigning one channel to each modality of data, thus cutting features from the channel dimension; then it encodes the data through the Linear layer; next, it merges the sliced ​​features; finally, it uses an attention mechanism in the channel dimension to promote the fusion of multimodal information between channels. Figure 5C middle, The sign indicating addition. To indicate the sign of the excitation function, The sign indicating multiplication. This is a symbol indicating positional encoding.

[0308] Further integration Figure 5C , Figure 5C The dashed box shows three Group layers from top to bottom, and the Linear layers connected to each Group layer. For ease of description, the topmost Group layer can be called the first Group layer, the topmost Linear layer can be called the first Linear layer, the middle Group layer can be called the second Group layer, the middle Linear layer can be called the second Linear layer, and the bottommost Group layer can be called the third Group layer, and the bottommost Linear layer can be called the third Linear layer.

[0309] The input of the first Linear layer is connected to the output of the first Group layer, the input of the second Linear layer is connected to the output of the first Linear layer and the output of the second Group layer, and the input of the third Linear layer is connected to the output of the second Linear layer and the output of the third Group layer.

[0310] from Figure 5C It can be observed that the output data from the top first group layer is encoded the most times, while the output data from the bottom third group layer is encoded the fewest times.

[0311] In this embodiment, high-quality features can be deployed in the lower Group layers, and low-quality features can be deployed in the upper Group layers. That is, features deployed in the topmost Group layer are relatively the worst, and features deployed in the bottommost Group layer are relatively the best. This way, features deployed in the bottommost Group layer require the fewest linear layers, i.e., the fewest encoding iterations. Features deployed in the topmost Group layer require the most linear layers, i.e., the most encoding iterations. In this way, a cross-channel attention encoder can combine the quality of each feature and encode each feature to different degrees, thereby achieving a better and more effective fusion of high-quality and low-quality features.

[0312] In this embodiment, after encoding and fusing high-quality and low-quality features to different degrees, the fused features can be transposed, and the transposed fused data can be input into the Features layer, which helps to achieve further effective fusion between features.

[0313] In this embodiment, the main function of the cross-channel attention encoder is to effectively fuse high-quality features with low-quality features to enhance the features, thereby improving the accuracy of the output features.

[0314] In addition, combined Figure 5A , Figure 5B and Figure 5C The cross-modal encoder adopts a lightweight design, which enables lightweight deployment.

[0315] (ii) Using a device owner template for device owner identification

[0316] In this embodiment of the application, the triggering conditions for the mobile phone to enter owner identification may include the following conditions 1-6. It is understood that this embodiment of the application only lists some of the triggering conditions, and other triggering conditions may exist as the actual application scenario changes.

[0317] Condition 1: Continuous authentication is enabled on your mobile phone.

[0318] In this embodiment, when the continuous authentication function is enabled, the mobile phone can automatically lock the target content during use. For example, when the continuous authentication function is enabled, the mobile phone can lock service cards containing the owner's private information.

[0319] Condition 2: The phone detects the action of opening the target application.

[0320] In this embodiment of the application, when an operation to open the target application is detected, it indicates that the user needs to open the target application. At this time, the phone can be triggered to enter the owner identification process. If the user is identified as the owner, the target application is opened; if the user is identified as not the owner, the target application is not opened. In this way, the security of the owner's private information can be guaranteed.

[0321] The target application is typically an application that may involve the user's privacy, such as WeChat, Alipay, or the photo album. Understandably, the target application can be set by the phone itself; for example, the phone can set an instant messaging application as the target application, or it can set an application that the user uses frequently above a certain frequency threshold as the target application. Of course, the target application can also be an application set by the user according to their own needs. For example, a user can set a photo album application that they consider to have high privacy as the target application. It is understood that this application does not specifically limit the target application.

[0322] Condition 3: The mobile phone receives a notification message.

[0323] In this embodiment, upon receiving a notification message from an application, the phone can be triggered to enter owner identification mode. If the user is identified as the owner, the notification message is previewed; if the user is not identified as the owner, the notification message is hidden and not previewed. This ensures the security of the owner's private information.

[0324] Condition 4: The phone needs to display a service card, and the service card must contain private information, such as bank card number, password, etc.

[0325] In this embodiment, the mobile phone can identify the owner before displaying the service card containing private information. If the owner is identified, the service card is displayed directly; if the owner is not identified, the private information in the service card is hidden. This ensures the security of the owner's private information.

[0326] Condition 5: The phone detects that the target function has been triggered.

[0327] The target function can be a password-free payment function, or it can be any other function that involves the security of the account holder's information or funds.

[0328] In this embodiment of the application, when the target function needs to be executed, the mobile phone can be triggered to enter the owner identification process. If the owner is identified, password-free payment is performed. If the non-owner is identified, password-free payment cannot be performed. For example, a password must be entered to complete the payment.

[0329] Condition 6: The phone detects that the target interface has been triggered.

[0330] The target interface can be an interface containing the owner's private information. For example, it could be the interface displaying the content of a short message.

[0331] In this embodiment of the application, when it is necessary to open the target interface, the mobile phone can be triggered to enter the owner recognition process. If the owner is recognized, the target interface is opened; if the owner is not recognized, the target interface is not opened, which can protect the security of the owner's privacy information.

[0332] If the current triggering conditions for owner identification are met, the phone will enter the owner identification process. The following will further explain how the phone performs owner identification.

[0333] Figure 6A The diagram illustrates the process of mobile phone owner identification based on multimodal fusion.

[0334] S601: The phone determines whether a face has been detected. If no face has been detected, proceed to S602. If a face has been detected, proceed to S605.

[0335] Here, the phone can analyze whether the detected image is a face image. For example, it can determine whether it is a face image by analyzing the image's pixel values. If it is a face image, the phone can combine it with a pre-recorded face template of the phone's owner to identify the owner's face and obtain a face recognition score. The face recognition score describes the degree of similarity between the tested face and the face template.

[0336] S602, if no face is detected, the phone determines whether a breathing sound is detected. If a breathing sound is detected, proceed to step S603; if no breathing sound is detected, it indicates that the detection result is unreliable, and in this case, proceed to step S604.

[0337] In this embodiment of the application, when no face is detected, the mobile phone can continue to collect the breathing audio data of the user to be tested and extract the breathing sound features in the breathing audio data to be tested. Then, by comparing the similarity between the breathing sound features to be tested and the breathing sound feature template in the owner template, it can be determined whether the user to be tested is the owner.

[0338] Among these features, the mobile phone can invoke a pre-stored audio processing module (see...). Figure 4A The audio processing module processes the respiratory audio data to be tested to extract the respiratory sound features from the respiratory audio data. The respiratory audio data to be tested is the respiratory audio data of the user being tested. The respiratory sound features to be tested are the respiratory sound characteristics within the respiratory audio data.

[0339] In this embodiment, the acquisition parameters for the respiratory audio data to be measured can be: single-channel acquisition, sampling rate of 16kHz, and acquisition duration of t seconds. The value of t ranges from 0.5 seconds to 5 seconds.

[0340] In practice, the acquisition duration of the respiratory audio data is not fixed. The acquisition duration is typically related to the usage scenario, and is generally positively correlated with the accuracy of the device identification process in that scenario, and inversely correlated with the speed of device identification. For example, in scenarios with very high accuracy requirements for device identification, the acquisition duration can be as long as possible; conversely, in scenarios with very high speed requirements for device identification, the acquisition duration can be as short as possible.

[0341] To illustrate further, in the scenario of user account login, where high accuracy in user identification is required, the collected breathing audio data could be 5 seconds long. In the scenario of viewing photos, if the user waits a long time to open a photo, such as 5 seconds, the user experience will obviously be poor. Therefore, this scenario requires high speed in user identification. In this case, to improve the speed of user identification, the collected breathing audio data could be 0.5 seconds long.

[0342] Because the respiratory sound recognition model in the audio processing module may identify that the collected respiratory audio data to be tested does not contain respiratory sounds, it is considered that no respiratory sounds were detected if the respiratory audio data to be tested does not contain respiratory sounds. Conversely, if the respiratory sound recognition model identifies that the respiratory audio data to be tested contains respiratory sounds, it is considered that respiratory sounds were detected.

[0343] S603 identifies the device owner based on the characteristics of the breath sounds to be tested and outputs the device owner identification result.

[0344] In this embodiment of the application, in order to distinguish the description, the breathing sound characteristics of the user to be tested can be denoted as the breathing sound characteristics to be tested.

[0345] Here, if S602 detects a breathing sound, it means that the breathing sound recognition model in the audio processing module recognizes that the breathing audio data to be tested contains a breathing sound, and then the audio processing module can finally output the breathing sound features to be tested.

[0346] Then, the phone can use the breath sound characteristics of the user being tested to identify whether the user is the owner of the phone.

[0347] Figure 6B A schematic diagram illustrating the process of a device that recognizes its owner based on breath sounds is shown. Combined with... Figure 6B The mobile phone can compare the breath sound characteristics of the user to be tested with the pre-stored breath sound characteristic template of the owner to determine whether the user to be tested is the owner of the phone.

[0348] In some optional implementations of the embodiments of this application, the mobile phone can determine whether the user to be tested is the owner of the phone in the following way: when the owner template includes a breathing sound feature template, the similarity between the breathing sound feature to be tested and the breathing sound feature template is determined. If the obtained similarity is greater than a preset second determination threshold, the user to be tested is determined to be the owner of the phone. Conversely, if the obtained similarity is less than or equal to the second determination threshold, the user to be tested is determined to be a non-owner of the phone.

[0349] The second judgment threshold mentioned above is usually a pre-set similarity value, for example, it can be 0.6, and the value range of the second judgment threshold is usually 0-1.

[0350] In practice, calculating the similarity between the breath sound feature to be tested and the breath sound feature template usually involves calculating the cosine similarity between the vector corresponding to the breath sound feature to be tested and the vector corresponding to the breath sound feature template.

[0351] In some optional implementations of the embodiments of this application, the mobile phone can also determine whether the user to be tested is the owner of the phone in the following way: when the owner template includes multiple breathing sound feature templates, the similarity between each breathing sound feature template and the breathing sound feature to be tested is determined respectively. When the average value of the multiple similarities obtained is greater than a preset fourth determination threshold, the user to be tested is determined to be the owner of the phone. Conversely, when the average value of the multiple similarities obtained is less than or equal to the fourth determination threshold, the user to be tested is determined to be a non-owner of the phone.

[0352] The fourth judgment threshold mentioned above is usually a pre-set similarity value, for example, it can be 0.3, and the value range of the fourth judgment threshold is 0-1.

[0353] To give a further example, if the host template contains three breath sound feature templates, namely breath sound feature template 1, breath sound feature template 2, and breath sound feature template 3, and the similarity between the breath sound feature to be tested and breath sound feature template 1 is S1, the similarity between the breath sound feature to be tested and breath sound feature template 2 is S2, and the similarity between the breath sound feature to be tested and breath sound feature template 3 is S3, then if the average of S1, S2, and S3 is greater than the fourth judgment threshold, the user to be tested is considered to be the host; otherwise, the user to be tested is considered to be a non-host.

[0354] It should be noted that when there are multiple breath sound feature templates in the device owner template, the accuracy of device owner identification can be improved by using the average similarity between the breath sound feature to be tested and each of the breath sound feature templates.

[0355] Optionally, if the device owner template includes multiple breath sound feature templates, the phone can also determine the similarity between each breath sound feature template and the breath sound feature to be tested. If the largest of the multiple similarities is greater than a second determination threshold, the user to be tested is determined to be the device owner. Conversely, if the largest of the multiple similarities is less than or equal to the second determination threshold, the user to be tested is determined to be a non-device owner.

[0356] In some optional implementations of the embodiments of this application, when multiple owner templates are entered in the mobile phone, the mobile phone can also determine whether the user to be tested is the owner in the following ways: For each owner template, if the owner template includes a breathing sound feature template, the similarity between the breathing sound feature to be tested and the breathing sound feature template is determined. If the obtained similarity is greater than a second determination threshold, the user to be tested is determined to be the owner, specifically the owner indicated by the owner template; if the owner template includes multiple breathing sound feature templates, the similarity between each breathing sound feature template in the owner template and the breathing sound feature to be tested is determined respectively. If the average of the obtained multiple similarities is greater than a fourth determination threshold, the user to be tested is determined to be the owner, specifically the owner indicated by the owner template.

[0357] Optionally, if the user under test is not the owner indicated by any owner template, the phone can determine that the user under test is not the owner.

[0358] In this embodiment, when the user under test is identified as the owner of the phone, the phone can continue to perform operations related to the owner's permissions, such as previewing notification messages and making password-free payments. When the user under test is identified as not the owner, the phone hides the owner's private information, such as hiding the content of communication messages and the content of service cards.

[0359] S604 indicates that the test result is unknown.

[0360] In this embodiment, if the mobile phone fails to detect either a face or a breathing sound during the owner identification process, the detection result is considered unreliable. In this case, the mobile phone can output "Detection result unknown," "Retry once," or "Try again," among other prompts.

[0361] Afterwards, the user to be tested can further cooperate with the phone to identify the owner, based on the prompts provided.

[0362] In S605, the phone inputs the face recognition score into the threshold decision unit to determine whether the user being tested is the owner of the phone.

[0363] Here, if the face recognition score is greater than a preset upper threshold, such as 0.65, the user being tested is considered the owner of the device. If the face recognition score is less than a preset lower threshold, such as 0.05, the user being tested is considered not the owner of the device. If the face recognition score is between the lower and upper thresholds, step S606 can continue.

[0364] The aforementioned upper threshold for face recognition can be the FAR (False Rejection Rate) threshold. The aforementioned lower threshold for face recognition can be the False Rejection Rate (FRR) threshold. In practice, a face recognition score between the aforementioned lower threshold and the aforementioned upper threshold usually indicates poor face detection results.

[0365] S606: If the face detection result is poor, the phone determines whether a breathing sound is detected. If no breathing sound is detected, proceed to S607; otherwise, if a breathing sound is detected, proceed to S608.

[0366] In this embodiment, if the face detection result is unsatisfactory, the mobile phone can continue to collect the breathing audio data of the user under test and call the pre-stored audio processing module (see...). Figure 4A The audio processing module processes the respiratory audio data to be tested in order to extract the respiratory sound features to be tested from the respiratory audio data.

[0367] In practice, if the breath sound recognition model detects that the breath audio data being tested does not contain breath sounds, it considers that no breath sounds have been detected. Conversely, if the breath sound recognition model detects that the breath audio data being tested contains breath sounds, it considers that breath sounds have been detected.

[0368] The S607 directly outputs face detection results even when no breathing sounds are detected.

[0369] Here, the phone uses a pre-recorded face template of the phone's owner to detect and recognize faces. During this process, it can use the face recognition score to make a preliminary judgment on whether the person is the owner. If no breathing sound is detected, the phone can output this preliminary judgment. For example, if the face recognition score is greater than a pre-set face recognition score threshold, the person is determined to be the owner; conversely, if the face recognition score is less than or equal to the threshold, the person is determined to be a non-owner. This face recognition score threshold can be a pre-set threshold, such as 0.5.

[0370] Optionally, in some application scenarios where the accuracy of owner recognition is very high, if the face detection result is poor and no breathing sound is detected, the phone can also output a prompt to detect again, such as "Please try again".

[0371] When a breathing sound is detected, the S608 combines the multimodal data of the user under test and calls the multimodal fusion owner recognition model to perform owner recognition.

[0372] The multimodal data may include face recognition scores, breath sound features to be tested, touchscreen data to be tested, and IMU data to be tested.

[0373] In this embodiment, only when facial recognition is ineffective, such as in situations with strong light, backlight, or low light, will the system further collect the user's breathing audio data, touchscreen data, and IMU data. This multimodal data is then combined to perform device owner identification, thereby improving the accuracy of owner recognition. Since facial recognition accuracy is already high enough most of the time, the use of multimodal data for owner identification is relatively rare. This approach allows for saving computational resources and reducing power consumption while maintaining accurate owner identification, thus improving the phone's usability and enhancing the user experience.

[0374] To distinguish the descriptions, the touchscreen data collected during the owner identification process can be recorded as the touchscreen data under test, and the IMU data collected can be recorded as the IMU data under test.

[0375] The process of collecting data from the touchscreen under test during owner identification is essentially the same as the process of collecting touchscreen data during owner template entry, and will not be elaborated here. Similarly, the process of collecting data from the IMU under test during owner identification is also essentially the same as the process of collecting IMU data during owner template entry.

[0376] Understandably, in some optional implementations, such as S608, when a breathing sound is detected, the phone can directly use the breathing sound characteristics of the user to be tested for owner identification. In this case, S608 can also be replaced with the following steps: when a breathing sound is detected, the phone identifies the owner based on the breathing sound characteristics and outputs the owner identification result. This operation is basically the same as the operation in S603 mentioned above, and will not be elaborated here.

[0377] Figure 6C A schematic diagram of the process of host identification by the multimodal fusion host identification model is shown.

[0378] S6081, calculate the cross-modal characteristics to be tested.

[0379] Among them, the cross-modal features to be tested are the cross-modal features corresponding to the users to be tested.

[0380] Here, the mobile phone can use a cross-modal encoder (see...). Figure 5A The data from the touchscreen under test, the data from the IMU under test, and the breath sound features under test are encoded to obtain the cross-modal features under test.

[0381] S6082, calculate the cross-modal similarity between the cross-modal feature to be tested and the cross-modal feature template.

[0382] Cross-modal similarity refers to the similarity between the cross-modal feature to be tested and the cross-modal feature template. In the embodiments of this application, cross-modal similarity can also be referred to as cross-modal recognition score.

[0383] Here, if the cross-modal feature template S m =E(I m ,T m ,R m ), and the cross-modal feature S to be tested t =E(I t ,T t ,R t ), where E(,,) represents the cross-modal encoder, I m For the owner's IMU data, T m For the owner's touchscreen data, R m As a template for the host's breathing sound characteristics, I t For the IMU data of the user under test, T t For the touchscreen data of the user under test, R t The breath sound characteristics of the user to be tested.

[0384] For example, cross-modal similarity can be calculated using the following formula (4):

[0385]

[0386] Among them, D c For cross-modal similarity, C(,) represents the cosine similarity comparator, and ||| is the L2 norm, used to calculate the magnitude of the vector.

[0387] S6083, calculate the similarity of the breath sounds between the breath sound feature to be tested and the breath sound feature template of the host.

[0388] Here, the breath sound similarity is the similarity between the breath sound feature to be tested and the breath sound feature template of the device owner. In the embodiments of this application, the breath sound similarity can also be referred to as the breath recognition score.

[0389] For example, the similarity of breath sounds can be calculated using the following formula (5):

[0390]

[0391] Among them, D r For breath sound similarity.

[0392] S6084 inputs cross-modal similarity, breath sound similarity, and face recognition score into the fusion decision model to obtain the multimodal recognition score.

[0393] The fusion decision model takes cross-modal similarity (also known as cross-modal recognition score), breath sound similarity (also known as breath recognition score), and face recognition score as inputs, and outputs a multimodal recognition score. The multimodal recognition score describes the comprehensive recognition results of the test user across multiple modalities.

[0394] In this embodiment of the application, the multimodal recognition score can be represented as follows: D = M(D c D r D f ), where D is the multimodal recognition score, D c For cross-modal similarity, D r For breath sound similarity, D f Let M(,,) represent the face recognition score, and M(,,) represent the fusion decision model.

[0395] The S6085, combined with multimodal recognition scores, determines whether the user under test is the owner of the device.

[0396] In this embodiment of the application, if the multimodal recognition score is greater than or equal to a preset fusion threshold, the user is identified as the owner; if the multimodal recognition score is less than the fusion threshold, the user is identified as a non-owner.

[0397] The aforementioned fusion threshold is usually a pre-set threshold, for example, it can be 0.6.

[0398] Combination Figure 6CIt can be observed that the multimodal fusion main recognition model employs a cross-modal encoder to perform feature-level fusion of IMU data, touchscreen data, and breathing sound features. Then, a decision fusion model performs decision-level fusion of the cross-modal recognition score, breathing recognition score, and face recognition score to obtain the multimodal recognition score. Notably, stronger modalities, such as the face modality, do not participate in feature-level fusion. This reduces data coupling and effectively utilizes the complementary advantages of multiple modalities, thereby improving the model's overall recognition success rate.

[0399] Considering that a phone may have one owner's template or multiple owner's templates registered in it. That is to say, a phone can have one owner's template or multiple owner's templates.

[0400] In some optional implementations of the embodiments of this application, when multiple owner templates are pre-recorded in the mobile phone, combined with the above... Figure 6C During the execution of S6081-S6084 on the mobile phone, a master template can be selected first, and the multimodal recognition score can be calculated using the selected master template.

[0401] In some application scenarios, the mobile phone can select the owner template with the highest face recognition score. For example, if there are three owner templates, namely owner template 1, owner template 2 and owner template 3, if the mobile phone executes S601 and identifies the user to be tested as the owner corresponding to owner template 1, then owner template 1 can be selected directly.

[0402] In this application scenario, during the execution of S6081-S6084, the phone can calculate a multimodal recognition score for owner template 1. The phone can then use this multimodal recognition score to further determine if the user is the owner. If the user is the owner, it means the user being tested is the owner indicated by owner template 1. It's easy to see that this application scenario only requires calculating a multimodal recognition score for one owner template, resulting in a relatively small computational load.

[0403] In some alternative implementations of this application, when multiple owner templates are pre-recorded in the mobile phone, the mobile phone can calculate a multimodal recognition score for each owner template. Then, the mobile phone can combine the obtained multiple multimodal recognition scores to determine whether the user to be tested is the owner, and to determine which specific owner it is.

[0404] To give a further example, if there are three owner templates, namely owner template 1, owner template 2, and owner template 3, the phone can use owner template 1 to perform the S6081-S6084 operation to obtain a multimodal recognition score, for example, 99 points. Using owner template 2 to perform the S6081-S6084 operation, another multimodal recognition score can be obtained, for example, 78 points. And using owner template 3 to perform the S6081-S6084 operation, yet another multimodal recognition score can be obtained, for example, 60 points. At this point, it can be determined that the user under test is the owner, specifically the owner indicated by owner template 1.

[0405] In this implementation, step S608, which employs a multimodal fusion model for further device owner identification, is only performed when face detection results are poor. If the face detection in step S601 is inaccurate, selecting a device owner template based solely on the face detection results may not be accurate enough. Therefore, in this implementation, when the face detection accuracy in S601 is not high, multimodal recognition scores are calculated separately for each device owner template to ensure the accuracy of device owner identification.

[0406] Considering that a host template may contain multiple feature templates. For example, it may contain multiple face templates; or multiple breathing sound feature templates; or multiple cross-modal feature templates.

[0407] In some optional implementations of the embodiments of this application, when the main template has multiple feature templates, during the execution of S6081-S6084 by the mobile phone, the average value of the multiple feature templates in the main template can be compared with the similarity of the feature to be tested in the same modality (e.g., the feature of the breath sound to be tested) to obtain the recognition score of the corresponding modality.

[0408] To give a further example, if the main template contains three breathing sound feature templates, the phone can calculate the similarity between the average of the three breathing sound feature templates and the breathing sound feature to be tested, and obtain the recognition score corresponding to the breathing modality, or breathing recognition score.

[0409] Optionally, the mobile phone can also calculate the similarity between each feature template in the master template and the test feature of the same modality, thereby obtaining multiple similarity scores. Then, the recognition score of the corresponding modality can be calculated based on the multiple similarity scores. For example, the maximum value among the multiple similarity scores can be used as the recognition score of the corresponding modality.

[0410] To give a further example, if the main template has three breathing sound feature templates, the phone can calculate the similarity between each breathing sound feature template and the breathing sound feature to be tested, and obtain three similarity scores. Then, the maximum value of the three similarity scores can be used as the recognition score corresponding to the breathing modality, or the breathing recognition score.

[0411] In some optional implementations of the embodiments of this application, when multiple owner templates are pre-recorded in the mobile phone, combined with the above... Figure 6C During the execution of S6085 on the mobile phone, since each owner template may correspond to a multimodal recognition score, the mobile phone can determine whether the user under test is the owner in the following way: For the multimodal recognition score corresponding to each owner template, if the multimodal recognition score is greater than the fusion threshold, the mobile phone can determine that the user under test is the owner. At this time, the user under test can be identified as the owner indicated by the owner template corresponding to the multimodal recognition score.

[0412] In some application scenarios, if multiple multimodal recognition scores simultaneously exceed the fusion threshold, the mobile phone can determine that the user being tested is the owner of the device. In this case, the user being tested can be identified as the owner indicated by the owner template corresponding to the highest of the multiple multimodal recognition scores.

[0413] The following is combined Figure 6D Further explanation of the fusion decision model in S6084. Figure 6D This is a schematic diagram of the fusion decision model provided in an embodiment of this application. (Combined with...) Figure 6C The input to the fusion decision model can be cross-modal recognition score, breathing recognition score, and face recognition score, and the output is multimodal recognition score.

[0414] Combination Figure 6D The fusion decision model includes a Stem layer, a 12-layer encoding layer, a Head layer, and a Sigmoid function. The encoding layer includes a Linear layer, an activation function, and a Linear function. Figure 6D In this model, the fusion decision model uses a Stem layer to expand the feature dimensions of cross-modal recognition scores, breathing recognition scores, and face recognition scores. These are then encoded using a 12-layer feedforward neural network with residual structures. Finally, a classification head layer and a sigmoid function are used to output the multimodal recognition scores. This decision fusion model employs a lightweight design, saving storage resources.

[0415] The training process of the multimodal fusion machine's main recognition model is described below. Figure 6CAs can be seen, the master recognition model of the multimodal fusion machine mainly includes a cross-modal encoder and a fusion decision model. In other words, training the master recognition model of the multimodal fusion machine is also training the cross-modal encoder and the fusion decision model.

[0416] During the training of the multimodal fusion machine master recognition model, difficult samples can be dynamically mined and used as negative samples, thereby further training the multimodal fusion machine master recognition model by pairing positive and negative samples.

[0417] Figure 6E The model training architecture based on positive and negative sample pairing using dynamic hard sample mining is shown. For example... Figure 6E As shown, the process of training the master recognition model of the multimodal fusion machine may include the following steps:

[0418] The first step is to obtain a batch of training samples from all the sample datasets corresponding to the sample users.

[0419] Combination Figure 6E The "m×" in the upper right corner of the dashed box indicates that m batches of training samples were sampled. Each batch of training samples includes 1 positive sample and n negative samples.

[0420] In this embodiment, each sample user corresponds to a sample dataset. Each sample data in the sample dataset may include sample IMU data, sample touchscreen data, sample breathing sound features, and sample face recognition score.

[0421] Here, we can start from sample user u i From all the sample datasets corresponding to i∈(1,k), obtain a positive sample. and n negative samples j∈(1,m). Where the sample label corresponding to the positive sample is... The sample label corresponding to the negative sample is Where k is the total number of sample users, and m is the total number of training batches.

[0422] For example, the multimodal fusion machine master recognition model can be represented as a nonlinear mapping. Where x is the input of the model and θ represents the parameters of the model, the output score of the multimodal fusion machine main recognition model can be expressed by the following formula (6):

[0423] f θ (x)=σ(φ θ (x)) (6)

[0424] Among them, f θ (x) represents the multimodal recognition score, and σ is the Sigmoid function.

[0425] The second step is to input the acquired batch of training samples into the multimodal fusion machine main recognition model to obtain the recognition probability of each sample, and to take the negative sample that was misclassified as a positive sample by the model and has the highest corresponding recognition probability as the hard sample (or dynamic hard sample).

[0426] For example, hard samples can be represented by the following formula (7):

[0427]

[0428] The third step is to use both hard samples and positive samples to calculate the loss for this training session.

[0429] For example, using hard samples and positive samples When jointly calculating the loss for this training session, the loss calculation formula (8) can be as follows:

[0430]

[0431] Among them, L total,ij Let L be the training loss for this model, and L(,) be the cross-entropy loss function.

[0432] The fourth step is to update the model parameters using the model training loss.

[0433] For example, the model parameters can be updated based on the following formula (9).

[0434]

[0435] Where, θ * These are the updated model parameters.

[0436] Based on the above training process, it can be found that in the process of training the multimodal fusion machine recognition model, for the same sample user, m training samples can be collected. That is, the model is trained using m sets of training data for each sample user, instead of just using one set of data. This can improve the accuracy of the model.

[0437] In addition, by identifying difficult samples as negative samples in each batch of training samples, and by calculating the model training loss together with the dynamic difficult samples and positive samples, and updating the model parameters based on the model training loss, the false recognition rate can be effectively reduced and the TAR index can be improved at a lower false recognition rate. In other words, the multimodal fusion machine owner recognition model trained can be more accurate and more stable in machine owner recognition.

[0438] The following further refers to Table 1, Figure 7 and Figure 8 This application describes the effectiveness of the multimodal fusion-based owner identification method provided in its embodiments.

[0439] As shown in Table 1, compared with the best-performing face modality in the single-modality method, the TAR (Take-Off Rate) of the multimodal fusion method is improved by 4.17% when the FAR (False Recognition Rate) is 1%, by 4.54% when the FAR is 1.5%, and by 5.83% when the FAR is 5%. In other words, the multimodal fusion method improves the TAR index at a lower false recognition rate.

[0440] Table 1 Performance evaluation results of the multimodal fusion-based owner identification method

[0441]

[0442] In Table 1, all values ​​are percentages. For example, the AUC for the respiratory single modality is 89.36, which means 89.36%, or 0.8936.

[0443] Figure 7 The ROC curves for the various modalities provided in the embodiments of this application are for the owner identification. Figure 8 The PR curves for the host identification corresponding to the various modalities provided in the embodiments of this application are shown.

[0444] Combined with Table 1, Figure 7 and Figure 8 It can be observed that, compared with the best-performing single-modality face modality, the multi-modal fusion method for host recognition improves the AUC by 3.42% and the AP by 2.20%.

[0445] Summary Table 1 Figure 7 and Figure 8 It can be observed that the multimodal fusion-based owner identification method can effectively improve the overall identification success rate by leveraging the complementary advantages between multiple modalities.

[0446] (III) Update the host template.

[0447] In practical applications, the characteristics of the device owner may change over time, such as changes in facial appearance and breathing sounds. Therefore, it is necessary to update the device owner template.

[0448] In this embodiment of the application, the timing for updating the owner template entered in the mobile phone can be: a certain update for face registration, and a periodic update for face unlock.

[0449] The update cycle for updating the host template can be one day or one week. This application embodiment does not specifically limit the update cycle.

[0450] It should be noted that the presence of face recognition registration indicates the registration of a new owner template, necessitating an update of the owner template on the phone. Once face unlock is enabled and the update cycle has arrived, the registered owner template can be updated.

[0451] Since the process of entering the host template has already been detailed in Part (I) above, the following mainly explains how to update the entered host template, and specifically how to update the cross-modal feature template and the respiratory sound feature template.

[0452] In practice, for any device owner, the face template in the owner's template can be updated synchronously with the cross-modal feature template and the breath sound feature template, or they can be updated asynchronously. For example, when the update conditions (or template update conditions) are met, the face template can be updated first, followed by the cross-modal feature template and the breath sound feature template. Alternatively, when updating the face template, if no breath sounds are detected in the collected breath audio data, the update of the cross-modal feature template and the breath sound feature template may be temporarily suspended.

[0453] In this embodiment, during the update of an existing owner template, if the update cycle is reached and face unlock occurs, the update of the corresponding owner's owner template can be triggered. For example, if the phone has already recorded owner templates for owners A, B, and C, after the update cycle is reached, if face unlock occurs, for example, if owner A is recognized during face unlock, the update of owner A's owner template can be triggered. At this time, the phone can collect the corresponding owner's IMU data, touchscreen data, and breathing audio data when face unlock is successful. Then, it can call the pre-stored audio processing module (see...). Figure 4A The audio processing module processes the breathing audio data to extract breathing sound features, which represent the current breathing sound features of the user. These current breathing sound features are those captured when face unlock is successful. Afterward, the phone can access the pre-stored cross-modal encoder (see...). Figure 5A The cross-modal encoder fuses breathing sound features, touchscreen data, and IMU data to obtain the current cross-modal features of the device owner. These current cross-modal features correspond to the cross-modal features of the multimodal data collected when face unlock is successful.

[0454] In some optional implementations of this application, to improve the accuracy of the master template update, the breath sound feature template can be updated when the confidence level of the breath sound corresponding to the current breath sound feature is greater than the third determination threshold c. Here, c > a, and a is the lower limit threshold used to determine whether the breath sound data includes a breath sound. In some scenarios of this application, a is also referred to as the first determination threshold.

[0455] In other words, in some application scenarios, the update conditions (or template update conditions) for updating the recorded owner template include, in addition to having face unlock and reaching the update cycle, the confidence level of the current breathing sound feature corresponding to the breathing sound is greater than c.

[0456] Understandably, mobile phones typically update the owner template identified when face unlock is successful. For example, if face unlock is successful and the identified owner is A, then, provided the template update conditions are met, the phone can update the owner template for owner A.

[0457] (1) Owner template update in a single feature template scenario

[0458] A single feature template typically refers to a host template that includes only one feature template. For example, a host template may include a breath sound feature template, or in other words, one host corresponds to one breath sound feature template.

[0459] In some optional implementations of the embodiments of this application, if the above-mentioned update conditions (or template update conditions) for updating the main template are met, the mobile phone can directly replace the original feature template with the current feature. The current feature may include the current breathing sound feature and the current cross-modal feature.

[0460] For example, the current breath sound feature can replace the original breath sound feature template, and the current cross-modal feature can replace the original cross-modal feature template. That is, the current breath sound feature is used as the breath sound feature template, and the current cross-modal feature is used as the cross-modal feature template.

[0461] In some optional implementations, if the template update conditions are met, the mobile phone can also determine the new feature template for the corresponding modality by taking the average of the feature template in the master template and the current feature of the same modality, and replace the old feature template of the corresponding modality with the new feature template.

[0462] For example, a mobile phone can determine a new breathing sound feature template by comparing the breathing sound feature template in the master template with the mean of the current breathing sound features, and replace the old breathing sound feature template with the new one. Simultaneously, it can determine a new cross-modal feature template by comparing the cross-modal feature template in the master template with the mean of the current cross-modal features, and replace the old cross-modal feature template with the new one.

[0463] It should be noted that using the average of the current feature and the old feature template as the new feature template can achieve accurate and stable updates to the host template, which helps to improve the accuracy and stability of host identification.

[0464] (2) Owner template update in multiple feature template scenarios

[0465] Multiple feature templates typically refer to a device owner template that includes multiple feature templates. For example, a device owner template may include multiple breath sound feature templates, or in other words, a single device owner may have multiple breath sound feature templates.

[0466] In some optional implementations of the embodiments of this application, when the template update conditions are met, the mobile phone can determine the average value of each feature template in the main template and the current feature of the same modality as the new feature template of the corresponding modality, and add the new feature template to the main template.

[0467] For example, a mobile phone can determine a new breathing sound feature template by comparing the breathing sound feature template in the main template with the mean of the current breathing sound features, and add the new breathing sound feature template to the main template. Simultaneously, it can determine a new cross-modal feature template by comparing the cross-modal feature template in the main template with the mean of the current cross-modal features, and add the new cross-modal feature template to the main template.

[0468] To give a further example, if the host template record has p-1 cross-modal feature templates, namely S m,1 ,S m,2 ,...,S m,p-1 The current cross-modal feature is S m,p Then the new cross-modal feature template S m It can be calculated using the following formula (10).

[0469]

[0470] Similarly, if the master template records q-1 breath sound feature templates, respectively R m,1 ,R m,2 ,...,R m,q-1 The current breath sound characteristic is R m,q The new breath sound feature template Rm It can be calculated using the following formula (11).

[0471]

[0472] It should be noted that using the average of the current feature and the old feature template as the new feature template can achieve accurate and stable updates to the host template, which helps to improve the accuracy and stability of host identification.

[0473] In some optional implementations of this application's embodiments, the mobile phone may first perform a feature quality assessment on the owner's current features, and determine whether to update the feature template based on the feature quality assessment result. If the feature quality assessment result is poor, the owner's current features can be deleted. If the feature quality assessment result is good, the current features can be used to update the owner's template.

[0474] As an example, if the feature quality assessment result of the current cross-modal feature of the device owner is not good, the phone can directly delete the current cross-modal feature.

[0475] As another example, if the current breathing sound characteristics of the device owner are assessed for quality and determined to be of good quality, then the current breathing sound characteristics can be used to update the device owner template. For instance, the current breathing sound characteristics can replace the original breathing sound characteristic template, or the current breathing sound characteristics can be directly added to the device owner template.

[0476] In some optional implementations of the embodiments of this application, when the template update condition is met, the mobile phone can determine the similarity between each feature template in the main template and the current feature of the same modality, and then select the similarity with the largest value from multiple similarities. For ease of description, it can be called the maximum similarity.

[0477] Update method 1: If the maximum similarity is greater than the fifth judgment threshold e, then the current feature replaces the feature template corresponding to the maximum similarity. The fifth judgment threshold is a pre-set threshold, with a value ranging from 0 to 1. In the scenario where update method 1 is used, the current feature replaces the feature template corresponding to the maximum similarity.

[0478] Update method 2: If the maximum similarity is less than or equal to e, and the number of feature templates for the corresponding modality stored in the current phone is less than a pre-set template quantity threshold, then the feature template corresponding to the maximum similarity is retained, and the current feature is added to the main template. In scenarios where update method 2 is used, the feature template in the main template is increased by one.

[0479] Update method 3: If the maximum similarity is less than or equal to e, and the number of feature templates stored in the current phone is equal to the template quantity threshold, then delete the feature template with the longest storage time (i.e., the one furthest from the present) from the owner template, and add the current feature to the owner template. In the scenario where update method 3 is used, the current feature replaces the feature template with the longest storage time.

[0480] The following examples will further illustrate this with the respiratory modality.

[0481] If the similarity between the current breath sound feature R0 and the existing breath sound feature templates R1, R2, and R3 in the host template is S1, S2, and S3 respectively, the maximum similarity is S3, and S3 corresponds to the breath sound feature template R3.

[0482] If S3 is greater than the threshold e, the mobile phone can replace the current breath sound feature template R3 with the current breath sound feature R0.

[0483] If S3 is less than or equal to the threshold e, and the number of breathing sound feature templates in the current phone is less than the template quantity threshold, then the phone can add the current breathing sound feature R0 to the main template. At this time, there are 4 breathing sound feature templates in the main template.

[0484] If S3 is less than or equal to the threshold e, and the number of breathing sound feature templates in the current phone is equal to the template quantity threshold, then the phone can first delete the breathing sound feature template that has been saved the longest. For example, breathing sound feature template R1 can be deleted. After that, the current breathing sound feature R0 can be added to the main template.

[0485] In this embodiment of the application, by combining the number of feature templates stored in the mobile phone and the storage time of the feature templates, the main template of the mobile phone can be updated, which can realize flexible and effective updates of each main template of the mobile phone.

[0486] Corresponding to the owner identification method in the above embodiment, Figure 9 A schematic diagram of the host identification device in the host provided in the embodiment of this application is shown. For ease of explanation, only the parts related to the embodiment of this application are shown.

[0487] Reference Figure 9 The main identification device of the machine includes:

[0488] The data acquisition unit 901 is used to collect user data of the user under test in response to the unlocking operation of the target content by the user under test. The target content is in a locked state. The user data includes at least one of the following: the face image of the user under test, the breathing audio data of the user under test, the IMU data of the user under test, and the touch screen data of the user under test.

[0489] The recognition execution unit 902 is used to perform owner recognition based on user data and pre-entered owner templates. The owner template includes at least one of the following: face template, breathing sound feature template, and cross-modal feature template. The cross-modal feature template is used to describe the comprehensive features of the owner in multiple modalities.

[0490] The unlocking execution unit 903 is used to unlock target content when the owner is identified.

[0491] As one embodiment of this application, the owner identification device can achieve the following: Figures 1A-8 The embodiments shown and other related method embodiments.

[0492] For details on how each module in the owner identification device provided in this application implements its respective function, please refer to the foregoing. Figures 1A-8 The descriptions of the embodiments shown and other related method embodiments are not repeated here.

[0493] It should be noted that the host identification and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.

[0494] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0495] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.

[0496] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0497] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."

[0498] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only for distinguishing descriptions and should not be construed as indicating or implying relative importance. It should also be understood that although the terms "first," "second," etc., are used in the text to describe various elements in some embodiments of this application, these elements should not be limited by these terms. These terms are merely used to distinguish one element from another. For example, a first status code may be named a second status code, and similarly, a second status code may be named a first status code, without departing from the scope of the various described embodiments. Both the first and second status codes are status codes, but they are not the same status code.

[0499] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0500] Figure 10 This is a schematic diagram of the structure of a terminal device provided in an embodiment of this application.

[0501] The terminal device may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a sensor module 180, buttons 190, a camera 193, a touch screen 194, and a SIM card interface 195, etc.

[0502] The audio module 170 may include a speaker and a microphone.

[0503] The sensor module 180 may include a pressure sensor, an IMU, a distance sensor, a proximity light sensor, a fingerprint sensor, a touch sensor, and an ambient light sensor (of course, the terminal device may also include other sensors, such as a temperature sensor, a bone conduction sensor, etc., which are not shown in the figure).

[0504] Processor 110 may include one or more processing units, such as: application processor (AP), modem processor, graphics processing unit (GPU), image signal processor (ISP), controller, memory, video codec, digital signal processor (DSP), baseband processor, and / or neural network processing unit (NPU). Different processing units may be independent devices or integrated into one or more processors. The controller may serve as the central nervous system and command center of the terminal device. The controller can generate operation control signals based on instruction opcodes and timing signals to control instruction fetching and execution.

[0505] The processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store instructions or data that the processor 110 has just used or that are used repeatedly. If the processor 110 needs to use the instruction or data again, it can retrieve it directly from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.

[0506] The processor 110 can run the owner identification method provided in the embodiments of this application to quickly and accurately create schedules and improve the user experience. The processor 110 may include different devices. For example, when integrating a CPU and a GPU, the CPU and GPU can cooperate to execute the owner identification method provided in the embodiments of this application. For example, some algorithms in the owner identification method are executed by the CPU, and other algorithms are executed by the GPU to achieve faster processing efficiency.

[0507] It should be understood that the illustrated terminal device is merely an example, and the terminal device may have more or fewer components than those shown in the figure, may combine two or more components, or may have different component configurations. The various components shown in the figure may be implemented in hardware, software, or a combination of hardware and software, including one or more signal processing and / or application-specific integrated circuits.

[0508] The software system of a terminal device can adopt a layered architecture, event-driven architecture, microkernel architecture, microservice architecture, or cloud architecture. This application uses the layered architecture Android system as an example to illustrate the software structure of the terminal device.

[0509] Figure 11This is a software structure block diagram of a terminal device according to an embodiment of this application.

[0510] A layered architecture divides software into several layers, each with a clear role and function. Layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers, from top to bottom: the application layer, the application framework layer, the Android runtime and system libraries, and the kernel layer.

[0511] The application layer can include a series of application packages.

[0512] like Figure 11 As shown, the application package can include multiple applications such as gallery, calendar, map, WeChat, QQ, and SMS. It can also include multiple models such as a breathing sound recognition model, a breathing sound feature extraction model, a multimodal fusion master recognition model, a cross-modal encoder, and a fusion decision model.

[0513] The application framework layer provides application programming interfaces (APIs) and a programming framework for applications in the application layer. The application framework layer includes some predefined functions.

[0514] In some implementations, multiple models, such as the breathing sound recognition model, the breathing sound feature extraction model, the multimodal fusion master recognition model, the cross-modal encoder, and the fusion decision model, can also be deployed at the application framework layer.

[0515] like Figure 11 As shown, the application framework layer may include a window manager, content provider, view system, phone manager, resource manager, notification manager, etc.

[0516] The window manager is used to manage windowed applications. It can retrieve screen size, determine the presence of a status bar, lock the screen, and capture screenshots, among other things.

[0517] Content providers store and retrieve data, making that data accessible to applications. This data can include videos, images, audio, phone calls made and received, browsing history and bookmarks, phone books, and more.

[0518] A view system includes visual controls, such as controls for displaying text and controls for displaying images. View systems can be used to build applications. A display interface can consist of one or more views. For example, a display interface including a text notification icon could include views for displaying text and views for displaying images.

[0519] A phone manager is used to provide communication functions for terminal devices. For example, it manages call status (including connection and disconnection).

[0520] The file explorer provides applications with various resources, such as localized strings, icons, images, layout files, video files, and more.

[0521] The notification manager allows applications to display notifications in the status bar. These notifications can be used to deliver informational messages and can disappear automatically after a short pause, requiring no user interaction. For example, the notification manager can be used to notify users of download completion or message alerts. The notification manager can also display notifications as icons or scrolling text in the top status bar, such as notifications from background applications, or as dialog boxes on the screen. Examples include displaying text messages in the status bar, emitting sounds, vibrating the device, and flashing indicator lights.

[0522] The Android Runtime consists of core libraries and a virtual machine. The Android Runtime is responsible for the scheduling and management of the Android system.

[0523] The core library consists of two parts: one part is the functionalities that need to be called by the Java language, and the other part is the Android core library.

[0524] The application layer and application framework layer run in a virtual machine. The virtual machine executes the Java files of the application layer and application framework layer as binary files. The virtual machine is used to perform functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection.

[0525] System libraries can include multiple functional modules. For example: surface manager, media libraries, 3D graphics processing libraries (e.g., OpenGL ES), 2D graphics engines (e.g., SGL), etc.

[0526] The Surface Manager is used to manage the display subsystem and provides the blending of 2D and 3D layers for multiple applications.

[0527] The media library supports playback and recording of various common audio and video formats, as well as still image files. It supports multiple audio and video encoding formats, such as MPEG4, H.164, MP3, AAC, AMR, JPG, and PNG.

[0528] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, compositing, and layer processing.

[0529] A 2D graphics engine is a graphics engine for 2D drawing.

[0530] The kernel layer is the layer between hardware and software. The kernel layer contains at least the display driver, camera driver, audio driver, and sensor driver.

[0531] Furthermore, those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. In the various embodiments of this application, each functional unit can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0532] This application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps described in the various method embodiments above.

[0533] This application provides a computer program product that, when run on a terminal device, enables the terminal device to implement the steps described in the various method embodiments.

[0534] This application also provides a chip system, which includes a processor coupled to a memory. The processor executes a computer program stored in the memory to implement the steps in the above-described method embodiments.

[0535] If an integrated module / unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable storage medium can include: any entity or device capable of carrying computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc.

[0536] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0537] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0538] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0539] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A method for identifying the owner of a device, characterized in that, Applied to a terminal device, the method includes: In response to the user's unlocking operation on the target content, user data of the user is collected, wherein the target content is in a locked state, and the user data includes at least one of the following: the face image of the user to be tested, the breathing audio data of the user to be tested, the IMU data of the user to be tested, and the touch screen data of the user to be tested. If a face is present in the face image to be tested, the owner is identified based on the face image to be tested and the face template in the pre-recorded owner template to obtain a face recognition score. The owner template includes at least one of the following: face template, breathing sound feature template, and cross-modal feature template. The cross-modal feature template is used to describe the comprehensive features of the owner in multiple modalities. If the face recognition score is greater than the preset lower threshold and less than the preset upper threshold, the test breathing sound features of the test user are extracted from the test breathing audio data. The breath sound features to be tested, the IMU data to be tested, the touch screen data to be tested, and the face recognition score are input into a pre-trained multimodal fusion device owner recognition model to obtain device owner indication information, which is used to indicate whether the user to be tested is the device owner; If the owner of the device is identified, the target content is unlocked.

2. The owner identification method according to claim 1, characterized in that, The target content includes any of the following: locked notification messages, locked service cards, locked applications, locked functions, and locked interfaces.

3. The owner identification method according to claim 1, characterized in that, The process of inputting the face template includes: In response to the user's face registration operation, a face registration interface is displayed, which includes a face capture frame; In response to the acquisition of a face image, the acquired face image is recorded as the face template.

4. The owner identification method according to claim 1, characterized in that, The process of inputting the cross-modal feature template includes: Under the condition of meeting the multimodal feature input conditions, the device owner's breathing audio data, the terminal device's IMU data, and the terminal device's touch screen data are collected. The multimodal feature input conditions include any one of the following: the device owner's face template is successfully entered, the device owner's face unlock is successful, and the terminal device has not yet entered the cross-modal feature template of the device owner currently performing the unlock. Extract respiratory sound features from the respiratory audio data; The IMU data, the touchscreen data, and the breathing sound features are encoded by a cross-modal encoder to obtain cross-modal features, and the cross-modal features are recorded as cross-modal feature templates. The cross-modal encoder is used to encode data from multiple modalities.

5. The owner identification method according to claim 4, characterized in that, The cross-modal encoder includes a cross-channel attention encoder; The cross-channel attention encoder includes a first slice encoding module, a second slice encoding module, and a third slice encoding module in parallel. The first slice encoding module includes a first Group layer and a first Linear layer connected in sequence; the second slice encoding module includes a second Group layer and a second Linear layer connected in sequence; the third slice encoding module includes a third Group layer and a third Linear layer connected in sequence. The input of the first Linear layer is connected to the output of the first Group layer, the input of the second Linear layer is connected to the outputs of the first Linear layer and the second Group layer, and the input of the third Linear layer is connected to the outputs of the second Linear layer and the third Group layer.

6. The owner identification method according to claim 4, characterized in that, The cross-modal encoder includes a cross-channel linear encoder, wherein the cross-channel linear encoder includes a sequentially connected Linear layer, an activation function, a Linear layer, a normal layer, a conv layer, and an activation function.

7. The owner identification method according to claim 1, characterized in that, The method also includes: If the face recognition score is greater than the face recognition upper limit threshold, the user to be tested is determined to be the owner of the device. If the face recognition score is less than the lower limit threshold for face recognition, the user to be tested is determined to be a non-owner of the device.

8. The owner identification method according to claim 1, characterized in that, The data processing procedure of the main recognition model of the multimodal fusion machine includes: The cross-modal features to be tested are obtained by encoding the IMU data to be tested, the touch screen data to be tested, and the breath sound features to be tested using a cross-modal encoder. Determine the cross-modal similarity between the cross-modal feature to be tested and the cross-modal feature template in the host template, and determine the respiratory sound similarity between the respiratory sound feature to be tested and the respiratory sound feature template in the host template; Based on the cross-modal similarity, the respiratory sound similarity, and the face recognition score, a multimodal recognition score is determined, wherein the multimodal recognition score is used to describe the comprehensive recognition result of the test user under multiple modalities; Based on the multimodal recognition score and the pre-set fusion threshold, it is determined whether the user to be tested is the owner of the device.

9. The owner identification method according to claim 8, characterized in that, When multiple host templates are entered into the terminal device, the multimodal recognition score can be multiple. The step of determining whether the user to be tested is the owner of the device based on the multimodal recognition score and a pre-set fusion threshold includes: The maximum score among multiple multimodal recognition scores is selected. If the selected multimodal recognition score is greater than the fusion threshold, the user to be tested is determined to be the owner of the device.

10. The owner identification method according to claim 1, characterized in that, The training process of the multimodal fusion machine master recognition model includes: From the sample dataset used to train the multimodal fusion machine master recognition model, m batches of training samples are selected. Each batch of training samples includes 1 positive sample and n negative samples, where n>1 and m>1. During the training process of the multimodal fusion machine master recognition model using each batch of training samples, the negative samples that are identified as positive samples and have the highest corresponding recognition probability are identified as difficult samples. The model training loss of the multimodal fusion machine master recognition model is determined based on the positive samples and the hard samples, and the model parameters of the multimodal fusion machine master recognition model are adjusted based on the determined model training loss.

11. The owner identification method according to any one of claims 1-10, characterized in that, The method further includes: If the template update conditions are met, the average value of the breathing sound feature template in the main device template and the current breathing sound feature is determined as the new breathing sound feature template, and the new breathing sound feature template is added to the main device template. The template update conditions include: successful face unlock and reaching the update cycle. The mean of the cross-modal feature template in the host template and the current cross-modal feature is determined as the new cross-modal feature template, and the new cross-modal feature template is added to the host template; The current breathing sound feature is the breathing sound feature in the breathing audio data collected when face unlock is successful.

12. The owner identification method according to any one of claims 1-11, characterized in that, The method further includes: If the template update conditions are met, the average value of the cross-modal feature template in the host template and the current cross-modal feature is determined as the new cross-modal feature template, and the new cross-modal feature template is added to the host template. The current cross-modal feature is the cross-modal feature corresponding to the multimodal data collected when face unlock is successful.

13. A terminal device, characterized in that, The terminal device includes a memory and a processor. The memory stores a computer program that can run on the processor. When the processor executes the computer program, it implements the owner identification method as described in any one of claims 1 to 12.