Method and device for unlocking electronic device

By performing two voiceprint feature comparisons on the wake-up word and command word input by the user, the problems of low accuracy and poor security of the voice unlocking solution in the existing technology are solved, and an efficient and secure automatic unlocking process is achieved.

CN114444042BActive Publication Date: 2025-09-23HUAWEI DEVICE CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202011188127.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-10-30
Publication Date
2025-09-23
Estimated Expiration
2040-10-30

AI Technical Summary

Technical Problem

Existing voice-based unlocking solutions have low accuracy, are easily affected by the environment, cannot prevent recording, speech synthesis and voice imitation attacks, and manual unlocking by users is cumbersome.

Method used

By performing two voiceprint feature comparisons on the wake-up word and command word entered by the user, the accuracy of voiceprint feature comparison is improved, and the user command is automatically executed after a successful match, reducing user manual operations.

Benefits of technology

It improves the security and convenience of voiceprint unlocking, prevents imitation and synthesis attacks, simplifies the unlocking process, and improves unlocking efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114444042B_ABST
    Figure CN114444042B_ABST
Patent Text Reader

Abstract

The present application provides a method and device for unlocking an electronic device, which relates to the field of smart terminal technology and is used to improve the unlocking security of terminal devices. In this method, the electronic device can receive the first voice data input by the user in the lock screen state, and extract the first voiceprint feature, and when the first voiceprint feature is consistent with the preset reference voiceprint feature comparison result, and the first voice data contains the specified text, it is still in the lock screen state. The electronic device can extract the second voiceprint feature of the second voice data, and when the second voiceprint feature is consistent with the preset reference voiceprint feature comparison result, the screen is unlocked and the control instruction is executed. In this way, when the user needs to unlock the electronic device, he can input voice data into the electronic device, and the electronic device can perform voiceprint feature comparison on the wake-up word and control command input by the user respectively, and unlock when the results of the two voiceprint feature comparisons indicate a successful match, thereby improving the security of voiceprint unlocking.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of smart terminal technology, and in particular to a method and device for unlocking an electronic device. Background Art

[0002] The industry currently offers solutions for unlocking terminal devices based on face, fingerprint, iris, or voice. However, most current voice-based unlocking solutions rely on voiceprint recognition based on a wake-up word entered by the user to unlock the terminal device. However, since the wake-up word is usually a short text, the accuracy of voiceprint recognition is low, failing to meet unlocking security requirements. Furthermore, these solutions are unable to protect against attacks such as recording, speech synthesis, and voice imitation, resulting in low security.

[0003] To improve unlocking security, the user can also manually unlock the terminal device after entering the wake-up word for voiceprint recognition. However, this solution requires manual user access, which loses the convenience of voice unlocking and is not smart enough. Summary of the Invention

[0004] The embodiments of the present application provide a method and apparatus for unlocking an electronic device, which are used to reduce the user's manual operations on the basis of improving the unlocking security of the terminal device, so as to improve the efficiency of the unlocking process.

[0005] In a first aspect, an embodiment of the present application provides a method for unlocking an electronic device. The method can be executed by the electronic device provided in the present application, or by a chip with similar electronic device functions. In the method, the electronic device can receive first voice data input by the user in a locked screen state and extract a first voiceprint feature of the first voice data. The electronic device can remain in a locked screen state when the first voiceprint feature is consistent with the preset reference voiceprint feature comparison result and the first voice data contains specified text. The electronic device can receive second voice data input by the user. The second voice data can contain a control instruction, which can be used to trigger at least one functional requirement of the user. The electronic device can extract the second voiceprint feature of the second voice data. When the second voiceprint feature is consistent with the preset reference voiceprint feature comparison result, the electronic device can unlock the screen and execute the control instruction to trigger the at least one functional requirement.

[0006] Based on the above solution, when a user needs to unlock an electronic device, he or she can input voice data into the electronic device. The electronic device can perform voiceprint feature comparison on the two voice data input by the user, that is, the wake-up word and the control command input by the user. When the results of both voiceprint feature comparisons indicate a successful match, the electronic device can be unlocked. Therefore, the accuracy of the voiceprint feature comparison results can be improved, and attacks such as imitation, recording, and synthesis can be effectively avoided, thereby improving the security of voiceprint unlocking. In addition, since both voiceprint feature comparisons are performed on the voice data input by the user, there is no need for manual operation by the user, which can provide convenience for the user. In addition, the control instructions carried in the second voice data input can enable the electronic device to execute the function required by the user after successful unlocking, thereby also improving the efficiency of the user in executing a certain requirement.

[0007] In one possible implementation, the electronic device may receive first registration voice data input by the user; the first registration voice data includes the specified text. The electronic device may obtain a first registration voiceprint feature of the first registration voice data; the electronic device may receive second registration voice data input by the user; the second registration voice data and the first registration voice data are from the same user; the electronic device may obtain a second registration voiceprint feature of the second registration voice data; and the electronic device may store the first registration voiceprint feature and the second registration voiceprint feature as the reference voiceprint feature.

[0008] Based on the above solution, the electronic device can obtain the user's voiceprint features as reference voiceprint features through the first registration voice data and the second registration voice data input by the user, so as to enable the user to unlock the electronic device by voiceprint.

[0009] In a possible implementation, the electronic device may combine the first registered voiceprint feature and the second registered voiceprint feature to obtain a third registered voiceprint feature; and the electronic device may store the third registered voiceprint feature as the reference voiceprint feature.

[0010] Based on the above solution, electronic devices can fuse two voiceprint features to improve the accuracy of voiceprint feature comparison and effectively prevent attacks such as voice imitation, voice synthesis and recording.

[0011] In a possible implementation, the electronic device may combine the second voiceprint feature and the first voiceprint feature to obtain a third voiceprint feature; and the electronic device may unlock the screen when the comparison result of the third voiceprint feature is consistent.

[0012] Based on the above scheme, the electronic device can merge the first voiceprint feature and the second voiceprint feature extracted based on the first voice data and the second voice data input by the user when performing voiceprint unlocking, and compare them with the stored reference voiceprint feature, which can improve the accuracy of voiceprint feature comparison.

[0013] In one possible implementation, the electronic device may use a pre-trained first voiceprint feature model to obtain the first voiceprint feature of the first voice data; the first voiceprint feature model is trained based on multiple first voice data labeled with speakers; the electronic device may use a pre-trained second voiceprint feature model to obtain the second voiceprint feature of the second voice data; the second voiceprint feature model is trained based on multiple second voice data labeled with speakers.

[0014] Based on the above solution, the electronic device can obtain the voiceprint features of the voice data input by the user based on the pre-trained voiceprint feature model, and can quickly and accurately extract the voiceprint features of the user.

[0015] An embodiment of the present application provides an electronic device, such as a foldable screen electronic device. The electronic device includes: one or more processors; a memory; multiple application programs; and one or more computer programs, wherein the one or more computer programs are stored in the memory and include instructions that, when executed by the electronic device, cause the electronic device to perform the technical solution of the first aspect and any possible implementation of the first aspect.

[0016] In a third aspect, an embodiment of the present application provides a chip that is coupled to a memory in an electronic device, and is used to call a computer program stored in the memory and execute the technical solution of the first aspect of the embodiment of the present application and any possible design of the first aspect thereof; in an embodiment of the present application, "coupling" refers to the direct or indirect combination of two components with each other.

[0017] In a fourth aspect, embodiments of the present application further provide a circuit system. The circuit system may be one or more chips, such as a system-on-a-chip (SoC). The circuit system includes: at least one processing circuit; the at least one processing circuit is configured to execute the technical solution of the first aspect of the embodiment of the present application and any possible design of the first aspect.

[0018] In the fifth aspect, an embodiment of the present application also provides an electronic device, which includes modules / units for executing the above-mentioned first aspect or any possible design method of the first aspect; these modules / units can be implemented through hardware, or corresponding software implementations can be executed through hardware.

[0019] In the sixth aspect, an embodiment of the present application also provides a computer-readable storage medium, which includes a computer program. When the computer program runs on an electronic device, the electronic device executes the technical solution of the first aspect of the embodiment of the present application and any possible design of the first aspect thereof.

[0020] In the seventh aspect, a program product in an embodiment of the present application includes instructions. When the program product runs on an electronic device, the electronic device executes the technical solution of the first aspect of the embodiment of the present application and any possible design of the first aspect thereof.

[0021] In addition, the beneficial effects of the second to seventh aspects can refer to the beneficial effects of the first aspect and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1A This is one of the schematic diagrams of recording the unlocking voice of an electronic device provided in an embodiment of the present application;

[0023] Figure 1B This is one of the schematic diagrams for unlocking an electronic device provided in an embodiment of the present application;

[0024] Figure 1C This is one of the schematic diagrams for unlocking an electronic device provided in an embodiment of the present application;

[0025] Figure 2 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application;

[0026] Figure 3 A schematic diagram of the software structure of the electronic device provided in an embodiment of the present application;

[0027] Figure 4 A schematic diagram of the process of training a voiceprint feature model provided in an embodiment of the present application;

[0028] Figure 5A This is one of the schematic diagrams of recording the unlocking voice of an electronic device provided in an embodiment of the present application;

[0029] Figure 5B This is one of the schematic diagrams of recording the unlocking voice of an electronic device provided in an embodiment of the present application;

[0030] Figure 6 A schematic diagram of the process of fusing voiceprint features provided in an embodiment of the present application;

[0031] Figure 7 One of the exemplary flow charts of the method for unlocking an electronic device provided in an embodiment of the present application;

[0032] Figure 8A This is one of the schematic diagrams for unlocking an electronic device provided in an embodiment of the present application;

[0033] Figure 8B This is one of the schematic diagrams for unlocking an electronic device provided in an embodiment of the present application;

[0034] Figure 9 This is one of the schematic diagrams for unlocking an electronic device provided in an embodiment of the present application;

[0035] Figure 10 This is one of the schematic diagrams for unlocking an electronic device provided in an embodiment of the present application;

[0036] Figure 11A This is one of the schematic diagrams for unlocking an electronic device provided in an embodiment of the present application;

[0037] Figure 11B A schematic diagram of locking an electronic device provided in an embodiment of the present application;

[0038] Figure 12 This is one of the schematic diagrams for unlocking an electronic device provided in an embodiment of the present application;

[0039] Figure 13 This is one of the schematic diagrams for unlocking an electronic device provided in an embodiment of the present application;

[0040] Figure 14 This is one of the schematic diagrams for unlocking an electronic device provided in an embodiment of the present application;

[0041] Figure 15 One of the exemplary flow charts of the method for unlocking an electronic device provided in an embodiment of the present application;

[0042] Figure 16 A schematic diagram of a scenario for unlocking an electronic device provided in an embodiment of the present application;

[0043] Figure 17 A block diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0044] Currently, terminal devices can be unlocked using face, iris, fingerprint, or voice. The process of voice unlocking is as follows:

[0045] See Figure 1A , the user can record an unlock voice in advance, and the terminal device can obtain the user's voiceprint characteristics based on the unlock voice recorded by the user and store the voiceprint characteristics. When the user needs to unlock the terminal device, refer to Figure 1B The user can input the aforementioned unlocking voice, and the terminal device can obtain the voiceprint feature based on the unlocking voice. The terminal device can compare the obtained voiceprint feature with the stored voiceprint feature. When the obtained voiceprint feature is consistent with the stored voiceprint feature, it determines to unlock and presents the main interface to the user. The main interface includes multiple application icons. For details, refer to Figure 1B shown.

[0046] However, in the above-mentioned voiceprint unlocking solution, since the unlocking voice is generally a voice with a specified short text content such as a wake-up word, for example, only a set voice such as "Please turn on the screen" can be input. Therefore, the voiceprint features that can be obtained based on the unlocking voice are relatively few. In addition, since the sound is greatly affected by the external environment, the voiceprint features obtained when the unlocking voice is recorded in advance may also be inaccurate. Therefore, the above-mentioned solution may fail to correctly unlock the terminal device, that is, the accuracy of the voice-based unlocking method of the terminal device is low. Not only that, since the unlocking voice is a voice with a specified short text content, the above-mentioned solution cannot solve attacks such as recording, speech synthesis and voice imitation, and the security of unlocking is also low.

[0047] In order to improve the security of the voice-based unlocking method of the terminal device, after the user enters the unlocking voice and passes the voiceprint feature comparison, if the unlocking fails, the user may need to manually enter the password or unlock fingerprint and other information for further verification. Figure 1C As shown, the terminal device may display "Voice unlock failed, please select fingerprint unlock or password unlock", at which point the user can choose to enter a numeric password through the numeric keypad or enter fingerprint information in the fingerprint recognition area for decoding. However, the above solution is not smart enough for users and is also relatively cumbersome, losing the convenience and efficiency of voiceprint unlocking.

[0048] Based on this, the present application proposes a new electronic device unlocking solution to avoid the above-mentioned problems and improve the security and efficiency of electronic devices based on voice unlocking. The embodiments of the present application can be applied to various types of electronic devices, such as electronic devices with curved screens, full screens, folding screens, etc. Electronic devices such as mobile phones, tablet computers, wearable devices (for example, watches, bracelets, smart helmets, etc.), vehicle-mounted equipment, smart homes, augmented reality (AR) / virtual reality (VR) devices, laptops, ultra-mobile personal computers (UMPC), netbooks, personal digital assistants (PDAs), etc. are not limited here in this application.

[0049] The method provided in the embodiments of this application considers that wake-up words are mostly short, specified text content. Therefore, the accuracy of voiceprint feature comparison based on wake-up words is insufficient. It also fails to address the low security issues caused by attacks such as recording, speech synthesis, or voice imitation. Therefore, it is proposed to improve the accuracy of voiceprint feature comparison by comparing multiple voiceprint features input by the user. In the embodiments of this application, a first voiceprint feature comparison can be performed on the wake-up word input by the user to determine whether the person who said the wake-up word is the same user as the user who registered. If the recognition result is the same person, the electronic device can perform a second voiceprint feature comparison on the command word input by the user to determine whether the person who said the command word is the same user as the user who registered. If the recognition result is the same person, the device can unlock and parse the text-related information contained in the command word, thereby executing the task associated with the command word. Therefore, the accuracy of voiceprint feature comparison can be improved, enhancing the security of unlocking the electronic device. Furthermore, since the voiceprint feature comparison process does not require manual user intervention, it provides convenience for the user, making the unlocking process more intelligent and efficient.

[0050] In order to fully understand the technical solutions provided by the embodiments of the present application, the terms appearing in the embodiments of the present application are first explained below.

[0051] 1) Wake-up word, which refers to the voice of a fixed short text content.

[0052] 2) Voiceprint recognition is a type of biometric technology, also known as speaker recognition, speaker identification, and speaker confirmation. It can convert a person's voice pattern into an electrical signal, and then use computer technology to extract the voice characteristics to identify the speaker.

[0053] 3) Voiceprint feature comparison refers to extracting voiceprint features from the voice input by the speaker and comparing them with the voiceprint features of the voice input by the user during registration.

[0054] The technical solutions in the embodiments of the present application will be described in detail below in conjunction with the drawings in the following embodiments of the present application.

[0055] The terms used in the following embodiments are only for the purpose of describing specific embodiments and are not intended to be limiting of the present application. As used in the specification of the present application and the appended claims, the singular expressions "a", "an", "said", "above", "the" and "this" are intended to also include expressions such as "one or more", unless there is a clear contrary indication in the context. It should also be understood that in the embodiments of the present application, "one or more" refers to one, two or more; "and / or" describes the association relationship of associated objects, indicating that three relationships may exist; for example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone, where A and B can be singular or plural. The character " / " generally indicates that the related objects before and after are in an "or" relationship.

[0056] References to "one embodiment" or "some embodiments" in this specification mean that a particular feature, structure, or characteristic described in conjunction with that embodiment is included in one or more embodiments of the present application. Thus, phrases such as "in one embodiment," "in some embodiments," "in other embodiments," and "in yet other embodiments" appearing in various places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more but not all embodiments," unless otherwise specifically emphasized. The terms "including," "comprising," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0057] The at least one involved in the embodiments of the present application includes one or more; wherein, more means greater than or equal to two. In addition, it should be understood that in the description of this application, the words "first" and "second" are only used for the purpose of distinguishing descriptions and cannot be understood as indicating or implying relative importance or order.

[0058] In the following embodiments, an electronic device is a mobile phone as an example for introduction. Various applications (applications, apps) can be installed in the mobile phone, which can be referred to as applications for short, to be software programs that can realize one or more specific functions. For example, applications include instant messaging applications, video applications, audio applications, image capture applications, and the like. Among them, instant messaging applications, for example, may include text messaging applications, WeChat (WeChat), WhatsApp Messenger, Line, photo sharing (Instagram), Kakao Talk, DingTalk, etc. Image capture applications, for example, may include camera applications (system cameras or third-party camera applications). Video applications, for example, may include Youtube, Twitter, Tik Tok, iQiyi, Tencent Video, etc. Audio applications, for example, may include Kugou Music, Xiami, QQ Music, etc. The applications mentioned in the following embodiments may be applications that are installed on the electronic device when it leaves the factory, or they may be applications that the user downloads from the Internet or obtains from other electronic devices during the use of the electronic device.

[0059] The present invention provides a method for unlocking an electronic device. The method can be applied to any electronic device. Figure 2 As shown in FIG, it is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application. The electronic device can be a mobile phone (folding screen mobile phone or non-folding screen mobile phone), a tablet computer (folding tablet computer or non-folding tablet computer), etc. Figure 2 As shown, the electronic device may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, an earphone interface 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc.

[0060] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU). The different processing units may be independent devices or integrated into one or more processors. The controller may serve as the nerve center and command center of the electronic device. The controller may generate operation control signals based on instruction opcodes and timing signals to control instruction fetching and execution. The processor 110 may also include memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a high-speed cache memory. This memory may store instructions or data that have just been used or are being recycled by the processor 110. If the processor 110 needs to use the instruction or data again, it can directly retrieve it from the memory. This avoids duplicate accesses, reduces the processor 110's waiting time, and thus improves system efficiency. The processor 110 may receive the voice data input by the audio module 170 and may obtain the voiceprint features of the voice data.

[0061] The USB interface 130 complies with USB standards and may be a Mini USB interface, a Micro USB interface, or a USB Type-C interface. The charging management module 140 receives charging input from a charger. The power management module 141 connects the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140 to power the processor 110, the internal memory 121, the external memory, the display 194, the camera 193, and the wireless communication module 160.

[0062] The wireless communication function of the electronic device can be implemented through antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modem processor and baseband processor. The mobile communication module 150 can provide solutions for wireless communications including 2G / 3G / 4G / 5G applied to electronic devices. The mobile communication module 150 may include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves from antenna 1, filter and amplify the received electromagnetic waves, and transmit them to the modem processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modem processor and convert it into electromagnetic waves for radiation through antenna 1.

[0063] The wireless communication module 160 can provide wireless communication solutions for electronic devices, including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared technology (IR), etc.

[0064] In some embodiments, the antenna 1 of the electronic device is coupled to the mobile communication module 150, and the antenna 2 is coupled to the wireless communication module 160, so that the electronic device can communicate with the network and other devices through wireless communication technology. The wireless communication technology may include global system for mobile communications (GSM), general packet radio service (GPRS), code division multiple access (CDMA), wideband code division multiple access (WCDMA), time-division code division multiple access (TD-SCDMA), long term evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technology. The GNSS may include global positioning system (GPS), global navigation satellite system (GLONASS), Beidou navigation satellite system (BDS), quasi-zenith satellite system (QZSS) and / or satellite based augmentation system (SBAS).

[0065] The display screen 194 is used to display the display interface of the application, etc. The display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode or an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), Miniled, MicroLed, Micro-oLed, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the electronic device may include 1 or N display screens 194, where N is a positive integer greater than 1. In some embodiments, the display screen 194 may light up when the electronic device enters wake-up mode, or may present the main interface to the user after the electronic device is unlocked.

[0066] The camera 193 is used to capture static images or videos. In some embodiments, the camera 193 may include at least one camera, such as a front camera and a rear camera.

[0067] The internal memory 121 can be used to store computer executable program code, which includes instructions. The processor 110 executes various functional applications and data processing of the electronic device by running the instructions stored in the internal memory 121. The internal memory 121 may include a program storage area and a data storage area. The program storage area may store an operating system and software code of at least one application (such as the iQiyi application, WeChat application, etc.). The data storage area may store data (such as images, videos, etc.) generated during the use of the electronic device. In addition, the internal memory 121 may include a high-speed random access memory and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc. In some embodiments, the internal memory 121 may store the voiceprint features of the user during registration and the model for extracting the voiceprint features of the voice data. For example, the internal memory 121 may store the first voiceprint feature model and the second voiceprint feature model in the embodiment of the present application.

[0068] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device. The external memory card communicates with the processor 110 via the external memory interface 120 to implement data storage functions. For example, files such as pictures and videos can be stored on the external memory card.

[0069] The electronic device can implement audio functions through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the headphone jack 170D, and the application processor. For example, the receiver 170B and the microphone 170C can receive voice data input by the user. The receiver 170B or the microphone 170C can be turned on when the electronic device is in sleep mode, or can also be turned on when the electronic device is in wake-up mode.

[0070] The sensor module 180 may include a fingerprint sensor 180H, a touch sensor 180K, and the like.

[0071] Fingerprint sensor 180H is used to collect fingerprints. Electronic devices can use the collected fingerprint characteristics to achieve fingerprint unlocking, access application locks, fingerprint photography, fingerprint answering calls, etc.

[0072] The touch sensor 180K is also called a "touch panel." The touch sensor 180K can be disposed on the display screen 194. The touch sensor 180K and the display screen 194 form a touch screen, also called a "touch screen." The touch sensor 180K is used to detect touch operations applied on or near the touch sensor. The touch sensor can transmit the detected touch operations to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through the display screen 194. In other embodiments, the touch sensor 180K can also be disposed on the surface of the electronic device, in a location different from that of the display screen 194.

[0073] The button 190 includes a power button, a volume button, etc. The button 190 can be a mechanical button. It can also be a touch button. The electronic device can receive button input and generate key signal input related to the user settings and function control of the electronic device. The motor 191 can generate a vibration prompt. The motor 191 can be used for incoming call vibration prompts, and can also be used for touch vibration feedback. For example, touch operations acting on different applications (such as taking pictures, audio playback, etc.) can correspond to different vibration feedback effects. The indicator 192 can be an indicator light, which can be used to indicate the charging status, power changes, and can also be used to indicate messages, missed calls, notifications, etc. The SIM card interface 195 is used to connect the SIM card. The SIM card can be inserted into the SIM card interface 195 or pulled out from the SIM card interface 195 to achieve contact and separation with the electronic device.

[0074] It is understandable that Figure 2 The components shown do not constitute a specific limitation on the mobile phone. The mobile phone may also include more or fewer components than shown in the figure, or combine some components, or split some components, or arrange the components differently. Figure 2 The electronic device shown is used as an example for introduction.

[0075] Figure 3 FIG1 shows a software structure block diagram of an electronic device provided by an embodiment of the present application. Figure 3 As shown, the software structure of the electronic device can be a layered architecture. For example, the software can be divided into several layers, each with a clear role and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers: from top to bottom, the application layer, the application framework layer (framework, FWK), the Android runtime (Android runtime) and system library, and the kernel layer.

[0076] The application layer can include a series of application packages. Figure 3 As shown, the application layer may include camera, settings, skin module, user interface (UI), third-party applications, etc. Among them, third-party applications may include WeChat, QQ, gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, video, short message, etc.

[0077] The application framework layer provides application programming interface (API) and programming framework for the application layer. The application framework layer may include some predefined functions. Figure 4 As shown, the application framework layer may include a window manager, a content provider, a view system, a phone manager, a resource manager, a notification manager, and the like.

[0078] The window manager manages windowed applications. It can determine the display size, determine whether a status bar is present, lock the screen, and take screenshots. Content providers store and retrieve data and make it accessible to applications. This data can include video, images, audio, incoming and outgoing calls, browsing history and bookmarks, and phone books.

[0079] The view system includes visual controls, such as those for displaying text and images. The view system is used to build applications. A display interface can consist of one or more views. For example, a display interface containing a text notification icon might include a view for displaying text and a view for displaying images.

[0080] The phone manager is used to provide communication functions of the electronic device 100, such as management of call status (including answering, hanging up, etc.).

[0081] The resource manager provides various resources for applications, such as localized strings, icons, images, layout files, video files, and so on.

[0082] The Notification Manager allows applications to display notifications in the status bar. These messages can be displayed briefly and then disappear automatically without user interaction. For example, the Notification Manager is used to notify users of completed downloads and message reminders. The Notification Manager can also display notifications in the top status bar of the system as icons or scrolling text, such as notifications from background applications, or as dialog windows on the screen. Examples include text messages in the status bar, beeps, vibrations on electronic devices, and flashing indicator lights.

[0083] The Android runtime includes the core library and the virtual machine. The Android runtime is responsible for scheduling and management of the Android system.

[0084] The core library consists of two parts: one containing the Java language's callable functions and the other the Android core library. The application layer and the application framework layer run in a virtual machine. The virtual machine executes the Java files in the application layer and application framework layer as binary files. The virtual machine manages object lifecycles, stack management, thread management, security and exception management, and garbage collection.

[0085] The system library can include multiple functional modules, such as a surface manager, media libraries, a 3D graphics processing library (such as OpenGL ES), and a 2D graphics engine (such as SGL).

[0086] The surface manager is used to manage the display subsystem and provide fusion of 2D and 3D layers for multiple applications.

[0087] The media library supports playback and recording of a variety of common audio and video formats, as well as static image files. The media library can support a variety of audio and video encoding formats, such as: MPEG4, H.264, MP3, AAC, AMR, JPG, PNG, etc.

[0088] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, compositing, and layer processing.

[0089] A 2D graphics engine is a drawing engine for 2D drawings.

[0090] In addition, the system library may also include display / rendering services and energy-saving display control services. The display / rendering service is used to determine the display digital stream, which includes display information for each pixel unit (hereinafter referred to as a pixel) on the display screen. Display information may include display brightness, display time, text information, image information, etc. The kernel layer is the layer between hardware and software. The kernel layer includes at least display drivers, camera drivers, audio drivers, and sensor drivers.

[0091] The hardware layer may include various sensors, such as the acceleration sensor, gyroscope sensor, touch sensor, etc. involved in the embodiments of the present application.

[0092] The following describes the software and hardware workflow of an electronic device in conjunction with the electronic device unlocking method according to an embodiment of the present application.

[0093] In the embodiment of the present application, the speaker's input voice needs to be verified twice, so the voiceprint feature acquisition model in the embodiment of the present application can be divided into a first voiceprint feature model and a second voiceprint feature model. The first voiceprint feature model is obtained by training based on the specified voice data input by the speaker, and the second voiceprint feature model is obtained by training based on the random voice data input by the speaker. The following first introduces the training method of the first voiceprint feature model in the embodiment of the present application, and then Figure 4 , which may include the following steps.

[0094] Step 401: Acquire first voice data of specified text content.

[0095] The first voice data of the specified text content here can be a specified short text voice data such as a wake-up word. For example, it can be a wake-up word such as "Hello Xiaoyi" that can wake up the electronic device. Among them, the first voice data input by different speakers can be obtained, and each obtained first voice data has been labeled with the speaker.

[0096] In one example, the first voice data input by the speaker in different environments may be obtained, for example, the first voice data may be input in a car, indoors, outdoors, or other environments.

[0097] Step 402: Using a deep neural network model, first speech data that has been labeled with a speaker is learned to obtain a first voiceprint feature model.

[0098] In an embodiment of the present application, each piece of first speech data labeled with a speaker can be used as input, and the voiceprint features of each piece of first speech data can be obtained through deep learning. Since each piece of first speech data is labeled with a speaker, the parameters of the first voiceprint feature model can be learned based on the correspondence between the acquired voiceprint features and the speaker, and the first voiceprint feature model can be constructed based on the parameters of the first voiceprint feature model. The first voiceprint feature model can extract the voiceprint features of the input first speech data.

[0099] Step 403: transplant the trained first voiceprint feature model to the electronic device.

[0100] Through steps 401 and 402 above, the first voiceprint feature model in the voiceprint feature model of the embodiment of the present application can be trained. Therefore, the trained first voiceprint feature model can be transplanted to an electronic device through step 403. The electronic device can then use the first voiceprint feature model to extract voiceprint features from the first voice data input by the user.

[0101] After introducing the training method of the first voiceprint feature model in the embodiment of the present application, the following introduces the training method of the second voiceprint feature model in the embodiment of the present application. Figure 4 , which may include the following steps:

[0102] Step 401: Acquire second voice data of random text content.

[0103] The second voice data here can be voice data containing random text content of a control command. For example, it can be a control command such as "Query today's weather" or "Play music." The second voice data input by different speakers can be obtained, and the speaker is labeled for each piece of the second voice data.

[0104] In one example, the second voice data of the speaker can be obtained in different application scenarios. For example, in-car scenarios, encyclopedia search scenarios, taxi scenarios, food delivery scenarios, etc. For example, in the in-car scenario, the speaker can enter the second voice data such as "Please open the navigation" or "Go to place A".

[0105] In another example, the second voice data of the speaker can be obtained in different environments, such as in a car, indoors, and outdoors. For example, the speaker can enter the second voice data such as "Please open the navigation" or "Please play music" in a car scene.

[0106] For another example, the speaker may enter the second voice data such as "How is the weather today" for an encyclopedia query scenario indoors.

[0107] Step 402: Using a deep neural network model, learn the second speech data that has been labeled with the speaker to obtain a second voiceprint feature model.

[0108] In an embodiment of the present application, each piece of second voice data labeled with a speaker can be used as input, and the voiceprint features of each piece of second voice data can be obtained through deep learning. Since each piece of second voice data is labeled with a speaker, the parameters of the second voiceprint feature model can be learned based on the correspondence between the acquired voiceprint features and the speaker, and then a second voiceprint feature model can be constructed based on the learned parameters of the second voiceprint feature model. The second voiceprint feature model can extract the voiceprint features of the input second voice data.

[0109] Step 403: transplant the trained second voiceprint feature model to the electronic device.

[0110] Through the above steps 401 and 402, the second voiceprint feature model in the voiceprint feature model of the embodiment of the present application can be trained. Therefore, the trained second voiceprint feature model can be transplanted to the electronic device through step 403. The electronic device can then use the second voiceprint feature model to extract voiceprint features from the second voice data input by the user.

[0111] After introducing the training method of the voiceprint feature model in the embodiment of the present application, the following describes the electronic device unlocking method provided in the embodiment of the present application in combination with the figure.

[0112] First, the user can pre-register voiceprint features on the electronic device. The electronic device can compare the user's pre-registered voiceprint features with the voiceprint features extracted from the first voice data through voiceprint recognition to determine whether the voiceprint recognition is successful. The following describes how users register voiceprint features.

[0113] See Figure 5A , users can register voiceprint features on electronic devices according to the prompts. Figure 5A As shown, after the voiceprint feature registration is activated, the electronic device can display "Please say, Hello Xiaoyi" on the display screen. The user can then say "Hello, Xiaoyi" according to the prompt. Therefore, the electronic device can obtain the first registered voice data of "Hello, Xiaoyi" input by the user through sensors such as the receiver and microphone. The electronic device can extract the first registered voiceprint feature from the first registered voice data input by the user through a pre-trained first voiceprint feature model. In this way, the electronic device can obtain the first registered voiceprint feature when the user says the first registered voice data of the specified short text content. The electronic device can store the first registered voiceprint feature in the memory.

[0114] See Figure 5BAs shown, the electronic device can display "Please say, Xiaoyi sends a message" on the display screen. The user can then say "Xiaoyi, send a message" according to the prompt. Therefore, the electronic device can obtain the second registered voice data of the user input "Xiaoyi, send a message" through sensors such as a receiver and a microphone. The electronic device can extract the second registered voiceprint feature from the second registered voice data input by the user using a pre-trained second voiceprint feature model. In this way, the electronic device can obtain the second registered voiceprint feature of the user when speaking the second registered voice data with random text content. The electronic device can store the second registered voiceprint feature in a memory.

[0115] Optionally, the electronic device may prompt the user to enter the second registration voice data multiple times via the display screen so that more second registration voiceprint features can be obtained, thereby improving the accuracy of voiceprint feature comparison. For example, after the user enters the second registration voice data such as "Xiaoyi, send a message," the electronic device may display prompts such as "Please tell me, Xiaoyi, how is the weather today?" or "Please tell me, Xiaoyi, play some music." The user can then enter the corresponding second registration voice data based on the prompts displayed on the display screen, so that the electronic device can extract the second registration voiceprint features of the second registration voice data.

[0116] In a possible implementation, the electronic device may combine the user's first registered voiceprint feature and the second registered voiceprint feature to obtain a third registered voiceprint feature. Figure 6 The electronic device obtains the user's first registered voiceprint feature and the second registered voiceprint feature through the first voiceprint feature model and the second voiceprint feature model, merges the user's first registered voiceprint feature and the second registered voiceprint feature in a specified manner, and stores the merged first registered voiceprint feature and the second registered voiceprint feature, that is, stores the third registered voiceprint feature.

[0117] Optionally, when the first registered voiceprint feature and the second registered voiceprint feature are merged, a simple merge can be performed. For example, if the first registered voiceprint feature is A and the second registered voiceprint feature is B, the first registered voiceprint feature and the second registered voiceprint feature can be merged to obtain a third registered voiceprint feature A+B. Alternatively, weights can be assigned to the first registered voiceprint feature and the second registered voiceprint feature, respectively, such as assigning a weight of 0.4 to the first registered voiceprint feature and a weight of 0.6 to the second registered voiceprint feature. The first registered voiceprint feature and the second registered voiceprint feature can then be multiplied by the corresponding weights and then merged.

[0118] See also Figure 7 , is a flow chart of a method for unlocking an electronic device provided in an embodiment of the present application, which can be performed by Figure 2 or Figure 3The electronic device shown in the figure performs the method, and the process includes:

[0119] 701: The electronic device receives first voice data input by a user and obtains a first voiceprint feature of the first voice data.

[0120] Among them, electronic equipment can be Figure 2 The receiver 170B or microphone 170C shown receives first voice data input by the user. The first voice data here can be a specified short text voice data, such as a wake-up word. For example, the user can say "Hello, Xiaoyi", and the electronic device can receive the user's voice data of "Hello, Xiaoyi" through the receiver 170B or microphone 170C.

[0121] In the embodiment of the present application, the user can input the first voice data when the display screen of the electronic device is not lit, that is, when the screen is black. It should be understood that even if the display screen of the electronic device is not lit, the receiver 170B or microphone 170C of the electronic device can be turned on. The user can say a specified short text voice data such as "Hello".

[0122] In order to save power consumption of the electronic device, the user can touch the display screen or press a button of the electronic device (such as Figure 2 The button 190 shown in the figure triggers the electronic device to turn on the receiver 170B or the microphone 170C. Figure 8A , after the screen of the electronic device lights up, the user can input the first voice data.

[0123] Optionally, after the screen of the electronic device lights up, the electronic device may display a prompt message on the display screen to prompt the user to input the first voice data. Figure 8B When the user touches the display screen or presses a button of the electronic device to light up the display screen of the electronic device, a prompt message such as "Please enter a voice password" may be displayed, prompting the user to enter voice data to unlock the electronic device.

[0124] In a possible implementation, if the external environment is noisy and the noise is large, the electronic device may not receive the first voice data input by the user. Therefore, the electronic device may also display a prompt message on the display screen to prompt the user to re-enter the first voice data. Figure 9 , the user inputs the first voice data, but the electronic device fails to receive the first voice data. Therefore, the electronic device can display a prompt message of "What did you say" to prompt the user to re-enter the first voice data. Optionally, if the electronic device fails to receive the first voice data input by the user for a specified number of times, the electronic device can prompt the user to unlock the phone by fingerprint or password, such as Figure 1CThe specified number of times here can be pre-set, for example, it can be set to 3 times, 4 times, etc., and this application does not make any specific limitation.

[0125] After receiving the first voice data input by the user, the electronic device can obtain the first voiceprint feature of the first voice data. Specifically, the electronic device can use the first voiceprint feature model obtained through steps 401 to 403 to extract the voiceprint feature of the first voice data to obtain the first voiceprint feature of the first voice data.

[0126] 702: The electronic device performs a voiceprint feature comparison on the first voice data according to the first voiceprint feature.

[0127] The electronic device may compare the first voiceprint feature with the stored first registered voiceprint feature of the user to determine whether the speaker is the same person.

[0128] In one example, the electronic device can determine whether the voiceprint features match by calculating the cosine distance d between the first voiceprint feature and the first registered voiceprint feature. Where d satisfies the following formula (1):

[0129] d=cos(x i ,x j )=x i T *x j Formula (1)

[0130] Among them, x i represents the first voiceprint feature extracted in step 701, x j represents the first registered voiceprint feature obtained during registration, and T represents the transpose of the matrix.

[0131] In the embodiment of the present application, a relationship between the cosine distance and the voiceprint feature comparison result can be maintained in advance. For example, when the cosine distance is less than a specified value, it indicates a successful match, and when the cosine distance is greater than or equal to the specified value, it indicates a failed match. Therefore, the cosine distance can be obtained by the above formula (1) and compared with the specified value to determine whether the match is successful, that is, whether the speaker is the same person.

[0132] 703: When the comparison result of the first voiceprint feature of the first voice data indicates a successful match, the electronic device performs text recognition on the first voice data.

[0133] When the electronic device determines that the first voiceprint feature successfully matches the first registered voiceprint feature, it can perform text recognition on the first voice data. The electronic device can also identify whether the first voice data contains specified text content. The specified text content here can be a short text wake-up word, etc.

[0134] The electronic device may pre-store specified text content in the memory. For example, the specified text content may include a short wake-up word such as "Hello," "Hello, Xiaoyi," or "Are you there?"

[0135] In a possible implementation, if the electronic device determines that the first voiceprint feature fails to match the first registered voiceprint feature, the electronic device may display a prompt message indicating the match failure on a display screen. Figure 10 , the electronic device may display a prompt message of "Unlock Failed" on the display screen. Figure 10 , the electronic device can prompt the user to re-enter the first voice data, so that the electronic device can perform voiceprint feature comparison based on the first voice data re-entered by the user.

[0136] In another possible implementation, if the electronic device performs voiceprint feature comparison on the first voice data and the number of failed matches is greater than or equal to a specified threshold, the electronic device may prompt the user to unlock the electronic device by fingerprint unlocking, face unlocking, or password unlocking. Figure 1C , the electronic device can display a password input interface or a fingerprint input interface on the display screen to prompt the user to unlock the electronic device by fingerprint unlocking or password unlocking. Optionally, after the user unlocks the electronic device by fingerprint unlocking or password unlocking, the voiceprint unlocking method of the electronic device can be restarted. Figure 11A After the user unlocks the electronic device by fingerprint unlocking or password unlocking, the user locks the electronic device. At this time, the voiceprint unlocking method of the electronic device is restarted, and the user can re-enter the first voice data so that the electronic device can perform voiceprint feature comparison on the first voice data.

[0137] In another possible implementation, if the electronic device performs voiceprint feature comparison on the first voice data and the number of matching failures is greater than or equal to a specified threshold, the electronic device may be locked for a specified time period. Figure 11B The electronic device can display the specified lock time on the display screen and can also display the remaining lock time of the electronic device in a countdown manner. When the lock time of the electronic device reaches the specified time, the user can unlock the electronic device by voiceprint unlocking, password unlocking, fingerprint unlocking or face unlocking.

[0138] 704: When the first voice data contains specified text content, the electronic device enters a wake-up mode.

[0139] When the electronic device determines that the first voice data contains the specified text content, it can enter the wake-up mode. In the wake-up mode, the electronic device can turn on the audio module, such as the receiver or microphone, to record. Figure 12 If the electronic device determines that the specified text content exists in the first voice recognition result, the electronic device can illuminate the display, turn on the audio module, and display a prompt such as "Please speak" or "I'm here" on the display to prompt the user to continue inputting control commands. It should be understood that the electronic device is still locked at this time.

[0140] When the electronic device determines that the specified text content does not exist in the first voice data, the electronic device may remain in a sleep state. For example, the electronic device may turn off the display screen, or the electronic device may turn off the audio module. The electronic device may display a prompt message such as "What did you say?" or "I don't understand" on the display screen to remind the user that the electronic device has not been awakened and that the first voice data needs to be re-entered so that the electronic device can perform voiceprint feature matching on the first voice data.

[0141] 705: The electronic device receives second voice data input by the user and obtains a second voiceprint feature of the second voice data.

[0142] After determining that the first voiceprint feature successfully matches the first registered voiceprint feature, the electronic device may enter a wake-up mode. In the wake-up mode, the electronic device may activate the audio module to receive second voice data input by the user. The second voice data may be random voice data. For example, the second voice data may include voice data containing a control command.

[0143] In a possible implementation, if the external environment is noisy and the noise is high, the electronic device may not receive the second voice data input by the user. Therefore, the electronic device may also display a prompt message on the display screen to prompt the user to re-enter the second voice data. Figure 13 , the user inputs the second voice data, but the electronic device fails to receive the second voice data. Therefore, the electronic device can display a prompt message "Please say it again" to prompt the user to re-enter the second voice data. Optionally, if the electronic device fails to receive the second voice data input by the user for a specified number of times, the electronic device can prompt the user to unlock the phone through fingerprint, face unlock or password, such as Figure 1C shown.

[0144] Optionally, if the electronic device fails to receive the second voice data a specified number of times, the electronic device may enter a sleep mode. For example, the electronic device may turn off the display screen or the audio module.

[0145] After receiving the second voice data input by the user, the electronic device can obtain the second voiceprint feature of the second voice data. Specifically, the electronic device can use the second voiceprint feature model obtained through steps 401 to 403 to extract the voiceprint feature of the second voice data to obtain the second voiceprint feature of the second voice data.

[0146] 706: The electronic device performs a voiceprint feature comparison on the second voice data based on the second voiceprint feature.

[0147] The electronic device may compare the second voiceprint feature with the stored second registered voiceprint feature of the user to determine whether the speaker is the same person.

[0148] In one example, the electronic device can calculate the pre-set distance d between the second voiceprint feature and the second registered voiceprint feature by using formula (1). i represents the second voiceprint feature extracted in step 705, x j represents the second registered voiceprint feature obtained during registration, and T represents the transpose of the matrix.

[0149] In another example, the electronic device may combine the first voiceprint feature obtained in step 701 with the second voiceprint feature obtained in step 705 in a specified manner. It should be understood that the method of combining the first voiceprint feature and the second voiceprint feature in this case should be the same as the method of combining the first registered voiceprint feature and the second registered voiceprint feature when the user registers the voiceprint feature. The electronic device may compare the third voiceprint feature obtained by combining the first voiceprint feature and the second voiceprint feature with the stored third registered voiceprint feature of the user to determine whether the speaker is the same person.

[0150] For example, when a user registers a voiceprint feature, the weight of the first registered voiceprint feature is 0.4, and the weight of the second registered voiceprint feature is 0.6. The first registered voiceprint feature and the second registered voiceprint feature are merged based on the weight of the first registered voiceprint feature and the weight of the second registered voiceprint feature, respectively, to obtain a third registered voiceprint feature. At this time, the weight of the first voiceprint feature obtained by the electronic device in step 701 is also 0.4, and the weight of the second voiceprint feature obtained in step 705 is also 0.6. The electronic device can merge the first voiceprint feature obtained in step 701 and the second voiceprint feature obtained in step 705 based on the weights of the first voiceprint feature and the second voiceprint feature to obtain a third voiceprint feature. The electronic device can compare the third voiceprint feature obtained by merging the first voiceprint feature and the second voiceprint feature with the stored third registered voiceprint feature.

[0151] The electronic device can calculate the pre-distance d between the third voiceprint feature and the third registered voiceprint feature by using the above formula (1).i Indicates the third voiceprint feature, x j represents the third registered voiceprint feature obtained during registration, and T represents the transpose of the matrix.

[0152] 707: When the comparison result of the second voiceprint feature of the second voice data indicates a successful match, the electronic device performs text recognition on the second voice data.

[0153] If the electronic device determines that the second voiceprint feature comparison result of the second voice data indicates a successful match, the electronic device can identify whether there is a control instruction in the second voice data. If there is a control instruction in the second voice data, the electronic device can respond to the control instruction.

[0154] See Figure 14 , the user inputs a second voice data containing the content "Play some music". When the electronic device determines that the second voice data has a successful match with the second voiceprint feature, the electronic device can perform text recognition on the second voice data. The electronic device can determine that the second voice data contains a control instruction, and therefore can respond to the control instruction by opening an application that can play music and randomly playing music.

[0155] In a possible implementation, if the electronic device compares the second voiceprint feature of the second voice data to a matching failure, the electronic device may display a matching failure prompt message on a display screen. Figure 10 , the electronic device can display a prompt message of "Unlock Failed" on the display screen to let the user know that the voiceprint unlocking has failed. Figure 10 , the electronic device may prompt the user to re-enter the second voice data. The user may re-enter the second voice data according to the prompt of the electronic device to unlock the electronic device.

[0156] In another possible implementation, if the electronic device fails to match the second voiceprint feature of the second voice data, the electronic device may terminate the voiceprint unlocking operation and turn off the display. At this point, the user may input the first voice data, and the electronic device may perform a voiceprint feature comparison on the first and second voice data input by the user using the method shown in steps 701-707 to determine whether to unlock the electronic device.

[0157] In another possible implementation, if the electronic device performs voiceprint feature comparison on the second voice data and the second voiceprint feature comparison result indicates that the number of matching failures is greater than or equal to a specified threshold, the electronic device may prompt the user to unlock the electronic device by fingerprint unlocking, face unlocking, or password unlocking. Figure 1C, the electronic device can display an interface for inputting a password and / or an interface for inputting a fingerprint on the display screen.

[0158] In another possible implementation, if the electronic device performs voiceprint feature comparison on the first voice data and the first voiceprint feature comparison result indicates that the number of matching failures is greater than or equal to a specified threshold, the electronic device may be locked for a specified time period. Figure 11B The electronic device can display the specified lock time on the display screen and can also display the remaining lock time of the electronic device in a countdown manner. When the lock time of the electronic device reaches the specified time, the user can unlock the electronic device by voiceprint unlocking, password unlocking, fingerprint unlocking or face unlocking.

[0159] Based on the above solution, when a user needs to unlock an electronic device, they can input voice data into the electronic device. The electronic device can then perform a voiceprint feature comparison on the two voice data input by the user, namely, the wake-up word and the control command input by the user. When the results of both voiceprint feature comparisons indicate a successful match, the electronic device can be unlocked. Therefore, the accuracy of the voiceprint feature comparison results can be improved, and attacks such as imitation, recording, and synthesis can be effectively avoided, thereby improving the security of voiceprint unlocking. In addition, since both voiceprint feature comparisons are performed on the voice data input by the user, no manual operation is required by the user, which provides convenience for the user.

[0160] The following describes a method for unlocking an electronic device according to an embodiment of the present application through specific examples.

[0161] Example 1

[0162] See Figure 15 , which is an exemplary flowchart of the electronic device unlocking method provided in an embodiment of the present application, may include the following steps.

[0163] Step 1501: The electronic device obtains a wake-up word input by the user through a microphone MIC. The wake-up word here can be a pre-set short text content, such as "Hello", "Hello, Xiaoyi", "Hi" or "Good morning".

[0164] Step 1502: The electronic device performs text recognition on the wake-up word.

[0165] The text recognition here is to determine whether the wake-up word is the same as the pre-set specified short text content.

[0166] Step 1503: The electronic device extracts the first voiceprint feature from the wake-up word input by the user.

[0167] Among them, the electronic device can extract the first voiceprint feature in the wake-up word through a pre-trained first voiceprint feature model.

[0168] Optionally, in the embodiment of the present application, step 1502 may be performed first and then step 1503, or step 1503 may be performed first and then step 1502, or step 1502 and step 1503 may be performed simultaneously.

[0169] Step 1504: The electronic device performs text comparison and voiceprint feature comparison.

[0170] Among them, the electronic device can determine whether the wake-up word input by the user is a preset designated short text content, and the electronic device can determine whether the first voiceprint feature extracted in step 1503 is the same as the voiceprint feature when the user registered.

[0171] If both the voiceprint feature comparison and the text comparison pass, step 1505 may be executed. If either the voiceprint feature comparison or the text comparison fails, the operation is terminated.

[0172] Step 1505: The electronic device enters the wake-up mode and turns on the screen. In the wake-up mode, the electronic device can start MIC recording to obtain the voice data input by the user.

[0173] Step 1506: The electronic device obtains the command word input by the user through the microphone MIC recording.

[0174] The command words here may be words for controlling the electronic device to perform corresponding operations, for example, words related to control commands such as "open navigation", "play some music" or "check the weather".

[0175] Step 1507: The electronic device extracts a second voiceprint feature from the command word input by the user.

[0176] Among them, the electronic device can extract the second voiceprint feature in the command word through a pre-trained second voiceprint feature model.

[0177] Step 1508: The electronic device combines the first voiceprint feature and the second voiceprint feature and verifies them.

[0178] The electronic device may combine the first voiceprint feature and the second voiceprint feature in a specified manner to obtain a third voiceprint feature. The electronic device may compare the combined third voiceprint feature with the third voiceprint feature obtained during registration to determine whether the speaker is the same person.

[0179] If the electronic device determines that the speaker is the same person, step 1509 may be executed; if the electronic device determines that the speaker is not the same person, this operation may be terminated.

[0180] Step 1509: The electronic device recognizes the command word and analyzes the task to be performed. The electronic device can perform ASR on the command word to determine the text-related information contained in the command word.

[0181] Step 1510: The electronic device is unlocked and executes the task corresponding to the command word.

[0182] Optionally, the electronic device may directly unlock the phone after determining in step 1508 that the speaker is the same person, that is, directly execute step 1509 .

[0183] Example 2

[0184] See Figure 16 User A is driving a vehicle and needs to turn on the navigation system on their electronic device. Therefore, user A can say, "Hello, please turn on the navigation system." "Hello" can be considered the first voice data, and "Please turn on the navigation system" can be considered the second voice data. The electronic device's microphone is on, so when user A says "Hello," the electronic device can capture the first voice data containing the user input "Hello" through the microphone. The electronic device can then turn off the microphone and extract a first voiceprint feature from the first voice data using a pre-trained first voiceprint feature model. This first voiceprint feature is then compared with the first registered voiceprint feature obtained during registration, thereby confirming that user A is the same person as the user who registered. The electronic device can perform text recognition on the first voice data and determine that the first voice data contains specified text. At this point, the electronic device can enter wake-up mode, turn on the microphone again, capture the second voice data containing user A's input "Please turn on the navigation system," and extract a second voiceprint feature using the pre-trained second voiceprint feature model. The electronic device can merge the first and second voiceprint features and compare them with the third registered voiceprint feature obtained during registration. This confirms again that user A is the same person as the user who registered. At this time, the electronic device can be unlocked and perform text recognition on the user input "Please open navigation". The electronic device recognizes that the user needs to open navigation, that is, the control command is "open navigation", so the electronic device can open the application that can navigate.

[0185] like Figure 17 As shown, some other embodiments of the present application disclose an electronic device 1700, which may include: one or more processors 1701; one or more memories 1702 and one or more display screens 1703; wherein the one or more memories 1702 store one or more computer programs, and the one or more computer programs include instructions. For example, Figure 17Schematically shows a processor 1701 and a memory 1702. When the instructions are executed by the one or more processors 1701, the electronic device 1700 performs the following steps:

[0186] In the lock screen state, receive the first voice data input by the user and extract the first voiceprint feature of the first voice data; when the comparison result of the first voiceprint feature is consistent with the preset reference voiceprint feature and the first voice data contains specified text, remain in the lock screen state; receive the second voice data input by the user, the second voice data contains a control instruction, and the control instruction is used to trigger at least one functional requirement of the user; extract the second voiceprint feature of the second voice data; when the comparison result of the second voiceprint feature is consistent with the preset reference voiceprint feature, unlock the screen and execute the control instruction to trigger the at least one functional requirement. The specified text and control instruction can be found in the following example. Figure 7 The relevant descriptions in the method embodiment shown are not repeated here.

[0187] In one design, the processor 1701 further performs the following steps: receiving the first registered voice data input by the user; the first registered voice data contains the specified text; obtaining the first registered voiceprint feature of the first registered voice data; receiving the second registered voice data input by the user; the second registered voice data and the first registered voice data are from the same user; obtaining the second registered voiceprint feature of the second registered voice data; the memory 1702 stores the first registered voiceprint feature and the second registered voiceprint feature as the reference voiceprint feature. The first registered voiceprint feature and the second registered voiceprint feature can be referred to as follows Figure 7 The relevant descriptions in the method embodiment shown are not repeated here.

[0188] In one design, the processor 1701 further performs the following steps: combining the first registered voiceprint feature and the second registered voiceprint feature to obtain a third registered voiceprint feature. The memory 1702 stores the third registered voiceprint feature as the reference voiceprint feature. The description of the third voiceprint feature can be found in Figure 7 The relevant descriptions in the method embodiment shown are not repeated here.

[0189] In one design, the processor 1701 is specifically configured to perform the following steps: combining the second voiceprint feature and the first voiceprint feature to obtain a third voiceprint feature; and unlocking the screen when the comparison result of the third voiceprint feature is consistent. The description of the third voiceprint feature can be found in Figure 7 The relevant descriptions in the method embodiment shown are not repeated here.

[0190] In one design, the processor 1701 is specifically configured to perform the following steps: using a pre-trained first voiceprint feature model to obtain the first voiceprint feature of the first voice data; and using a pre-trained second voiceprint feature model to obtain the second voiceprint feature of the second voice data. The description of the first voiceprint feature model and the second voiceprint feature model can be found in FIG. Figure 7 The relevant descriptions in the method embodiment shown are not repeated here.

[0191] It should be noted that the division of units in the embodiments of the present application is schematic and is only a logical functional division. There may be other division methods in actual implementation. The functional units in the embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. For example, in the above embodiment, the first acquisition unit and the second acquisition unit can be the same unit or different units. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0192] As used in the above embodiments, the term “when…” may be interpreted to mean “if…” or “after…” or “in response to determining…” or “in response to detecting…”, depending on the context. Similarly, the phrases “upon determining…” or “if (stated condition or event) is detected” may be interpreted to mean “if determining…” or “in response to determining…” or “upon detecting (stated condition or event)” or “in response to detecting (stated condition or event)”, depending on the context.

[0193] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrations. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state hard disk).

[0194] For the purpose of explanation, the foregoing description is described with reference to specific embodiments. However, the above exemplary discussion is not intended to be exhaustive, nor is it intended to limit the present application to the precise forms disclosed. In light of the above teachings, many modifications and variations are possible. The embodiments are selected and described in order to fully illustrate the principles of the present application and its practical application, so that others skilled in the art can take full advantage of the present application and various embodiments with various modifications suitable for the specific purposes contemplated.

Claims

1. A method for unlocking an electronic device, characterized in that: The method comprises: The electronic device receives first voice data input by a user when the display screen is not lit, and extracts a first voiceprint feature of the first voice data; When the comparison result of the first voiceprint feature and the preset reference voiceprint feature is consistent and the first voice data contains the specified text, the electronic device remains in the locked screen state; The electronic device receives second voice data input by the user when a comparison result of the first voiceprint feature is consistent with a preset reference voiceprint feature and the first voice data includes specified text, where the second voice data includes a control instruction, and the control instruction is used to trigger at least one functional requirement of the user; extracting, by the electronic device, a second voiceprint feature of the second voice data; When the comparison result of the second voiceprint feature and the preset reference voiceprint feature is consistent, the electronic device unlocks the screen and executes the control instruction to trigger the at least one functional requirement.

2. The method according to claim 1, characterized in that Also includes: The electronic device receives first registration voice data input by the user; The first registered voice data includes the specified text; The electronic device obtains a first registration voiceprint feature of the first registration voice data; The electronic device receives second registration voice data input by the user; The second registered voice data and the first registered voice data are from the same user; The electronic device obtains a second registration voiceprint feature of the second registration voice data; The electronic device stores the first registered voiceprint feature and the second registered voiceprint feature as the reference voiceprint feature.

3. The method according to claim 2, characterized in that Also includes: The electronic device combines the first registered voiceprint feature and the second registered voiceprint feature to obtain a third registered voiceprint feature; The electronic device stores the third registered voiceprint feature as the reference voiceprint feature.

4. The method according to claim 2, characterized in that The electronic device unlocks the screen when a comparison result of the second voiceprint feature and the preset reference voiceprint feature is consistent, including: The electronic device combines the second voiceprint feature and the first voiceprint feature to obtain a third voiceprint feature; The electronic device unlocks the screen when the comparison result of the third voiceprint feature is consistent.

5. The method according to any one of claims 1 to 4, characterized in that: The electronic device uses a pre-trained first voiceprint feature model to obtain the first voiceprint feature of the first voice data; the first voiceprint feature model is trained based on a plurality of first voice data labeled with speakers; The electronic device uses a pre-trained second voiceprint feature model to obtain the second voiceprint feature of the second voice data; The second voiceprint feature model is trained based on a plurality of second speech data labeled with speakers.

6. An electronic device, characterized in that: include: one or more processors; Memory; Multiple applications; and one or more computer programs, wherein the one or more computer programs are stored in the memory, and the one or more computer programs include instructions that, when executed by the electronic device, cause the electronic device to perform the following steps: When the display screen is not lit, receiving first voice data input by a user and extracting a first voiceprint feature of the first voice data; When the comparison result of the first voiceprint feature is consistent with the preset reference voiceprint feature and the first voice data contains the specified text, the screen remains locked; If the comparison result of the first voiceprint feature is consistent with a preset reference voiceprint feature and the first voice data contains specified text, receiving second voice data input by the user, where the second voice data contains a control instruction, and the control instruction is used to trigger at least one functional requirement of the user; extracting a second voiceprint feature of the second voice data; When the comparison result of the second voiceprint feature is consistent with the preset reference voiceprint feature, the screen is unlocked and the control instruction is executed to trigger the at least one functional requirement.

7. The electronic device according to claim 6, wherein: When the instruction is executed by the electronic device, the electronic device further performs the following steps: receiving first registration voice data input by the user; the first registration voice data includes the specified text; Obtaining a first registration voiceprint feature of the first registration voice data; receiving second registration voice data input by the user, wherein the second registration voice data and the first registration voice data are from the same user; Obtaining a second registration voiceprint feature of the second registration voice data; The first registered voiceprint feature and the second registered voiceprint feature are stored as the reference voiceprint feature.

8. The electronic device according to claim 7, wherein: When the instruction is executed by the electronic device, the electronic device further performs the following steps: Combining the first registered voiceprint feature and the second registered voiceprint feature to obtain a third registered voiceprint feature; The third registered voiceprint feature is stored as the reference voiceprint feature.

9. The electronic device according to claim 7, wherein: When the instruction is executed by the electronic device, the electronic device specifically performs the following steps: Combining the second voiceprint feature and the first voiceprint feature to obtain a third voiceprint feature; When the comparison result of the third voiceprint feature is consistent, the screen is unlocked.

10. The electronic device according to any one of claims 6 to 9, characterized in that: When the instruction is executed by the electronic device, the electronic device specifically performs the following steps: Acquiring the first voiceprint feature of the first speech data using a pre-trained first voiceprint feature model; the first voiceprint feature model is trained based on a plurality of first speech data labeled with speakers; Acquire the second voiceprint feature of the second speech data using a pre-trained second voiceprint feature model; The second voiceprint feature model is trained based on a plurality of second speech data labeled with speakers.

11. A computer-readable storage medium comprising instructions, characterized in that: When the instruction is executed on an electronic device, the electronic device is caused to execute the method according to any one of claims 1 to 5.

12. A computer program product comprising instructions, characterized in that When the computer program product is run on an electronic device, the electronic device is enabled to perform the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Electronic device control method and device, storage medium and electronic device

    CN108847242A