Voice input method, electronic equipment and storage medium

By adopting a two-stage chip processing method in electronic devices combined with voiceprint verification methods, the problems of inconvenience and false triggering of voice input are solved, and more efficient and safe voice input is achieved.

CN120276699APending Publication Date: 2025-07-08HONOR DEVICE CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202311853279.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-28
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

In the prior art, the voice input method of electronic devices is not convenient enough, especially when holding with one hand, it is difficult to click the top and bottom buttons at the same time, resulting in frequent occurrence of falsely triggering voice input.

Method used

The two-stage chip processing method is adopted, and the first-level wake-up judgment and the second-level calibration judgment are combined with voiceprint verification to reduce false triggering. The specific steps include: obtaining trigger data, performing device posture, user proximity situation and audio data detection, and if passed, performing first-level wake-up judgment; after the first-level wake-up judgment is passed, performing voice/noise recognition, recording and playback attack recognition, sound source angle recognition and directional sound pickup, and finally performing voiceprint verification to confirm the voice data of the designated user.

Benefits of technology

It effectively reduces the situation of accidentally triggering voice input, improves the convenience and security of voice input, and is suitable for mid- and low-end electronic devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120276699A_ABST
    Figure CN120276699A_ABST
Patent Text Reader

Abstract

The invention provides a voice input method, electronic equipment and a storage medium, and is applied to the technical field of information, the method comprises the following steps: obtaining trigger data of the electronic equipment, the trigger data comprising audio data; performing first-level wake-up judgment according to the trigger data; under the condition that the first-level awakening judgment is passed, performing second-level verification judgment according to the trigger data; under the condition that the second-level verification judgment is passed, voice-to-text input is carried out; wherein the second-level verification according to the trigger data comprises the step of carrying out voiceprint verification according to the audio data. In combination with first-level wake-up judgment and second-level verification judgment, a voice false triggering condition of a non-specified user is eliminated in a voiceprint verification mode, so that the condition of false triggering voice input is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of information technology, and particularly to a voice input method, an electronic device, and a storage medium. Background Art

[0002] With the development of speech recognition technology, voice input scenarios in electronic devices have become increasingly common. For example, various third-party applications such as search engines, social platforms, shopping software, and input methods have started to add voice input scenarios to support users in inputting text with voice instead, and subsequent processing such as voice recognition is performed by the application background.

[0003] Currently, switching to voice input in various application scenarios requires manual operation. However, on the one hand, as Figure 1-1 and Figure 1-2 shown, the position of the button for switching to voice input is difficult to click; on the other hand, users often use electronic devices in a single-handed holding manner, but the touch range of single-handed holding is difficult to cover both the banner notification button at the top and the voice input button at the bottom at the same time, as Figure 1-3 and Figure 1-4 shown (OKAY indicates that it can be touched with single-handed holding, where EAST&ACCURATE and EASY indicate that it can be simply and / or accurately touched with single-handed holding; other areas indicate that it is difficult to touch). Therefore, a more convenient voice input method is needed to support the voice input scenario of electronic devices.

[0004] In order to perform voice input more conveniently, the breath wake-up function has emerged. After the electronic device collects the voice data corresponding to the user's speech through the microphone and the gesture data corresponding to the user lifting the electronic device through the inertial detection sensor, it sends them to the low-power ADSP (Digital Signal Processor). The ADSP detects the voice data and gesture data through the breath wake-up model. When the ADSP detects that the voice data meets the preset voice conditions, for example, it is human speech, or the sound intensity is greater than the preset sound intensity, or the similarity with the preset wake-up breath data is greater than the first threshold, and the similarity between the gesture data and the preset wake-up gesture data is greater than the second threshold, the ADSP sends the voice data to the breath wake-up software module. The breath wake-up software module controls the application program to start, thus realizing voice input. However, using the above method, false triggering of voice input is likely to occur. Summary of the Invention

[0005] The purpose of the embodiments of this application is to provide a voice input method, an electronic device, and a storage medium to reduce the situation of false triggering of voice input. The specific technical solutions are as follows:

[0006] In a first aspect, the embodiments of this application provide a voice input method, including:

[0007] Obtain trigger data of an electronic device, where the trigger data includes audio data;

[0008] Perform a primary wake-up determination based on the trigger data;

[0009] When the primary wake-up determination passes, perform a secondary verification determination based on the trigger data;

[0010] When the secondary verification determination passes, perform voice-to-text input;

[0011] Among them, the performing the secondary verification based on the trigger data includes:

[0012] Perform voiceprint verification based on the audio data.

[0013] In the embodiment of the present application, in combination with the primary wake-up determination and the secondary verification determination, the voiceprint verification method is used to exclude the voice mis-triggering situation of non-designated users, so as to reduce the situation of mis-triggered voice input.

[0014] In a possible implementation manner, the method is applied to an electronic device, and the electronic device includes a primary processing chip and a secondary processing chip. Among them, the primary wake-up determination is executed by the primary processing chip, and the secondary verification determination is executed by the secondary processing chip.

[0015] In the embodiment of the present application, the primary processing chip executes the primary wake-up determination, and the secondary processing chip executes the secondary verification determination. By using the processing method of two-level chips, the processing pressure of a single chip can be reduced, and the speed of voice output can be improved.

[0016] In a possible implementation manner, the primary processing chip is a digital signal processor ADSP, and the secondary processing chip is an application processor AP.

[0017] In the embodiment of the present application, using the ADSP to execute the primary wake-up determination and using the AP to execute the secondary verification determination can reduce the memory occupation of the ADSP, reduce the computing power requirement for the ADSP, and is applicable to mid- to low-end electronic devices.

[0018] In a possible implementation manner, the method is applied to an electronic device, and the electronic device includes an inertial measurement unit, a microphone, and a proximity sensor;

[0019] The obtaining the trigger data of the electronic device includes:

[0020] Obtain inertial measurement data collected by the inertial measurement unit, audio data collected by the microphone, and proximity light data collected by the proximity sensor;

[0021] Performing a first-level wake-up determination based on the trigger data includes:

[0022] Performing device attitude detection of the electronic device based on the inertial measurement data;

[0023] Detecting the user's approach based on the proximity light data;

[0024] Performing audio data detection based on the audio data;

[0025] Among them, when the device attitude detection, the user approach situation, and the audio data detection all pass, it is determined that the first-level wake-up determination passes.

[0026] In the embodiment of the present application, the first-level wake-up determination includes device attitude detection, user approach situation detection, and audio data detection. When all three detections pass, it is determined that the first-level wake-up determination passes, thereby reducing the situation of mis-triggering voice input.

[0027] In a possible implementation manner, performing a second-level verification determination based on the trigger data includes:

[0028] Performing voice / noise recognition on the audio data to obtain voice data;

[0029] Performing recognition of recording playback attacks on the voice data;

[0030] When the voice data is not a recording, determining the sound source angle of the user according to the trigger data;

[0031] Performing directional sound pickup on the voice data according to the sound source angle to obtain user voice data;

[0032] Performing voiceprint verification based on the audio data includes:

[0033] Determining whether the voiceprint of the user voice data is the voiceprint of a specified user; among them, if it is the voiceprint of the specified user, it is determined that the second-level verification determination passes.

[0034] In the embodiment of the present application, the second-level verification determination includes voice / noise recognition, recording playback attack recognition, sound source angle recognition, directional sound pickup, and voiceprint verification, which can reduce the situation of mis-triggering voice input; and the recording playback attack recognition can increase the security of voice input.

[0035] In a possible implementation manner, before determining whether the voiceprint of the user voice data is the voiceprint of a specified user, the method further includes: performing voice enhancement on the user voice data.

[0036] In the embodiments of the present application, by enhancing the user's voice data, the accuracy of subsequent voiceprint recognition can be increased, thereby reducing the situation of mis-triggered voice input.

[0037] In a possible implementation manner, the method further includes:

[0038] Pre-record the fixed wake-up word voice of the specified user, extract the voiceprint features of the fixed wake-up word voice, and obtain the voiceprint of the specified user;

[0039] The determination of whether the voiceprint of the user voice data is the voiceprint of the specified user includes:

[0040] Calculate the similarity between the voiceprint of the user voice data and the voiceprint of the specified user. If the similarity is greater than a preset threshold, it is determined that the voiceprint of the user voice data is the voiceprint of the specified user.

[0041] In the embodiments of the present application, voiceprint detection is realized by pre-recording the voiceprint of the specified user, and the voiceprint detection method is simple.

[0042] In a possible implementation manner, the method further includes:

[0043] Obtain the text-independent voice of the specified user, and use the text-independent voice of the specified user to train the voiceprint recognition model to obtain the trained voiceprint recognition model;

[0044] The determination of whether the voiceprint of the user voice data is the voiceprint of the specified user includes:

[0045] Use the voiceprint recognition model to determine whether the voiceprint of the user voice data is the voiceprint of the specified user.

[0046] In the embodiments of the present application, using the voiceprint recognition model for voiceprint detection can be applicable to the situation where the user does not input a fixed wake-up word, and the applicable range is wider.

[0047] In a second aspect, the embodiments of the present application provide an electronic device, including:

[0048] A memory for storing a computer program;

[0049] A processor for implementing the following method when executing the program stored in the memory:

[0050] Obtain the trigger data of the electronic device, where the trigger data includes audio data;

[0051] Perform a first-level wake-up judgment according to the trigger data;

[0052] When the first-level wake-up judgment passes, perform a second-level verification judgment based on the trigger data;

[0053] When the second-level verification judgment passes, perform voice-to-text input;

[0054] Among them, the second-level verification according to the trigger data includes:

[0055] Perform voiceprint verification according to the audio data.

[0056] In a possible implementation manner, the processor includes a first-level processing chip and a second-level processing chip;

[0057] The first-level processing chip is used to perform a first-level wake-up judgment according to the trigger data;

[0058] The second-level processing chip is used to perform a second-level verification judgment according to the trigger data when the first-level wake-up judgment passes; and perform voice-to-text input when the second-level verification judgment passes.

[0059] In a possible implementation manner, the first-level processing chip is a digital signal processor ADSP, and the second-level processing chip is an application processor AP.

[0060] In a possible implementation manner, the electronic device further includes an inertial measurement unit, a microphone, and a proximity sensor;

[0061] The inertial measurement unit is used to collect inertial measurement data;

[0062] The microphone is used to collect audio data;

[0063] The proximity sensor is used to collect proximity light data;

[0064] The first-level processing chip is specifically used for: detecting the device posture of the electronic device according to the inertial measurement data; detecting the user's proximity according to the proximity light data; performing audio data detection according to the audio data, wherein when the device posture detection, the user's proximity situation, and the audio data detection all pass, it is determined that the first-level wake-up judgment passes.

[0065] In a possible implementation, the secondary processing chip is specifically configured to: perform speech / noise recognition on the audio data to obtain speech data; perform recognition of recording playback attacks on the speech data; when the speech data is not a recording, determine the sound source angle of the user according to the trigger data; perform directional sound pickup on the speech data according to the sound source angle to obtain user speech data; determine whether the voiceprint of the user speech data is the voiceprint of a specified user; wherein, if it is the voiceprint of the specified user, it is determined that the secondary verification judgment passes.

[0066] In a possible implementation, the secondary processing chip is further configured to: perform speech enhancement on the user speech data.

[0067] In a possible implementation, the secondary processing chip is further configured to: pre-enter the fixed wake-up word voice of the specified user, extract the voiceprint features of the fixed wake-up word voice to obtain the voiceprint of the specified user;

[0068] The secondary processing chip is specifically configured to: calculate the similarity between the voiceprint of the user speech data and the voiceprint of the specified user, wherein, if the similarity is greater than a preset threshold, it is determined that the voiceprint of the user speech data is the voiceprint of the specified user.

[0069] In a possible implementation, the secondary processing chip is further configured to: obtain the text-independent speech of the specified user, and use the text-independent speech of the specified user to train a voiceprint recognition model to obtain a trained voiceprint recognition model;

[0070] The secondary processing chip is specifically configured to: use the voiceprint recognition model to determine whether the voiceprint of the user speech data is the voiceprint of a specified user.

[0071] In a third aspect, an embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the above-mentioned voice input method is implemented.

[0072] Of course, implementing any product or method of the present application does not necessarily require achieving all of the above advantages at the same time. BRIEF DESCRIPTION OF THE DRAWINGS

[0073] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application, and those of ordinary skill in the art can also obtain other embodiments based on these drawings.

[0074] Figure 1-1 It is an example diagram of the button position for the first type of voice input provided by the embodiment of this application;

[0075] Figure 1-2 It is an example diagram of the button position for the second type of voice input provided by the embodiment of this application;

[0076] Figure 1-3 It is an example diagram of the touch range for the first type of single - hand holding provided by the embodiment of this application;

[0077] Figure 1-4 It is an example diagram of the touch range for the second type of single - hand holding provided by the embodiment of this application;

[0078] Figure 1-5 It is a schematic diagram of a breath - wake - up event provided by the embodiment of this application;

[0079] Figure 1-6 It is a schematic diagram for introducing a scenario of the application of the breath - wake - up engine provided by the embodiment of this application;

[0080] Figure 2 It is the first schematic diagram of the structure of the electronic device provided by the embodiment of this application;

[0081] Figure 3 It is the second schematic diagram of the structure of the electronic device provided by the embodiment of this application;

[0082] Figure 4 It is a software structure block diagram of the electronic device provided by the embodiment of this application;

[0083] Figure 5 It is a schematic diagram of the electronic device implementing the voice - input method provided by the embodiment of this application;

[0084] Figure 6 It is another schematic diagram of the electronic device implementing the voice - input method provided by the embodiment of this application;

[0085] Figure 7 It is a schematic flowchart of the voice - input method provided by the embodiment of this application;

[0086] Figure 8 It is another schematic flowchart of the voice - input method provided by the embodiment of this application. Detailed implementation manners

[0087] Next, the technical solutions in the embodiments of this application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art based on this application belong to the scope protected by this application.

[0088] Since in the related art, the method of using voice input on an electronic device is not convenient enough, to solve this problem, the embodiments of the present application provide a voice input method, apparatus, electronic device, and storage medium.

[0089] The following will describe in detail the implementation manners of the embodiments of the present application with reference to the accompanying drawings.

[0090] The voice input method provided by the embodiments of the present application can be applied to any electronic device with voice input capability. The electronic device can be an electronic device with a display screen hardware and corresponding software support. For example, the electronic device can be a mobile phone, a tablet computer, a wearable device, a vehicle-mounted device, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), a home device, etc. The present application does not impose any restrictions on the specific type of the electronic device.

[0091] In a possible embodiment, to more clearly illustrate the electronic device provided by the embodiments of the present application, the following will give an exemplary description of a possible application scenario of the electronic device provided by the embodiments of the present application. It can be understood that the following examples are only a possible application scenario of the electronic device provided by the embodiments of the present application. In other possible embodiments, the electronic device provided by the embodiments of the present application can also be applied to other possible application scenarios, and the following examples do not impose any restrictions on this.

[0092] The voice input method provided by the embodiments of the present application is applicable to any user who uses an electronic device and the voice assistant therein, and is applicable to any scenario where a user needs to use an electronic device and the voice assistant therein. Specifically, it can be a scenario where the electronic device cannot be directly operated, or a scenario where the voice assistant cannot be awakened by voice. For example, when the user is in a relatively quiet public scenario such as a coffee shop, a western restaurant, or a high-speed rail / airport lounge, when the user is in a place such as a subway station, an airport, a railway station security check, a scenic spot ticket gate, etc., and has a lot of luggage in hand, when the user is in a bustling place such as a mall or a supermarket, and needs to select goods in hand, when the user is in an outdoor dog-walking scenario, when the user is driving into / leaving a parking lot, a toll station, a community / park, etc., and has just left the steering wheel... and so on.

[0093] As Figure 1-5 shown, Figure 1-5 A schematic diagram of a scenario of a method for waking up an application program provided by the embodiments of the present application is shown. The user can lift the electronic device, bring the bottom of the electronic device close to the mouth, and speak towards the microphone.

[0094] When the electronic device collects the corresponding voice data when the user is speaking through the microphone, and the gesture data corresponding to the user lifting the electronic device through the inertial detection sensor (by determining the posture of the electronic device to determine the gesture data of the user), it is sent to the low-power ADSP, and the ADSP performs a first-level wake-up judgment. The ADSP detects the voice data and gesture data through the breath wake-up model. When the ADSP detects that the voice data meets the preset voice conditions, for example, it is a human voice, or the sound intensity is greater than the preset sound intensity, or the similarity with the preset wake-up breath data is greater than the first threshold, and the similarity between the gesture data and the preset wake-up gesture data is greater than the second threshold, the ADSP reports a breath wake-up event to the AP (application processing chip), and the AP performs a second-level verification, including voiceprint verification. After the second-level verification passes, the AP sends the user's voice to the breath wake-up software module. The breath wake-up software module controls the application program to start.

[0095] Among them, the distance between the electronic device and the user's mouth can be maintained within 0-7 cm, and the angle between the mouth and the MIC (microphone) is within 60°, which is convenient for the microphone of the electronic device to accurately collect the user's voice data.

[0096] In addition, the above gesture data can be the gesture data of raising the wrist.

[0097] Among them, Figure 1-5 From the gesture state indicated by the box A part in the left figure to the gesture state indicated by the box B part in the right figure, it can be seen that the user can lift the electronic device through the gesture of raising the wrist, and the inertial detection sensor can collect the gesture data of raising the wrist.

[0098] Figure 1-5 From the right figure, it can be seen that after the user lifts the electronic device through the gesture of raising the wrist, the mouth is close to the microphone of the electronic device to speak, and the breath indicated by part C can be emitted, and the microphone can collect the breath and the corresponding voice data.

[0099] As Figure 1-6 shown, it shows a schematic diagram of the scenario introduction of an application of a breath wake-up engine provided by an embodiment of the present application. It is used to prompt that the user can have a natural conversation and answer questions immediately through breath wake-up. When the electronic device is a mobile phone, when the user lifts the mobile phone, brings the bottom of the mobile phone close to the mouth (within a distance of 5 cm), and aligns it with the bottom microphone, the conversation journey of the user (addressing the user as "you" to enhance the sense of interaction) will be started, which also means that the breath wake-up engine application will detect a breath wake-up event in this case.

[0100] It should be understood that the above is an example of the scenario, and does not limit the scenario of the present application in any way.

[0101] First, in the first aspect of the embodiment of the present application, an electronic device is provided, such asFigure 2 As shown, the electronic device includes:

[0102] A memory 201 for storing computer programs;

[0103] A processor 202, when executing the program stored on the memory 201, implements the following steps:

[0104] Obtain trigger data of the electronic device, where the trigger data includes audio data;

[0105] Perform a primary wake-up determination based on the trigger data;

[0106] When the primary wake-up determination passes, perform a secondary verification determination based on the trigger data;

[0107] When the secondary verification determination passes, perform voice-to-text input;

[0108] Wherein, the performing the secondary verification based on the trigger data includes:

[0109] Perform voiceprint verification based on the audio data.

[0110] In one example, the above-mentioned electronic device may further include a communication bus and / or a communication interface, and the processor 202, the communication interface, and the memory 201 complete communication with each other through the communication bus.

[0111] The communication bus mentioned in the above-mentioned electronic device may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, only a thick line is shown in the figure, but it does not mean that there is only one bus or one type of bus.

[0112] The communication interface is used for communication between the above-mentioned electronic device and other devices.

[0113] The memory may include a Random Access Memory (RAM), or may also include a Non-Volatile Memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.

[0114] It may include one or more processing units. For example, the processor 202 may include an application processor (AP), a digital signal processor (e.g., ADSP), a modem processor, a graphics processor, an image signal processor (ISP), a controller, a memory, a video stream codec, a digital signal processor, a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors 202. In one example, the above-mentioned processor 202 may include a primary processing chip and a secondary processing chip; the above-mentioned primary processing chip is used to make a primary wake-up judgment according to the above-mentioned trigger data; the above-mentioned secondary processing chip is used to make a secondary verification judgment according to the above-mentioned trigger data when the primary wake-up judgment passes; and when the secondary verification judgment passes, perform voice-to-text input.

[0115] Specifically, the above-mentioned primary processing chip may be a low-power processor such as ADSP, and the above-mentioned secondary processing chip may be a processor with higher computing power such as AP.

[0116] For ease of explanation, a mobile phone is used as an example of the electronic device for illustration.

[0117] As Figure 3 shown, in some embodiments, the electronic device 300 may include a processor 301, a communication module 302, etc.

[0118] Among them, the processor 301 is the same as the Figure 2 processor 202, and may include one or more processing units. For example, the processor 301 may include an application processor (AP), a digital signal processor (e.g., ADSP), a modem processor, a graphics processor, an image signal processor (ISP), a controller, a memory, a video stream codec, a digital signal processor, a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors 301.

[0119] The controller may be the nerve center and command center of the electronic device 300. The controller may generate operation control signals according to the instruction operation code and timing signal to complete the control of fetching instructions and executing instructions.

[0120] A memory can also be provided in the processor 301 for storing instructions and data.

[0121] In some embodiments, the memory in the processor 301 is a cache memory. This memory can store the instructions or data that the processor 301 has just used or recycled. If the processor 301 needs to use the instruction or data again, it can directly call it from the memory. This avoids repeated accesses, reduces the waiting time of the processor 301, and thus improves the efficiency of the system.

[0122] In some embodiments, the processor 301 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.

[0123] The communication module 302 may include antenna 1, antenna 2, a mobile communication module, and / or a wireless communication module.

[0124] As Figure 3 shown, in some embodiments, the electronic device 300 may further include an external memory interface 305, an internal memory 304, a USB interface 306, a charge management module 307, a power management module 308, a battery 309, and a sensor module 303, etc.

[0125] The NPU is a neural-network (NN) computing processor. By learning from the structure of biological neural networks, such as the transmission mode between human brain neurons, it can quickly process input information and can also continuously self-learn. Through the NPU, applications such as intelligent cognition of the electronic device 300 can be realized, such as: image recognition, face recognition, speech recognition, text understanding, etc.

[0126] The charging management module 307 is used to receive charging input from a charger. Herein, the charger can be a wireless charger or a wired charger.

[0127] In some embodiments of wired charging, the charging management module 307 can receive the charging input from a wired charger through the USB interface 306.

[0128] In some embodiments of wireless charging, the charging management module 307 can receive wireless charging input through the wireless charging coil of the electronic device 300. While charging the battery 309, the charging management module 307 can also supply power to the electronic device 300 through the power management module 308.

[0129] The power management module 308 is used to connect the battery 309, the charging management module 307, and the processor 301. The power management module 308 receives the inputs from the battery 309 and / or the charging management module 307 and supplies power to the processor 301, the internal memory 304, the external memory, the communication module 302, etc. The power management module 308 can also be used to monitor parameters such as the battery capacity, the number of battery cycles, and the battery health status (leakage, impedance).

[0130] In some other embodiments, the power management module 308 can also be disposed in the processor 301.

[0131] In some other embodiments, the power management module 308 and the charging management module 307 can also be disposed in the same device.

[0132] The external memory interface 305 can be used to connect an external memory card, such as a Micro SD card, to implement the storage capacity expansion of the electronic device 300. The external memory card communicates with the processor 301 through the external memory interface 305 to implement the data storage function. For example, files such as music and video streams are saved in the external memory card.

[0133] The internal memory 304 can be used to store computer-executable program codes, and the executable program codes include instructions. The processor 301 executes various functional applications and data processing of the electronic device 300 by running the instructions stored in the internal memory 304. The internal memory 304 can include a program storage area and a data storage area. Among them, the program storage area can store an operating system and applications required for at least one function (such as a sound playback function, an image playback function, etc.). The data storage area can store data created during the use of the electronic device 300 (such as audio data, phone book, etc.). In addition, the internal memory 304 can include high-speed random access memory and can also include non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc.

[0134] The sensor module 303 in the electronic device 300 may include an Inertial Measurement Unit (IMU), a microphone, and a proximity sensor. The above-mentioned inertial measurement unit is used to collect inertial measurement data. The above-mentioned microphone is used to collect audio data. The above-mentioned proximity sensor is used to collect proximity light data.

[0135] In one example, the sensor module 303 may further include components such as an image sensor, a touch sensor, a pressure sensor, a gyroscope sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a distance sensor, a proximity light sensor, an ambient light sensor, a fingerprint sensor, a temperature sensor, a bone conduction sensor, etc., to implement the function of sensing and / or acquiring different signals.

[0136] Optionally, the electronic device 300 may further include peripheral devices, such as a mouse, buttons, an indicator light, a keyboard, a speaker, a microphone, etc.

[0137] The buttons include a power-on button, a volume button, etc. The buttons can be mechanical buttons or touch buttons. The electronic device 300 can receive button inputs and generate key signal inputs related to the user settings and function controls of the electronic device 300.

[0138] The indicator can be an indicator light, which can be used to indicate the charging state and the change of battery power, and can also be used to indicate messages, missed calls, and notifications, etc.

[0139] It can be understood that the structure illustrated in this embodiment does not constitute a specific limitation on the electronic device 300.

[0140] In other embodiments, the electronic device 300 may include more or fewer components than those illustrated, or combine certain components, or split certain components, or have different component arrangements. The illustrated components can be implemented in hardware, software, or a combination of software and hardware.

[0141] In an electronic device, its software system can be divided into several layers, such as Figure 4 shown Figure 4 This is a software structure block diagram of the electronic device provided by the embodiment of the present application. The layered architecture divides the software system of the electronic device into several layers, and each layer has a clear role and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the application system of the electronic device can be divided into an application layer, a framework layer (fwk), a hardware abstraction layer (HAL), and a kernel layer / chip layer.

[0142] The application layer may include a series of application packages, and the application layer runs applications by calling the application programming interfaces (APIs) provided by the application framework layer. The application packages may include multiple applications, such as a breath wake-up engine application, a voice assistant application, an input method application, an AIPlugin Engine (intelligent plug-in engine) application, a call application, and applications such as maps. Among them, an automatic speech recognition algorithm may be run in the intelligent plug-in engine application.

[0143] The framework layer provides APIs and programming frameworks for the applications in the application layer. The application framework layer includes some predefined functions. As Figure 4 shown, the application framework layer may include an audio trigger module (Sound Trigger), an audio service module (Audio Server), etc. Among them, the audio trigger module can be used to control the startup of applications in the application layer; the audio service module is used to send a startup notification to the audio trigger module in response to the startup notification sent by the applications in the application layer. The audio service module can also be used to control services such as the storage (record) and deletion of audio data.

[0144] In addition, the framework layer may also include a sound trigger module, an audio policy service module, etc. Additionally, the application framework layer may also include a sound trigger service module, an audio flinger module, etc. The sound trigger module is used to send a wake-up event to the audio trigger module and send a notification to the audio driver in the driver layer to indicate the stop or start of the breath wake-up model. The audio policy service module is used to establish a voice recognition channel with the voice assistant application in the application layer.

[0145] In addition, the audio trigger module is also used to send a start notification to the audio trigger service module in response to the start notification sent by the audio service module. The audio trigger service module is used to send a start notification to the sound trigger module in response to the start notification sent by the audio trigger module. The sound trigger module is also used to send a notification to start running the breath wake-up model to the audio policy service module in response to the start notification sent by the audio trigger module. The audio policy service module responds to the notification to start running the breath wake-up model and starts running the breath wake-up model. The audio indication module is used to send a notification to load the breath wake-up model to the audio driver in the driver layer in response to the notification sent by the audio policy service module. The audio driver in the driver layer is used to send a notification for instructing to start running the breath wake-up model to the digital signal processor of the electronic device.

[0146] The driver layer is the layer between the hardware and the software and is used to drive the hardware so that the hardware works. Multiple drivers for driving the hardware to work can be installed in the driver layer. For example, the audio driver (sound trigger hal), the primary driver (Primary hal), the sound trigger driver (Sound Trigger driver), and the audio stream input driver (AudioStreamin driver).

[0147] The kernel layer / chip layer includes a digital signal processor (ADSP) and an application processor AP, which are used to determine the existence of a breath wake-up event based on the breath wake-up algorithm, and can obtain the event identifier, that is, the event ID (identity, identification) corresponding to the wake-up event, for distinguishing the breath wake-up event and the voice wake-up event. This event ID is preset by the developer. In one example, after detecting the breath wake-up event, the breath wake-up event including the event ID is reported to the driver layer; the driver layer then reports the breath wake-up event to the framework layer; the framework layer reports the breath wake-up event to the breath wake-up engine application, and the breath wake-up engine application processes the breath wake-up event.

[0148] In one example, taking voice input in an input method application as an example, see Figure 5 , the user uses the input service (Input Service) to input trigger data. In one example, for example Figure 6 as shown, the trigger data includes inertial measurement data collected by the inertial measurement unit, audio data collected by the microphone, and proximity light data collected by the proximity sensor.

[0149] The breath wake-up engine application may include a breath wake-up setting module, a voice wake-up setting module, a cloning / upgrading module, a data uploading module, a wake-up service module, a recording management module, a voiceprint management module, a secondary verification module, a Java / JNI (Java Native Interface) wake-up module, and a Voice Kit module. Among them, the Java / JNI wake-up module can be used to call secondary verification algorithms, text-independent voiceprint training algorithms, voiceprint clearing algorithms, etc. in the algorithm library. Among them, the voice wake-up setting module is used to set the relevant settings of voice wake-up in the breath wake-up engine application, the cloning / upgrading module is used for cloning and upgrading the breath wake-up engine application, and the data uploading module is used for uploading relevant data generated by the breath wake-up engine application, such as error logs, etc.

[0150] The breath wake-up setting module sends a start notification to the audio trigger module, and the audio data collected by the microphone triggers the audio trigger module to call the audio driver, and the audio driver is used to send a notification indicating the start of running the breath wake-up algorithm to the ADSP of the electronic device. For example Figure 6 As shown, the ADSP obtains the trigger data. The IMU module in the ADSP is used to perform device posture detection on the electronic device according to the inertial measurement data (IMU data) (that is, the device posture corresponds to the user's gesture, so it can also be regarded as the detection of gesture data). When the similarity between the current device posture and the preset device posture is greater than the preset posture threshold, it is determined that the device posture detection passes. The voice module in the ADSP is used to perform audio data detection according to the audio data. When the audio data meets the preset voice conditions, for example, it is a human voice, or the sound intensity is greater than the preset sound intensity, or the similarity between the audio data and the preset wake-up breath data is greater than the preset audio threshold, it is determined that the audio data detection passes. The proximity light module in the ADSP is used to detect the user's proximity according to the proximity light data. Among them, when the proximity light data indicates that the user has a proximity action and the distance from the electronic device is within the preset distance range, it is determined that the user's proximity detection passes. When the device posture detection passes, the audio data detection passes, and the user's proximity detection passes, it is determined that the primary wake-up judgment passes; otherwise, it is determined that the primary wake-up judgment fails. If the primary wake-up judgment fails, the judgment of this breath wake-up is ended, and the trigger data is deleted.

[0151] When the first-level wake-up judgment passes, the ADSP reports a breath wake-up event. The audio driver forwards the breath wake-up event of the ADSP to the audio trigger module. The audio trigger module sends the breath wake-up event to the wake-up service module. The wake-up service module calls the recording management module, and the recording management module calls the audio service module to record, and obtains the audio data obtained from the recording of the audio service module. Among them, the audio service module can implement functions such as storage and deletion of audio data based on the Primary hal. After obtaining the audio data, the recording management module calls the voiceprint management module and the secondary verification module to implement the secondary verification judgment. The secondary verification judgment can be implemented through the computing resources of the AP.

[0152] The secondary verification module running in the AP calls the secondary verification algorithm in the algorithm library (such as the SO library) through the Java / JNI wake-up module, for example Figure 6 As shown, the secondary verification algorithm can include voice / noise recognition algorithm, recording playback attack recognition algorithm, sound source angle recognition algorithm, and directional sound pickup algorithm. The AP performs voice / noise recognition on the audio data through the voice / noise recognition algorithm, removes the noise in the audio data, and extracts the voice in the audio data to obtain voice data. The AP uses the recording playback attack recognition algorithm to identify the recording playback attack on the voice data and determine whether the voice data is a recording. The recognition algorithm for whether the voice data is a recording can refer to the prior art and is not specifically limited in this application. When the voice data is not a recording, subsequent operations are continued. When the voice data is a recording, the current voice input is stopped and the stored audio data is deleted. The AP determines the sound source angle of the user according to the trigger data through the sound source angle recognition algorithm. According to the sound source angle, for example, the device posture of the electronic device can be determined according to the inertial measurement data, and the sound source angle of the user can be determined according to the device posture. For example, at least two microphones can be included in the electronic device, and the audio data collected by the two microphones respectively is used to determine the sound source angle of the user. The AP performs directional sound pickup on the voice data according to the sound source angle through the directional sound pickup algorithm to obtain the user voice data. After determining the sound source angle of the user, the user voice data can be obtained through the directional sound pickup algorithm. The specific implementation process of the directional sound pickup algorithm can refer to the prior art and is not specifically limited in this application. In one example, the secondary verification algorithm can also include the NN (Neural Network) front-end enhancement algorithm. The AP can perform voice enhancement on the user voice data or the voice data through the NN front-end enhancement algorithm, thereby improving the accuracy of subsequent voiceprint recognition.

[0153] The voiceprint management module running in the AP can call the voiceprint or voiceprint recognition model (large model) of a specified user in the algorithm library through the Java / JNI wake-up module, so as to determine whether the voiceprint of the user voice data is the voiceprint of the specified user; if it is the voiceprint of the specified user, it is determined that the secondary verification passes, otherwise it is determined that the secondary verification fails.

[0154] In one example, the voiceprint corresponding to the fixed wake-up word voice of the specified user is pre-recorded in the algorithm library. The fixed wake-up word can be custom-set according to the actual situation. For example, it can be "enable voice input", etc. The voiceprint management module running in the AP can calculate the similarity between the voiceprint of the user voice data and the voiceprint of the specified user. Among them, if the similarity is greater than the preset threshold, it is determined that the voiceprint of the user voice data is the voiceprint of the specified user.

[0155] In one example, without pre-recording the voiceprint corresponding to the fixed wake-up word voice of the specified user, the voiceprint management module can train the voiceprint recognition model through the text-independent voiceprint training algorithm; it can also remove the training influence of the specified voiceprint sample on the voiceprint recognition model through the voiceprint clearing algorithm. Specifically, during daily use or during the previous breath wake-up process, the text-independent voice of the specified user can be obtained, and the text-independent voice of the specified user is used to train the voiceprint recognition model to obtain the trained voiceprint recognition model. The AP calls the trained voiceprint recognition model to determine whether the voiceprint of the current user voice data is the voiceprint of the specified user. The text-independent voice does not limit the specific content of the voice, as long as it is the voice of the specified user. The specified user is the actual user of the electronic device. Through the voiceprint recognition of the specified user, the situation where other users accidentally trigger voice input can be reduced.

[0156] After it is determined that the secondary verification passes, the voice suite module in the breath wake-up engine application can call the intelligent plugin engine application, and the intelligent plugin engine application converts the user voice data into text data through the ASR algorithm, and realizes the input of the text data through the input method application.

[0157] It should be noted that other contents may also be included in the application layer, application framework layer, driver layer, and kernel layer / chip layer, which are not specifically limited here.

[0158] It can be seen that by using the electronic device in the embodiment of the present application, combining the primary wake-up judgment and the secondary verification judgment, and performing voice-to-text input when the voiceprint verification passes, the situation of accidental triggering of voice input can be reduced.

[0159] To solve the technical problem that the method of using voice input in the electronic device is not convenient enough, in the second aspect of the embodiment of the present application, as Figure 7 shown Figure 7Schematic flowchart of a voice input method provided by an embodiment of the present application, including:

[0160] Step S701: Obtain trigger data of the electronic device, where the trigger data includes audio data;

[0161] Step S702: Perform a first-level wake-up determination based on the trigger data;

[0162] Step S703: When the first-level wake-up determination passes, perform a second-level verification determination based on the trigger data;

[0163] Step S704: When the second-level verification determination passes, perform voice-to-text input;

[0164] Among them, the second-level verification based on the trigger data includes:

[0165] Perform voiceprint verification based on the audio data.

[0166] In the embodiment of the present application, by combining the first-level wake-up determination and the second-level verification determination, and performing voice-to-text input when the voiceprint verification passes, the situation of mis-triggering voice input can be reduced.

[0167] The voice input method in the embodiment of the present application can be implemented by one or more processors. Specifically, the processor may include an AP, an ADSP, a modulation and demodulation processor, a graphics processor, an ISP, a controller, a digital signal processor, a baseband processor, and / or an NPU, etc. Among them, different processors may be independent devices or integrated in one or more processors. In a possible implementation manner, the voice input method is applied to an electronic device, and the electronic device includes a first-level processing chip and a second-level processing chip. Among them, the first-level wake-up determination is executed by the first-level processing chip, and the second-level verification determination is executed by the second-level processing chip. In one example, the first-level processing chip is an ADSP, and the second-level processing chip is an AP.

[0168] The first-level wake-up determination may include at least one of device attitude detection (also known as user gesture detection), user proximity detection, and audio data detection. In a possible implementation manner, the electronic device further includes an inertial measurement unit, a microphone, and a proximity sensor; the obtaining of the trigger data of the electronic device includes: obtaining inertial measurement data collected by the inertial measurement unit, audio data collected by the microphone, and proximity light data collected by the proximity sensor.

[0169] The above first-level wake-up determination based on the above trigger data includes: performing device attitude detection of the above electronic device according to the above inertial measurement data; detecting the user's approach according to the above proximity light data; performing audio data detection according to the above audio data; wherein, when the above device attitude detection, the above user approach, and the above audio data detection all pass, it is determined that the first-level wake-up determination passes.

[0170] The secondary verification determination may include at least one of processes such as voice / noise recognition, recording playback attack recognition, sound source angle recognition, directional sound pickup, and voiceprint detection. In a possible implementation manner, the above secondary verification determination based on the above trigger data includes:

[0171] Step A, performing voice / noise recognition on the above audio data to obtain voice data;

[0172] Step B, performing recognition of recording playback attacks on the above voice data;

[0173] Step C, when the above voice data is not a recording, determining the sound source angle of the user according to the above trigger data;

[0174] Step D, performing directional sound pickup on the above voice data according to the above sound source angle to obtain user voice data;

[0175] The above voiceprint verification based on the above audio data includes:

[0176] Step E, determining whether the voiceprint of the above user voice data is the voiceprint of a specified user; wherein, if it is the voiceprint of the specified user, it is determined that the secondary verification determination passes.

[0177] To further improve the accuracy of voiceprint detection, voice enhancement can also be performed on the user voice data. In a possible implementation manner, before determining whether the voiceprint of the above user voice data is the voiceprint of a specified user, the above method further includes: performing voice enhancement on the above user voice data. For example, voice enhancement can be achieved through the NN front-end enhancement algorithm.

[0178] In some scenarios, the voiceprint corresponding to the fixed wake-up word of the specified user is pre-stored in the database. In a possible implementation manner, the above method further includes: pre-recording the fixed wake-up word voice of the above specified user, and extracting the voiceprint features of the above fixed wake-up word voice to obtain the voiceprint of the above specified user;

[0179] The above determination of whether the voiceprint of the above user voice data is the voiceprint of a specified user includes:

[0180] Calculate the similarity between the voiceprint of the above user voice data and the voiceprint of the above specified user. If the similarity is greater than a preset threshold, determine that the voiceprint of the above user voice data is the voiceprint of the above specified user.

[0181] In some scenarios, it is not necessary for the user to pre-enter the voiceprint corresponding to the fixed wake-up word. In a possible implementation, the above method further includes: obtaining the text-independent voice of the above specified user, and using the text-independent voice of the above specified user to train a voiceprint recognition model to obtain a trained voiceprint recognition model.

[0182] The above determination of whether the voiceprint of the above user voice data is the voiceprint of the specified user includes: using the above voiceprint recognition model to determine whether the voiceprint of the above user voice data is the voiceprint of the specified user.

[0183] See Figure 8 , the implementation process of the voice input method in the embodiments of the present application can also be as Figure 8 shown, including:

[0184] 1-1: Primary wake-up judgment.

[0185] The user triggers breath wake-up, the inertial measurement unit collects inertial measurement data, the microphone collects audio data, and the proximity sensor collects proximity light data. The ADSP uses the chip algorithm ability to perform a primary wake-up judgment based on the inertial measurement data, audio data, and proximity light data.

[0186] 2-1: Report of breath wake-up event.

[0187] After the primary wake-up judgment is passed, report the breath wake-up event to the wake-up service in the voice assistant.

[0188] 2-2: Record audio from the specified source (data source).

[0189] The wake-up service starts the recording service to record audio from the specified source to obtain audio data.

[0190] 2-3: Secondary verification of audio data.

[0191] The secondary verification management module obtains the audio data from the recording service and uses the chip algorithm ability of the AP to perform secondary verification on the audio data, which may include voice / noise recognition, recording playback attack recognition, sound source angle recognition, directional pickup, NN front-end enhancement, and voiceprint recognition, etc.

[0192] 2-4: Return of secondary verification result.

[0193] The secondary verification management module obtains the secondary verification result.

[0194] 2-5: Judgment based on the secondary verification result.

[0195] If the secondary verification passes, execute 2-6; if the secondary verification fails, stop recording and clear the cache through the recording service.

[0196] 2-6: After the recording is started, the audio data is synchronously written into the buffer.

[0197] Start recording through the recording service and synchronously write the audio data into the audio buffer.

[0198] 3-1: Perform voiceprint training on the audio data.

[0199] The voiceprint management module uses the audio data in the audio buffer to train the voiceprint recognition model to further improve the accuracy of the voiceprint recognition model.

[0200] 3-2: Perform ASR recognition on the audio data.

[0201] The middleware engine calls the ASR algorithm to perform ASR recognition on the audio data to obtain the ASR recognition result.

[0202] 3-3: Return the ASR recognition result.

[0203] Return the ASR recognition result to the interface management module.

[0204] 3-4: Pull up the floating ball and display the text by jumping.

[0205] The interface management module pulls up the floating ball and displays the text by jumping for the ASR recognition result, so as to present the input result of speech-to-text in the UI (user interface) of the voice assistant.

[0206] The embodiment of the present application also provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the method described in any one of the present application is implemented.

[0207] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from a website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).

[0208] It should be noted that, in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover a non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.

[0209] Each embodiment in this specification is described in a related manner. Each embodiment focuses on the differences from other embodiments. For the same and similar parts among the embodiments, reference can be made to each other.

[0210] The above are only the preferred embodiments of the present application and are not intended to limit the protection scope of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application are included in the protection scope of the present application.

Claims

1. A voice input method, characterized in that, The method includes: Obtaining trigger data of an electronic device, where the trigger data includes audio data; Performing a first-level wake-up determination based on the trigger data; When the first-level wake-up determination passes, performing a second-level verification determination based on the trigger data; When the second-level verification determination passes, performing voice-to-text input; Wherein, the performing the second-level verification based on the trigger data includes: Performing voiceprint verification based on the audio data.

2. The method according to claim 1, characterized in that The method is applied to an electronic device, and the electronic device includes a first-level processing chip and a second-level processing chip. Among them, the first-level wake-up determination is executed by the first-level processing chip, and the second-level verification determination is executed by the second-level processing chip.

3. The method according to claim 2, wherein The first-level processing chip is a digital signal processor ADSP, and the second-level processing chip is an application processor AP.

4. The method according to claim 1, wherein The method is applied to an electronic device, and the electronic device includes an inertial measurement unit, a microphone, and a proximity sensor; The obtaining the trigger data of the electronic device includes: Obtaining inertial measurement data collected by the inertial measurement unit, audio data collected by the microphone, and proximity light data collected by the proximity sensor; The performing the first-level wake-up determination based on the trigger data includes: Performing device attitude detection of the electronic device based on the inertial measurement data; Detecting the user's approach based on the proximity light data; Performing audio data detection based on the audio data; Among them, when the device attitude detection, the user's approach, and the audio data detection all pass, it is determined that the first-level wake-up determination passes.

5. The method according to claim 1, characterized in that, The performing the second-level verification determination based on the trigger data includes: Performing voice / noise recognition on the audio data to obtain voice data; Performing recognition of recording playback attacks on the voice data; When the voice data is not a recording, determining the sound source angle of the user based on the trigger data; Performing directional sound pickup on the voice data according to the sound source angle to obtain user voice data; The performing the voiceprint verification based on the audio data includes: Determining whether the voiceprint of the user voice data is the voiceprint of a specified user; among them, if it is the voiceprint of the specified user, it is determined that the second-level verification determination passes.

6. The method according to claim 5, characterized in that, Before determining whether the voiceprint of the user voice data is the voiceprint of a specified user, the method further includes: Performing voice enhancement on the user voice data.

7. The method according to claim 5, characterized in that The method further includes: Pre-recording the fixed wake-up voice of the specified user, and extracting the voiceprint features of the fixed wake-up voice to obtain the voiceprint of the specified user; The determining whether the voiceprint of the user voice data is the voiceprint of a specified user includes: Calculating the similarity between the voiceprint of the user voice data and the voiceprint of the specified user. Among them, if the similarity is greater than a preset threshold, it is determined that the voiceprint of the user voice data is the voiceprint of the specified user.

8. The method according to claim 5, characterized in that, The method further includes: Obtaining text-independent speech of the specified user, and training a voiceprint recognition model using the text-independent speech of the specified user to obtain a trained voiceprint recognition model; The determining whether the voiceprint of the user voice data is the voiceprint of a specified user includes: Use the voiceprint recognition model to determine whether the voiceprint of the user voice data is the voiceprint of a specified user.

9. An electronic device, characterized in that, It includes: A memory for storing computer programs; A processor for implementing the following method when executing the program stored on the memory: Obtain the trigger data of the electronic device, where the trigger data includes audio data; Perform a first-level wake-up judgment based on the trigger data; When the first-level wake-up judgment passes, perform a second-level verification judgment based on the trigger data; When the second-level verification judgment passes, perform voice-to-text input; Among them, the second-level verification based on the trigger data includes: Perform voiceprint verification based on the audio data.

10. The electronic device according to claim 9, characterized in that, The processor includes a first-level processing chip and a second-level processing chip; The first-level processing chip is used to perform a first-level wake-up judgment based on the trigger data; The second-level processing chip is used to perform a second-level verification judgment based on the trigger data when the first-level wake-up judgment passes; When the second-level verification judgment passes, perform voice-to-text input.

11. The electronic device according to claim 10, characterized in that, The first-level processing chip is a digital signal processor ADSP, and the second-level processing chip is an application processor AP.

12. The electronic device according to claim 10, wherein The electronic device further includes an inertial measurement unit, a microphone, and a proximity sensor; The inertial measurement unit is used to collect inertial measurement data; The microphone is used to collect audio data; The proximity sensor is used to collect proximity light data; The first-level processing chip is specifically used for: detecting the device posture of the electronic device according to the inertial measurement data; detecting the user's approach according to the proximity light data; performing audio data detection according to the audio data, where when the device posture detection, the user's approach, and the audio data detection all pass, it is determined that the first-level wake-up judgment passes.

13. The electronic device according to claim 10, wherein The second-level processing chip is specifically used for: performing voice / noise recognition on the audio data to obtain voice data; performing recognition of recording playback attacks on the voice data; when the voice data is not a recording, determining the sound source angle of the user according to the trigger data; performing directional sound pickup on the voice data according to the sound source angle to obtain user voice data; determining whether the voiceprint of the user voice data is the voiceprint of a specified user; among them, if it is the voiceprint of the specified user, it is determined that the second-level verification judgment passes.

14. The electronic device according to claim 13, wherein The second-level processing chip is further used for: performing voice enhancement on the user voice data.

15. The electronic device according to claim 13, characterized in that, The second-level processing chip is further used for: pre-recording the fixed wake-up voice of the specified user, extracting the voiceprint features of the fixed wake-up voice to obtain the voiceprint of the specified user; The second-level processing chip is specifically used for: calculating the similarity between the voiceprint of the user voice data and the voiceprint of the specified user, where if the similarity is greater than a preset threshold, it is determined that the voiceprint of the user voice data is the voiceprint of the specified user.

16. The electronic device according to claim 13, wherein The second-level processing chip is further used for: obtaining the text-independent voice of the specified user, and using the text-independent voice of the specified user to train the voiceprint recognition model to obtain the trained voiceprint recognition model; The secondary processing chip is specifically configured to: use the voiceprint recognition model to determine whether the voiceprint of the user voice data is the voiceprint of a specified user.

17. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1-8 is implemented.

Citation Information

Patent Citations

  • Voice wake-up method and electronic equipment

    CN115588435A

  • Voiceprint recognition method and device, electronic equipment and storage medium

    CN115995232A

  • Voice wake-up method and electronic equipment

    CN117116258A