Interaction method and device, electronic equipment and readable storage medium

Through a two-level recognition mechanism and multiple audio processing technologies, the wake-up process of the voice assistant is simplified, the false wake-up rate is reduced, it is suitable for portable electronic devices, and the user experience is improved.

CN119007717BActive Publication Date: 2026-01-09HONOR DEVICE CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202310575797.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-19
Publication Date
2026-01-09
Estimated Expiration
2043-05-19

AI Technical Summary

Technical Problem

The wake-up process of existing voice assistants is cumbersome and has a high false wake-up rate, which hinders their further promotion and popularization.

Method used

A two-level recognition mechanism is adopted. The audio digital signal processor performs the first recognition of the specified state, and triggers the application processor to perform the second recognition of the cached audio data. Combined with technologies such as voice/wind noise recognition, front-end enhancement, directional sound pickup, and breath recognition, the performance requirements of ADSP are reduced, and the convenience and accuracy of wake-up are improved.

Benefits of technology

It simplifies the voice assistant wake-up process, reduces the false wake-up rate, is suitable for low-spec electronic devices, and improves the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119007717B_ABST
    Figure CN119007717B_ABST
Patent Text Reader

Abstract

The application discloses an interactive method, device, electronic equipment and readable storage medium. The method comprises the following steps: performing first identification on a specified state of the electronic equipment by using an ADSP, and obtaining a first identification result; when the first identification result is matched with a state of using a wake-up word to wake up a voice assistant, triggering an application processor (AP) to perform second identification on first audio data, and obtaining a second identification result; the first audio data is audio data obtained from a microphone and stored in a cache; and when the second identification result is that the first audio data comprises audio data that needs to be responded by the voice assistant, the voice assistant is woken up. When the embodiment is used for interaction, the voice assistant can be woken up by using the wake-up word, and the operation is simple and convenient. In addition, by using the second identification, a part of the identification task is completed by the AP, so that the requirement on the performance of the ASDP is reduced; the cooperation of the ASDP and the AP in identification is also beneficial to reducing the false wake-up rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech recognition technology, and in particular to an interaction method, device, electronic device, and readable storage medium. Background Technology

[0002] With the development of speech recognition technology, voice assistants are becoming increasingly widely used in portable electronic devices such as mobile phones, smartwatches, and wearable devices. Voice assistants allow users to control functions and query information via voice. In actual use, users need to wake up the voice assistant by saying a preset wake-up word.

[0003] It should be noted that activating a voice assistant using a wake word requires prior registration. The registration process is cumbersome. For example, registering the wake word "Hello YOYO" requires finding a multi-level menu entry point and then saying "Hello YOYO" three times in a quiet environment at a preset distance (e.g., 30cm). This cumbersome process hinders the further promotion and widespread adoption of the feature, and in actual use, the voice assistant experiences a relatively high rate of false wake-ups.

[0004] Therefore, how to conveniently wake up the voice assistant and reduce the false wake-up rate is an urgent technical problem to be solved. Summary of the Invention

[0005] This application provides an interaction method, device, electronic device, and readable storage medium that can improve the convenience of waking up a voice assistant and reduce the false wake-up rate.

[0006] Firstly, an interaction method is provided for use in an electronic device, the method comprising:

[0007] An audio digital signal processor (ADSP) is used to perform a first identification on a specified state of the electronic device to obtain a first identification result; the first identification result includes: determining whether the specified state of the electronic device matches the state when the voice assistant of the electronic device is woken up with a wake-up word;

[0008] When the first recognition result matches the specified state with the state when waking up the voice assistant with a wake-up word, the application processor (AP) is triggered to perform a second recognition on the first audio data to obtain a second recognition result; the first audio data is audio data obtained from the microphone stored in the cache; the second recognition result includes: determining whether the first audio data includes audio data that the voice assistant needs to respond to;

[0009] wake up the voice assistant when the second recognition result is that the first audio data includes audio data that the voice assistant needs to respond to.

[0010] When the interaction is performed by using this embodiment, two-stage recognition is adopted, first, the ASDP is used to perform first recognition on the specified state, to determine whether the specified state of the electronic device matches the state when the voice assistant is woken up by using the free wake-up word, and when the specified state matches the state, the AP is triggered to perform second recognition on the audio data stored in the cache and acquired from the microphone, to determine whether the audio data includes audio data that the voice assistant needs to respond to, and when the audio data stored in the cache and acquired from the microphone includes audio data that the voice assistant needs to respond to, the voice assistant is woken up. When the interaction is performed by using this embodiment, the voice assistant is woken up by using the free wake-up word, which is simple and convenient to operate. In addition, by using the second recognition, part of the recognition task is performed on the AP side, which reduces the requirement on the performance of the ASDP, and is conducive to applying the technical solution to electronic devices with low ASDP configuration. The recognition is performed by the cooperation of the ASDP and the AP, which is conducive to reducing the false wake-up rate.

[0011] It should be noted that the space of the cache is limited, and the audio data acquired from the microphone is circularly cached in the cache. When the first recognition result is that the specified state matches the state when the voice assistant is woken up by using the free wake-up word, the latest stored audio data (for example, 2 seconds of data can be stored each time, or data of a predetermined size can be stored) in the cache is acquired as the first audio data, and the first audio data is recognized to determine whether the first audio data includes audio data that the voice assistant needs to respond to. In some possible embodiments, if the first audio data is a voice uttered by a target user (if the electronic device is a mobile phone, the target user can be the owner of the mobile phone), the voice uttered by the target user can be regarded as audio data that the voice assistant needs to respond to. The audio data corresponding to the recording, the voice uttered by other users, and the wind sound can not be regarded as audio data that the voice assistant needs to respond to.

[0012] It should be noted that the first audio data can be audio data with a preset time length or audio data with a capacity of a preset byte, which can be determined according to needs and / or experience, and the present application is not limited in this regard.

[0013] By way of an example of the present application, the specified state includes at least one of the following: a motion state of the electronic device, a spatial posture of the electronic device, a distance between a microphone of the electronic device and a sound source, and an intensity of proximity light. Where the electronic device is exemplified by a mobile phone, when the voice assistant is woken up using the wake-up word-free, the user will usually have the action of picking up the electronic device, that is, the electronic device has a moving motion state, the bottom microphone of the mobile phone is close to the mouth (that is, the distance between the sound source and the microphone is close, such as within 5 cm), the horizontal angle of the electronic device is within a certain range (such as -60°~60° relative to the horizontal plane), and the light intensity around the mobile phone is usually greater than a certain value (if the light intensity is low, it may be in a closed space, such as in a backpack, a pocket, etc.). In specific embodiments, the specified state can be set as needed. Where the motion state of the electronic device can be detected using an acceleration sensor or the like detection module to determine whether the electronic device has a moving state. The spatial posture of the electronic device can be the angle between the display interface (such as the display screen of the mobile phone) of the electronic device and the horizontal plane, which can be detected by a gyroscope, an acceleration sensor or the like module; the distance between the sound source and the microphone can be identified using the intensity of the sound source detected by different microphones (in actual use, one, two, or three or the like number of microphones can be used), and the intensity of the proximity light can be identified using a proximity light sensor.

[0014] By way of an example of the present application, the second identification includes: pre-processing the first audio data to obtain second voice data; the pre-processing includes at least one of the following operations: voice / wind noise identification, front-end enhancement, directional sound pickup; and identifying the second voice data using a pre-set breath identification model to obtain a first confidence; determining a second identification result according to the first confidence. The pre-set breath identification model can be a pre-trained neural network model, which is used to identify the second voice data. When the voice assistant is woken up using the wake-up word-free, the user usually speaks close to the microphone. The breath identification model is used to identify the breath features generated by the user when speaking to the microphone to obtain the first confidence. The second identification result is determined according to the first confidence. For example, if the first confidence is high or exceeds a threshold, it can be determined that the second identification result is that the first audio data includes audio data that needs to be responded by the voice assistant. It can be understood that when the second identification is performed, if other parameters are identified, the identification results of the other parameters can be fused to determine the second identification result. Specifically, how to fuse can be determined according to the influence of each parameter on the second identification result, or can be determined according to experience.

[0015] The voice / wind noise recognition preprocessing can identify whether there is wind noise in the first audio data, and if there is wind noise, the wind noise is removed from the first audio data. The front-end enhancement preprocessing can identify the position of the sound source, determine the end close to the sound source as the front end, and determine the other end as the rear end. For example, when the mobile phone is held horizontally, the sound source close to the direction of the microphone at the bottom of the mobile phone is determined as the front-end sound source, and the voice in this direction can be amplified, which is beneficial to improve the wake-up rate. Further, the sound in a specific angle or a specific angle range can be picked up, that is, the directional sound pickup preprocessing. Which preprocessing operation is used can be determined according to actual conditions.

[0016] As an example of the present application, the second identification further includes: performing sound source angle identification on the first audio data, determining a third identification result according to the identified sound source angle and a preset sound source angle threshold, the preset sound source angle threshold being an angle range corresponding to the audio data in the first audio data that the voice assistant needs to respond to; and determining the second identification result according to the first confidence level, including: determining the second identification result according to the first confidence level and the third identification result. The angle of the sound source can be the inclination angle of the mobile phone display screen relative to the horizontal plane, and the third identification result obtained can be fused with other module identification results to obtain a further identification result.

[0017] As an example of the present application, determining the second identification result according to the first confidence level and the third identification result includes: performing fusion processing on the first confidence level and the third identification result according to a preset fusion rule to determine the second identification result.

[0018] As an example of the present application, the second identification can further include: identifying the first audio data according to a text-independent voiceprint identification module to obtain a fourth identification result, the fourth identification result including: determining whether the first audio data includes voiceprint information of a target user, and the second identification result being obtained after fusing the fourth identification result.

[0019] As an example of the present application, the second identification further includes: identifying the first audio data according to a recording playback attack defense module to obtain a fifth identification result, the fifth identification result including: determining whether the first audio data is recording data. The second identification result is obtained after fusing the fifth identification result. The recording data includes: synthesized / spliced recording data. This embodiment only responds to real human speech, and does not respond to recording playback, recording splicing, and synthesized recording, which is beneficial to protect privacy and security.

[0020] In a second aspect, the present application provides an interaction device applied to an electronic device, the interaction device comprising: an audio digital signal processor (ADSP), an application processor (AP), and a wake-up module, wherein

[0021] The ADSP is configured to perform first identification on a specified state of the electronic device to obtain a first identification result, the first identification result comprising: determination of whether the specified state of the electronic device matches a state when a voice assistant is woken up by a free wake-up word.

[0022] The AP is configured to perform second identification on first audio data to obtain a second identification result when the first identification result indicates that the specified state matches the state when the voice assistant is woken up by the free wake-up word, the first audio data being audio data obtained from a microphone and stored in a cache, and the second identification result comprising: determination of whether the first audio data includes audio data that needs to be responded to by the voice assistant.

[0023] The wake-up module is configured to wake up the voice assistant when the second identification result indicates that the first audio data includes the audio data that needs to be responded to by the voice assistant.

[0024] As an example of the present application, the specified state comprises at least one of the following: a motion state of the electronic device, a spatial posture of the electronic device, a distance between a microphone of the electronic device and a sound source, and a light intensity of a proximity light.

[0025] As an example of the present application, the second identification performed by the AP can specifically comprise: pre-processing the first audio data to obtain second voice data, the pre-processing comprising at least one of the following operations: voice / wind noise identification, front-end enhancement, directional sound pickup; and identifying the second voice data using a preset breath identification model to obtain a first confidence degree; and determining the second identification result according to the first confidence degree.

[0026] As an example of the present application, when performing the second identification, the AP is further configured to: perform sound source angle identification on the first audio data, determine a third identification result according to a recognized sound source angle and a preset sound source angle threshold value, the preset sound source angle threshold value being a corresponding angle range when the first audio data includes the audio data that needs to be responded to by the voice assistant, and the determination of the second identification result according to the first confidence degree comprising: determination of the second identification result according to the first confidence degree and the third identification result.

[0027] As an example of the present application, when performing the second identification, the AP is further configured to: perform fusion processing on the first confidence degree and the third identification result according to a preset fusion rule to determine the second identification result.

[0028] As an example of the present application, the AP in the second identification is further used for: identifying the first audio data according to a text-independent voiceprint identification module to obtain a fourth identification result, the fourth identification result including: determining whether the voiceprint information of the target user is included in the first audio data; and the second identification result being obtained after fusing the fourth identification result.

[0029] As an example of the present application, the AP in the second identification is further used for: the second identification further including:

[0030] identifying the first audio data according to a recording playback attack defense module to obtain a fifth identification result, the fifth identification result including: determining whether the first audio data is recording data, and the second identification result being obtained after fusing the fifth identification result.

[0031] In a third aspect, the present application provides an electronic device, comprising: a memory and one or more processors, the memory being coupled with the processor; wherein the memory stores computer program code, the computer program code including computer instructions, when the computer instructions are executed by the processor, causing the electronic device to perform the interaction method of the first aspect or any possible implementation manner of the first aspect.

[0032] In a fourth aspect, the present application provides a computer readable storage medium, comprising computer instructions, when the computer instructions run on an electronic device, causing the electronic device to perform the interaction method of the first aspect or any possible implementation manner of the first aspect.

[0033] The technical effects obtained by the above-mentioned second aspect, third aspect, and fourth aspect are similar to the technical effects obtained by the corresponding technical means in the above-mentioned first aspect, which will not be repeated here. BRIEF DESCRIPTION OF DRAWINGS

[0034] Figure 1 is a schematic diagram of a microphone pickup process in the prior art;

[0035] Figure 2 is a schematic diagram of a process of an interaction method of the present application;

[0036] Figure 3 is a schematic diagram of a microphone pickup process in the present application;

[0037] Figure 4A is a schematic diagram of an embodiment architecture of the present application;

[0038] Figure 4B is a schematic diagram of ADSP and AP interaction in an embodiment of the present application;

[0039] Figure 5 is a schematic diagram of interaction between ADSP and AP in an embodiment of the present application;

[0040] Figure 6A is a schematic diagram of processing flow at AP side in an embodiment of the present application;

[0041] Figure 6B is a schematic diagram of processing flow at AP side in an embodiment of the present application;

[0042] Figure 6C is a schematic diagram of processing flow at AP side in an embodiment of the present application;

[0043] Figure 7 is a schematic diagram of structure of an electronic device provided in an embodiment of the present application;

[0044] Figure 8 is a block diagram of software system of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0045] In order to make the objects, technical solutions and advantages of the present application clearer, the embodiments of the present application will be further described in detail below with reference to the drawings.

[0046] It should be understood that the "multiple" mentioned in the present application refers to two or more than two. In the description of the present application, unless otherwise specified, " / " represents the meaning of or, for example, A / B can represent A or B; "and / or" in the present application is only a description of the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B can represent the three cases of A alone, A and B together, and B alone. In addition, in order to clearly describe the technical solutions of the present application, the same items or similar items with basically the same functions and roles are distinguished by using "first", "second", etc. The skilled in the art can understand that "first", "second", etc. do not limit the quantity and execution order, and "first", "second", etc. also do not necessarily mean different.

[0047] In the present application, the reference "one embodiment" or "some embodiments" means that in one or more embodiments of the present application, the specific features, structures or characteristics described in connection with the embodiment are included. Therefore, the statements "in one embodiment", "in some embodiments", "in other some embodiments", "in other some embodiments" and the like appearing in the present specification are not necessarily all referring to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized. The terms "include", "contain", "have" and their variants mean "include but not limited to", unless otherwise specifically emphasized.

[0048] Voice assistants are increasingly being used in portable electronic devices such as mobile phones, smartwatches, and wearable devices. They enable functions such as voice control and information retrieval.

[0049] like Figure 1 In the existing technology shown, a fixed wake word is used to wake up the voice assistant. To avoid accidental wake-ups, Figure 1 The existing technology described above employs a two-level low-power wake-up recognition system: a first-level low-power voice wake-up recognition system and a second-level low-power voice wake-up recognition system. For example, if the wake-up word is "Hello YOYO," and the first-level low-power voice wake-up recognition module receives the audio data of "Hello YOYO" through the microphone, if the wake-up word matches a preset wake-up word, the second-level low-power voice wake-up recognition module retrieves cached data from the audio data buffer to further identify whether the wake-up word is the target user's voiceprint information. If it is the target user's voiceprint information, the second-level low-power voice wake-up recognition passes, waking up the voice assistant. The voice recognition and processing module then retrieves real-time audio data from the microphone, and the voice assistant performs recognition on the retrieved real-time audio data. It should be noted that the existing technology requires wake-up word registration, which is cumbersome. For example, registering the wake-up word "Hello YOYO" requires finding a multi-level menu entry point and then saying "Hello YOYO" three times in a quiet environment at a preset distance (e.g., 30cm). This cumbersome operation hinders the further promotion and popularization of the function, and in actual use, the voice assistant has a high false wake-up rate. In addition, voiceprint information is extracted by reading a specified paragraph or from a fixed wake word repeatedly spoken during registration, so the accuracy of voiceprint is not high.

[0050] See Figure 2 , Figure 2 This is a flowchart illustrating an interaction method provided in an embodiment of this application. The electronic device using this interaction method in this embodiment is described using a mobile phone as an example. Figure 2 As shown, the interaction method includes steps 201 to 203.

[0051] Among them, 201, the audio digital signal processor (ADSP) is used to perform a first identification of the specified state of the electronic device to obtain a first identification result.

[0052] The first recognition result includes: determining whether a specified state of the electronic device matches the state when the voice assistant is woken up with a wake-up word.

[0053] In some possible embodiments, the specified state includes at least one of the following states: the motion state of the electronic device, the spatial orientation of the electronic device, the distance between the microphone of the electronic device and the sound source, and the light intensity of the proximity light.

[0054] 202、in a case where the first recognition result matches a specified state and a state of waking up the voice assistant with the wake-up word, triggering the application processor (AP) to perform second recognition on the first audio data to obtain a second recognition result.

[0055] In some possible embodiments, the first audio data is audio data obtained from a microphone and stored in a cache; and the second recognition result includes determining whether the first audio data includes audio data that needs to be responded to by the voice assistant.

[0056] The second recognition includes: pre-processing the first audio data to obtain second voice data; the pre-processing includes at least one of the following operations: voice / wind noise recognition, front-end enhancement, directional sound pickup; and identifying the second voice data by using a preset breath recognition model to obtain a first confidence degree.

[0057] In some possible embodiments, the second recognition can further include: performing sound source angle recognition on the first audio data, and determining a third recognition result according to a recognized sound source angle and a preset sound source angle threshold, the preset sound source angle threshold being a corresponding angle range in a case where the first audio data includes audio data that needs to be responded to by the voice assistant.

[0058] In some possible embodiments, the second confidence degree is obtained by performing fusion processing on the first confidence degree and the third recognition result according to a preset fusion rule; and a fourth recognition result is obtained according to the second confidence degree.

[0059] In some possible embodiments, the second recognition further includes: identifying the first audio data according to a text-independent voiceprint recognition module to obtain a fifth recognition result, the recognition result including determining whether the first audio data includes voiceprint information of a target user.

[0060] In some possible embodiments, the second recognition further includes: identifying the first audio data according to a recording playback attack defense module to obtain a sixth recognition result, the sixth recognition result including determining whether the first audio data is recording data.

[0061] 203、in a case where the second recognition result indicates that the first audio data includes audio data that needs to be responded to by the voice assistant, waking up the voice assistant.

[0062] When the implementation is used for interaction, two-stage recognition is used, first, the ASDP is used to perform first recognition on the specified state, to determine whether the specified state of the electronic device matches the state when the voice assistant is woken up by the wake-up word, when the states match, the AP is triggered to perform second recognition on the audio data stored in the cache and obtained from the microphone, to determine whether the audio data includes audio data that needs to be responded to by the voice assistant, when the audio data stored in the cache and obtained from the microphone includes audio data that needs to be responded to by the voice assistant, the voice assistant is woken up. When the implementation is used for interaction, the voice assistant is woken up by the wake-up word, which is simple and convenient to operate. In addition, through the second recognition, part of the recognition task is performed on the AP side, which reduces the requirement on the performance of the ASDP, and is conducive to applying the technical solution to electronic devices with low ASDP configuration. The cooperation of the ASDP and the AP for recognition is conducive to reducing the false wake-up rate.

[0063] In Figure 2 The processing flow of the pickup flowchart of the microphone in the interaction method shown in FIG. 8 can be as shown in FIG. 9. Figure 3 It should be noted that in some embodiments, other states can also be included, such as the proximity light can be recognized, and the preset specified state includes the angle (such as within ±60°) between the display screen of the mobile phone and the horizontal plane: the proximity light detected by the proximity light sensor exceeds the threshold, and the proximity light recognition is set, which is to avoid possible misoperations in the scenario where the mobile phone is placed in the pocket. It should be noted that the distance between the mobile phone and the user's mouth, that is, the distance between the mobile phone and the sound source, when two microphones are provided on the mobile phone, the distance between the mobile phone and the sound source can be determined by the difference in breath strength of the voice information obtained by the two microphones. It can be understood that in some possible embodiments, other methods can also be used to determine the distance between the mobile phone and the sound source.

[0064] When the user expects to use the wake-up word to wake up the voice assistant, the mobile phone can be held close to the mouth, and then a voice instruction such as "What's the weather today" can be spoken. When the current state of the mobile phone matches the preset specified state, the voice recognition and processing module obtains the audio data obtained from the microphone from the audio data circular buffer and performs recognition thereon, and when it is determined through the recognition that the audio data includes audio data that needs to be responded to by the voice assistant, the voice assistant is woken up, and the voice assistant responds to the audio data.

[0065] Specifically, as shown in Figure 4A and Figure 4B The embodiment of the present application uses the ADSP and the AP to perform one-time recognition, respectively. As shown in Figure 4AAs shown, the inertial measurement unit (IMU) data acquired by the sensor is transmitted to the wake-up module without the wake-up word, and the distance between the microphone and the sound source can be determined according to the IMU data, and it needs to be noted that Figure 4A The embodiment shown is compatible with the fixed wake-up word wake-up function and the function without the wake-up word, and it can be understood that in some possible implementations, only the function of the wake-up without the wake-up word can be included, and both are feasible.

[0066] The ADSP side identifies whether the state of the mobile phone matches the specified state according to the voice data acquired from the microphone 1 and the microphone 2 and the sensor signal acquired from the sensor, such as Figure 4B As shown, the IMU module and the voice module can be used to identify the specified state, and when the identification result is that the specified state matches the state of the wake-up voice assistant without the wake-up word, the AP side further identifies, such as Figure 4B The AP side can perform post-processing, and the large model can be to identify the audio data acquired from the audio data circular buffer by using the breath recognition module, and the post-processing can further set some identification parameters to identify whether the state matches the voice recognition wake-up. It needs to be noted that the ADSP side can only perform simple identification of the IMU and the voice, and most of the identification tasks can be performed on the AP side, which is beneficial to ensure the identification performance and reduce the dependence on the performance of the ADSP side. For some electronic devices with low ADSP configuration, the technology provided in the present application can also be used for wake-up without the wake-up word, which is beneficial to improve the application range of the present application.

[0067] It needs to be noted that which functions are implemented by the ADSP side and the AP side and which parameters are identified can be set according to experience or needs. For example, Figure 5 As shown, the ADSP side can use the IMU module, the voice module, the proximity light module, etc. to identify the mobile state (whether the mobile phone has moved), the distance between the microphone and the user's mouth, simple voice recognition, proximity light intensity, etc. In specific implementation, when any parameter detected by the ADSP does not match the specified state, subsequent operations are not performed, and when all the identified parameters match the specified state, the AP further identifies and processes other parameters. For example, the AP side can further use the voice / wind noise recognition module, the front-end enhancement module, the directional sound pickup module, the sound source angle module, the recording playback attack defense module, the breath recognition module, the text-independent voiceprint module, and the post-processing module to identify, and finally obtain a judgment result. According to the judgment result, it is determined whether to wake up the voice assistant to respond to the audio data acquired by the microphone.

[0068] After the identification on the ADSP side is passed, the AP side further identifies, please refer toFigure 6A-C is a flowchart of the process handled by the AP side in different embodiments.

[0069] As shown in Figure 6A , the AP side obtains double-channel voice data from the audio data circular buffer, and then obtains a discrimination result through steps such as voice / wind noise recognition, neural network (NN) front-end enhancement, directional sound pickup, sound source angle determination, angle adaptive threshold module to determine whether the sound source angle meets the preset requirement, large model, fusion judgment, post-processing, etc. Among them, the voice / wind noise recognition removes wind sound from the obtained voice data, the NN front-end enhancement performs front-end signal enhancement on the voice data from which the wind sound is removed, such as enhancing the intensity of the voice obtained close to the bottom microphone of the mobile phone, the directional sound pickup obtains the voice signal of a specific direction, the sound source angle is used to identify the inclination angle of the mobile phone screen and the horizontal plane, the angle adaptive threshold module judges whether the angle of the sound source matches the preset sound source angle, if the sound source angle is within the preset sound source angle threshold range, it is considered that the sound source angle matches the preset sound source angle threshold, the large model can perform breath recognition on the data after the front-end enhancement and the directional sound pickup, the post-processing can perform recognition of other parameters defined by the user, and finally the discrimination result is obtained.

[0070] It should be noted that the results output by each module can be confidence (high, medium, low, or numerical value, etc.), and the results of different modules can be fused through fusion for fusion judgment. How to fuse is not limited here. For example, if the result output by the large model is a confidence score, if the result output by the angle adaptive threshold module is a match, the result of the fusion judgment is a high confidence score, and if the result of the post-processing is also positive, the discrimination result can be positive, so as to wake up the voice assistant and trigger the voice assistant to recognize and respond to the audio data saved in the audio data circular buffer and obtained by the microphone recently. For example, if the user holds the mobile phone and says "What's the weather today" to the microphone, the voice assistant can reply according to the real-time weather condition.

[0071] Please refer to Figure 6B , the difference between this embodiment and the embodiment shown in Figure 6A is that it further includes recognition of text-independent voiceprint. The text-independent voiceprint is relative to the prior art, which requires the user to read a specified sentence or paragraph to obtain the voiceprint, and the voiceprint feature of the target user is identified through the read sentence or paragraph. The text-independent voiceprint feature in this embodiment does not require the user to read a specified sentence or paragraph. This embodiment uses artificial intelligence to learn and identify the voiceprint of the user when the user uses the mobile phone. The longer the mobile phone is used, the higher the matching degree of the obtained voiceprint information and the voiceprint of the mobile phone user, which can avoid the false wake-up of the voice assistant when other people get the user's mobile phone.

[0072] Please refer toFigure 6C This embodiment is similar to Figure 6A The difference in the illustrated embodiment is that it also includes the identification of the recording playback attack defense function. It should be noted that the voice recognized by the voice assistant should be the voice spoken by the phone owner in real time, not a recording of what the phone owner said. It should also be noted that there are some signals in the recording file that are different from the real-time voice, such as noise such as current noise. By setting up the recording playback attack defense module, the situation of waking up the voice assistant through recording can be avoided.

[0073] This application also provides an interactive device applied to an electronic device. The interactive device includes: an audio digital signal processor (ADSP), an application processor (AP), and a wake-up module. The ADSP is used to perform a first recognition on a specified state of the electronic device to obtain a first recognition result. The first recognition result includes determining whether the specified state of the electronic device matches the state when a voice assistant is woken up using a wake-up word. The AP is used to perform a second recognition on first audio data when the first recognition result shows that the specified state matches the state when a voice assistant is woken up using a wake-up word, to obtain a second recognition result. The first audio data is audio data acquired from a microphone stored in a cache. The second recognition result includes determining whether the first audio data includes audio data that the voice assistant needs to respond to. The wake-up module is used to wake up the voice assistant when the second recognition result shows that the first audio data includes audio data that the voice assistant needs to respond to. Possible implementations of each module can be found in the descriptions of the previous embodiments, and will not be repeated here.

[0074] Next, the electronic devices involved in the embodiments of this application will be described.

[0075] Figure 7 This is a structural schematic diagram of an electronic device 700 provided in an embodiment of this application, which can specifically be a mobile phone, smartwatch, portable wearable device, etc. See also... Figure 7The electronic device 700 can include a processor 710, an external memory interface 720, an internal memory 721, a universal serial bus (USB) interface 730, a charging management module 740, a power management module 741, a battery 742, an antenna 1, an antenna 2, a mobile communication module 750, a wireless communication module 760, an audio module 770, a speaker 770A, a receiver 770B, a microphone 770C, a headset interface 770D, a sensor module 780, a key 790, a motor 791, an indicator 792, a camera 793, a screen 794, and a subscriber identification module (SIM) card interface 795, etc. The sensor module 780 can include a pressure sensor 780A, a gyro sensor 780B, a barometric sensor 780C, a magnetic sensor 780D, an acceleration sensor 780E, a distance sensor 780F, a proximity light sensor 780G, a fingerprint sensor 780H, a temperature sensor 780J, a touch sensor 780K, an ambient light sensor 780L, a bone conduction sensor 780M, etc.

[0076] The processor 710 can include one or more processing units, such as: the processor 710 can include an application processor (AP), an audio digital signal processor (ADSP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Different processing units can be independent devices, or can be integrated in one or more processors.

[0077] The controller can be the nerve center and command center of the electronic device 700. The controller can generate operation control signals according to instruction operation codes and timing signals, and complete the control of fetching and executing instructions.

[0078] The processor 710 can also be provided with a memory for storing instructions and data. In some embodiments, the memory in the processor 710 is a cache memory. The memory can save instructions or data that the processor 710 has just used or repeatedly uses. If the processor 710 needs to use the instructions or data again, it can directly call from the memory. This avoids repeated access and reduces the waiting time of the processor 710, thus improving the efficiency of the system.

[0079] The electronic device 700 implements a display function through a GPU, a screen 794, and an application processor, etc. The GPU is a microprocessor for image processing, connected to the screen 794 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 710 can include one or more GPUs that execute program instructions to generate or change display information.

[0080] The screen 794 is used to display images, videos, etc. The screen 794 includes a display panel. The display panel can adopt a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flex light-emitting diode (FLED), a Miniled, a MicroLed, a Micro-OLED, a quantum dot light emitting diode (QLED), etc. In some embodiments, the electronic device 700 can include 1 or N screens 794, N being an integer greater than 1.

[0081] The electronic device 700 can implement a shooting function through an ISP, a camera 793, a video codec, a GPU, a screen 794, and an application processor, etc.

[0082] The ISP is used to process data fed back by the camera 793. For example, when taking a photo, the shutter is opened, the light is transmitted to the camera photosensitive element through the lens, the light signal is converted into an electric signal, and the camera photosensitive element transmits the electric signal to the ISP for processing to convert it into an image visible to the naked eye. The ISP can also optimize the algorithm for noise, brightness, and skin color of the image. The ISP can also optimize the exposure, color temperature, and other parameters of the shooting scene. In some embodiments, the ISP can be disposed in the camera 793.

[0083] The camera 793 is configured to capture still images or videos. An object projects an optical image through a lens to a photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, which is then passed to an ISP for conversion into a digital image signal. The ISP outputs the digital image signal to a DSP for processing. The DSP converts the digital image signal into a standard RGB, YUV, or the like format image signal. In some embodiments, the electronic device 700 can include one or N cameras 793, where N is an integer greater than one.

[0084] The external memory interface 720 can be configured to connect with an external memory card, such as a Micro SD card, to extend the storage capability of the electronic device 700. The external memory card communicates with the processor 710 via the external memory interface 720 to perform data storage functions. For example, music, video, and the like files can be saved in the external memory card.

[0085] The internal memory 721 can be configured to store computer executable program code including instructions. The processor 710 performs various functional applications and data processing of the electronic device 700 by running the instructions stored in the internal memory 721. The internal memory 721 can include a program storage area and a data storage area. The program storage area can store an operating system, at least one application program required by a function (such as a sound playing function, an image playing function, and the like), and the like. The data storage area can store data (such as audio data, a phone book, and the like) created during use of the electronic device 700, and the like. In addition, the internal memory 721 can include a high-speed random access memory, and can further include a non-volatile memory such as at least one magnetic disk storage device, a flash memory device, a universal flash storage (UFS), and the like.

[0086] The acceleration sensor 780E can detect the magnitude of acceleration of the electronic device 700 in each direction (generally, three axes). When the electronic device 700 is stationary, the acceleration sensor 780E can detect the magnitude and direction of gravity. The acceleration sensor 780E can also be used to identify the posture of the electronic device 700, and can be applied to a landscape / portrait screen switching function, a pedometer, and the like. Of course, the acceleration sensor 780E can also be combined with the gyro sensor 780B to identify the posture of the electronic device 700, and can be applied to a landscape / portrait screen switching function.

[0087] The gyroscope sensor 780B can be used to determine the motion posture of the electronic device 700. In some embodiments, the angular velocity of the electronic device 700 around three axes (i.e., x, y and z axes) can be determined by the gyroscope sensor 780B. The gyroscope sensor 780B can be used for anti-shake shooting. For example, when the shutter is pressed, the gyroscope sensor 780B detects the angle of shaking of the electronic device 700, calculates the distance that the lens module needs to compensate according to the angle, and makes the lens offset the shaking of the electronic device 700 by reverse movement to achieve anti-shake. The gyroscope sensor 780B can also be used for landscape / portrait switching, navigation, and motion sensing game scenarios.

[0088] It can be understood that the structure illustrated in the embodiments of the present application does not constitute a specific limitation on the electronic device 700. In some other embodiments of the present application, the electronic device 700 can include more or fewer components than those illustrated, or combine certain components, or split certain components, or different arrangement of components. The illustrated components can be implemented in hardware, software or a combination of software and hardware.

[0089] The electronic device provided by the embodiments of the present application can be a user equipment (UE), such as a mobile terminal (e.g., a mobile phone), a tablet computer, a handheld computer, a smart watch, a personal digital assistant (PAD), and the like.

[0090] In addition, an operating system runs on the above components. For example, an Android open source operating system developed by Google Company, etc.

[0091] The software system of the electronic device 700 can adopt a layered architecture, an event-driven architecture, a microkernel architecture, a microservice architecture, or a cloud architecture. In order to more clearly illustrate the display optimization method when the screen is rotated provided by the embodiments of the present application, the software system of the electronic device 700 is exemplarily illustrated by taking the layered architecture of the Android system as an example.

[0092] Figure 8 is a block diagram of a software system of an electronic device 700 provided by the embodiments of the present application. Referring to Figure 8 , the electronic device can include a hardware layer and a software layer, wherein the layered architecture of the Android system can include an application layer, an application framework layer, a system library layer and a kernel layer. In some optional embodiments, the system of the electronic device can also include a layer not mentioned in the above technical architecture, such as Android runtime (Android Runtime).

[0093] The application program layer can include a series of application program packages, such as navigation applications, music applications and video applications, etc. For example, Figure 8As shown, the application package can include video, chat, etc. applications, and a system user interface (System UI).

[0094] The video, chat, voice, etc. applications are used to provide corresponding services for users. For example, a user uses a video application to watch a video, uses a chat application to chat with other users, uses a music application to listen to music, uses a voice assistant to respond to voice instructions of the user, etc.

[0095] The System UI is used to manage a user interface (UI) of the electronic device, and in the embodiments of the present application, the System UI is used to manage an image after image synthesis to be displayed on the screen.

[0096] The application framework layer provides an application programming interface (API) and a programming framework for the applications of the application layer. The application framework layer includes some pre-defined functions. For example, Figure 8 As shown, the application framework layer can include a window management service module (WMS), a display rotation module (also referred to as DisplayRotation), an activity management service module (AMS), and an input management module (also referred to as Input), etc.

[0097] The WMS is used to manage a window program. The window manager can obtain a screen size, determine whether there is a status bar, perform a cutout on an image on the screen, etc. In the embodiments of the present application, the WMS can create and manage a window corresponding to an application.

[0098] The display rotation module is used to control the screen to rotate, so that the screen presents a portrait or landscape layout. For example, when it is determined that the screen needs to be rotated, the Surfaceflinger is notified to switch the portrait or landscape of the application interface.

[0099] The AMS is used to start a specific application according to the operation of the user. For example, when the image completes the synthesis operation, the image is triggered to be displayed on the screen, after the image is displayed, the image that needs to be executed for the cutout operation is triggered to perform the cutout operation, and an application stack corresponding to the video application is created, so that the video application can normally run.

[0100] The system library layer can include a plurality of functional modules, such as a sensor module (also referred to as sensor) and a SurfaceFlinger.

[0101] The sensor module is configured to acquire data collected by the sensor, such as ambient light collected under the screen. The gravity direction information of the electronic device is collected. Alternatively, the sensor module can also adjust the brightness of the screen according to the ambient light, and determine the landscape / portrait screen state information of the electronic device according to the gravity direction information of the electronic device, the landscape / portrait screen state information being used to indicate whether the electronic device is in a landscape screen state or a portrait screen state.

[0102] The Surfaceflinger is a system service configured to create, control and manage layers.

[0103] In addition, the system library layer can further include a surface manager, media libraries, a three-dimensional graphics processing library (such as OpenGL ES), a 2D graphics engine (such as SGL), and the like. The surface manager is configured to manage the display subsystem, and provide a fusion of 2D and 3D layers for multiple applications. The media libraries support playback and recording of multiple common audio and video formats, as well as static image files, and the like. The media libraries can support multiple audio and video encoding formats, such as MPEG4, H.264, MP3, AAC, AMR, JPG, PNG, and the like. The three-dimensional graphics processing library is configured to implement three-dimensional graphics drawing, image rendering, synthesis, and layer processing, and the like. The 2D graphics engine is a drawing engine for 2D drawing.

[0104] The kernel layer is a layer between hardware and software. In the embodiment of the present application, the kernel layer at least includes a touch driving module and a display driving module.

[0105] The display driving module is configured to display a synthesized image in the screen according to image data provided by the modules of the application framework layer and the application programs of the application layer. For example, a video application delivers one frame of image data of a video to the display driving module, and the display driving module displays one frame of image of the video on the touch screen according to the image data. The SystemUI delivers image data to the display driving module, and the display driving module displays the synthesized image in the screen.

[0106] The touch driving module is configured to monitor the capacitance values of each area of the touch screen. When a user clicks or slides on the touch screen, the capacitance values of the clicked or slid area will change, and the touch driving module can monitor the changes of the capacitance values of each area of the touch screen, and send a capacitance value change message to the input management module, the capacitance value change message carrying information such as the change amplitude of the capacitance values (or capacitance sampling values) of each area of the touch screen and the time when the change occurs.

[0107] The input management module can determine the touch operation according to the reported capacitance value change message, and then send the recognized touch operation to other modules. The touch operation here can include a click operation, a drag operation, and a specific gesture operation (such as an up-slip gesture operation, a horizontal-slip gesture operation, etc.).

[0108] The hardware layer includes a screen and an ambient light sensor, etc., and the ambient light sensor is used to detect the ambient light information under the screen, etc. When the ambient light sensor has a processing function, the image information corresponding to the matting operation can be obtained, the real ambient light information is determined according to the image information of the image obtained by the matting operation and the detected ambient information under the screen, and the adjustment signal for adjusting the screen brightness is generated according to the real ambient light information.

[0109] The above technical architecture lists the modules and devices that may be involved in the present application in the electronic device. In actual application, the electronic device can include all or part of the modules and devices of the above technical architecture, and other modules and devices not mentioned in the above technical architecture, of course, it can also only include the modules and devices of the above technical architecture, and the present embodiment does not limit this.

[0110] In order to facilitate the understanding of the interactive method provided by the embodiments of the present application, first, the technical architecture of the electronic device shown in Figure 8 is taken as an example, and the implementation manner of the interactive method provided by the present application is explained.

[0111] In the process of Figure 7 , the processor 710 includes an ADSP and an AP, the ADSP performs first identification on a specified state of the electronic device to obtain a first identification result; the first identification result includes: determining whether the specified state matches the state when the voice assistant is awakened by the free wake-up word; the specified state may be, for example: the user picks up the mobile phone and holds the bottom of the mobile phone close to the mouth (within 5 cm).

[0112] When the first identification result is that the specified state matches the state when the voice assistant is awakened by the free wake-up word, the application processor AP is triggered to perform second identification on the first audio data to obtain a second identification result; the first audio data is the audio data stored in the cache and obtained from the microphone; the second identification result includes: determining whether the first audio data includes audio data that needs to be responded by the voice assistant; when the second identification result is that the first audio data includes audio data that needs to be responded by the voice assistant, the voice assistant is awakened.

[0113] In some possible implementations, the specified state includes at least one of the following states: the mobile phone has moved, the angle between the display screen of the mobile phone and the horizontal plane is within a preset range, the distance between the microphone of the mobile phone and the sound source is within a preset distance, and the light intensity of the proximity light is greater than a preset value.

[0114] In some possible implementation manners, the second recognition comprises: pre-processing the first audio data to obtain second speech data; the pre-processing comprises at least one of the following operations: speech / wind noise recognition, front-end enhancement, directional sound pickup; and recognizing the second speech data by using a preset breath recognition model to obtain a first confidence, and determining the second recognition result according to the first confidence.

[0115] In some possible implementation manners, the second recognition further comprises: performing sound source angle recognition on the first audio data, determining a third recognition result according to a recognized sound source angle and a preset sound source angle threshold, the preset sound source angle threshold being an angle range corresponding to audio data in the first audio data that the voice assistant needs to respond to, and the determining the second recognition result according to the first confidence comprises: determining the second recognition result according to the first confidence and the third recognition result.

[0116] In some possible implementation manners, the determining the second recognition result according to the first confidence and the third recognition result comprises: performing fusion processing on the first confidence and the third recognition result according to a preset fusion rule to determine the second recognition result.

[0117] In some possible implementation manners, the second recognition further comprises: recognizing the first audio data according to a text-independent voiceprint recognition module to obtain a fourth recognition result, the fourth recognition result comprising: determining whether voiceprint information of a target user is included in the first audio data, and the second recognition result being obtained after the fourth recognition result is fused. For example, if the fourth recognition result is no, the second recognition result is that the first audio data does not include audio data that the voice assistant needs to respond to, and the voice assistant is not woken up.

[0118] In some possible implementation manners, the second recognition further comprises: recognizing the first audio data according to a recording playback attack defense module to obtain a fifth recognition result, the fifth recognition result comprising: determining whether the first audio data is recording data, and the second recognition result being obtained after the fifth recognition result is fused. For example, if the fifth recognition result is yes, the second recognition result is that the first audio data does not include audio data that the voice assistant needs to respond to, and the voice assistant is not woken up.

[0119] When interaction is performed by using the embodiment, two-stage recognition is adopted, first, the ASDP is used to perform first recognition on the specified state, to determine whether the specified state of the electronic device matches the state when the voice assistant is woken up by using the wake-up word, when the states match, the AP is triggered to perform second recognition on the audio data stored in the cache and acquired from the microphone, to determine whether the audio data includes audio data that needs to be responded by the voice assistant, when the audio data stored in the cache and acquired from the microphone includes the audio data that needs to be responded by the voice assistant, the voice assistant is woken up. When interaction is performed by using the embodiment, the voice assistant is woken up by using the wake-up word, and the operation is simple and convenient. In addition, by using the second recognition, part of the recognition task is performed on the AP side, the requirement on the performance of the ASDP is reduced, and the technical solution is beneficial to application to the electronic device with low ASDP configuration. The recognition is performed by cooperation of the ASDP and the AP, and the false wake-up rate is reduced.

[0120] The embodiment of the present application further provides a computer readable storage medium, which stores a computer program, and when the computer program is executed by a processor, steps in each of the method embodiments can be implemented.

[0121] The embodiment of the present application provides a computer program product, which comprises a computer program, and when the computer program is executed by a processor, steps in each of the method embodiments can be implemented.

[0122] The computer program can be stored in a computer readable storage medium. When the computer program is executed by the processor, the steps of the above-mentioned method embodiments can be implemented. The computer program includes computer program code. The computer program code can be in a form of source code, object code, executable file, or some intermediate form. The computer readable medium can include at least any entity or device that can carry the computer program code to the photographing device / electronic device, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium. For example, U disk, mobile hard disk, magnetic disk or optical disk, etc. In some jurisdictions, according to legislation and patent practice, the computer readable medium can not be an electrical carrier signal and a telecommunication signal.

[0123] In the above embodiments, the description of each embodiment has its own focus. The parts not described or recorded in detail in a certain embodiment can be referred to the related description of other embodiments.

[0124] Those skilled in the art can appreciate that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be realized in electronic hardware, or a combination of computer software and electronic hardware. Whether the functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0125] In the embodiments provided in the present application, it should be understood that the disclosed method and electronic device can be implemented in other ways. For example, the above-described device / network device embodiments are only schematic, for example, the division of the modules or units is only a logical function division, and there can be another division manner in actual implementation, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual coupling or direct coupling or communication connection between each other can be indirect coupling or communication connection through some interface, device or unit, and can be electrical, mechanical or in other forms.

[0126] The units described as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, that is, they can be located in one place, or can be distributed on a plurality of network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiments.

[0127] It should be understood that when used in the specification and the appended claims of the present application, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or sets thereof.

[0128] It should also be understood that the term "and / or" used in the specification and the appended claims of the present application refers to any combination of one or more of the associated listed items and all possible combinations, and includes these combinations.

[0129] Reference throughout this specification to "one embodiment" or "an embodiment" means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the application. Thus, the appearances of the phrases "in one embodiment" or "in an embodiment" or "in a various embodiment" or "in some embodiments" in various places throughout this specification are not necessarily referring to the same embodiment, but can, but not necessarily, refer to different embodiments. The terms "including," "comprising," "having" and variations thereof are meant to encompass the items listed thereafter, but do not exclude other items from also being present. The term "consisting of" is meant to exclude any item not specified, but "consisting essentially of" does not exclude other items that do not "materially" affect the essential characteristics of the application.

[0130] The above-described embodiments are merely intended to illustrate the technical solutions of the present application, but not to limit the same; even though the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacements to some technical features therein; and these modifications or replacements do not make the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application.

Claims

1. An interaction method, characterized in that, The method is applied to an electronic device, and the method comprises: A first identification is performed on a specified state of the electronic device by using an audio digital signal processor (ADSP), and a first identification result is obtained; the first identification result comprises: determining whether the specified state of the electronic device matches a state when a voice assistant of the electronic device is woken up by using a hotword; the specified state comprises at least one of the following states: a motion state of the electronic device, a spatial posture of the electronic device, a distance between a microphone of the electronic device and a sound source, and a light intensity of a proximity light; When the first identification result is that the specified state matches the state when the voice assistant is woken up by using the hotword, a second identification is performed on first audio data by using an application processor (AP), and a second identification result is obtained; the first audio data is audio data obtained from a microphone and stored in a cache; the second identification result comprises: determining whether the first audio data comprises audio data that needs to be responded to by the voice assistant; The second identification comprises: performing preprocessing on the first audio data to obtain second voice data; performing identification on the second voice data by using a preset breath identification model to obtain a first confidence degree; performing sound source angle identification on the first audio data, and determining a third identification result according to a sound source angle obtained by identification and a preset sound source angle threshold value, the preset sound source angle threshold value being a corresponding angle range when the first audio data comprises the audio data that needs to be responded to by the voice assistant; determining the second identification result according to the first confidence degree and the third identification result; performing identification on the first audio data according to a recording playback attack defense module to obtain a fifth identification result, the fifth identification result comprising: determining whether the first audio data is recording data, and the second identification result being obtained after the fifth identification result is fused; When the second identification result is that the first audio data comprises the audio data that needs to be responded to by the voice assistant, the voice assistant is woken up.

2. The method of claim 1, wherein, The preprocessing comprises at least one of the following operations: voice / wind noise identification, front-end enhancement, and directional sound pickup.

3. The method of claim 1, wherein, The determination of the second identification result according to the first confidence degree and the third identification result comprises: The first confidence degree and the third identification result are fused according to a preset fusion rule, and the second identification result is determined.

4. The method according to any one of claims 1 to 3, characterized in that, The second identification further comprises: The first audio data is identified according to a text-independent voiceprint identification module to obtain a fourth identification result, the fourth identification result comprising: determining whether the first audio data comprises voiceprint information of a target user; and the second identification result being obtained after the fourth identification result is fused.

5. An interactive device, characterized by The interactive device is applied to an electronic device, and the interactive device comprises: an audio digital signal processor (ADSP), an application processor (AP), and a wake-up module, wherein The ADSP is configured to perform first identification on a specified state of the electronic device to obtain a first identification result. The first identification result includes a determination of whether the specified state of the electronic device matches a state when a voice assistant is woken up by a free wake-up word. The specified state includes at least one of a motion state of the electronic device, a spatial posture of the electronic device, a distance between a microphone of the electronic device and a sound source, and a light intensity of a proximity light. The AP is configured to perform second identification on first audio data to obtain a second identification result when the first identification result indicates that the specified state matches the state when the voice assistant is woken up by the free wake-up word. The first audio data is audio data obtained from a microphone and stored in a cache. The second identification result includes a determination of whether the first audio data includes audio data that needs to be responded to by the voice assistant. The second identification includes pre-processing the first audio data to obtain second voice data, identifying the second voice data by using a preset breath identification model to obtain a first confidence degree, performing sound source angle identification on the first audio data, determining a third identification result based on a sound source angle obtained by the identification and a preset sound source angle threshold value, the preset sound source angle threshold value being a corresponding angle range when the first audio data includes the audio data that needs to be responded to by the voice assistant, determining the second identification result based on the first confidence degree and the third identification result, and identifying the first audio data by using a recording playback attack defense module to obtain a fifth identification result, the fifth identification result including a determination of whether the first audio data is recording data, and the second identification result being obtained after the fifth identification result is fused. The wake-up module is configured to wake up the voice assistant when the second identification result indicates that the first audio data includes the audio data that needs to be responded to by the voice assistant.

6. An electronic device, comprising: The electronic device includes a memory and one or more processors coupled to the memory. The memory stores computer program code including computer instructions that, when executed by the processor, cause the electronic device to perform the method of any one of claims 1-4. The computer program product includes computer instructions that, when executed on an electronic device, cause the electronic device to perform the method of any one of claims 1-4.

7. A computer readable storage medium characterized in that, ​

Citation Information

Patent Citations

  • Voice interaction method and related electronic equipment

    CN115881118A