Voice Interaction Method, Electronic Device, and Storage Medium
The multi-microphone channel module collects environmental audio and analyzes the voice frame energy to determine the speaker's orientation, solving the shortcomings of existing equipment in voice enhancement and scene adaptability, and achieving a more efficient voice interaction experience.
Patent Information
- Application Number
- CN202111517080.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-08
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2041-12-08
AI Technical Summary
Existing voice intelligent interactive devices have problems such as poor noise cancellation capabilities or weak scene adaptability in voice enhancement, which makes it impossible to achieve a better voice interactive experience.
The microphone module of the multi-microphone channel collects environmental audio files, extracts speaker audio, and determines a microphone close to the speaker's orientation by analyzing the speech frame energy in each microphone channel, thereby performing voice interaction.
There is no need to know the prior knowledge such as the location of the signal source. The speech frame energy analysis of different channels is used to locate the speaker's orientation, effectively avoid the interference of environmental noise, and adapt in a wider range of scenarios, improving the reliability and experience of voice interaction.
Smart Images

Figure CN114120984B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of Internet technologies, and particularly relates to a voice interaction method, an electronic device, and a storage medium. Background Art
[0002] With the continuous development of voice technologies, various voice interaction devices have been integrated into all aspects of people's lives, such as voice ticket vending machines, voice chatbots, and so on.
[0003] In the process of voice interaction between a voice device and a speaker, how the device accurately identifies the position of the speaker and enhances the speech of the speaker is the key to ensuring the accuracy of speech recognition and the reliability of voice interaction.
[0004] Currently, although there are some voice intelligent interaction devices on the market, they have problems such as poor noise cancellation ability or weak scene adaptability in terms of voice enhancement, resulting in an inability to achieve a better voice interaction experience.
[0005] In response to the above problems, the industry has not yet provided a better solution for the time being. Summary of the Invention
[0006] Embodiments of the present invention provide a voice interaction method, an electronic device, and a storage medium, which are used to solve at least one of the above technical problems.
[0007] In a first aspect, an embodiment of the present invention provides a voice interaction method, including: collecting an environmental audio file based on a microphone module having multiple microphone channels; each of the microphone channels is respectively configured with a corresponding microphone orientation; extracting a speaker audio from the environmental audio file; for the environmental audio file, determining the voice component energy of the speaker audio in each microphone channel, and determining the microphone orientation of the microphone channel corresponding to the maximum voice component energy as the speaker orientation; and performing a voice interaction operation based on the speaker orientation.
[0008] In a second aspect, an embodiment of the present invention provides an electronic device, including: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the steps of the above method.
[0009] In a third aspect, an embodiment of the present invention provides a storage medium, in which one or more programs including execution instructions are stored, and the execution instructions can be read and executed by an electronic device (including but not limited to a computer, a server, or a network device, etc.) to execute the steps of the above method of the present invention.
[0010] Fourthly, an embodiment of the present invention further provides a computer program product, which includes a computer program stored on a storage medium. The computer program includes program instructions that, when executed by a computer, cause the computer to execute the steps of the above method.
[0011] The beneficial effects of the embodiments of the present invention are as follows:
[0012] The electronic device uses a microphone module with multiple microphone channels to collect an environmental audio file, extracts the speaker audio from the environmental audio file, determines the microphone close to the speaker's orientation based on the energy analysis results of the speech frames in each microphone channel, thereby obtaining the speaker's orientation, and performs a voice interaction operation. Thus, without prior knowledge such as the signal source position, the speaker's orientation is located by analyzing the energy of speech frames in different channels, which can effectively avoid the interference of environmental noise and can be adapted in a wider range of scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for the description of the embodiments. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0014] Figure 1 Shows a flowchart of an example of the voice interaction method according to an embodiment of the present invention;
[0015] Figure 2 Shows a flowchart of an example in the multimodal interaction process of the voice interaction method according to an embodiment of the present invention;
[0016] Figure 3 Is a schematic structural diagram of an embodiment of the electronic device of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0017] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following clearly and completely describes the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0018] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.
[0019] The present invention may be described in the general context of computer-executable instructions, such as program modules, executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform particular tasks or implement particular abstract data types. The present invention may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communications network. In a distributed computing environment, program modules may be located in both local and remote computer storage media including storage devices.
[0020] In the present invention, terms such as "module", "system", etc. refer to related entities applied to a computer, such as hardware, a combination of hardware and software, software, or software in execution. Specifically, for example, a component may, but is not limited to, be a process running on a processor, a processor, an object, an executable component, an execution thread, a program, and / or a computer. Also, an application program or a script program running on a server, and the server can both be components. One or more components may be in an execution process and / or thread, and the components may be localized on one computer and / or distributed between two or more computers, and can be run by various computer-readable media. The components can also communicate through local and / or remote processes according to a signal having one or more data packets, for example, a signal from data that interacts with another component in a local system, a distributed system, and / or interacts with other systems through a network on the Internet.
[0021] Finally, it should also be noted that in this article, the terms "comprising" and "including" not only include those elements, but also include other elements not explicitly listed, or also include elements inherent to such a process, method, article, or device. Without further limitation, an element defined by the statement "comprising..." does not exclude the existence of additional identical elements in the process, method, article, or device including the said element.
[0022] It should be noted that in the current related technologies, there are various multi-modal intelligent interaction devices, and most of their voice enhancement parts use beamforming for voice enhancement and calculate the azimuth angle through the MUSIC (Multiple Signal Classification) method.
[0023] (1) Perform spatial filtering based on the beamforming method
[0024] Beamforming is to use spatial information for spatial filtering. The microphone array has at least two microphones and can distinguish the direction of the incoming wave to a certain extent. The interfering speech or other non-stationary noises in the non-expected direction can be linearly attenuated.
[0025] (2) MUSIC locates the speaker's azimuth angle based on the matrix feature space decomposition method;
[0026] a. The array signal contains signal and noise, and its covariance matrix is calculated;
[0027] b. Perform eigendecomposition on the covariance matrix to obtain the signal subspace and noise subspace;
[0028] c. Because the signal and noise are orthogonal to each other, when performing an angle search in the entire space, the angle corresponding to the maximum value of the power spectrum is the DOA (Direction-of-arrival estimation, or angle of arrival).
[0029] However, both beamforming technology and MUSIC technology have some defects, such as poor noise cancellation capability for co-directional noise interference; the shape and layout of the microphone array are strictly restricted and need to be strongly bound to the algorithm itself, and changing the structure will no longer be applicable. For example, it is suitable for flat microphones, but not effective for lateral microphones; and the embedded platform support is not friendly.
[0030] Specifically, beamforming requires the position information of the microphone array to perform time delay or phase compensation and amplitude weighting on the output of each array element to form a beam pointing in a specific direction. Therefore, the spatial position information of the microphone array must be strictly adapted to the algorithm. In addition, the performance of the beamforming algorithm depends on the azimuth information of the target signal. If the target signal and the interference signal are in the same beam, beamforming cannot distinguish them, spatial filtering cannot be performed, and naturally they cannot be eliminated.
[0031] The MUSIC (Multiple Signal Classification) algorithm is used to calculate the Direction of arrival (DOA). The SVD decomposition has high requirements for calculation accuracy. In addition, some embedded platforms are not friendly to the boundary value representation of floating-point numbers. In the case of extremely high precision, data processing exceptions may occur, which affects the overall audio processing effect.
[0032] Figure 1 The flowchart of an example of a voice interaction method according to an embodiment of the present invention is shown. It should be noted that the execution subject of the method embodiment of the present invention can be various electronic devices, such as various mobile terminals or electronic devices with voice interaction functions.
[0033] like Figure 1 As shown, in step 110, an ambient audio file is collected based on a microphone module having multiple microphone channels. Here, each microphone channel is respectively configured with a corresponding microphone orientation, and the ambient audio file contains audio information corresponding to each different microphone channel.
[0034] In step 120, speaker audio is extracted from the ambient audio file. Here, various speech frame extraction techniques can be adopted for the extraction operation.
[0035] It should be understood that in addition to the speaker audio (i.e., speech audio) in the ambient audio file, there may also be other types of audio information, such as background noise, etc.
[0036] In step 130, for the ambient audio file, the speech component energy of the speaker audio in each microphone channel is determined, and the microphone orientation of the microphone channel with the maximum speech component energy is determined as the speaker orientation.
[0037] Since the position of the speaker is specific, when the speaker's voice is collected by microphone channels in different orientations, the corresponding energies will also be different. For example, when the speaker shouts on the left side of the electronic device, obviously, the energy received by the left microphone of the electronic device will be much higher than that of the right microphone.
[0038] In step 140, based on the speaker orientation, a voice interaction operation is performed.
[0039] Exemplarily, when the speaker orientation is determined, audio enhancement can be performed on the microphone channel with a matching orientation, or the collected audio of other microphone channels can be suppressed to obtain high-quality speech frames, and then corresponding speech recognition and interaction operations can be performed. Or, the orientation of the interaction module (such as a touch screen) of the mobile terminal can be adjusted to further enrich the voice interaction experience.
[0040] Regarding the implementation details of the above step 120, in some examples of the embodiments of the present invention, the speaker audio can be extracted from the ambient audio file based on a preset BSS (Blind Source Separation) algorithm, which separates certain source signals from the observed mixed signals based on statistical methods and does not require prior knowledge such as the signal source position. Therefore, there is no strict restriction on the shape and layout of the microphones, so the algorithm has better adaptability to the hardware. Since the BSS algorithm model is not sensitive to the spatial position information of the sound source, it is also superior to the beamforming scheme for co-directional noise interference.
[0041] In some examples of the embodiments of the present invention, the voice interaction function operation is performed using a display screen, that is, some interactive functions are carried out between the user and the display screen. Specifically, the electronic device can adjust the position of the display screen according to the speaker's orientation, and when the display screen is successfully adjusted to the speaker's orientation, perform voice interaction operations based on the display screen. In this way, the orientation of the display screen can be intelligently adjusted so that the display screen faces the user, ensuring that the user has a high-quality voice operation experience.
[0042] Figure 2 The flowchart of an example in the multimodal interaction process of the voice interaction method according to the embodiments of the present invention is shown.
[0043] As Figure 2 In the multimodal interaction process shown, in step 210, environmental audio information is collected based on a microphone module with multiple microphone channels.
[0044] Here, the microphone channels in the microphone module are arranged in a ring to fully pick up audio signals from various directions. For example, a ring-shaped 6-microphone voice interaction scheme is adopted, both planar microphones and side microphones can be adapted, and a single camera is configured for image interaction to achieve a reliable voice interaction process.
[0045] In step 220, the speaker audio is extracted from the environmental audio file based on the BSS algorithm.
[0046] Furthermore, before using the BSS algorithm, the AEC (Acoustic Echo Cancelling) algorithm can be adopted for preprocessing to optimize the voice frame quality.
[0047] In step 230, it is detected whether the speaker audio meets the preset voice wake-up condition.
[0048] Specifically, the speaker audio can be subjected to speech recognition to determine whether it contains a specific voice wake-up keyword, such as "Xiaobu Xiaobu".
[0049] In step 240, when the speaker audio meets the voice wake-up condition, for the environmental audio file, the voice component energy of the speaker audio in each microphone channel is determined, and the microphone orientation of the microphone channel with the maximum voice component energy is determined as the speaker orientation.
[0050] Here, through the voice wake-up function, the voice frames before blind source separation 1 to 2 seconds after waking up are rolled back and the energy is calculated, the energy sizes of each microphone channel are compared, and the position of the microphone channel with the largest energy is taken as the preliminary target person orientation information.
[0051] In step 250, when the speaker's audio does not meet the voice wake-up condition, it indicates that the speaker's intention is not for voice interaction, and the operation can be directly ended.
[0052] In some examples of the embodiments of the present invention, the microphone channels in the microphone module are arranged in a ring. In this way, through the above energy analysis of the voice component energy of different microphone channels, preliminary speaker localization can be achieved. However, in order to achieve a more accurate speaker localization effect, a single camera can also be triggered for object confirmation to improve the reliability of voice interaction.
[0053] In step 260, an environmental image corresponding to the speaker's orientation is collected. For example, the camera's shooting function is activated according to the speaker's orientation to collect the corresponding environmental image.
[0054] In step 270, it is identified whether there is target object information in the environmental image. Here, the target object information can be the face information of the target user.
[0055] Specifically, the above target detection scheme can adopt the yolov3 (You only look once, detect object categories and positions at once) algorithm scheme, so that a lightweight inference framework darknet can be deployed on the embedded platform to achieve better real-time processing capabilities.
[0056] In step 280, when target object information is recognized in the environmental image, a voice interaction operation is triggered.
[0057] In some cases, pixel-level analysis can also be performed on the target object information in the environmental image to extract the orientation information of the target object, and the position of the microphone that matches the orientation is fine-tuned accordingly to focus on the speaker's voice. Thus, by fusing the voice frame energy analysis and localization scheme of different channels with the image target recognition and localization scheme, and adopting the combined image and voice scheme for azimuth angle localization, linking the camera and the microphone can significantly improve the accuracy of target object localization and ensure the reliability of the microphone's sound pickup function.
[0058] In step 290, when target object information is not recognized in the environmental image, the speaker's orientation is calibrated based on the orientations of the microphones in each microphone channel. Specifically, the voice frame energy of each microphone channel can be recalculated, and the speaker's orientation can be calibrated using the recalculated energy results. In addition, the orientation of the microphone with the second-highest voice frame energy in the previous energy calculation can also be used to calibrate the speaker's orientation, and the camera can be adjusted to integrate the target recognition scheme to accurately locate the speaker's position.
[0059] In some examples of the embodiments of the present invention, the camera may be a front camera of the display screen. When the display screen is adjusted to a specified position and a target face is captured, the interactive display screen is correspondingly aligned with the interactive person, and the voice interaction function of the electronic device can be directly triggered.
[0060] Through the embodiments of the present invention, the fusion of target detection of images and microphone voice signal processing is applied to improve the accuracy and robustness of user positioning, thereby bringing a better user experience to the interactive person.
[0061] Through the embodiments of the present invention, by using voice frame energy analysis and BSS algorithm, the layout adaptation of the microphone is more flexible, which is convenient for structural design. The anti-noise ability is stronger in the case of co-directional interference. Improving the wake-up rate can quickly respond, which helps to improve the user experience. In addition, the application of the multi-modal fusion algorithm of target detection of images and voice signal processing to improve the accuracy and robustness of DOA makes the user experience better.
[0062] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of actions combined. However, those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention. In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0063] In some embodiments, the embodiments of the present invention provide a non-volatile computer-readable storage medium. One or more programs including execution instructions are stored in the storage medium. The execution instructions can be read and executed by an electronic device (including but not limited to a computer, a server, or a network device, etc.) to execute any one of the above voice interaction methods of the present invention.
[0064] In some embodiments, the embodiments of the present invention further provide a computer program product. The computer program product includes a computer program stored on a non-volatile computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer is made to execute any one of the above voice interaction methods.
[0065] In some embodiments, the embodiments of the present invention further provide an electronic device, which includes: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute a voice interaction method.
[0066] Figure 3 is a schematic hardware structure diagram of an electronic device for executing a voice interaction method provided in another embodiment of the present application, as Figure 3 shown, the device includes:
[0067] One or more processors 310 and a memory 320, Figure 3 Taking one processor 310 as an example.
[0068] The device for executing the voice interaction method may further include: an input device 330 and an output device 340.
[0069] The processor 310, the memory 320, the input device 330, and the output device 340 may be connected through a bus or other means, Figure 3 Taking connection through a bus as an example.
[0070] The memory 320, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as program instructions / modules corresponding to the voice interaction method in the embodiments of the present application. The processor 310 executes various functional applications and data processing of the server by running the non-volatile software programs, instructions, and modules stored in the memory 320, that is, implements the voice interaction method in the above method embodiments.
[0071] The memory 320 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the voice interaction device, etc. In addition, the memory 320 may include a high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the memory 320 may optionally include a memory remotely provided relative to the processor 310, and these remote memories can be connected to the voice interaction device through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0072] The input device 330 can receive input digital or character information, and generate signals related to user settings and function control of the voice interaction device. The output device 340 may include a display device such as a display screen.
[0073] The one or more modules are stored in the memory 320 and, when executed by the one or more processors 310, perform the voice interaction method in any of the above method embodiments.
[0074] The above product can execute the method provided in the embodiments of the present application, and has functional modules and beneficial effects corresponding to the execution of the method. For technical details not described in detail in this embodiment, reference may be made to the method provided in the embodiments of the present application.
[0075] The electronic devices in the embodiments of the present application exist in various forms, including but not limited to:
[0076] (1) Mobile communication devices: These devices are characterized by having mobile communication functions and mainly aim to provide voice and data communication. Such terminals include: smart phones, multimedia phones, functional phones, and low-end phones, etc.
[0077] (2) Ultra-mobile personal computer devices: These devices belong to the category of personal computers, have computing and processing functions, and generally also have the characteristic of mobile Internet access. Such terminals include: PDAs, MIDs, and UMPC devices, etc.
[0078] (3) Portable entertainment devices: These devices can display and play multimedia content. Such devices include: audio and video players, handheld game consoles, e-books, and smart toys and portable vehicle navigation devices.
[0079] (4) Other on-board electronic devices with data interaction functions, such as in-vehicle device installed on a vehicle.
[0080] The device embodiments described above are merely illustrative, where the units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0081] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a general hardware platform, and of course, can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the related technology, can be embodied in the form of a software product, which can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0082] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than limiting them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present application.
Claims
1. A voice interaction method, comprising: collecting an environmental audio file based on a microphone module having multiple microphone channels; each of the microphone channels is respectively configured with a corresponding microphone orientation; extracting a speaker audio from the environmental audio file based on a preset blind source separation algorithm; detecting whether the speaker audio meets a preset voice wake-up condition; when the speaker audio meets the voice wake-up condition, rolling back the speech frames before blind source separation, calculating the speech component energies of the speaker audio in each microphone channel, and determining the microphone orientation of the microphone channel corresponding to the maximum speech component energy as the speaker orientation; performing a voice interaction operation based on the speaker orientation; wherein, performing a voice interaction operation based on the speaker orientation includes: collecting an environmental image corresponding to the speaker orientation based on a camera; identifying whether there is target object information in the environmental image; when target object information is identified in the environmental image, triggering a voice interaction operation; when no target object information is identified in the environmental image, calibrating the speaker orientation based on the microphone orientations of each microphone channel, including: determining the microphone orientation of the microphone channel corresponding to the second largest speech component energy as the speaker orientation, and adjusting the camera to re-perform target recognition to locate the speaker position.
2. The method according to claim 1, wherein, performing a voice interaction operation based on the speaker orientation includes: adjusting the position of a display screen according to the speaker orientation; when the display screen is successfully adjusted to the speaker orientation, performing a voice interaction operation based on the display screen.
3. The method according to claim 1, wherein, before extracting the speaker audio from the environmental audio file based on a preset blind source separation algorithm, the method further includes: preprocessing the environmental audio file by using an echo cancellation algorithm.
4. The method according to any one of claims 1-3, wherein, the microphone channels in the microphone module are arranged in a ring.
5. An electronic device, which comprises: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the steps of the method according to any one of claims 1-4.
6. A storage medium, on which a computer program is stored, characterized in that, when the program is executed by a processor, it implements the steps of the method according to any one of claims 1-4.
Citation Information
Patent Citations
Intelligent interaction method and system based on circle microphone array
CN104936091A
Voice waking up method and voice recognition device in man-machine interaction
CN105912092A
Natural human-computer voice interaction method and system
CN107230476A
Conference speaker tracking method and device, computer equipment and storage medium
CN112040119A
Man-machine conversation method and device, robot, computer equipment and storage medium
CN112309395A