Method and electronic device for signal processing
By combining a microphone array and a camera, the direction of the target sound source is determined, and audio signal processing is performed using a speech enhancement model combined with lip video information. This solves the problem of low speech clarity in far-field sound pickup technology and achieves a significant improvement in speech recognition efficiency.
Patent Information
- Application Number
- CN202011065346.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-09-30
- Publication Date
- 2026-01-02
- Estimated Expiration
- 2040-09-30
AI Technical Summary
Existing far-field pickup technology suffers from low speech intelligibility in noisy and reverberant environments, leading to a decrease in speech recognition rate.
By combining a microphone array and a camera, the direction of the target sound source is determined, and an audio signal processing method is used in conjunction with lip video information to enhance the audio signal in the direction of the target sound source and suppress noise and interference from other directions.
It significantly improves speech recognition efficiency, reduces the impact of environmental noise and reverberation on speech recognition, and enhances the clarity of audio signals.
Smart Images

Figure CN114333831B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of acoustics, and more particularly, to a signal processing method and an electronic device. BACKGROUND
[0002] Currently, smart devices such as smart televisions, smart sound boxes, smart electric lamps, etc. can all perform far-field sound pickup. For example, a user says a command of "turn off the light" five meters away, and the smart device picks up the voice and recognizes the voice, and controls the electric lamp to perform the corresponding action of turning off the light.
[0003] A commonly used far-field sound pickup technology is to use a microphone array to pick up an audio signal, and to use a beamforming technology and an echo cancellation algorithm to suppress environmental noise and echo, so as to obtain a relatively clear audio signal. However, in an actual environment, there can be various noises and interferences, such as cooking and dishwashing noises in a kitchen, television program noises, interference noises from family members chatting, etc., and in some homes, the rooms are spacious or the walls are decorated with materials having a large sound reflection coefficient, resulting in a large reverberation and a blurred sound. All these unfavorable factors can greatly reduce the clarity of the sound picked up by the microphone array, thereby greatly reducing the voice recognition rate.
[0004] Therefore, it is necessary to provide a technology that can greatly improve the voice recognition efficiency. SUMMARY
[0005] Embodiments of the present application provide a signal processing method and an electronic device. The target sound source direction of a user who is performing voice interaction with the electronic device is determined based on an audio signal and a video obtained by a camera. Then, the audio signal picked up is subjected to voice enhancement processing based on the user's lip video in the target sound source direction obtained by the camera and a preset voice enhancement model, so that a relatively clear audio signal is obtained or recovered, and the voice recognition efficiency can be greatly improved.
[0006] In a first aspect, a signal processing method is provided. The method is applied to an electronic device, and the electronic device includes a microphone array and a camera. The method includes:
[0007] performing sound source positioning on a first audio signal obtained by the microphone array to obtain sound source direction information;
[0008] processing a first video obtained by the camera to obtain user direction information;
[0009] determining a target sound source direction according to the sound source direction information and the user direction information;
[0010] obtaining a user's lip video in the target sound source direction by the camera;
[0011] obtaining a second audio signal through the microphone array;
[0012] obtaining a third audio signal through a speech enhancement model according to the second audio signal and the user lip video, the speech enhancement model comprising a correspondence between pronunciation and lip shape.
[0013] The sound source direction information comprises at least one sound source direction, and the at least one sound source direction comprises a target sound source direction. The user direction information comprises some directions related to the user, and exemplarily comprises at least one type of direction related to the user. The target sound source direction is a direction in which a target user who is in voice interaction with the electronic device is located, i.e., a direction from which the target user emits sound.
[0014] The user lip video records a plurality of lip shapes in the process of the user speaking, and the lip shapes have a correspondence with pronunciation, i.e., one lip shape can correspond to one or more pronunciations. When the user does not speak, the lips are in a static state. The user lip video in the target sound source direction can also be understood as a lip video of the target user in practice.
[0015] The purpose of the speech enhancement model is to perform pickup enhancement processing on the audio signal, enhance the audio signal in the target sound source direction, and suppress or eliminate the audio signal generated by other directions including the speaker or background noise, so as to obtain or restore a relatively clear audio signal. The speech enhancement model of the embodiment of the present application fuses the audio and video information and integrates the correspondence between pronunciation and lip shape, i.e., one or more pronunciations can correspond to one lip shape.
[0016] Exemplarily, the camera is a rotatable camera, and after the target sound source direction is determined, the camera can be rotated to the target sound source direction to shoot the user lip video in the target sound source direction.
[0017] The signal processing method of the embodiment of the present application can greatly improve the estimation accuracy of the target sound source direction by obtaining the first video through the camera, combining the first audio signal obtained through the microphone array, determining the target sound source direction, avoiding the false sound source interference with the determination of the target sound source direction due to the strong reflected sound when the target sound source direction is determined only through the audio signal, and performing speech enhancement processing on the second audio signal obtained through the microphone array by using the user lip video in the target sound source direction obtained through the camera and the preset speech enhancement model. Since the speech enhancement model integrates the correspondence between pronunciation and lip shape, the relatively clean third audio signal can be restored by combining the user lip video and the speech enhancement model, and finally the speech recognition efficiency can be effectively improved.
[0018] In combination with the first aspect, in some implementations of the first aspect, the electronic device further comprises a directional microphone, and the method further comprises:
[0019] obtain a fourth audio signal in the target sound source direction through the directional microphone; and
[0020] obtain a third audio signal through a speech enhancement model according to the second audio signal and a user lip video in the target sound source direction, including:
[0021] obtain the third audio signal through the speech enhancement model according to the second audio signal, the fourth audio signal and the user lip video.
[0022] In some embodiments, the directional microphone can be fixed on the camera. In this way, after the target sound source direction is determined, the directional microphone is rotated in the process of rotating the camera, and finally rotated to the target sound source direction, the camera shoots the user lip video in the target sound source direction, and the directional microphone picks up the fourth audio signal in the target sound source direction.
[0023] The signal processing method of the embodiments of the present application, after the target sound source direction is determined, obtains a fourth audio signal in the target sound source direction through a directional microphone. Since the directional microphone has a certain inhibitory effect on reverberation, interference outside the target sound source direction, and echo of the display screen itself, and further inhibitory effect on the echo residue after echo cancellation, therefore, the embodiments of the present application use the fourth audio signal obtained in the target sound source direction by the directional microphone, in combination with the second audio signal obtained by the microphone array, and take the two audio signals as audio input, which can greatly improve the effect of sound pickup enhancement, so as to improve the speech recognition efficiency.
[0024] In combination with the first aspect, in some implementations of the first aspect, the user direction information includes at least one of the following types of directions:
[0025] The first type of direction includes a direction in which at least one lip in an active state is located;
[0026] The second type of direction includes a direction in which at least one user is located;
[0027] The third type of direction includes a direction in which at least one user is looking at the electronic device.
[0028] The signal processing method of the embodiments of the present application can effectively exclude, for the first type of direction determination target sound source direction, a scenario in which a person in the video is speaking, and can also exclude, for an electronic device with a display screen, a scenario in which the user is speaking is interfered; for the second type of direction determination target sound source direction, the user appearing in the first video detection frame can effectively exclude other non-user emitted interference signals, for example, interference signals emitted by a sound box; for the third type of direction determination target sound source direction, whether the user in the first video detection frame is looking at the electronic device, in general, especially for an electronic device with a display screen, if the user has an interaction intention with the electronic device, in most cases, the user will issue a voice instruction to the electronic device so that the electronic device can well receive the voice instruction, and the user can also quickly know whether the electronic device executes the instruction or obtains some feedback from the electronic device, for example, the user issues a voice instruction to inquire about the weather state, and the user needs to see the weather state displayed on the electronic device.
[0029] With reference to the first aspect, in some implementations of the first aspect, the sound source direction information includes at least one sound source direction, and
[0030] The determining, according to the sound source direction information and the user direction information, of the target sound source direction includes:
[0031] The at least one sound source direction and the at least one type of direction are combined to obtain at least one combined direction;
[0032] The target sound source direction is determined from the at least one direction.
[0033] The signal processing method of the embodiments of the present application can simplify the calculation by combining the at least one sound source direction and the at least one type of direction to determine the target sound source direction.
[0034] With reference to the first aspect, in some implementations of the first aspect, the determining, from the at least one direction, of the target sound source direction includes:
[0035] The target sound source direction is determined from the at least one direction according to at least one parameter;
[0036] The at least one parameter includes:
[0037] a sum of frequencies at which each direction in the at least one direction is detected in the sound source direction and the at least one type of direction;
[0038] whether the electronic device has successfully performed voice interaction with the user in a preset time period and a preset angle range corresponding to each direction, the preset time period being a time period between a current time and a historical time;
[0039] an included angle between each direction and a direction perpendicular to a display screen of the electronic device.
[0040] For the parameter of "a sum of frequencies in which each direction is detected in the sound source direction and the at least one type of direction", it can be understood that the more the sum of frequencies in which a direction is detected, the more likely the direction is the target sound source direction. Ideally, the direction is substantially the target sound source direction.
[0041] For the parameter of "whether the electronic device has successfully performed voice interaction with the user in a preset time period and a preset angle range corresponding to each direction", the angle of the preset angle range corresponding to each direction can include not only an angle corresponding to the direction, but also an angle near the angle. The parameter can be understood as whether the electronic device has successfully performed voice interaction with the user in a preset time period near an angle corresponding to a certain direction.
[0042] For the parameter of "an included angle between each direction and a direction perpendicular to a display screen of the electronic device", it is more suitable for an electronic device with a display screen. The parameter can be understood as whether the user is near a certain specific direction defined when a preset scene is used for the electronic device.
[0043] The signal processing method of the embodiments of the present application sets different parameters in combination with specific scenes, determines the target sound source direction from the at least one direction through the at least one parameter, and can further effectively improve the estimation accuracy of the target sound source direction for a specific electronic device (for example, a smart television), so as to improve the voice recognition efficiency.
[0044] In combination with the first aspect, in some implementations of the first aspect, the determining the target sound source direction from the at least one direction according to the at least one parameter comprises:
[0045] determining a confidence degree of each direction according to the at least one parameter;
[0046] determining a direction corresponding to a maximum value of the confidence degrees in the at least one direction as the target sound source direction.
[0047] In combination with the first aspect, in some implementations of the first aspect, the obtaining the second audio signal through the microphone array comprises:
[0048] obtaining the second audio signal in the target sound source direction based on a beam forming technology through the microphone array.
[0049] The method for signal processing of the embodiment of the application obtains the second audio signal in the direction of the target sound source through the beamforming technology, enhances the sound pickup effect, and effectively reduces the influence of the interference signals in other directions on the speech recognition efficiency.
[0050] With reference to the first aspect, in some implementations of the first aspect, the first audio signal is a wake-up signal.
[0051] In a second aspect, an electronic device is provided, which includes a microphone array, a camera, and a processor, the processor being configured to:
[0052] perform sound source positioning on a first audio signal obtained through the microphone array to obtain sound source direction information;
[0053] perform processing on a first video obtained through the camera to obtain user direction information;
[0054] determine a target sound source direction according to the sound source direction information and the user direction information;
[0055] obtain a user lip video in the target sound source direction through the camera;
[0056] obtain a second audio signal through the microphone array;
[0057] obtain a third audio signal through a speech enhancement model according to the second audio signal and the user lip video, the speech enhancement model including a corresponding relationship between pronunciation and lip shape.
[0058] With reference to the second aspect, in some implementations of the first aspect, the electronic device further includes a directional microphone, and the processor is further configured to:
[0059] obtain a fourth audio signal in the target sound source direction through the directional microphone; and
[0060] the processor is specifically configured to:
[0061] obtain the third audio signal through the speech enhancement module according to the second audio signal, the fourth audio signal, and the user lip video.
[0062] With reference to the second aspect, in some implementations of the first aspect, the directional microphone is fixedly connected with the camera.
[0063] With reference to the second aspect, in some implementations of the first aspect, the user direction information includes at least one type of direction as follows:
[0064] a first type of direction, the first type of direction including a direction in which at least one lip is active;
[0065] a second type of direction, the second type of direction including a direction in which at least one user is located;
[0066] a third type of direction, the third type of direction including a direction in which at least one user is gazing at the electronic device.
[0067] With reference to the second aspect, in some implementations of the first aspect, the sound source direction information includes at least one sound source direction, and
[0068] The processor is specifically configured to:
[0069] merge the at least one sound source direction and the at least one type of direction to obtain at least one merged direction;
[0070] determine the target sound source direction from the at least one direction.
[0071] With reference to the second aspect, in some implementations of the first aspect, the processor is specifically configured to:
[0072] determine the target sound source direction from the at least one direction according to at least one parameter;
[0073] The at least one parameter includes:
[0074] a sum of frequencies at which each direction in the at least one direction is detected in the sound source direction and the at least one type of direction;
[0075] whether the electronic device and a user have successfully performed voice interaction within a preset time period and a preset angle range corresponding to the each direction, the preset time period being a time period between a current time and a historical time;
[0076] an included angle between the each direction and a direction perpendicular to a display screen of the electronic device.
[0077] With reference to the second aspect, in some implementations of the first aspect, the processor is specifically configured to:
[0078] determine a confidence level of the each direction according to the at least one parameter;
[0079] determine a direction corresponding to a maximum value of the confidence levels in the at least one direction as the target sound source direction.
[0080] With reference to the second aspect, in some implementations of the first aspect, the processor is specifically configured to:
[0081] The second audio signal is obtained in the direction of the target sound source based on a beamforming technique through the microphone array.
[0082] With reference to the second aspect, in some implementations of the first aspect, the first audio signal is a wake-up signal.
[0083] With reference to the second aspect, in some implementations of the first aspect, the electronic device is a smart television.
[0084] In a third aspect, a chip is provided, including a processor configured to invoke and run instructions stored in a memory, so that an electronic device installed with the chip performs the method of the first aspect.
[0085] In a fourth aspect, a computer storage medium is provided, including a processor coupled with a memory, the memory configured to store programs or instructions, when the programs or instructions are executed by the processor, causing the apparatus to perform the method of the first aspect.
[0086] In a fifth aspect, the present application provides a computer program product, when the computer program product is run on an electronic device, causing the electronic device to perform the method of any one of the first aspect.
[0087] It can be understood that the electronic device, chip, computer storage medium and computer program product provided above are all used to perform the corresponding method provided above, and thus the beneficial effects achieved thereby can refer to the beneficial effects of the corresponding method provided above, which will not be described here again. BRIEF DESCRIPTION OF DRAWINGS
[0088] Figure 1 is a schematic structural diagram of an electronic device provided by an embodiment of the present application.
[0089] Figure 2 is a schematic structural diagram of an electronic device provided by another embodiment of the present application.
[0090] Figure 3 is a schematic scene diagram of a camera shooting a video provided by an embodiment of the present application.
[0091] Figure 4 is an exemplary block diagram of an electronic device provided by an embodiment of the present application.
[0092] Figure 5 is a schematic scene diagram provided by an embodiment of the present application.
[0093] Figure 6 is a schematic flowchart of a signal processing method provided by an embodiment of the present application.
[0094] Figure 7 is a schematic flowchart of a method of signal processing provided by another embodiment of the present application.
[0095] Figure 8 is a schematic flowchart of a method of determining a target sound source direction by an electronic device provided by another embodiment of the present application.
[0096] Figure 9 is a schematic flowchart of a method of signal processing provided by another embodiment of the present application. DETAILED DESCRIPTION
[0097] The technical solutions in the present application will be described below with reference to the accompanying drawings.
[0098] The method of signal processing provided by the embodiments of the present application determines the direction (denoted as the target sound source direction) in which the user (denoted as the target user) who is performing voice interaction with the electronic device is located based on an audio signal and a video obtained based on a camera, and then performs voice enhancement processing on the picked-up audio signal based on the user's lip video in the direction obtained based on the camera and a preset voice enhancement model, to obtain or restore a clearer audio signal, which can greatly improve the voice recognition efficiency.
[0099] For ease of description, the embodiments of the present application define some terms, which will be introduced as follows.
[0100] The target user is a person who is performing voice interaction with the electronic device, and the target user is issuing a voice instruction to the electronic device to perform a certain action. The target user can also be understood as the actual speaker.
[0101] The target sound source direction is the direction in which the target user is located, i.e., the direction of the sound source of the target user. Due to the influence of various interference signals in the environment, the electronic device can pick up audio signals of multiple sound source directions, so the direction in which the target user is located is defined as the target sound source direction.
[0102] The user's lip video records the lip shapes (denoted as lip shapes) of the user in the process of speaking. When the user speaks, the lips will make various lip shape actions, and the lip video can record multiple lip shapes. The lip shape has a corresponding relationship with the pronunciation, i.e., one lip shape can correspond to one or more pronunciations, for example, "w" (the Chinese character for "nest"), "wo" (the Chinese character for "I") and "shou" (the Chinese character for "hold") represent three different pronunciations, but correspond to one lip shape. When the user does not speak, the lips are in a static state. In the embodiments of the present application, the user's lip video in the target sound source direction can also be understood as the lip video of the target user.
[0103] The voice enhancement model aims to perform pickup enhancement processing on the audio signal, enhance the audio signal in the direction of the target sound source, suppress or eliminate the audio signal generated by other directions including the speaker or background noise, etc., to obtain or restore a clearer audio signal. The voice enhancement model of the embodiments of the present application fuses the audio and video information, integrates the correspondence between pronunciation and lip shape, and one or more pronunciations can correspond to one lip shape. In the embodiments of the present application, the audio signal and the user's lip video are taken as the input of the voice enhancement model, and the voice enhancement model can perform voice enhancement processing on the audio signal based on the correspondence between pronunciation and lip shape and the input user's lip video, to obtain a clearer audio signal for voice recognition.
[0104] Exemplarily, the voice enhancement module can perform noise reduction processing, echo residual elimination processing, dereverberation processing, etc. on the audio signal.
[0105] The signal processing method of the embodiments of the present application can be applied to any electronic device capable of recognizing voice. In an example, the electronic device can be a smart television (also known as a smart screen) or the like voice control device. In another example, the electronic device can be a mobile phone, a computer or the like voice call device.
[0106] Hereinafter, the voice enhancement model will be described in combination with Figures 1 to 3 The electronic device of the embodiments of the present application will be described taking a smart television as an example.
[0107] Reference Figure 1 The electronic device 10 includes a housing 110, a display screen 120, a microphone array 130, and a camera 140. The display screen 120, the microphone array 130, and the camera 140 are installed in the housing 110.
[0108] The display screen 120 is used to display images, videos, etc. The display screen 120 includes a display panel. The display panel can adopt a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light emitting diode (AMOLED), a flexible light-emitting diode (FLED), a Miniled, a MicroLed, a Micro-oLed, a quantum dot light emitting diode (QLED), etc.
[0109] The microphone array 130 is used to pick up audio signals, and includes a plurality of microphones that can pick up audio signals in multiple directions. Exemplarily, the microphones in the microphone array 130 can be omnidirectional microphones, directional microphones, or a combination of omnidirectional microphones and directional microphones, and the present application does not make any limitation.
[0110] The microphone array 130 can be arranged at any position of the housing 110, and the present application does not make any limitation on the embodiments.
[0111] In an example, as shown in FIG. 1, the microphone array 130 is arranged in the housing 110 and located at a region on one side of the display screen 120, and the sound outlet holes of the microphone array 130 are arranged on the front surface of the housing 110, and the orientation of the sound outlet holes is the same as the orientation of the display screen 120. The front surface of the housing 110 can be understood as the surface with the same orientation as the display screen 120, or the front surface of the housing 110 can be understood as the surface that faces the user in the normal use condition. Figure 1 The microphone array 130 can be arranged in the housing 110 and located at a region on any side of the display screen 120. Assuming that the microphone array 130 shown in FIG. 2 is arranged in the housing 110 and located at a region on the top side of the display screen 120, the microphone array 130 can also be arranged in the housing 110 and located at a region on the other side (for example, the left side, the right side, or the bottom side) of the display screen 120. Figure 1
[0112] In another example, the microphone array 130 can be arranged in the housing 110 and located at a region on the top side of the display screen 120, and the sound outlet holes of the microphone array 130 are arranged on the top surface of the housing 110 (not shown in the figure), and the top surface of the housing 110 is connected to the front surface of the housing 110, and the orientation of the sound outlet holes is perpendicular to the orientation of the display screen 120.
[0113] In another example, the microphone array 130 can also be arranged on the back side of the display screen 120, and the sound outlet holes of the microphone array 130 are arranged on the display screen 120 (not shown in the figure).
[0114] In another example, the microphone array 130 can also be arranged on the back side of the display screen 120, and the sound outlet holes of the microphone array 130 are arranged on the front surface of the housing 110.
[0115] The microphone array 130 can be arranged as shown in FIG. 3. Figure 1 The linear structure arrangement shown can also be other structure arrangements, and embodiments of the present application do not make any limitation. For example, the microphone array 130 can be arranged in a circular structure or a rectangular structure, etc.
[0116] The camera 140 is used to capture still images or videos. An object generates an optical image through a lens and projects the optical image to a photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, and then transmits the electrical signal to an ISP to convert into a digital image signal. The ISP outputs the digital image signal to a DSP for processing. The DSP converts the digital image signal into an image signal in a standard RGB, YUV, etc. format. In some embodiments, the electronic device 10 can include one or N cameras 130, where N is a positive integer greater than 1.
[0117] In embodiments of the present application, the camera 140 can rotate within a preset angle range to capture a video within a certain angle range, which can be used to determine the direction of a target sound source; and after the electronic device 10 determines the direction of the target sound source, the camera 140 can rotate to the direction of the target sound source so that the camera 140 is directly opposite to the direction of the target sound source, and the target user is as possible as possible displayed in the center of the screen, so as to better capture a video in the direction of the target sound source, and the obtained user lip video can be used to process an audio signal of an input speech enhancement model to output a clearer audio signal for speech recognition.
[0118] In an example, referring to Figure 1 , the camera 140 is arranged on the top surface of the housing 110 and protrudes from the top surface to better realize the rotation of the camera 140. In embodiments in which the microphone array 130 is located in the area on the top side of the display screen 120, the camera 140 can be located above the microphone array 130.
[0119] In another example, referring to Figure 2 , the camera 140 can be arranged on the front surface of the housing 110 and located in the area on the top side of the display screen 120.
[0120] The camera 140 can rotate within a preset angle range, which can be any range of angles. Referring to Figure 3In the embodiment where the electronic device 10 is a smart television, the camera 140 can rotate within an angle range less than or equal to 180°, and the angle range can be 120°, for example. The camera 140 can rotate within an angle range of 120° in front of the display screen 120, and in combination with the field of view of the camera 140, the camera 140 can capture all pictures within an angle range of 180° in front of the smart television.
[0121] In some embodiments, referring to Figure 1 and Figure 2 , the electronic device 10 further includes a directional microphone 150, which can rotate to pick up audio signals in a specific direction. After the electronic device 10 determines the target sound source direction, the directional microphone 150 can rotate to the target sound source direction to pick up audio signals in the target sound source direction.
[0122] Since the directional microphone 150 can pick up audio signals in the target sound source direction without distortion, it can have certain inhibitory effect on interference and reverberation, and the directional microphone 150 picks up audio signals in front, which can also have good inhibitory effect on echo. Therefore, in the embodiments of the present application, the audio signals obtained by the directional microphone 150 and the audio signals obtained by the microphone array 130 can be used as audio inputs of a speech enhancement model, and clearer audio signals can be obtained or recovered.
[0123] In combination with the embodiment where the camera 140 can rotate to the target sound source direction to capture video in the target sound source direction, in an example, referring to Figure 1 and Figure 2 , the directional microphone 150 can be arranged on the camera 140, and the directional microphone 150 is fixedly connected to the camera 140, for example. When the camera 140 rotates to the target sound source direction, the directional microphone 150 also rotates to the target sound source direction, which is simple and convenient. The electronic device 10 further includes a processor (not shown in the figure), and the display screen 120, the microphone array 130, the camera 140, and the directional microphone 150 are all connected to the processor, for inputting signals collected by the components to the processor for further processing. The processor runs instructions to implement the signal processing method of the embodiments of the present application to obtain clearer audio signals sent by the user, and after speech recognition of the audio signals, the corresponding components can be controlled to execute instructions corresponding to the audio signals.
[0124] It should be understood that the structure of the electronic device 10 described above by taking a smart television as an example is only illustrative, and the electronic device 10 can have more or fewer components.
[0125] In some embodiments, the electronic device 10 can include the microphone array 130, the camera 140, and optionally, the directional microphone 150, but the electronic device 10 can not include the display screen 120.
[0126] In some other embodiments, the electronic device 10 can include the directional microphone 150 and the camera 140, but the electronic device 10 does not include the microphone array 130. In this embodiment, the audio signal picked up by the directional microphone 150 and the video taken by the camera 140 can be used to determine the target sound source direction, and the video taken by the camera 140 in the target sound source direction and the audio signal picked up by the directional microphone 150 in the target sound source direction can be used to recover a clearer audio signal through the video in the target sound source direction and the speech enhancement model. For example, the directional microphone 150 can rotate to collect the audio signal all the time before the target sound source direction is determined.
[0127] In some other embodiments, the electronic device 10 can include more components in addition to the microphone array 130, the camera 140, and the directional microphone 150. For example, the electronic device 10 can be a mobile phone or a computer.
[0128] Figure 4 FIG. 4 is an exemplary block diagram of the electronic device 10 provided by the embodiments of the present application. The electronic device 10 can include the display screen 120, the microphone array 130, the directional microphone 150, and the camera 140 shown in FIG. 3. For example, the electronic device 10 can further include one or more of the following components: a processor 160, a wireless communication module 171, an audio module 172, a speaker 173, a touch sensor 174, a key 175, and an internal memory 176.
[0129] The wireless communication module 171 can provide a solution for wireless communication, including wireless local area networks (WLAN) (e.g., wireless fidelity (Wi-Fi) network), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared (IR) technology, etc. The wireless communication module 171 can be one or more devices that integrate at least one communication processing module. The wireless communication module 171 receives electromagnetic waves via an antenna, frequency-modulates and filters the electromagnetic wave signals, and transmits the processed signals to the processor. The wireless communication module 171 can also receive signals to be transmitted from the processor, frequency-modulate them, amplify them, and radiate them as electromagnetic waves via the antenna.
[0130] The audio module 172 is configured to convert digital audio information into analog audio signals for output, and to convert analog audio input into digital audio signals. The audio module 172 can also be configured to encode and decode audio signals. In some embodiments, the audio module 172 can be disposed in the processor 160, or some of the functional modules of the audio module 172 can be disposed in the processor 160.
[0131] The speaker 173, also referred to as a "loudspeaker", is configured to convert audio electrical signals into sound signals. The electronic device 10 can listen to music or sound in a video through the speaker 173. In embodiments in which the electronic device 10 is a mobile phone, the speaker 173 can also be used for listening to a hands-free call.
[0132] The touch sensor 174, also referred to as a "touch panel", can be disposed on the display screen 120. The touch sensor 174 and the display screen 120 together form a touch screen, also referred to as a "touch screen". The touch sensor 174 is configured to detect a touch operation applied to or near the touch sensor 174. The touch sensor 174 can transmit the detected touch operation to the processor 160 to determine the type of touch event. The display screen 120 can be used to provide visual output related to the touch operation. In other embodiments, the touch sensor 174 can be disposed on the surface of the electronic device 10, rather than on the display screen 120.
[0133] The keys 175 include a power key, a volume key, etc. The keys 175 can be mechanical keys. Alternatively, the keys 175 can be touch keys. The electronic device 10 can receive key 175 inputs and generate key signal inputs related to user settings and function control of the electronic device 10.
[0134] The internal memory 176 is configured to store computer-executable program codes including instructions. The processor 160 performs various function applications and data processing of the electronic device 10 by running the instructions stored in the internal memory. The internal memory 176 can include a program storage area and a data storage area. The program storage area can store an operating system, at least one application program (such as a sound playing function, an image playing function, etc.) required by a function, etc. The data storage area can store data (such as audio data, a phone book, etc.) created during use of the electronic device 10, etc. In addition, the internal memory 176 can include a high-speed random access memory, and can further include a non-volatile memory such as at least one magnetic disk storage device, a flash memory device, a universal flash storage (UFS), etc.
[0135] Figure 5 is a schematic scenario provided by an embodiment of the present application. Taking a smart television as an example, the target user is watching television and says to the smart television “Xiao Yi, Xiao Yi, I want to watch a variety show”, the smart television receives and recognizes the instruction to tune the smart television to a variety show. Figure 5
[0136] In the embodiments of the present application, in order to facilitate description, a certain direction can be expressed by an angle, and a reference direction can be defined, and a certain direction can be expressed by an angle between the certain direction and the reference direction. It should be understood that the reference direction can be arbitrary, and the embodiments of the present application do not make any limitation.
[0137] For example, the direction extending along the left side (the direction indicated by the arrow of the 0° direction in the figure) in the length direction (for example, the x direction) of the electronic device 10 can be recorded as a reference direction, and the angle corresponding to the reference direction is 0°. The target user is facing the electronic device 10, and the angle between the target sound source direction of the target user and the reference direction is 90°. Figure 5
[0138] In the following, the method of signal processing of the embodiments of the present application is described, which can be performed by the electronic device 10. The electronic device 10 includes a microphone array 130, a camera 140 and a processing unit 160. The processing unit 160 can include a target sound source direction determination module 161 and a voice enhancement module 162, and the electronic device 10 can further include a directional microphone 150. Figures 6 to 9
[0139] is a schematic flowchart of the method of signal processing provided by an embodiment of the present application. Referring to Figure 6 Figure 6 The general process of the embodiments of the present application is as follows:
[0140] S210, the target user starts to issue a voice instruction to the electronic device, and the microphone array 130 picks up a first audio signal.
[0141] S220, the camera 140 shoots a video to obtain a first video.
[0142] S230, the processing unit 160 performs sound source positioning on the first audio signal to obtain sound source direction information including at least one sound source direction, and the processing unit 160 processes the first video to obtain user direction information. This step can be performed by the target sound source direction determination module 161 in the processing unit 160.
[0143] S240, the processing unit 160 determines the target sound source direction in which the target user is located according to the sound source direction information and the user direction information. This step can be performed by the target sound source direction determination module 161 in the processing unit 160.
[0144] S250, the processing unit 160 controls the camera 140 to rotate to the target sound source direction, and the camera 140 shoots a video in the target sound source direction to obtain a user lip video in the target sound source direction.
[0145] S260, the microphone array 130 continues to pick up a second audio signal, which is actually required for voice recognition.
[0146] S270, in the embodiment in which the electronic device includes the directional microphone 150, the processing unit 160 can also control the directional microphone 150 to rotate to the target sound source direction, and the directional microphone 150 picks up a fourth audio signal in the target sound source direction.
[0147] In the embodiment in which the directional microphone 150 is arranged on the camera 140, the processing unit 160 controls the camera 140 and the directional microphone 150 to rotate to the target sound source direction together.
[0148] S280, taking the second audio signal and the user lip video in the target sound source direction as inputs, the processing unit 160 performs voice enhancement processing on the second audio signal through a voice enhancement model to obtain a third audio signal that is clearer after enhancement. This step can be performed by the voice enhancement module 162 in the processing unit 160.
[0149] In the embodiment in which the electronic device includes the directional microphone 150, in S280, taking the second audio signal, the fourth audio signal, and the user lip video in the target sound source direction as inputs, the processing unit 160 performs voice enhancement processing on the second audio signal and the fourth audio signal through the voice enhancement model to obtain the third audio signal.
[0150] It should be understood that the size of the sequence number of each process described above does not mean the order of execution in various embodiments of the method 200 of the present application, and the execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application. For example, step S210 and step S220 can be executed simultaneously, step S250 and step S260 can be executed simultaneously, and step S250, step S260 and step S270 can be executed simultaneously. For another example, step S250 can be executed before step S260, or can be executed after step S260.
[0151] The signal processing method of the embodiments of the present application can obtain a first video through a camera, combine a first audio signal obtained by a microphone array, determine a target sound source direction of a target user who is performing voice interaction with the electronic device, and can greatly improve the estimation accuracy of the target sound source direction, avoid false sound source interference with the determination of the target sound source direction due to strong reflected sound when only the audio signal is used to determine the target sound source direction, and perform voice enhancement processing on a second audio signal obtained by the microphone array through a camera-acquired user lip video in the target sound source direction and a preset voice enhancement model. Since the voice enhancement model integrates the corresponding relationship between pronunciation and lip shape, the user lip video and the voice enhancement model can be combined to restore a relatively clean third audio signal, and finally the voice recognition efficiency can be effectively improved.
[0152] In addition, the directional microphone has a certain inhibitory effect on reverberation, interference outside the target sound source direction, and echo of the display screen itself, and has a further inhibitory effect on residual echo after echo cancellation. The embodiments of the present application use the fourth audio signal picked up by the directional microphone in the target sound source direction, combine the second audio signal obtained by the microphone array, and use the two audio signals as audio input, which can greatly improve the effect of sound pickup enhancement and improve the voice recognition efficiency.
[0153] Figure 7 The signal processing method 300 provided by another embodiment of the present application is a schematic flow chart of the method, which can be executed by the processing unit 160 of the electronic device 10.
[0154] In step S310, sound source positioning is performed on the first audio signal obtained by the microphone array 130 to obtain sound source direction information, and the sound source direction information includes at least one sound source direction. The at least one sound source direction includes a target sound source direction.
[0155] The user issues a voice instruction to the electronic device, and the microphone array 130 picks up an audio signal. This step can be used for sound source positioning, and the first audio signal can be a small part of the content of the voice instruction issued by the user, which does not substantially affect the subsequent content for voice recognition.
[0156] Exemplarily, the first audio signal can be a wake-up signal. For example, the voice instruction issued by the user is “Xiao Yi, I want to watch a variety show”, and then the first audio signal can be one or more words of “Xiao Yi” or multiple “Xiao Yis”. The one or more words of “Xiao Yi” or multiple “Xiao Yis” can be understood as a wake-up signal. The electronic device detects “Xiao Yi”, can determine that the user can need the electronic device to execute the voice instruction, the microphone array 130 performs sound source positioning, and continuously picks up subsequent audio signals.
[0157] Of course, in a voice instruction without a wake-up signal, the first audio signal can be the first few words in the voice instruction. Generally, the microphone array 130 can detect the audio signal as long as the user issues one or two words. For example, the voice instruction issued by the user is “I want to watch a variety show”, and the first audio signal can be “I”.
[0158] The microphone array 130 performs sound source positioning on the first audio signal, aiming to determine a target sound source direction of a target user, that is, a sound source direction actually issuing the voice instruction. However, the microphone array 130 can pick up audio signals in various directions, and due to the influence of various interference sounds in the environment, the finally determined sound source direction can not be accurate, and at least one sound source direction can be obtained, which includes the target sound source direction and can also include a direction of an interference signal. For example, the target user issues a voice instruction to the electronic device, a sound box is playing music, and another user (referred to as an interference user) is speaking, and it is assumed that the above three sounds can be picked up by the microphone array 130. Then, the microphone array 130 can determine three or two or one sound source directions, which is not accurate, so the target sound source direction needs to be further determined in combination with a video.
[0159] Exemplarily, the sound source positioning technology of the embodiment of the present application can be a controllable beamforming technology based on maximum output power, a high-resolution spectrum estimation technology, or a sound source positioning technology based on time-delay estimation (TDE), which is not limited in the embodiment of the present application.
[0160] As described above, a certain direction of the embodiment of the present application can be expressed by an angle, and here, the sound source direction can be expressed by an angle θ, and the angle of any direction in the at least one sound source direction can be denoted as θi. ii = 1, 2, …, I, I is the number of sound source directions included in at least one sound source direction.
[0161] In step S320, the first video obtained by the camera 140 is processed to obtain user direction information. The user direction information includes some directions related to the user. For example, the user direction information includes at least one type of direction related to the user.
[0162] In some embodiments, the user issues a voice instruction to the electronic device, the camera 140 can capture a video, and the electronic device can determine the target sound source direction based on the obtained first video.
[0163] For example, the voice instruction issued by the user can be used as a trigger condition for the camera 140 to capture a video. The electronic device detects the voice instruction issued by the user and controls the camera 140 to start capturing a video. Since the camera 140 of the embodiment of the present application can rotate, in some examples, the camera 140 can capture a video while rotating to obtain a picture in a larger angle range.
[0164] In other embodiments, the camera 140 can capture a video during the operation of the electronic device. After the electronic device receives a voice instruction issued by the user, a video of a period of time is used as the first video for determining the target sound source direction.
[0165] The electronic device processes the first video captured by the camera 140, detects the content related to the user in the first video, and obtains user direction information including at least one type of direction related to the user. In this way, based on the sound source direction information and in combination with the user direction information, other non-user generated interference signals can be effectively excluded, for example, interference signals generated by a sound box can be excluded.
[0166] It should be understood that the user involved in the user direction information not only includes the target user who is in voice interaction with the electronic device, but also includes other users. As long as the user is detected in the first video, it can be included, but relative to the target user, the other user can be understood as an interference user.
[0167] The user direction information includes at least one type of direction related to the user, and each type of direction includes at least one direction.
[0168] In some embodiments, the at least one type of direction includes at least one of the following:
[0169] The first type of direction includes at least one direction in which a lip in an active state is located;
[0170] The second type of direction includes at least one direction in which the user is located;
[0171] The third type of direction includes at least one direction in which a user who is gazing at the electronic device is located.
[0172] For the first type of direction, by detecting whether the lips of a person in the first video are moving, that is, by detecting whether a person is speaking, a scenario in which a person in the video is speaking can be effectively excluded. For an electronic device 10 with a display screen, a scenario in which a user is speaking can also be excluded to some extent. For example, a target user is watching television and giving voice instructions to the television, and user 1 is also speaking, but is looking down and doing housework and is not speaking to the television. In this case, the lips of user 1 are not detected as moving in the first video, and only the lips of the target user are detected as moving. Therefore, user 1 is an interfering user and can be effectively excluded.
[0173] If there are multiple users (including the target user) in the environment who are speaking within the field of view of the camera 140, the lips of multiple users can be detected as moving, and multiple directions in which the lips of multiple users are moving can be obtained. In a normal case, the first type of direction includes the direction of the target sound source.
[0174] Here, the first type of direction is represented by an angle γ, and the angle corresponding to any direction in the first type of direction can be denoted as γ l , l = 1, 2,..., L, and L is the number of directions included in the first type of direction.
[0175] For the second type of direction, by detecting the users appearing in the first video, interfering signals emitted by other non-users can be effectively excluded, for example, interfering signals emitted by a sound box can be excluded.
[0176] If there are multiple users (including the target user) in the environment, multiple users can be detected in the first video, and multiple directions in which the users are located can be obtained. It should be understood that, in a normal case, the second type of direction includes the direction of the target sound source.
[0177] For ease of distinction, the second type of direction can be represented by an angle α, and the angle corresponding to any direction in the second type of direction can be denoted as α j , j = 1, 2,..., J, and J is the number of directions included in the second type of direction.
[0178] For the third type of direction, by detecting whether a user is gazing at the electronic device in the first video, in general, especially for electronic devices with a display, if the user has an intention to interact with the electronic device, the user will most likely issue a voice instruction to the electronic device so that the electronic device can better receive the voice instruction, and the user can also quickly know whether the electronic device executes the instruction or obtains some feedback from the electronic device. For example, the user issues a voice instruction to inquire about the weather state, and the user needs to see the weather displayed on the electronic device. Therefore, by detecting the user gazing at the electronic device, the scenario in which the user is speaking can be effectively excluded. For example, the target user is watching TV and issuing a voice instruction to the TV, and user 1 is speaking to the target user but not gazing at the TV. In this case, the first video can not detect that user 1 is gazing at the electronic device, and can only detect that the target user is gazing at the electronic device. Therefore, user 1 is an interfering user and can be effectively excluded.
[0179] If there are multiple users (including the target user) in the environment, multiple users gazing at the electronic device can be detected in the first video, and multiple directions of the users gazing at the electronic device can be obtained. In a normal case, the third type of direction includes the target sound source direction.
[0180] For convenience of distinction, the third type of direction can be represented by an angle β, and the angle corresponding to any direction in the third type of direction can be denoted as β k. k k = 1, 2,..., K, and K is the number of directions included in the third type of direction.
[0181] It should be understood that the user direction information can include one type, two types, or three types of the above-mentioned three types of directions, and the embodiments of the present application do not make any limitation. Of course, the more types of directions included in the user direction information, the more conducive to improving the accuracy of determining the target sound source direction.
[0182] It should also be understood that in addition to the above-mentioned three types of directions, the user direction information can also include other directions related to the user, and the embodiments of the present application do not make any limitation. For example, the user direction information can include other directions related to the user's behavior.
[0183] In step S330, the target sound source direction is determined according to the sound source direction information and the user direction information.
[0184] The target sound source direction is the direction of the target user who is in voice interaction with the electronic device.
[0185] It can be understood that the sound source direction in the sound source direction information can be regarded as a type of direction, and at least one type of direction related to the user is combined to determine the target sound source direction.
[0186] Figure 8 is a schematic flowchart of a method 230 for determining a target sound source direction of an electronic device according to another embodiment of the present application.
[0187] In some embodiments, with reference to Figure 8 , the electronic device can determine the target sound source direction in the following manner:
[0188] In step S331, at least one sound source direction in the sound source direction information and at least one type of direction in the user direction information are combined to obtain at least one combined direction.
[0189] In step S332, the target sound source direction is determined from the at least one combined direction.
[0190] For ease of description, in the following, the sound source direction and the above-mentioned three types of directions are taken as examples, and the manner of obtaining the at least one combined direction is first described.
[0191] In the combining process, in order to simplify the calculation, if the deviation between the angles corresponding to multiple directions is less than a threshold value, a direction can be determined based on the multiple directions, and logically the multiple directions can be considered as the same direction. The finally determined direction can be any one of the multiple directions, or an average value of the multiple directions, and the present embodiment does not make any limitation. The threshold value can be reasonably designed based on the actual application scenario. For example, the threshold value can be 5°.
[0192] Suppose that the sound source direction information includes four sound source directions, and the corresponding angles are 30°, 60°, 95° and 120°, respectively. The first type of direction includes one direction, and the corresponding angle is 93°. The second type of direction includes two directions, and the corresponding angles are 63° and 95°, respectively. The third type of direction includes one direction, and the corresponding angle is 95°.
[0193] List all the angles corresponding to the directions in ascending order: 30°, 60°, 63°, 93°, 95°, 95°, 95°, 120°. 60° is close to or the same as 63°, and 93° is close to or the same as 95°. For example, in the manner of taking an average value of the two directions, the combined angles are: 30°, 61.5°, 94.5° and 120°. That is, the fourth type of direction obtained by combining includes four directions, and the target sound source direction is one of the four directions. In fact, the direction corresponding to 94.5° is the target sound source direction, and the target user is basically facing the electronic device and interacting with the electronic device through voice.
[0194] The electronic device determines the target sound source direction from the at least one combined direction after obtaining the at least one combined direction.
[0195] In the embodiments of the present application, some parameters can be set based on the specific scene of the far-hang pickup, and the target sound source direction is determined based on the parameters.
[0196] In some embodiments, in step S332, the electronic device can determine the target sound source direction from the at least one direction according to at least one parameter, wherein the at least one parameter includes:
[0197] The sum of the frequencies at which each direction is detected in the sound source direction and at least one type of direction;
[0198] Whether the electronic device has successfully performed voice interaction with the user within a preset time period and a preset angle range corresponding to each direction, the preset time period being a time period between the current time and a historical time, and the preset angle range including an angle corresponding to each direction;
[0199] The included angle between each direction and a direction perpendicular to the display screen of the electronic device.
[0200] Taking the at least one type of direction as an example, the above four sound source directions and the angles corresponding to the three types of directions are taken as examples, and each parameter is described.
[0201] The four sound source directions correspond to angles of 30°, 60°, 95°, and 120°, the first type of direction includes one direction corresponding to an angle of 93°, the second type of direction includes two directions corresponding to angles of 63° and 95°, the third type of direction includes one direction corresponding to an angle of 95°, and the four directions obtained by merging the processing correspond to angles of 30°, 61.5°, 94.5°, and 120°.
[0202] The first parameter: the sum of the frequencies at which each direction is detected in the sound source direction and at least one type of direction.
[0203] The frequencies at which 30° is detected in the sound source direction, the first type of direction, the second type of direction, and the third type of direction are 1, 0, 0, and 0, respectively, and the sum of the frequencies is 1; the frequencies at which 61.5° is detected in the sound source direction, the first type of direction, the second type of direction, and the third type of direction are 1, 0, 1, and 0, respectively, and the sum of the frequencies is 2; the frequencies at which 94.5° is detected in the sound source direction, the first type of direction, the second type of direction, and the third type of direction are 1, 1, 1, and 1, respectively, and the sum of the frequencies is 4; the frequencies at which 120° is detected in the sound source direction, the first type of direction, the second type of direction, and the third type of direction are 1, 0, 0, and 0, respectively, and the sum of the frequencies is 1. It can be seen that the sum of the frequencies at which 94.5° is detected in the sound source direction and at least one type of direction is the most.
[0204] It can be understood that the more the sum of frequencies of which the direction is detected, the more likely the direction is the target sound source direction. Ideally, the direction is substantially the target sound source direction.
[0205] The second parameter: whether the electronic device and the user have successfully performed voice interaction within a preset time period and a preset angle range corresponding to each direction, the preset time period being a time period between a current time and a historical time, and the preset angle range including an angle corresponding to each direction.
[0206] The angle of the preset angle range corresponding to each direction can include not only the angle corresponding to the direction, but also angles near the angle, for example, the angle corresponding to a certain direction is 30°, and the preset angle range can be 25°-35°. It should be understood that the smaller the preset angle range, the more accurate the target sound source direction determined by using the parameter.
[0207] The preset time period is a time period between a current time and a historical time, and the historical time is a time located before the current time. Generally, the length of the preset time period should not be too long, which is conducive to accurately determining the target sound source direction. For example, the length of the preset time period can be set to 1 minute, 5 minutes, 10 minutes, etc. Assuming that the current time is 10:30, and the length of the preset time period is 10 minutes, the historical time is 10:20, and the preset time period is a time period between 10:20 and 10:30.
[0208] For the second parameter, in other words, it can be understood that whether the electronic device and the user have successfully performed voice interaction within a preset angle range corresponding to a certain direction.
[0209] In actual scenarios, the user is likely to use the electronic device continuously within a certain time period, especially for electronic devices with a display screen, such as a smart television. When the user is watching television, the user is unlikely to frequently move positions. Therefore, within a preset time period and a preset angle range corresponding to a certain direction, if the electronic device and the user have successfully performed voice interaction, it means that the direction is more likely to be the target sound source direction, and vice versa, the direction is less likely to be the target sound source direction. Further, the more the frequency of successful voice interaction between the electronic device and the user, the more likely the direction is the target sound source direction, and vice versa.
[0210] The third parameter: an included angle between each direction and a direction perpendicular to a display screen of the electronic device.
[0211] The third parameter is more suitable for electronic devices with a display screen, and the direction perpendicular to the display screen of the electronic device can be understood as the thickness direction of the electronic device.
[0212] In actual scenarios, when a user watches a video, the user will face the electronic device (or the display screen) in front of the electronic device to have a better watching experience. Therefore, if an included angle between a certain direction and a direction perpendicular to the display screen of the electronic device is smaller, it means that the user is likely to face the electronic device to watch the video, and thus the user is more likely to issue a voice instruction, and therefore, the direction is more likely to be the target sound source direction, and vice versa. In other words, if a certain direction is closer to the direction perpendicular to the display screen, the direction is more likely to be the target sound source direction.
[0213] For the third parameter, in other words, it can be understood that whether the user is in a certain specific direction defined when the user uses the preset scene for the electronic device.
[0214] It should be understood that the electronic device can determine the target sound source direction based on one or two or three of the above parameters, and embodiments of the present application do not make any limitation, which is described below.
[0215] In some embodiments, the at least one parameter includes the first parameter, that is, the at least one parameter includes: a sum of frequencies detected in the sound source direction and the at least one type of direction for each direction. Exemplarily, as a principle, a direction with the largest sum of frequencies detected in the sound source direction and the at least one type of direction can be determined as the target sound source direction.
[0216] In other embodiments, the at least one parameter includes the second parameter, that is, the at least one parameter includes: whether the electronic device and the user have successfully performed voice interaction within a preset period of time and a preset angle range corresponding to each direction. Exemplarily, as a principle, a direction corresponding to an angle for which the electronic device and the user have successfully performed voice interaction within the preset period of time and the preset angle range is determined as the target sound source direction.
[0217] In other embodiments, the at least one parameter includes the third parameter, that is, the at least one parameter includes: an included angle between each direction and a direction perpendicular to the display screen of the electronic device. Exemplarily, as a principle, a direction with the smallest included angle between the direction perpendicular to the display screen of the electronic device can be determined as the target sound source direction.
[0218] In other embodiments, the at least one parameter includes any two or three parameters, and exemplarily, for each parameter, a candidate sound source direction can be obtained based on the principle in the above corresponding example, and a direction with the highest repetition rate among the candidate sound source directions is determined as the target sound source direction.
[0219] For example, the at least one parameter includes a first parameter and a second parameter, for the first parameter, a direction in which a sum of frequencies detected in the sound source direction and the at least one type of direction is maximum is taken as a candidate sound source direction, assuming that the candidate sound source direction is 94.5°, for the second parameter, a direction corresponding to an angle in which the electronic device and the user successfully perform voice interaction within a preset time period and a preset angle range is taken as another candidate sound source direction, assuming that the candidate sound source direction is 94.5°, then a target sound source direction obtained based on the two candidate sound source directions is 94.5°.
[0220] In some other embodiments, the electronic device can determine a confidence degree of each direction according to the at least one parameter, and determine a direction corresponding to a confidence degree with a maximum value in the at least one direction as the target sound source direction. The confidence degree of each direction can also be referred to as the reliability of each direction, which represents a probability that the direction is the target sound source direction, and the greater the confidence degree, the greater the possibility that the direction corresponding to the confidence degree is the target sound source direction.
[0221] For example, the at least one parameter includes three parameters, and a manner of determining the target sound source direction based on the confidence degree is described. It should be understood that the manner of determining the target sound source direction based on the confidence degree in the embodiment in which the at least one parameter includes one or two parameters is similar to the embodiment in which the at least one parameter includes three parameters, and reference can be made to the description below, and subsequent descriptions are omitted.
[0222] For example, a weighted value can be configured for each parameter according to a priority of the three parameters, and the target sound source direction is determined based on a confidence degree of each direction. For example, a direction corresponding to a confidence degree with a maximum value in the at least one direction is determined as the target sound source direction.
[0223] For example, the priority of the three parameters is in a descending order as follows: the priority of the first parameter > the priority of the second parameter > the priority of the third parameter, and correspondingly, the weighted value of the first parameter > the weighted value of the second parameter > the weighted value of the third parameter.
[0224] For example, the at least one parameter includes a first parameter and a second parameter, for the first parameter, a direction in which a sum of frequencies detected in the sound source direction and the at least one type of direction is maximum is taken as a candidate sound source direction, assuming that the candidate sound source direction is 94.5°, for the second parameter, a direction corresponding to an angle in which the electronic device and the user successfully perform voice interaction within a preset time period and a preset angle range is taken as another candidate sound source direction, assuming that the candidate sound source direction is 94.5°, then a target sound source direction obtained based on the two candidate sound source directions is 94.5°.
[0225] Assuming that the first parameter has a weighting value of 0.5, the second parameter has a weighting value of 0.3, and the third parameter has a weighting value of 0.2, for the first parameter, if each direction after merging is detected in the sound source direction and the three types of directions, the score of each direction detected once is 10 points, for the second parameter, if the electronic device has successfully performed voice interaction with the user within a preset time period and a preset angle range corresponding to a certain direction, the score of the direction is also 10 points. For the third parameter, if the angle between a certain direction and the direction perpendicular to the display screen is less than a threshold value, the score of the direction is also 10 points, for example, the threshold value is 10°.
[0226] Among them, 4 sound source directions correspond to angles of 30°, 60°, 95°, and 120°, the first type of direction includes 1 direction corresponding to an angle of 93°, the second type of direction includes 2 directions corresponding to angles of 63° and 95°, and the third type of direction includes 1 direction corresponding to an angle of 95°. The angles corresponding to the 4 directions obtained by merging processing are 30°, 61.5°, 94.5°, and 120°.
[0227] In 30°, for the first parameter, only in the sound source direction, 1 10 points can be obtained, for the second parameter and the third parameter, the conditions are not met, the score is 0, and the confidence is 10*0.5=5.
[0228] In 61.5°, for the first parameter, in the sound source direction and the second type of direction, 2 10 points can be obtained, i.e. 20 points, for the second parameter, in the direction corresponding to 61.5°, the electronic device has successfully performed voice interaction once, 1 10 points can be obtained, for the third parameter, the condition is not met, the score is 0, and the confidence is 20*0.5+10*0.3=13.
[0229] In 94.5°, for the first parameter, in the sound source direction and the three types of directions, 4 10 points can be obtained, i.e. 40 points, for the second parameter, in the direction corresponding to 94.5°, the electronic device has successfully performed voice interaction once, 1 10 points can be obtained, for the third parameter, 94.5°-90°=4.5°, 4.5° is less than 10°, the condition is met, and 1 10 points can also be obtained, therefore, the confidence is 40*0.5+10*0.3+10*0.2=25.
[0230] In 10°, for the first parameter, only in the third type of direction, 1 10 points can be obtained, for the second parameter and the third parameter, the conditions are not met, the score is 0, and the confidence is 10*0.5=5.
[0231] In summary, the confidence value of 94.5° is the highest, and then the direction corresponding to 94.5° is determined as the target sound source direction.
[0232] In step S340, the user's lip video in the target sound source direction is obtained by the camera 140.
[0233] After determining the target sound source direction, the electronic device rotates the camera 140 to the target sound source direction, and the camera 140 shoots a video in the target sound source direction, which includes the user's lip video of the target user in the target sound source direction.
[0234] In step S350, the second audio signal is obtained by the microphone array 130.
[0235] It should be understood that the second audio signal is a signal for indicating the actual voice command. For example, assuming that the voice instruction issued by the target user throughout the process is "Xiao Yi, I want to watch a variety show", then the second audio signal can be used to indicate the voice instruction "I want to watch a variety show".
[0236] In order to improve the pickup effect, in some embodiments, the second audio signal is obtained in the target sound source direction by the microphone array 130 based on the beamforming technology.
[0237] In step S350, according to the second audio signal and the user's lip video, a third audio signal is obtained by a voice enhancement model, and the voice enhancement model includes a correspondence relationship between a plurality of pronunciations and a plurality of lip shapes.
[0238] The purpose of the voice enhancement model is to perform pickup enhancement processing on the audio signal, enhance the audio signal in the target sound source direction, and suppress or eliminate the audio signal in other directions to obtain or restore a clearer audio signal. The voice enhancement model integrates the audio-video information and the correspondence relationship between the pronunciation and the lip shape, that is, one or more pronunciations correspond to one lip shape. The second audio signal is used as the audio input, and the lip information in the target sound source direction is used as the video input. The voice enhancement model can perform enhancement processing on the audio signal based on the correspondence relationship between the pronunciation and the lip shape and the input user's lip video, to obtain or restore a clearer third audio signal for voice recognition. Compared with the way of processing the audio signal based on only the audio information, the voice enhancement model in the embodiment of the present application processes the audio signal based on the audio-video information, which can obtain a relatively clean audio signal and greatly improve the pickup enhancement effect.
[0239] For example, the voice enhancement module can perform noise reduction processing, echo residual elimination processing, dereverberation processing, etc. on the second audio signal.
[0240] Figure 9is a schematic flowchart of a method 400 of signal processing provided by another embodiment of the present application, which can be performed by the processing unit 160 of the electronic device 10.
[0241] In step S410, sound source localization is performed on the first audio signal obtained by the microphone array 130 to obtain sound source direction information, which includes at least one sound source direction. The at least one sound source direction includes the target sound source direction.
[0242] For detailed description of step S410, reference can be made to the relevant description of step S310 above.
[0243] In step S420, the first video obtained by the camera 140 is processed to obtain user direction information, which includes at least one type of direction related to the user.
[0244] For detailed description of step S420, reference can be made to the relevant description of step S320 above.
[0245] In step S430, the target sound source direction is determined according to the sound source direction information and the user direction information, which is the direction of the target user who is in voice interaction with the electronic device.
[0246] For detailed description of step S430, reference can be made to the relevant description of step S330 above.
[0247] In step S440, the user lip video in the target sound source direction is obtained by the camera 140.
[0248] For detailed description of step S440, reference can be made to the relevant description of step S340 above.
[0249] In step S450, the second audio signal is obtained by the microphone array 130.
[0250] For detailed description of step S450, reference can be made to the relevant description of step S350 above.
[0251] In step S460, the fourth audio signal in the target sound source direction is obtained by the directional microphone 150.
[0252] After the electronic device determines the target sound source direction, the electronic device can control the directional microphone 150 to rotate to the target sound source direction to pick up the fourth audio signal in the target sound source direction.
[0253] In the embodiment in which the directional microphone 150 is disposed at the camera 140, the electronic device can control the camera 140 and the directional microphone 150 to rotate to the target sound source direction together.
[0254] In step S470, a third audio signal is obtained by a speech enhancement model according to the second audio signal, the fourth audio signal, and the user lip video.
[0255] In this step, the second audio signal picked up by the microphone array 130 and the fourth audio signal picked up by the directional microphone 150 are taken as audio inputs of the speech enhancement model, and the user lip video is taken as a video input. The speech enhancement model is used to process the input audio signals to obtain a clearer third audio signal.
[0256] Since the directional microphone 150 can pick up sound without distortion in the direction of the target sound source, it can have a certain inhibitory effect on interference and reverberation. In addition, the directional microphone 150 picks up sound forward, which can also effectively suppress echo. Therefore, the fourth audio signal obtained by the directional microphone 150 and the second audio signal obtained by the microphone array 130 are taken as audio inputs of the speech enhancement model, which can obtain or restore a clearer third audio signal.
[0257] It should be understood that, similar to the method 200 described above, in various embodiments of the methods 300 and 400 described above, the sequence numbers of the processes do not mean the order of execution. The execution order of the processes should be determined according to their functions and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0258] The embodiments of the present application also provide an electronic device, which can be Figure 4 The electronic device shown in the figure includes a microphone array 130, a rotatable camera 140, and a processor 160. The processor 160 is configured to:
[0259] perform sound source positioning on a first audio signal obtained by the microphone array 130 to obtain sound source direction information;
[0260] process a first video obtained by the camera 140 to obtain user direction information;
[0261] determine a target sound source direction according to the sound source direction information and the user direction information;
[0262] obtain a user lip video in the target sound source direction by the camera 140;
[0263] obtain a second audio signal by the microphone array 130;
[0264] obtain a third audio signal by a speech enhancement model according to the second audio signal and the user lip video, the speech enhancement model including a corresponding relationship between pronunciation and lip shape.
[0265] Optionally, the electronic device further comprises a directional microphone 150, and the processor 160 is further configured to:
[0266] obtain a fourth audio signal in the target sound source direction through the directional microphone 150; and
[0267] The processor 160 is specifically configured to:
[0268] obtain the third audio signal through the speech enhancement module according to the second audio signal, the fourth audio signal and the user lip video.
[0269] Optionally, the directional microphone 150 is fixedly connected with the camera 140.
[0270] Optionally, the user direction information comprises at least one type of direction, and the at least one type of direction comprises:
[0271] a first type of direction, the first type of direction comprising a direction in which at least one lip in an active state is located;
[0272] a second type of direction, the second type of direction comprising a direction in which at least one user is located;
[0273] a third type of direction, the third type of direction comprising a direction in which at least one user is looking at the electronic device.
[0274] Optionally, the sound source direction information comprises at least one sound source direction, and the processor 160 is specifically configured to:
[0275] merge the at least one sound source direction and the at least one type of direction to obtain at least one merged direction;
[0276] determine the target sound source direction from the at least one direction.
[0277] Optionally, the processor 160 is specifically configured to:
[0278] determine the target sound source direction from the at least one direction according to at least one parameter;
[0279] wherein the at least one parameter comprises:
[0280] a sum of frequencies at which each direction in the at least one direction is detected in the sound source direction and the at least one type of direction;
[0281] whether the electronic device and the user have successfully performed speech interaction within a preset period of time and a preset angle range corresponding to each direction, the preset period of time being a period of time between a current time and a historical time;
[0282] an included angle between the each direction and a direction perpendicular to a display screen of the electronic device.
[0283] Optionally, the processor 160 is specifically configured to:
[0284] determine a confidence degree of the each direction according to the at least one parameter;
[0285] determine a direction corresponding to a maximum confidence degree in the at least one direction as the target sound source direction.
[0286] Optionally, the processor 160 is specifically configured to:
[0287] obtain the second audio signal in the target sound source direction based on a beam forming technology through the microphone array 130.
[0288] Optionally, the first audio signal is a wake-up signal.
[0289] Optionally, the electronic device is a smart television.
[0290] It should be understood that, in the embodiments of the present application, unless otherwise explicitly specified and limited, the terms "connection", "fixed connection" and the like should be understood in a broad sense. For those skilled in the art, the specific meanings of the above-mentioned various terms in the embodiments of the present application can be understood according to the specific circumstances.
[0291] Exemplarily, for "connection", it can be various connection modes such as fixed connection, rotary connection, flexible connection, movable connection, one-piece forming, electrical connection, etc. It can be directly connected, or it can be indirectly connected through an intermediate medium, or it can be the internal communication of two elements or the interaction relationship between two elements.
[0292] Exemplarily, for "fixed connection", one element can be directly or indirectly fixed to another element. Fixed connection can include mechanical connection, welding and adhesive bonding and the like. The mechanical connection can include riveting, bolt connection, threaded connection, key pin connection, buckle connection, lock connection, plug-in and the like. The adhesive bonding can include adhesive bonding and solvent bonding and the like.
[0293] It should also be understood that the "parallel" or "perpendicular" described in the embodiments of the present application can be understood as "approximately parallel" or "approximately perpendicular".
[0294] It should also be understood that the terms "length", "width", "thickness", "upper", "lower", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer" and the like, indicating the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, are only for the convenience of describing the present application and simplifying the description, and do not indicate or imply that the device or element referred to must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as a limitation on the present application.
[0295] It should be noted that the terms "first", "second" are only for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the technical features indicated. The features defined with "first", "second" can explicitly or implicitly include one or more of the features.
[0296] In the embodiments of the present application, "at least one" means one or more, and "multiple" means two or more. "At least part of the element" means part or all of the element. "And / or" describes the association relationship of the associated objects, which means that there can be three kinds of relationships, for example, A and / or B can represent the following three cases: A exists alone, A and B exist together, B exists alone, where A and B can be singular or plural. The character " / " generally represents that the front and rear associated objects are in an "or" relationship.
[0297] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be realized in electronic hardware, or a combination of computer software and electronic hardware. Whether the functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. A person skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0298] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the system, device and unit described above can refer to the corresponding process in the foregoing method embodiments, which will not be repeated here.
[0299] In several embodiments provided in the present application, it should be understood that the disclosed system, device and method can be implemented in other manners. For example, the described device embodiments are merely illustrative. For example, the division of the units is only a logical function division. There can be another division manner for the actual implementation, for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections can be indirect couplings or communication connections through some interfaces, devices or units, and can be in electrical, mechanical or other forms.
[0300] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, i.e., can be located in one place, or can be distributed on multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the embodiment.
[0301] In addition, each functional unit in the various embodiments of the present application can be integrated into a processing unit, or each unit can be a physically independent unit, or two or more units can be integrated into a unit.
[0302] If the functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.
[0303] The same or similar parts between the various embodiments in this application can be referred to mutually. In the various embodiments of this application, and in the various implementation methods / methods / implementations within each embodiment, unless otherwise specified or logically conflicting, the terminology and / or descriptions between different embodiments and between the various implementation methods / methods / implementations within each embodiment are consistent and can be mutually referenced. The technical features in different embodiments and the various implementation methods / methods / implementations within each embodiment can be combined according to their inherent logical relationships to form new embodiments, implementation methods, methods, or implementation approaches. The above-described embodiments of this application do not constitute a limitation on the scope of protection of this application.
[0304] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method of signal processing, characterized by, The method is applied to an electronic device including a microphone array and a camera, and comprises: sound source localization on a first audio signal obtained by the microphone array to obtain sound source direction information; processing a first video obtained by the camera to obtain user direction information; determining a target sound source direction according to the sound source direction information and the user direction information; obtaining a user lip video in the target sound source direction by the camera; obtaining a second audio signal by the microphone array; obtaining a third audio signal by a speech enhancement model according to the second audio signal and the user lip video, the speech enhancement model including a correspondence between pronunciation and lip shape; wherein the user direction information includes at least one of the following types of directions: a first type of direction including a direction in which at least one lip in an active state is located; a second type of direction including a direction in which at least one user is located; a third type of direction including a direction in which at least one user is looking at the electronic device; wherein the sound source direction information includes at least one sound source direction, and determining the target sound source direction according to the sound source direction information and the user direction information includes: merging processing the at least one sound source direction and the at least one type of direction to obtain at least one merged direction; determining the target sound source direction from the at least one direction; wherein determining the target sound source direction from the at least one direction includes: determining the target sound source direction from the at least one direction according to at least one parameter; wherein the at least one parameter includes: a sum of frequencies at which each direction in the at least one direction is detected in the sound source direction and the at least one type of direction; whether the electronic device and a user have successfully performed voice interaction within a preset period and a preset angle range corresponding to each direction, the preset period being a period between a current time and a historical time; an included angle between each direction and a direction perpendicular to a display screen of the electronic device.
2. The method of claim 1, wherein, The electronic device further includes a directional microphone, and the method further includes: obtaining a fourth audio signal in the target sound source direction by the directional microphone; and obtaining a third audio signal by a speech enhancement model according to the second audio signal and the user lip video in the target sound source direction includes: obtaining the third audio signal by a speech enhancement module according to the second audio signal, the fourth audio signal and the user lip video.
3. The method according to claim 1 or 2, characterized in that, determining the target sound source direction from the at least one direction according to the at least one parameter includes: determining a confidence degree of each direction according to the at least one parameter; determining a direction corresponding to a maximum value of the confidence degrees in the at least one direction as the target sound source direction.
4. The method according to claim 1 or 2, characterized in that, obtaining a second audio signal by the microphone array includes: obtaining the second audio signal in the target sound source direction based on a beam forming technology by the microphone array.
5. The method according to claim 1 or 2, characterized in that, The first audio signal is a wake-up signal.
6. An electronic device, comprising: The electronic device comprises a microphone array, a camera, and a processor, wherein the processor is configured to: perform sound source positioning on a first audio signal obtained by the microphone array to obtain sound source direction information; process a first video obtained by the camera to obtain user direction information; determine a target sound source direction according to the sound source direction information and the user direction information; obtain a user lip video in the target sound source direction through the camera; obtain a second audio signal through the microphone array; obtain a third audio signal through a speech enhancement model according to the second audio signal and the user lip video, wherein the speech enhancement model comprises a corresponding relationship between pronunciation and lip shape; wherein the user direction information comprises at least one of the following types of directions: a first type of direction, which comprises a direction in which at least one lip in an active state is located; a second type of direction, which comprises a direction in which at least one user is located; a third type of direction, which comprises a direction in which at least one user is looking at the electronic device; wherein the sound source direction information comprises at least one sound source direction, and the processor is specifically configured to: merge the at least one sound source direction and the at least one type of direction to obtain at least one merged direction; determine the target sound source direction from the at least one direction; wherein the processor is specifically configured to: determine the target sound source direction from the at least one direction according to at least one parameter; wherein the at least one parameter comprises: a sum of frequencies at which each direction in the at least one direction is detected in the sound source direction and the at least one type of direction; whether the electronic device and a user have successfully performed voice interaction within a preset time period and a preset angle range corresponding to each direction, wherein the preset time period is a time period between a current time and a historical time; an included angle between each direction and a direction perpendicular to a display screen of the electronic device.
7. The electronic device of claim 6, wherein, The electronic device further comprises a directional microphone, and the processor is further configured to: obtain a fourth audio signal in the target sound source direction through the directional microphone; and the processor is specifically configured to: obtain the third audio signal through a speech enhancement module according to the second audio signal, the fourth audio signal, and the user lip video.
8. The electronic device of claim 7, wherein, The directional microphone is fixedly connected with the camera.
9. The electronic device of any of claims 6-8, wherein, The processor is specifically configured to: determine a confidence degree of each direction according to the at least one parameter; determine a direction corresponding to a maximum confidence degree in the at least one direction as the target sound source direction.
10. The electronic device of any of claims 6-8, wherein, The processor is specifically configured to: obtain the second audio signal in the target sound source direction based on a beamforming technology through the microphone array.
11. The electronic device of any of claims 6-8, wherein, The first audio signal is a wake-up signal.
12. The electronic device of any of claims 6-8, wherein, The electronic device is a smart television.
13. A computer storage medium, comprising, comprises: a processor coupled with the memory, the memory to store a program or instructions that, when executed by the processor, cause the method of any one of claims 1 to 5 to be performed.
14. A computer program product comprising instructions, characterized in that, The computer program product, when run on an electronic device, causes the electronic device to perform the method of any one of claims 1 to 5.
Citation Information
Patent Citations
Speech signal processing method, apparatus and electronic device
CN107146614A
Sound processing method and device and electronic equipment
CN107993671A