Audio processing method, terminal and storage medium
By combining voiceprints and face features matching in audio and video calls, the target user is identified and interfered with audio, and the problem of noise and interfering with human voice in public places is solved, audio noise reduction is achieved, and user experience is improved.
Patent Information
- Application Number
- CN202411354994.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-25
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2044-09-25
AI Technical Summary
When making audio and video calls in public places, the audio data collected by the terminal contains ambient noise and disturbing voices, resulting in a decline in user experience.
By obtaining the to-processed audio and face images during the audio call, using voiceprint features and face features matching to identify the target user, suppressing interfering audio other than the target user, combining audio and video preprocessing to improve feature extraction accuracy, adjusting the difference threshold to deal with voiceprint changes, and performing lip motion detection to confirm user consistency.
Improves the accuracy of suppression of interfering audio, improves user experience, and ensures clear transmission of target audio.
Smart Images

Figure CN120455583A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to an audio processing method, terminal, and storage medium. Background Art
[0002] Users can make audio and video calls using devices (such as mobile phones, tablets, and PCs). If users are in public places, such as offices, cafes, restaurants, airports, and train stations, the audio data collected by the devices may contain ambient noise and interfering human voices, which can degrade the user experience.
[0003] Therefore, in order to allow the other party of the video call to better hear the user's voice, the terminal needs to perform noise reduction processing on the environmental noise and interfering human voices in the collected audio data to improve the user experience. Summary of the Invention
[0004] The purpose of the embodiments of the present application is to provide an audio processing method, terminal, and storage medium to improve the accuracy of suppressing interfering audio, achieve audio noise reduction, and improve the user experience. The specific technical solution is as follows:
[0005] In a first aspect, to achieve the above-mentioned objectives, an embodiment of the present application provides an audio processing method, which is applied to a first terminal and includes:
[0006] During an audio call between the first terminal and the second terminal, obtaining collected audio to be processed and a facial image to be processed;
[0007] For each template voiceprint feature, if the voiceprint feature of the audio to be processed matches the template voiceprint feature, the user to which the template voiceprint feature belongs is determined to be the user to which the audio to be processed belongs;
[0008] For each template facial feature, if the facial feature of the face image to be processed matches the template facial feature, determining that the user to which the template facial feature belongs is the user to which the face image to be processed belongs;
[0009] If the user to which the audio to be processed belongs and the user to which the facial image to be processed belongs are the same target user, interference audio other than the audio to be processed of the target user is suppressed to obtain target audio, and the target audio is sent to the second terminal.
[0010] From the above, it can be seen that the technical solution provided by this embodiment, when the target user's to-be-processed audio and the template voiceprint features, as well as the to-be-processed facial image and the template facial features are successfully matched, indicates that the target user is a user making an audio call, that is, the target user's to-be-processed audio is not interference audio. In this case, the interference audio other than the target user's to-be-processed audio is suppressed, which can improve the accuracy of suppressing the interference audio, achieve audio noise reduction, and improve the user experience.
[0011] In one embodiment of the present application, for each template voiceprint feature, if the voiceprint feature of the audio to be processed matches the template voiceprint feature, before determining that the user to which the template voiceprint feature belongs is the user to which the audio to be processed belongs, the method further includes:
[0012] For each audio to be processed, extract the features of the audio to be processed to obtain the voiceprint features of the audio to be processed;
[0013] For each template voiceprint feature, if the difference between the voiceprint feature of the audio to be processed and the template voiceprint feature is less than a preset difference threshold, it is determined that the voiceprint feature of the audio to be processed matches the template voiceprint feature.
[0014] In one embodiment of the present application, after determining that the voiceprint feature of the audio to be processed matches the template voiceprint feature if the difference between the voiceprint feature of the audio to be processed and the template voiceprint feature is less than a preset difference threshold, the method further includes:
[0015] If the user to which the audio to be processed belongs and the user to which the facial image to be processed belongs are not the same target user, increasing the preset difference threshold from the first value to the second value;
[0016] For each template voiceprint feature, if the difference between the voiceprint feature of the audio to be processed and the template voiceprint feature is less than the adjusted difference threshold, it is determined that the voiceprint feature of the audio to be processed matches the template voiceprint feature, and the user to whom the template voiceprint feature that matches the voiceprint feature of the audio to be processed belongs is determined to be the user to whom the audio to be processed belongs;
[0017] If the user to which the audio to be processed belongs and the user to which the facial image to be processed belongs are the same target user, interference audio other than the audio to be processed of the target user is suppressed to obtain target audio, and the target audio is sent to the second terminal.
[0018] As can be seen from the above, the technical solution provided by this embodiment avoids the problem of recognition errors caused by changes in the target user's voiceprint characteristics by adjusting the preset difference threshold, further improves the accuracy of suppressing interfering audio, and improves the user experience.
[0019] In one embodiment of the present application, after determining that the user to whom the template voiceprint feature that matches the voiceprint feature of the audio to be processed belongs is the user to whom the audio to be processed belongs, the method further includes:
[0020] If the user to which the audio to be processed belongs and the user to which the facial image to be processed belongs are not the same user, it is determined that the audio to be processed is interference audio.
[0021] As can be seen from the above, the technical solution provided by this embodiment determines that the audio to be processed whose facial image features do not match the voiceprint features is interference audio, can accurately identify the interference audio, and suppress the interference audio, further improving the accuracy of suppressing the interference audio and improving the user experience.
[0022] In one embodiment of the present application, if the user to which the audio to be processed belongs and the user to which the facial image to be processed belong are not the same target user, after increasing the preset difference threshold from the first value to the second value, the method further includes:
[0023] If the difference between the voiceprint feature of the audio to be processed and the voiceprint features of each template is greater than the adjusted preset difference threshold, it is determined that the audio to be processed is interference audio.
[0024] From the above, it can be seen that the technical solution provided by this embodiment determines that the audio to be processed whose voiceprint features do not match the voiceprint features of each template is interference audio, can accurately identify the interference audio, and suppress the interference audio, further improve the accuracy of suppressing the interference audio, and improve the user experience.
[0025] In one embodiment of the present application, if the user to which the audio to be processed belongs and the user to which the facial image to be processed belongs are the same target user, before suppressing interference audio other than the audio to be processed of the target user to obtain target audio, and sending the target audio to the second terminal, the method further includes:
[0026] Performing lip movement detection on the face image to be processed;
[0027] If lip movement of the user in the face image to be processed is detected, it is detected whether the user to whom the audio to be processed belongs and the user to whom the face image to be processed belongs are the same target user.
[0028] As can be seen from the above, the technical solution provided by this embodiment is that if the user in the face image to be processed has lip movements, indicating that the user within the camera field of view of the first terminal has made a sound, then the detection of whether the user making the sound is the same user as the user within the camera field of view of the first terminal is continued, thereby avoiding the problem of audio suppression errors caused by mistakenly identifying the sounds made by other users as the sounds made by the target user, further improving the accuracy of suppressing interfering audio and improving the user experience.
[0029] In one embodiment of the present application, obtaining the collected audio to be processed and the facial image to be processed includes:
[0030] Obtaining audio captured by a microphone of the first terminal, performing audio preprocessing on the acquired audio to obtain original audio, extracting a vocal portion from the original audio to obtain audio to be processed; wherein the audio preprocessing includes acoustic echo cancellation, automatic gain control, and noise suppression;
[0031] A video image containing a user image captured by a camera is obtained, and image preprocessing is performed on the obtained video image to obtain an original video image, and a face image in the original video image is extracted to obtain a face image to be processed; wherein the image preprocessing includes: automatic exposure, automatic white balance and automatic focus.
[0032] From the above, it can be seen that the technical solution provided in this embodiment performs audio preprocessing on the collected audio, which can improve the accuracy of the subsequently extracted voiceprint features, and performs video preprocessing on the collected video, which can improve the accuracy of the subsequently extracted facial features, further improve the accuracy of suppressing interfering audio, and improve the user experience.
[0033] In one embodiment of the present application, for each template voiceprint feature, if the voiceprint feature of the audio to be processed matches the template voiceprint feature, before determining that the user to which the template voiceprint feature belongs is the user to which the audio to be processed belongs, the method further includes:
[0034] Obtain the collected audio and facial image of the user to be registered;
[0035] Performing feature extraction on the audio to be registered to obtain a template voiceprint feature of the audio to be registered, and performing feature extraction on the facial image to be registered to obtain a template facial feature of the facial image to be registered;
[0036] Performing lip movement detection on the face image to be registered;
[0037] If lip movement is detected in the face image to be registered of the user to be registered, the template voiceprint features of the audio to be registered, the template facial features of the face image to be registered, and the corresponding user identification of the user to be registered are recorded in a preset storage area locally on the first terminal.
[0038] As can be seen from the above, the technical solution provided by this embodiment can record the template voiceprint features of the audio to be registered, the template facial features of the facial image to be registered, and the corresponding user ID of the user to be registered. Subsequently, if the target user's to-be-processed audio successfully matches the template voiceprint features, and the to-be-processed facial image successfully matches the template facial features, indicating that the target user is the user making an audio call, that is, the target user's to-be-processed audio is not interference audio. In this case, interference audio other than the target user's to-be-processed audio is suppressed, which can improve the accuracy of interference audio suppression, achieve audio noise reduction, and enhance the user experience.
[0039] In one embodiment of the present application, the preset storage area is a secure storage area of the first terminal.
[0040] As can be seen from the above, the technical solution provided by this embodiment presets the storage area as a local secure storage area of the first terminal, which can improve the security of the registered template voiceprint features and template facial features and protect the user's privacy.
[0041] In a second aspect, an embodiment of the present application further provides a terminal, including:
[0042] one or more processors and memory;
[0043] The memory is coupled to the one or more processors, and is used to store computer program code, where the computer program code includes computer instructions. The one or more processors call the computer instructions to enable the terminal to execute any one of the above-mentioned audio processing methods.
[0044] In a third aspect, an embodiment of the present application further provides a computer-readable storage medium, comprising a computer program, which, when executed on a terminal, enables the terminal to execute any of the above-described audio processing methods.
[0045] In a fourth aspect, an embodiment of the present application further provides a computer program product, which includes executable instructions. When the executable instructions are executed on a terminal, the terminal executes any of the above-mentioned audio processing methods.
[0046] In a fifth aspect, an embodiment of the present application further provides a chip system, which is applied to a terminal. The chip system includes one or more processors, and the processor is used to call computer instructions so that the terminal executes any of the above-mentioned audio processing methods.
[0047] The beneficial effects of the solutions provided by the embodiments in the second, third, fourth and fifth aspects can be referred to the beneficial effects of the solutions provided by the embodiments in the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0049] Figure 1 A structural diagram of a terminal provided in an embodiment of the present application;
[0050] Figure 2 A software structure diagram of a terminal provided in an embodiment of the present application;
[0051] Figure 3 A schematic diagram of an application scenario of the first audio processing method provided in an embodiment of the present application;
[0052] Figure 4 A schematic diagram of an application scenario of the second audio processing method provided in an embodiment of the present application;
[0053] Figure 5 A schematic diagram of an application scenario of the third audio processing method provided in an embodiment of the present application;
[0054] Figure 6 A flowchart of the first audio processing method provided in an embodiment of the present application;
[0055] Figure 7 A flowchart of a second audio processing method provided in an embodiment of the present application;
[0056] Figure 8 A flowchart of a third audio processing method provided in an embodiment of the present application;
[0057] Figure 9 A flowchart of a fourth audio processing method provided in an embodiment of the present application;
[0058] Figure 10 A flowchart for obtaining audio data provided in an embodiment of the present application;
[0059] Figure 11 A flowchart for obtaining image data provided in an embodiment of the present application;
[0060] Figure 12 A flowchart of the first audio data and image data transmission provided in an embodiment of the present application;
[0061] Figure 13 A flowchart of a registration template voiceprint feature and template face feature provided in an embodiment of the present application;
[0062] Figure 14 A flowchart for obtaining template voiceprint features provided in an embodiment of the present application;
[0063] Figure 15 A flowchart for obtaining template facial features provided in an embodiment of the present application;
[0064] Figure 16 A flowchart of the second audio data and image data transmission provided in an embodiment of the present application;
[0065] Figure 17 A structural diagram of a chip system provided in an embodiment of the present application. DETAILED DESCRIPTION
[0066] In order to better understand the technical solution of the present application, the embodiments of the present application are described in detail below with reference to the accompanying drawings.
[0067] In order to facilitate the clear description of the technical solutions of the embodiments of the present application, in the embodiments of the present application, words such as "first" and "second" are used to distinguish between identical or similar items with substantially the same functions and effects. For example, the first instruction and the second instruction are intended to distinguish different user instructions and do not limit their order. Those skilled in the art will understand that words such as "first" and "second" do not limit the quantity and execution order, and words such as "first" and "second" do not necessarily limit them to be different.
[0068] It should be noted that, in this application, words such as "exemplarily" or "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design described in this application as "exemplary" or "for example" should not be construed as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplarily" or "for example" is intended to present the relevant concepts in a concrete manner.
[0069] The audio processing method provided in the embodiments of the present application is applied to a terminal. The terminal may be a mobile phone, tablet computer, personal digital assistant (PDA), smart watch, wearable terminal, augmented reality (AR) device, virtual reality (VR) device, robot, smart glasses, or other device capable of recording audio and capturing images. In this way, the terminal can suppress the interfering audio during the user's audio and video call according to the audio processing method provided in the embodiments of the present application, improve the accuracy of suppressing the interfering audio, achieve audio noise reduction, and improve the user experience.
[0070] For example, Figure 1 The structure of the terminal 100 is shown. The terminal 100 may include a processor 110, a display screen 120, a camera 130, an internal memory 140, a Subscriber Identification Module (SIM) card interface 150, a Universal Serial Bus (USB) interface 160, a charging management module 170, a battery management module 171, a battery 172 with a cell and a battery protection device, a sensor module 180, a mobile communication module 190, a wireless communication module 200, an antenna 1, and an antenna 2. The sensor module 180 may include a pressure sensor 180A, a fingerprint sensor 180B, a touch sensor 180C, an ambient light sensor 180D, and the like.
[0071] It should be understood that the structures illustrated in the embodiments of the present application do not constitute a specific limitation on the terminal 100. In other embodiments of the present application, the terminal 100 may include more or fewer components than shown, or may combine or separate certain components, or arrange the components differently. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0072] The processor 110 may include one or more processing units. For example, the processor 110 may include a central processing unit (CPU), an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU). The different processing units may be independent components or integrated into one or more processors. In some embodiments, the terminal 100 may also include one or more processors 110. The controller may generate operation control signals based on instruction opcodes and timing signals to control instruction fetching and execution. In other embodiments, the processor 110 may also include a memory for storing instructions and data. For example, the memory in the processor 110 may be a cache memory. This memory may store instructions or data that have just been used or are being recycled by the processor 110. If the processor 110 needs to use the instruction or data again, it can directly call it from the memory. This avoids repeated access, reduces the waiting time of the processor 110, and thus improves the efficiency of the terminal 100 in processing data or executing instructions.
[0073] In some embodiments, the processor 110 may include one or more interfaces. The interfaces may include an Inter-Integrated Circuit (I2C) interface, an Inter-Integrated Circuit Sound (I2S) interface, a Pulse Code Modulation (PCM) interface, a Universal Asynchronous Receiver / Transmitter (UART) interface, a Mobile Industry Processor Interface (MIPI), a General-Purpose Input / Output (GPIO) interface, a SIM card interface, and / or a USB interface. The USB interface 160 is an interface that complies with USB standards and may be a Mini USB interface, a Micro USB interface, a USB Type-C interface, or the like. The USB interface 160 may be used to connect a charger to charge the terminal 100, or to transfer data between the terminal 100 and peripheral devices. The USB interface 160 may also be used to connect headphones to play audio through the headphones.
[0074] It is understood that the interface connection relationship between the modules illustrated in the embodiments of the present application is for illustrative purposes only and does not constitute a structural limitation on the terminal 100. In other embodiments of the present application, the terminal 100 may also adopt different interface connection methods from the above embodiments, or a combination of multiple interface connection methods.
[0075] The wireless communication function of the terminal 100 can be implemented through the antenna 1, the antenna 2, the mobile communication module 190, the wireless communication module 200, the modem processor, and the baseband processor.
[0076] Antenna 1 and Antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in terminal 100 can be used to cover a single or multiple communication frequency bands. Different antennas can also be reused to improve antenna utilization. For example, antenna 1 can be reused as a diversity antenna for a wireless local area network. In other embodiments, the antennas can be used in conjunction with a tuning switch.
[0077] Terminal 100 implements display functions through a GPU, display screen 120, and an application processor. The GPU is a microprocessor for image processing that connects display screen 120 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. Processor 110 may include one or more GPUs that execute program instructions to generate or modify display information.
[0078] The display screen 120 is used to display images, videos, etc. The display screen 120 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode or an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a mini-LED, a micro-LED, a micro-o-LED, or a quantum dot light-emitting diode (QLED). In some embodiments, the terminal 100 may include one or more display screens 120.
[0079] In some embodiments of the present application, when the display panel adopts materials such as OLED, AMOLED, FLED, etc., the above Figure 1 The display screen 120 can be bent. Here, the display screen 120 can be bent to any angle at any position and can be maintained at that angle. For example, the display screen 120 can be folded in half from the middle to the left or right. It can also be folded in half from the middle to the top or bottom.
[0080] The display screen 120 of the terminal 100 may be a flexible screen. Currently, flexible screens have attracted much attention due to their unique characteristics and huge potential. Compared with traditional screens, flexible screens are more flexible and bendable, which can provide users with a new way of interaction based on the bendable characteristics, and can meet more user demands for the terminal. For terminals equipped with a foldable display, the foldable display on the terminal can be switched between a small screen in a folded form and a large screen in an unfolded form at any time. Therefore, users are using the split-screen function on terminals equipped with a foldable display more and more frequently.
[0081] The terminal 100 can implement a shooting function through an ISP, a camera 130, a video codec, a GPU, a display screen 120, and an application processor, wherein the camera 130 includes a front camera and a rear camera.
[0082] The ISP processes data fed back by the camera 130. For example, when shooting, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, which is then passed to the ISP for processing and converted into a visible image. The ISP can perform algorithmic optimization on image noise, brightness, and color. The ISP can also optimize parameters such as exposure and color temperature of the captured scene. In some embodiments, the ISP can be located within the camera 130.
[0083] The camera 130 is used to take photos or videos. The object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal oxide semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, and then passes the electrical signal to the ISP for conversion into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in a standard red, green, blue (RGB), YUV, or other format. In some embodiments, the terminal 100 may include 1 or N cameras 130, where N is a positive integer greater than 1.
[0084] The digital signal processor is used to process digital signals. In addition to processing digital image signals, it can also process other digital signals. For example, when the terminal 100 selects a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy.
[0085] Video codecs are used to compress or decompress digital video. Terminal 100 may support one or more video codecs. This allows terminal 100 to play or record videos in various encoding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, and MPEG4.
[0086] The NPU is a neural network (NN) computing processor. Drawing on the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, it rapidly processes input information and can continuously self-learn. The NPU enables intelligent cognitive applications in the terminal 100, such as image recognition, face recognition, speech recognition, and text comprehension.
[0087] The internal memory 140 can be used to store one or more computer programs, each of which includes instructions. The processor 110 can execute the above instructions stored in the internal memory 140, thereby causing the terminal 100 to perform the image generation method provided in some embodiments of the present application, as well as various applications and data processing. The internal memory 140 may include a program storage area and a data storage area. The program storage area may store an operating system; the program storage area may also store one or more applications (such as a gallery, contacts, etc.). The data storage area may store data created during the use of the terminal 100 (such as photos, contacts, etc.). In addition, the internal memory 140 may include a high-speed random access memory and may also include a non-volatile memory, such as one or more disk storage components, a flash memory component, a universal flash memory (UFS), etc. In some embodiments, the processor 110 can execute the audio processing method provided in the embodiments of the present application, as well as other applications and data processing by executing the instructions stored in the internal memory 140 and / or the instructions stored in the memory provided in the processor 110.
[0088] The internal memory 140 can be used to store the relevant programs of the audio processing method provided in the embodiment of the present application, and the processor 110 can be used to call the relevant programs of the audio processing method stored in the internal memory 140 when displaying information to execute the audio processing method of the embodiment of the present application.
[0089] The sensor module 180 may include a pressure sensor 180A, a fingerprint sensor 180B, a touch sensor 180C, an ambient light sensor 180D, and the like.
[0090] Pressure sensor 180A is used to sense pressure signals and convert them into electrical signals. In some embodiments, pressure sensor 180A can be located on display screen 120. There are many types of pressure sensors 180A, including resistive, inductive, and capacitive pressure sensors. A capacitive pressure sensor may comprise at least two parallel plates made of conductive material. When force is applied to pressure sensor 180A, the capacitance between the electrodes changes, and terminal 100 determines the intensity of the pressure based on the change in capacitance. When a touch operation is applied to display screen 120, terminal 100 detects the touch operation based on pressure sensor 180A. Terminal 100 can also calculate the touch location based on the detection signal from pressure sensor 180A. In some embodiments, touch operations applied to the same touch location but with different touch operation intensities can correspond to different operation instructions. For example, when a touch operation with an intensity less than a first pressure threshold is applied to a short message application icon, a command to view short messages is executed; when a touch operation with an intensity greater than or equal to the first pressure threshold is applied to a short message application icon, a command to create a new short message is executed.
[0091] The fingerprint sensor 180B is used to collect fingerprints. The terminal 100 can use the collected fingerprint characteristics to implement functions such as unlocking, accessing application locks, taking photos, and answering calls.
[0092] Touch sensor 180C, also known as a touch-sensitive device, can be provided on display screen 120. Touch sensor 180C and display screen 120 form a touch screen, also known as a touchscreen. Touch sensor 180C is used to detect touch operations applied thereto or in the vicinity thereof. Touch sensor 180C can transmit the detected touch operations to an application processor to determine the type of touch event. Visual output related to the touch operations can be provided via display screen 120. In other embodiments, touch sensor 180C can also be provided on the surface of terminal 100, in a different location from display screen 120.
[0093] Ambient light sensor 180D is used to sense ambient light brightness. Terminal 100 can adaptively adjust the brightness of display screen 120 based on the perceived ambient light brightness. Ambient light sensor 180D can also be used to automatically adjust white balance during photography. Ambient light sensor 180D can also transmit information about the device's environment to the GPU.
[0094] The ambient light sensor 180D is also used to obtain the brightness, light ratio, color temperature, etc. of the acquisition environment in which the camera 130 captures images.
[0095] Figure 2The present invention is a software structure block diagram of a terminal applicable to an embodiment of the present application. The terminal software system can adopt a layered architecture, an event-driven architecture, a micro-kernel architecture, a microservice architecture, or a cloud architecture. The layered architecture divides the terminal software system into several layers, each with a clear role and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the software system can be divided into three layers, namely, the application layer (applications), the application framework layer (application framework) and the driver layer (hardware abstract layer, HAL).
[0096] The application layer can include a series of application packages, and the application layer runs applications by calling the application programming interface (API) provided by the application framework layer. Figure 2 As shown, the application package may include multiple applications, such as camera, clock, browser, music, etc. It is understandable that the port of each of the above applications may be used to receive data.
[0097] The application framework layer provides API and programming framework for the applications in the application layer. The application framework layer includes some predefined functions. Figure 2 As shown, the application framework layer may include a window manager, a content provider, a view system, a resource manager, a notification manager, and a Dynamic Host Configuration Protocol (DHCP) module.
[0098] The driver layer is the layer between hardware and software, responsible for driving the hardware and making it work. The driver layer can contain multiple drivers for driving the hardware. For example, there are drivers for cameras, displays, audio, and sensors.
[0099] In addition, the terminal also includes a hardware layer, which may include a camera, a speaker, a display, a battery, etc. The hardware layer is connected to the driver layer.
[0100] The following describes the application scenarios of the audio processing method provided in the embodiments of the present application.
[0101] When users use terminals (such as mobile phones, tablets, and PCs) to make audio and video calls in public places such as offices, cafes, restaurants, airports, and train stations, the audio data collected by the terminals will contain environmental noise and interfering human voices, which will reduce the user experience.
[0102] The following takes the mobile phone as an example, combined with Figure 3The application scenarios of the embodiments of this application are described. Figure 3 This is only an example description and does not limit the audio processing method provided in this embodiment to only be applicable to Figure 3 The application scenario shown.
[0103] The target user is using the instant messaging app on their phone to make a voice call with contact 1. In addition to the target user's voice, the interfering user's voice is also making sounds, as is the alarm clock in the environment. The audio collected by the phone contains the target user's audio, the interfering user's audio, and the ambient noise. Therefore, to ensure the caller can better hear the target user's voice, the phone needs to suppress the interfering user's audio and the ambient noise in the collected audio. This is called noise reduction processing for the target user's audio, improving the user experience.
[0104] Currently, terminals can perform voiceprint noise reduction on environmental noise and interfering human voices. However, there are still many problems with current voiceprint noise reduction. For example, recognition anomalies caused by changes in the target user's voiceprint and incorrect learning during the voiceprint learning stage result in low accuracy of voiceprint noise reduction.
[0105] The following describes the reasons why voiceprint noise reduction is less accurate, based on the principles of audio noise reduction performed by the terminal.
[0106] Before voiceprint noise reduction, the target user's voiceprint features must be registered with the terminal. That is, the terminal saves the target user's voiceprint features as a template voiceprint feature. Accordingly, in the subsequent noise reduction phase, if the voiceprint features of the audio collected by the terminal match the saved template voiceprint features, it indicates that the collected audio is the target user's audio, and other audio except the target user's audio can be suppressed.
[0107] However, if Figure 4 As shown in , during the registration phase, in addition to the target user making a sound, the interfering user also makes a sound, and the terminal may mistakenly save the interfering user's voiceprint features as the template voiceprint features. Figure 5 As shown, if the target user's audio voiceprint characteristics do not match the template voiceprint characteristics, the terminal will perform noise reduction on the target user's audio. Alternatively, if the target user's voiceprint characteristics change, such as due to a cold, the target user's audio voiceprint characteristics will also not match the template voiceprint characteristics, and the terminal will also perform noise reduction on the target user's audio. Consequently, the terminal will output silent audio to the other party in the voice call, reducing the user experience.
[0108] The audio processing method provided in the embodiment of the present application, when the target user's to-be-processed audio and the template voiceprint features, as well as the to-be-processed facial image and the template facial features are successfully matched, indicates that the target user is a user making an audio call, that is, the target user's to-be-processed audio is not interference audio. In this case, interference audio other than the target user's to-be-processed audio is suppressed, which can improve the accuracy of suppressing the interference audio, achieve audio noise reduction, and improve the user experience.
[0109] Next, the audio processing method provided in the embodiment of the present application is described in detail through specific examples.
[0110] See also Figure 6 , Figure 6 This is a flowchart of an audio processing method provided in an embodiment of the present application. The method is applied to a terminal and includes the following steps:
[0111] S601: During an audio call between a first terminal and a second terminal, acquired audio to be processed and a facial image to be processed are obtained.
[0112] S602: For each template voiceprint feature, if the voiceprint feature of the audio to be processed matches the template voiceprint feature, determine that the user to whom the template voiceprint feature belongs is the user to whom the audio to be processed belongs.
[0113] S603: For each template facial feature, if the facial feature of the face image to be processed matches the template facial feature, determine that the user to which the template facial feature belongs is the user to which the face image to be processed belongs.
[0114] S604: If the user to which the audio to be processed belongs and the user to which the facial image to be processed belongs are the same target user, suppress the interference audio except for the audio to be processed of the target user to obtain the target audio, and send the target audio to the second terminal.
[0115] From the above, it can be seen that the technical solution provided by this embodiment, when the target user's to-be-processed audio and the template voiceprint features, as well as the to-be-processed facial image and the template facial features are successfully matched, indicates that the target user is a user making an audio call, that is, the target user's to-be-processed audio is not interference audio. In this case, the interference audio other than the target user's to-be-processed audio is suppressed, which can improve the accuracy of suppressing the interference audio, achieve audio noise reduction, and improve the user experience.
[0116] For step S601, when the user uses the first terminal to make an audio call with the second terminal, the microphone of the first terminal collects the user's audio (i.e., the audio to be processed), and the sensor of the first terminal collects the user's image (i.e., the facial image to be processed).
[0117] The sensor can be a camera of the first terminal, or an ultrasonic detector (such as a radar) of the first terminal. The first terminal can capture the user's image by shooting with a camera, or can capture the user's image by performing ultrasonic detection using an ultrasonic detector. The method of capturing the user's image is not limited in the embodiments of the present application. The following embodiments are described using the method of capturing the user's image by shooting with a camera as an example.
[0118] In one scenario, a user uses an instant messaging application installed in a first terminal to make a video call with a second terminal. During the call, the microphone of the first terminal collects the user's audio to be processed, and the camera of the first terminal collects the user's facial image to be processed.
[0119] In another scenario, a user uses an instant messaging app installed on a first terminal to conduct a voice call with a second terminal. Since the user does not use the video call function, the microphone of the first terminal collects the user's audio to be processed during the call. If the user authorizes the camera to capture images during the voice call, the camera of the first terminal collects the user's facial image to be processed.
[0120] In some embodiments, step S601 may include the following steps: obtaining audio captured by a microphone of the first terminal, performing audio preprocessing on the obtained audio to obtain original audio, extracting the vocal portion of the original audio to obtain audio to be processed. The audio preprocessing includes acoustic echo cancellation, automatic gain control, and noise suppression. Obtaining a video image captured by a camera containing a user image, performing image preprocessing on the obtained video image to obtain an original video image, and extracting a facial image from the original video image to obtain a facial image to be processed. The image preprocessing includes automatic exposure, automatic white balance, and autofocus.
[0121] The audio quality captured by the first terminal is low, potentially containing noise and other issues. Therefore, the first terminal preprocesses the captured audio using the 3A call algorithm to obtain the original audio. The first terminal then extracts the vocal portion of the original audio to obtain the processed audio, which is the sound emitted by the user during the audio call between the first terminal and the second terminal.
[0122] When there is only one user making a sound, the audio to be processed of that user is extracted. When there are multiple users making a sound, the audio to be processed of each user can be extracted separately, and each audio to be processed is processed according to the method provided in the embodiment of the present application.
[0123] The 3A call algorithm includes: Acoustic Echo Cancellation (AEC), Automatic Gain Control (AGC) and Noise Suppression (NS).
[0124] The quality of the video image collected by the first terminal can also be low, such as low clarity, etc. Therefore, the first terminal uses the 3A video algorithm to perform image preprocessing on the collected video image to obtain the original video image. The 3A video algorithm includes: Auto Exposure (AE), Auto White Balance (AWB) and Auto Focus (AF). Then, the facial image in the original video image is extracted to obtain the facial image to be processed. The facial image to be processed is the facial image of the user who is within the field of view of the camera of the first terminal and facing the camera of the first terminal during the audio call between the first terminal and the second terminal.
[0125] From the above, it can be seen that the technical solution provided in this embodiment performs audio preprocessing on the collected audio, which can improve the accuracy of the subsequently extracted voiceprint features, and performs video preprocessing on the collected video, which can improve the accuracy of the subsequently extracted facial features, further improve the accuracy of suppressing interfering audio, and improve the user experience.
[0126] Regarding step S602 and step S603, the template voiceprint feature and template face feature are stored locally on the first terminal during the registration phase. The registration process of the template face feature and template voiceprint feature is described in the subsequent embodiments.
[0127] In some embodiments, Figure 6 Based on Figure 7 Before step S602, the method may further include the following steps:
[0128] S605: For each audio to be processed, extract features of the audio to be processed to obtain a voiceprint feature of the audio to be processed.
[0129] S606: For each template voiceprint feature, if the difference between the voiceprint feature of the audio to be processed and the template voiceprint feature is less than a preset difference threshold, it is determined that the voiceprint feature of the audio to be processed matches the template voiceprint feature.
[0130] After obtaining the audio to be processed, the first terminal extracts the voiceprint features of the audio to be processed and compares the voiceprint features of the audio to be processed with the locally stored template voiceprint features. If the voiceprint features of the audio to be processed match any template voiceprint features, it is determined that the user to whom the template voiceprint features belong is the user to whom the audio to be processed belongs (which can be called the first user), that is, the audio to be processed is the sound emitted by the first user.
[0131] In some embodiments, voiceprint features may include: the fundamental frequency of the audio being processed, the pitch period of the audio being processed, the harmonic components of the audio being processed, and the Mel Frequency Cepstrum Coefficients (MFCCs) of the audio being processed. The fundamental frequency (i.e., pitch frequency) refers to the frequency at which the vocal cords open and close each time. The pitch period refers to the period of vocal cord vibration.
[0132] For each template voiceprint feature, if the difference between the voiceprint feature of the audio to be processed and the template voiceprint feature is less than the preset difference threshold, it indicates that the voiceprint feature of the audio to be processed is relatively similar to the template voiceprint feature, then it can be determined that the voiceprint feature of the audio to be processed matches the template voiceprint feature.
[0133] If the differences between the voiceprint features of the audio to be processed and the voiceprint features of each template are greater than the preset difference threshold, it indicates that the audio to be processed is not the voice of the registered user. It may be that the voiceprint features of the target user have changed, causing the audio to be processed to be identified as the audio of another user, such as the voiceprint features of the target user changing due to a cold, etc., then the preset difference threshold is increased from the first value to the second value.
[0134] The voiceprint features of the audio to be processed are again compared with the voiceprint features of each template. If the difference between the voiceprint features of the audio to be processed and any of the template voiceprint features is less than the adjusted preset difference threshold, the voiceprint features of the audio to be processed are determined to match the template voiceprint features, and the user to which the template voiceprint features that match the voiceprint features of the audio to be processed belong is determined to be the user to which the audio to be processed belongs (i.e., the first user). In other words, the user to which the template voiceprint features are matched is determined to be the user who made the sound. Furthermore, if the first user to which the audio to be processed belongs and the second user to which the facial image to be processed belongs are the same target user, indicating that the user making the sound is the same user who is within the camera field of view of the first terminal and facing the camera of the first terminal, i.e., the target user is using the first terminal to conduct an audio call with a user using the second terminal. In this case, all other audio in the original audio except the audio to be processed by the target user is interference audio. The first terminal then suppresses the interference audio except the audio to be processed by the target user to obtain the target audio, which is noise-reduced and of higher quality. The first terminal then sends the target audio to the second terminal.
[0135] After the first terminal obtains the facial image to be processed, since the facial image to be processed is the facial image of a user who was within the camera field of view of the first terminal during the audio call between the first terminal and the second terminal, the first terminal extracts facial features from the facial image to be processed and compares the facial image to be processed with locally stored template facial features. If the facial features of the facial image to be processed match any of the template facial features, it is determined that the user to which the template facial features belong is the user to which the facial image to be processed belongs (which can be referred to as the second user), that is, the second user is the facial image of the user who was within the camera field of view of the first terminal during the audio call between the first terminal and the second terminal. Facial features can include key points in the facial image, such as the 68 key points obtained by face detection.
[0136] Furthermore, it is determined whether the first user to whom the audio to be processed belongs and the second user to whom the image to be processed belongs are the same user (i.e., the target user). That is, it is detected whether, during the audio call between the first terminal and the second terminal, the user who makes the sound and the user who is within the camera field of view of the first terminal and facing the camera of the first terminal are the same user.
[0137] In some embodiments, Figure 6 Based on Figure 8 Before step S604, the method may further include the following steps:
[0138] S607: Perform lip movement detection on the face image to be processed.
[0139] S608: If lip movement of the user in the face image to be processed is detected, it is detected whether the user to whom the audio to be processed belongs and the user to whom the face image to be processed belongs are the same target user.
[0140] The first terminal performs lip movement detection on the facial image to be processed. For example, the facial image to be processed is input into a pre-trained lip movement detection model to obtain a detection result of whether the user in the facial image to be processed has made lip movements. If the user in the facial image to be processed has not made lip movements, it indicates that the user within the field of view of the camera of the first terminal has not made any sound. There is no need to further detect whether the first and second users are the same target user, and the audio to be processed is directly suppressed.
[0141] If the user in the face image to be processed has lip movements, indicating that the user within the camera field of view of the first terminal makes a sound, then continue to detect whether the user who makes the sound is the same user within the camera field of view of the first terminal, so as to avoid mistakenly identifying the sound made by other users as the sound made by the target user, which leads to audio suppression errors, further improve the accuracy of suppressing interfering audio, achieve audio noise reduction, and improve user experience.
[0142] For step S604, if the first user to whom the audio to be processed belongs and the second user to whom the facial image to be processed belongs are the same target user, it indicates that the user emitting the sound and the user within the camera field of view of the first terminal are the same target user, that is, the target user uses the first terminal to conduct an audio call with the user using the second terminal. In this case, all the audio in the original audio except the audio to be processed of the target user is interference audio. The first terminal suppresses the interference audio except the audio to be processed of the target user to obtain the target audio. The target audio is noise-reduced and of higher quality. The first terminal then sends the target audio to the second terminal. Interference audio includes: ambient noise and sounds emitted by users other than the target user.
[0143] In some embodiments, Figure 7 Based on Figure 9 After detecting whether the user to which the audio to be processed belongs and the user to which the facial image to be processed belongs are the same user, the method may further include the following steps:
[0144] S609: If the user to which the audio to be processed belongs and the user to which the facial image to be processed belong are not the same target user, the preset difference threshold is increased from the first value to the second value.
[0145] S610: For each template voiceprint feature, if the difference between the voiceprint feature of the audio to be processed and the template voiceprint feature is less than the adjusted preset difference threshold, it is determined that the voiceprint feature of the audio to be processed matches the template voiceprint feature, and the user to whom the template voiceprint feature that matches the voiceprint feature of the audio to be processed belongs is determined to be the user to whom the audio to be processed belongs.
[0146] S611: If the user to which the audio to be processed belongs and the user to which the facial image to be processed belongs are the same target user, suppress the interference audio except for the audio to be processed of the target user, obtain the target audio, and send the target audio to the second terminal.
[0147] If the first user to whom the audio to be processed belongs and the second user to whom the facial image to be processed belongs are not the same target user, it indicates that the user who makes the sound and the user within the camera field of view of the first terminal are not the same target user. It may be that the voiceprint characteristics of the target user have changed, causing the audio to be processed to be identified as audio of another user, such as the voiceprint characteristics of the target user changing due to a cold, etc., then the preset difference threshold is increased from the first value to the second value.
[0148] The voiceprint features of the audio to be processed are again compared with the voiceprint features of each template. If the difference between the voiceprint features of the audio to be processed and any of the template voiceprint features is less than the adjusted preset difference threshold, the voiceprint features of the audio to be processed are determined to match the template voiceprint features, and the user to whom the template voiceprint features that match the voiceprint features of the audio to be processed belong is determined to be the user to whom the audio to be processed belongs (i.e., the first user). In other words, the user to whom the template voiceprint features are matched is determined to be the user who made the sound. Furthermore, if the first user to whom the audio to be processed belongs and the second user to whom the facial image to be processed belongs are the same target user, indicating that the user making the sound is the same user within the camera field of view of the first terminal, i.e., the target user is using the first terminal to conduct an audio call with a user using the second terminal. In this case, all other audio in the original audio except the audio to be processed of the target user is interference audio. The first terminal then suppresses the interference audio except the audio to be processed of the target user to obtain the target audio, which is noise-reduced and of higher quality. The first terminal then sends the target audio to the second terminal.
[0149] The first and second values of the preset difference threshold can be determined by a technician based on experimental results. For example, the unchanged voiceprint features and the changed voiceprint features of the same user can be collected, and the two collected voiceprint features can be compared to determine the preset difference threshold at which the two collected voiceprint features can be identified as the voiceprint features of the same user.
[0150] As can be seen from the above, in the technical solution provided by this embodiment, by adjusting the preset difference threshold, the problem of recognition errors caused by changes in the target user's voiceprint characteristics is avoided, the accuracy of suppressing interfering audio is further improved, audio noise reduction is achieved, and the user experience is improved.
[0151] In some embodiments, if the difference between the voiceprint features of the audio to be processed and the voiceprint features of each template is greater than the adjusted preset difference threshold, it indicates that the audio to be processed is not the sound emitted by a registered user, nor is the change in the voiceprint features of the target user causing the audio to be processed to be identified as audio of other users. The first terminal determines that the audio to be processed is interference audio.
[0152] As can be seen from the above, the technical solution provided by this embodiment determines that the audio to be processed whose voiceprint features do not match the voiceprint features of each template is interference audio, can accurately identify the interference audio, and suppress the interference audio, further improve the accuracy of suppressing the interference audio, achieve audio noise reduction, and improve the user experience.
[0153] In some embodiments, if the first user to whom the audio to be processed is determined again and the second user to whom the facial image to be processed belongs are still not the same user, it indicates that the change in the voiceprint characteristics of the target user causes the audio to be processed to be identified as audio of another user, and the first terminal determines that the audio to be processed is interference audio.
[0154] As can be seen from the above, the technical solution provided by this embodiment determines that the audio to be processed whose facial image features do not match the voiceprint features is interference audio, can accurately identify the interference audio, and suppress the interference audio, further improving the accuracy of suppressing the interference audio, realizing audio noise reduction, and improving the user experience.
[0155] The following combination Figure 10 、 Figure 11 and Figure 12 The following describes the data flow when performing audio noise reduction. Figure 10 and Figure 11 For the process simultaneously executed by the first terminal, in order to clearly explain the audio data stream and the image data stream, Figure 10 and Figure 11 Introduce them separately.
[0156] See also Figure 10 , Figure 10 A flowchart for obtaining audio data is provided in an embodiment of the present application.
[0157] S1001: APP initiates an audio or video call.
[0158] In this step, the APP is an application in the application layer that provides the user with an audio call function, such as an instant messaging application, a conference application, a social application, etc. The APP initiates an audio and video call after receiving the user's instruction.
[0159] S1002: APP calls audio service.
[0160] In this step, the application framework layer includes various services of the first terminal, such as audio service (ie, audio FWR), camera service, sensor service, etc. When the APP initiates an audio or video call according to the user's instructions, it calls the audio FWR in the application framework layer.
[0161] S1003: The audio FWR calls the audio HAL to start recording.
[0162] In this step, the driver layer includes drivers for various hardware components in the first terminal, such as the audio driver (i.e., audio HAL), camera driver, and sensor driver. When the app initiates an audio or video call according to the user's instructions, it calls the audio service in the application framework layer. Upon receiving the app's recording request, the audio FWR calls the audio HAL to initiate the recording function.
[0163] S1004: The audio HAL activates the DSP.
[0164] In this step, the audio HAL detects the call of the audio FWR and activates the DSP.
[0165] S1005: The audio HAL calls the microphone to record.
[0166] In this step, after the audio HAL detects the call of the audio FWR, it calls the microphone to record to start the recording service.
[0167] S1006: DSP loads the audio preprocessing algorithm.
[0168] In this step, after the audio HAL activates the DSP, the DSP loads the audio pre-processing algorithm (ie, acoustic echo cancellation, automatic gain control, and noise suppression in the above embodiment).
[0169] In the embodiment of the present application, the execution order of the above-mentioned step S1005 and step S1006 is not limited, and step S1005 and step S1006 can be executed at the same time.
[0170] S1007: The microphone transmits the collected audio to the DSP.
[0171] In this step, after the microphone is started, it starts recording and transmits the recorded audio to the DSP. Figure 12 In step ①, the microphone transmits the audio data to the DSP.
[0172] S1008: DSP performs audio preprocessing.
[0173] In this step, the DSP pre-processes the audio transmitted by the microphone according to the loaded audio pre-processing algorithm to obtain the pre-processed audio (i.e. the original audio in the above embodiment), i.e. Figure 12 In the call, DSP uses the 3A call algorithm to process the audio data transmitted by the microphone.
[0174] S1009: The DSP transmits the pre-processed audio to the audio HAL.
[0175] In this step, DSP transmits the pre-processed audio to the audio HAL, i.e. Figure 12In step ②, the DSP transmits the audio processed by the 3A call algorithm to the audio HAL.
[0176] S1010: The audio HAL performs voiceprint noise reduction on the pre-processed audio.
[0177] In this step, the audio HAL extracts the human voice portion of the preprocessed audio (i.e., the audio to be processed) and its voiceprint features. The voiceprint features of the audio to be processed are compared with the locally stored template voiceprint features. The user whose template voiceprint features match the voiceprint features of the audio to be processed is determined, thereby obtaining the user to whom the audio to be processed belongs.
[0178] Furthermore, when the user to whom the audio to be processed belongs and the user to whom the facial image to be processed belongs are the same target user, the audio HAL suppresses the interference audio in the preprocessed audio except for the audio to be processed of the target user to obtain the noise-reduced audio (i.e., the target audio).
[0179] That is Figure 12 In step ③, the audio HAL compares the voiceprint features of the audio to be processed with the template voiceprint features stored in the secure OS (i.e., the preset storage area). The secure OS stores N voiceprint features, namely, the voiceprint of user 1, user 2, and so on, in the voiceprint management. If the voiceprint of user 1 matches the voiceprint features of the audio to be processed, user 1 is determined to be the user to whom the audio to be processed belongs. Furthermore, if the user to whom the facial image to be processed belongs is also user 1, the audio HAL suppresses the interfering audio in the preprocessed audio except for the audio of user 1, obtaining the noise-reduced audio (i.e., the target audio).
[0180] S1011: The audio HAL transmits the noise-reduced audio to the audio FWR.
[0181] In this step, the audio HAL transmits the noise-reduced audio to the audio FWR, i.e. Figure 12 In step ④, the audio HAL transmits the noise-reduced audio to the audio FWR.
[0182] S1012: The audio FWR performs audio encoding on the noise-reduced audio.
[0183] In this step, the audio FWR encodes the noise-reduced audio transmitted by the audio HAL, that is, Figure 12 In the process, the audio FWR performs audio encoding on the noise-reduced audio transmitted by the audio HAL to obtain a recording file (AudioRecord).
[0184] S1013: The audio FWR transmits the encoded audio to the APP.
[0185] In this step, the audio FWR transmits the encoded audio to the APP, that is, Figure 12 In step ⑤, the audio FWR transmits the AudioRecord to the APP, and the APP subsequently uses the AudioRecord to make an audio call with the second terminal.
[0186] See also Figure 11 , Figure 11 A flowchart for obtaining image data is provided in an embodiment of the present application.
[0187] S1101: The APP initiates an audio or video call.
[0188] In this step, the APP is an application in the application layer that provides audio and video call functions for users, such as instant messaging applications, conference applications, social applications, etc. The APP initiates the audio and video call after receiving the user's instruction.
[0189] S1102: APP calls camera service.
[0190] In this step, the application framework layer includes various services of the first terminal, such as audio service, camera service (ie, camera FWR), sensor service, etc. When the APP initiates an audio or video call according to the user's instructions, it calls the audio FWR in the application framework layer.
[0191] S1103: The camera FWR calls the camera HAL to start shooting.
[0192] In this step, the driver layer includes drivers for various hardware components in the first terminal, such as the audio driver, camera driver (i.e., camera HAL), and sensor drivers. When the app initiates an audio or video call according to the user's instructions, it calls the camera FWR in the application framework layer. Upon receiving the app's capture request, the camera FWR calls the camera HAL to initiate the capture function.
[0193] S1104: The camera HAL activates the Compute Digital Signal Processing (CDSP).
[0194] In this step, the camera HAL detects the call of the camera FWR and activates the CDSP.
[0195] S1105: The camera HAL calls the camera to shoot.
[0196] In this step, after the camera HAL detects the call of the camera FWR, it calls the camera to shoot to start the video shooting service.
[0197] S1106: The CDSP loads the video preprocessing algorithm.
[0198] In this step, after the camera HAL activates the CDSP, the CDSP loads the video pre-processing algorithm (ie, the automatic exposure, automatic white balance, and automatic focus in the above embodiment).
[0199] In the embodiment of the present application, the execution order of the above-mentioned step S1105 and step S1106 is not limited, and step S1105 and step S1106 can be executed at the same time.
[0200] S1107: The camera transmits the captured video to the DSP.
[0201] In this step, after the camera is started, it starts shooting and transmits the captured video to the CDSP. Figure 12 In step ①, the camera transmits the image data to the CDSP.
[0202] S1108: CDSP performs video preprocessing.
[0203] In this step, the CDSP pre-processes the image transmitted by the camera according to the loaded video pre-processing algorithm to obtain the pre-processed video image (i.e. the original video image in the aforementioned embodiment), i.e. Figure 12 In the image processing, CDSP performs image processing.
[0204] S1109: CDSP transmits the pre-processed video to the camera HAL.
[0205] In this step, CDSP transmits the pre-processed video to the camera HAL, i.e. Figure 12 In step ②, CDSP transmits the video after image processing to the camera HAL.
[0206] S1110: The camera HAL performs face recognition and lip movement detection on the pre-processed video.
[0207] In this step, the camera HAL extracts a facial image from the preprocessed video (i.e., the face image to be processed) and performs facial recognition on it. Specifically, facial features are extracted from the face image to be processed. The facial features of the face image to be processed are compared with the facial features of a locally stored template. The user whose template facial features match the facial features of the face image to be processed is determined, thereby obtaining the user to whom the face image to be processed belongs.
[0208] Furthermore, the camera HAL performs lip movement detection on the face image to be processed. When lip movement of the user in the face image to be processed is detected, it is detected whether the user to whom the audio to be processed belongs and the user to whom the face image to be processed belongs are the same target user.
[0209] That is Figure 12In step ③, the camera HAL compares the facial features of the facial image to be processed with the template facial features stored in the secure OS. The secure OS stores N facial features, namely the facial features of user 1, user 2, and so on, in the face recognition management. If the facial features of user 1 match the facial features of the facial image to be processed, user 1 is determined to be the user to whom the facial image to be processed belongs. Furthermore, when the user to whom the facial image to be processed belongs and the user to whom the audio to be processed belongs are both user 1, the audio HAL suppresses interfering audio in the preprocessed audio, excluding the audio of user 1, to obtain the noise-reduced audio (i.e., the target audio).
[0210] S1111: The camera HAL transmits the processed video to the camera FWR.
[0211] In this step, after the face recognition obtains the user to whom the face image to be processed belongs, the camera HAL transmits the original video image to the camera FWR, i.e. Figure 12 In step ④, the camera HAL transmits the raw video image to the camera FWR.
[0212] S1112: The camera FWR performs video encoding on the original video image.
[0213] In this step, the camera FWR encodes the original video image transmitted by the camera HAL, that is, Figure 12 In the process, the camera FWR performs video encoding on the original video image transmitted by the camera HAL to obtain a video file.
[0214] S1113: The camera FWR transmits the encoded video to the APP.
[0215] In this step, the camera FWR transmits the encoded video to the APP, which is Figure 12 In step ⑤, the camera FWR transmits the video file to the APP, and the APP subsequently uses the video file to make a video call with the second terminal.
[0216] From the above, it can be seen that the technical solution provided by this embodiment, when the target user's to-be-processed audio and the template voiceprint features, as well as the to-be-processed facial image and the template facial features are successfully matched, indicates that the target user is a user making an audio call, that is, the target user's to-be-processed audio is not interference audio. In this case, the interference audio other than the target user's to-be-processed audio is suppressed, which can improve the accuracy of suppressing the interference audio, achieve audio noise reduction, and improve the user experience.
[0217] The following describes how to register template voiceprint features and template facial features on the first terminal. Figure 13 , Figure 13A flowchart of a registration template voiceprint feature and template face feature is provided in an embodiment of the present application. The method includes the following steps:
[0218] S1301: Acquire the collected audio and facial image of the user to be registered.
[0219] S1302: Extract features of the audio to be registered to obtain template voiceprint features of the audio to be registered, and extract features of the face image to be registered to obtain template face features of the face image to be registered.
[0220] S1303: Perform lip movement detection on the face image to be registered.
[0221] S1304: If lip movement is detected in the face image of the user to be registered, the template voiceprint features of the audio to be registered, the template facial features of the face image to be registered, and the corresponding user ID of the user to be registered are recorded in a preset storage area locally on the first terminal.
[0222] As can be seen from the above, the technical solution provided by this embodiment can record the template voiceprint features of the audio to be registered, the template facial features of the facial image to be registered, and the corresponding user ID of the user to be registered. Subsequently, if the target user's to-be-processed audio successfully matches the template voiceprint features, and the to-be-processed facial image successfully matches the template facial features, indicating that the target user is the user making an audio call, that is, the target user's to-be-processed audio is not interference audio. In this case, interference audio other than the target user's to-be-processed audio is suppressed, which can improve the accuracy of interference audio suppression, achieve audio noise reduction, and enhance the user experience.
[0223] During the registration phase, the first terminal acquires audio captured by a microphone, performs audio preprocessing on the acquired audio, extracts the vocal portion of the preprocessed audio, and obtains the audio to be registered of the user to be registered. For details on the audio preprocessing method, refer to the relevant description of the preceding embodiment. Furthermore, the first terminal acquires a user image of the user to be registered captured by a camera, performs video preprocessing on the acquired user image, and extracts the facial image of the preprocessed user image to obtain the facial image to be registered. For details on the video preprocessing method, refer to the relevant description of the preceding embodiment.
[0224] In one application scenario, when a first terminal is not in an audio or video call with another terminal, the first terminal displays preset text according to a user's instructions, prompting the user to read the preset text aloud. Accordingly, the first terminal obtains the sound produced by the user while reading the preset text, extracts the vocal portion, and obtains the audio to be registered. Furthermore, the first terminal obtains an image of the user while reading the preset text, extracts the facial image, and obtains the facial image to be registered.
[0225] In another application scenario, during an audio or video call between a first terminal and another terminal, the first terminal obtains audio captured during a preset time period and extracts the human voice portion to obtain the audio to be registered. It also obtains user images captured during the preset time period and extracts the facial image to obtain the facial image to be registered. The length of the preset time period is set as needed.
[0226] Furthermore, the first terminal performs feature extraction on the audio to be registered to obtain a template voiceprint feature of the audio to be registered, and performs feature extraction on the facial image to be registered to obtain a template facial feature of the facial image to be registered. Furthermore, the first terminal performs lip movement detection on the facial image to be processed. If lip movement is detected in the facial image to be registered, indicating that the user within the field of view of the terminal camera is making a sound, that is, the user within the field of view of the terminal camera is performing voiceprint registration and facial image registration, the first terminal then binds the template voiceprint feature and template facial feature of the user to be registered, that is, records the template voiceprint feature of the audio to be registered, the template facial feature of the facial image to be registered, and the corresponding user ID of the user to be registered in a local preset storage area. The user ID of the user to be registered can be a unique identifier such as a user name or user number.
[0227] The preset storage area is a local secure storage area of the first terminal, such as a secure OS. The secure OS is a secure storage area in the operating system of the first terminal. Developers can access this secure storage area through a designated encrypted interface. Other personnel cannot access this secure storage area, which can improve the security of the registered template voiceprint features and template facial features and protect the user's privacy.
[0228] The following combination Figure 14 、 Figure 15 and Figure 16 The data flow when performing audio registration is described. Figure 14 and Figure 15 For the process simultaneously executed by the first terminal, in order to clearly explain the audio data stream and the image data stream, Figure 14 and Figure 15 Introduce them separately.
[0229] See also Figure 14 , Figure 14 A flowchart of obtaining template voiceprint features provided in an embodiment of the present application.
[0230] S1401: Receive user instructions.
[0231] In this step, the APP may be a setting application. When the first terminal is not in an audio or video call with another terminal, the setting application of the first terminal displays a preset text according to the user's instruction, that is, the setting application receives an instruction to collect audio.
[0232] Alternatively, the APP is an application in the application layer that provides the user with an audio call function, such as an instant messaging application, a conference application, a social application, etc. The APP initiates an audio or video call after receiving an instruction from the user.
[0233] S1402: APP calls audio service.
[0234] In this step, the application framework layer includes various services of the first terminal, such as audio service (ie, audio FWR), camera service, sensor service, etc. After receiving the instruction to collect audio, the APP calls the audio FWR in the application framework layer.
[0235] S1403: The audio FWR calls the audio HAL to start recording.
[0236] In this step, the driver layer includes drivers for various hardware components in the first terminal, such as the audio driver (i.e., audio HAL), camera driver, and sensor driver. When the app initiates an audio or video call according to the user's instructions, it calls the audio service in the application framework layer. Upon receiving the app's recording request, the audio FWR calls the audio HAL to initiate the recording function.
[0237] S1404: The audio HAL activates the DSP.
[0238] In this step, the audio HAL detects the call of the audio FWR and activates the DSP.
[0239] S1405: The audio HAL calls the microphone to record.
[0240] In this step, after the audio HAL detects the call of the audio FWR, it calls the microphone to record to start the recording service.
[0241] S1406: The DSP loads the audio preprocessing algorithm.
[0242] In this step, after the audio HAL activates the DSP, the DSP loads the audio pre-processing algorithm (ie, acoustic echo cancellation, automatic gain control, and noise suppression in the above embodiment).
[0243] In the embodiment of the present application, the execution order of the above-mentioned step S1405 and step S1406 is not limited, and step S1405 and step S1406 can be executed at the same time.
[0244] S1407: The microphone transmits the collected audio to the DSP.
[0245] In this step, after the microphone is started, it starts recording and transmits the recorded audio to the DSP. Figure 16 In step ①, the microphone transmits the audio data to the DSP.
[0246] S1408: DSP performs audio preprocessing.
[0247] In this step, the DSP preprocesses the audio transmitted by the microphone according to the loaded audio preprocessing algorithm to obtain the preprocessed audio, i.e. Figure 16 In the call, DSP uses the 3A call algorithm to process the audio data transmitted by the microphone.
[0248] S1409: The DSP transmits the pre-processed audio to the audio HAL.
[0249] In this step, DSP transmits the pre-processed audio to the audio HAL, i.e. Figure 16 In step ②, the DSP transmits the audio processed by the 3A call algorithm to the audio HAL.
[0250] S1410: The audio HAL extracts voiceprint features from the preprocessed audio.
[0251] In this step, the audio HAL extracts the human voice part of the pre-processed audio (ie, the audio to be registered), and extracts the voiceprint features of the audio to be registered.
[0252] S1411: The audio HAL saves the voiceprint feature as a template voiceprint feature.
[0253] In this step, when lip movements of the user to be registered are detected in the face image to be registered of the user to be registered, the audio HAL saves the voiceprint features of the audio to be registered as template voiceprint features, that is, records the template voiceprint features of the audio to be registered, the template face features of the face image to be registered, and the corresponding user identification of the user to be registered in the local preset storage area.
[0254] That is Figure 16 In step ③, the audio HAL saves the voiceprint features of the audio to be registered as template voiceprint features in the secure OS. The secure OS stores N voiceprint features, namely the voiceprint of user 1, user 2, ..., user N in the voiceprint management.
[0255] In one application scenario, when the APP is a setting application, the registration process ends after recording the template voiceprint features of the audio to be registered, the template facial features of the facial image to be registered, and the corresponding user identification of the user to be registered in the local preset storage area.
[0256] In another application scenario, when the APP is an application in the application layer that provides the user with an audio call function, after recording the template voiceprint features of the audio to be registered, the template facial features of the facial image to be registered, and the corresponding user ID of the user to be registered in the local preset storage area, the first terminal executes the subsequent audio noise reduction process. Audio noise reduction process reference Figure 10 、 Figure 11 and Figure 12 Related introduction.
[0257] Figure 16 The processing of audio data stream in step ④ and step ⑤ is the same as Figure 12 The processing of the audio data stream in step ④ and step ⑤ is similar, and reference is made to the relevant introduction of the aforementioned embodiment.
[0258] See also Figure 15 , Figure 15 A flowchart for obtaining template facial features is provided in an embodiment of the present application.
[0259] S1501: Receive user instructions.
[0260] In this step, the APP may be a setting application. When the first terminal is not in an audio or video call with another terminal, the setting application of the first terminal displays a preset text according to the user's instruction, that is, the setting application receives an instruction to collect audio.
[0261] Alternatively, the APP is an application in the application layer that provides the user with an audio call function, such as an instant messaging application, a conference application, a social application, etc. The APP initiates an audio or video call after receiving an instruction from the user.
[0262] S1502: APP calls camera service.
[0263] In this step, the application framework layer includes various services of the first terminal, such as audio service, camera service (ie, camera FWR), sensor service, etc. When the APP initiates an audio or video call according to the user's instructions, it calls the audio FWR in the application framework layer.
[0264] S1503: The camera FWR calls the camera HAL to start shooting.
[0265] In this step, the driver layer includes drivers for various hardware components in the first terminal, such as the audio driver, camera driver (i.e., camera HAL), and sensor drivers. When the app initiates an audio or video call according to the user's instructions, it calls the camera FWR in the application framework layer. Upon receiving the app's capture request, the camera FWR calls the camera HAL to initiate the capture function.
[0266] S1504: The camera HAL activates the CDSP.
[0267] In this step, the camera HAL detects the call of the camera FWR and activates the CDSP.
[0268] S1505: The camera HAL calls the camera to shoot.
[0269] In this step, after the camera HAL detects the call of the camera FWR, it calls the camera to shoot to start the video shooting service.
[0270] S1506: The CDSP loads the video preprocessing algorithm.
[0271] In this step, after the camera HAL activates the CDSP, the CDSP loads the video pre-processing algorithm (ie, the automatic exposure, automatic white balance, and automatic focus in the above embodiment).
[0272] In the embodiment of the present application, the execution order of the above-mentioned step S1505 and step S1506 is not limited, and step S1505 and step S1506 can be executed at the same time.
[0273] S1507: The camera transmits the captured video to the DSP.
[0274] In this step, after the camera is started, it starts shooting and transmits the captured video to the CDSP. Figure 16 In step ①, the camera transmits the image data to the CDSP.
[0275] S1508: CDSP performs video preprocessing.
[0276] In this step, CDSP preprocesses the image transmitted by the camera according to the loaded video preprocessing algorithm to obtain the preprocessed video image, that is, Figure 16 In the image processing, CDSP performs image processing.
[0277] S1509: The CDSP transmits the pre-processed video to the camera HAL.
[0278] In this step, CDSP transmits the pre-processed video to the camera HAL, i.e. Figure 16 In step ②, CDSP transmits the video after image processing to the camera HAL.
[0279] S1510: The camera HAL performs face detection, feature extraction, and lip movement detection on the preprocessed video.
[0280] In this step, the camera HAL performs face detection on the preprocessed video image and extracts the face image (i.e., the face image to be registered) from the preprocessed video image. It then extracts facial features from the face image to be registered. Furthermore, the camera HAL performs lip movement detection on the face image to be registered.
[0281] S1511: The camera HAL saves the facial features as template facial features.
[0282] In this step, when the camera HAL detects lip movement of the user to be registered in the face image to be registered, the camera HAL saves the facial features as template facial features.
[0283] That is Figure 16 In step 3, the camera HAL saves the facial features of the face image to be registered as template facial features in the secure OS. The secure OS stores N facial features, namely, the facial features of user 1, user 2, and so on, in the face recognition management.
[0284] In one application scenario, when the APP is a setting application, the registration process ends after recording the template voiceprint features of the audio to be registered, the template facial features of the facial image to be registered, and the corresponding user identification of the user to be registered in the local preset storage area.
[0285] In another application scenario, when the APP is an application in the application layer that provides the user with an audio call function, after recording the template voiceprint features of the audio to be registered, the template facial features of the facial image to be registered, and the corresponding user ID of the user to be registered in the local preset storage area, the first terminal executes the subsequent audio noise reduction process. Audio noise reduction process reference Figure 10 、 Figure 11 and Figure 12 Related introduction.
[0286] Figure 16 The processing of audio data stream in step ④ and step ⑤ is the same as Figure 12 The processing of the audio data stream in step ④ and step ⑤ is similar, and reference is made to the relevant introduction of the aforementioned embodiment.
[0287] As can be seen from the above, the technical solution provided by this embodiment can record the template voiceprint features of the audio to be registered, the template facial features of the facial image to be registered, and the corresponding user ID of the user to be registered. Subsequently, if the target user's to-be-processed audio successfully matches the template voiceprint features, and the to-be-processed facial image successfully matches the template facial features, indicating that the target user is the user making an audio call, that is, the target user's to-be-processed audio is not interference audio. In this case, interference audio other than the target user's to-be-processed audio is suppressed, which can improve the accuracy of interference audio suppression, achieve audio noise reduction, and enhance the user experience.
[0288] In a specific implementation, the present application also provides a terminal, which includes one or more processors and a memory; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code includes computer instructions, and the one or more processors call the computer instructions to enable the terminal to execute some or all of the steps in the above method embodiments.
[0289] The present application also provides a computer-readable storage medium including a computer program. When the computer program is executed on a terminal, the terminal executes some or all of the steps in the above method embodiment. The above storage medium may be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).
[0290] In a specific implementation, an embodiment of the present application further provides a computer program product, which includes executable instructions. When the executable instructions are executed on a terminal, the terminal executes some or all of the steps in the above method embodiment.
[0291] like Figure 17 As shown, the present application also provides a chip system, which is applied to a terminal. The chip system includes one or more processors 1701. The processor 1701 is used to call computer instructions so that the terminal inputs the data to be processed into the chip system. The chip system processes the data based on the audio processing method provided in the embodiment of the present application and outputs the processing results.
[0292] In one possible implementation, the chip system further includes input and output interfaces for inputting and outputting data.
[0293] The various embodiments of the mechanisms disclosed in this application can be implemented in hardware, software, firmware, or a combination of these implementation methods. The embodiments of the present application can be implemented as a computer program or program code executed on a programmable system, which includes at least one processor, a storage system (including volatile and non-volatile memory and / or storage elements), at least one input device, and at least one output device.
[0294] Program code can be applied to input instructions to perform the functions described herein and generate output information. The output information can be applied to one or more output devices in a known manner. For the purposes of this application, a processing system includes any system having a processor such as, for example, a digital signal processor (DSP), a microcontroller, an application specific integrated circuit (ASIC), or a microprocessor.
[0295] Program code can be implemented with a high-level programming language or an object-oriented programming language to communicate with the processing system. Where necessary, program code can also be implemented in assembly language or machine language. In fact, the mechanism described in this application is not limited to the scope of any particular programming language. In either case, the language can be a compiled language or an interpreted language.
[0296] In some cases, the disclosed embodiments can be implemented in hardware, firmware, software or any combination thereof. The disclosed embodiments can also be implemented as instructions carried or stored on one or more temporary or non-temporary machine-readable (e.g., computer-readable) storage media, which can be read and executed by one or more processors. For example, instructions can be distributed over a network or by other computer-readable media. Therefore, machine-readable media can include any mechanism for storing or transmitting information in a machine (e.g., computer) readable form, including but not limited to, floppy disks, optical disks, optical disks, compact disc read-only memories (Compact Disc Read Only Memory, CD-ROMs), magneto-optical disks, read-only memories, random access memories, erasable programmable read-only memories (Erasable Programmable Read Only Memory, EPROM), electrically erasable programmable read-only memories (Electrically Erasable Programmable Read Only Memory, EEPROM), magnetic cards or optical cards, flash memory, or tangible machine-readable memories for transmitting information (e.g., carrier waves, infrared signal digital signals, etc.) using the Internet in electrical, optical, acoustic or other forms of propagation signals. Accordingly, machine-readable media includes any type of machine-readable media suitable for storing or transmitting electronic instructions or information in a form readable by a machine (eg, a computer).
[0297] In the accompanying drawings, some structural or method features may be shown in a particular arrangement and / or order. However, it should be understood that such a particular arrangement and / or order may not be required. Rather, in some embodiments, these features may be arranged in a manner and / or order different from that shown in the accompanying drawings. In addition, the inclusion of a structural or method feature in a particular figure does not imply that such a feature is required in all embodiments, and in some embodiments, such features may not be included or may be combined with other features.
[0298] It should be noted that the units / modules mentioned in the various device embodiments of the present application are all logical units / modules. Physically, a logical unit / module can be a physical unit / module, or a part of a physical unit / module, or can be implemented as a combination of multiple physical units / modules. The physical implementation of these logical units / modules themselves is not the most important. The combination of functions implemented by these logical units / modules is the key to solving the technical problems raised by this application. In addition, in order to highlight the innovative part of this application, the above-mentioned device embodiments of this application do not introduce units / modules that are not closely related to solving the technical problems raised by this application. This does not mean that other units / modules do not exist in the above-mentioned device embodiments.
[0299] It should be noted that in the examples and description of this patent, relational terms such as first and second, etc. are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device that includes a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "including a" does not exclude the presence of other identical elements in the process, method, article or device that includes the element.
[0300] Although the present application has been shown and described with reference to certain preferred embodiments thereof, it will be understood by those skilled in the art that various changes in form and details may be made therein without departing from the spirit and scope of the application.
Claims
1. An audio processing method, characterized in that: The method is applied to a first terminal, and includes: During an audio call between the first terminal and the second terminal, obtaining collected audio to be processed and a facial image to be processed; For each template voiceprint feature, if the voiceprint feature of the audio to be processed matches the template voiceprint feature, the user to which the template voiceprint feature belongs is determined to be the user to which the audio to be processed belongs; For each template facial feature, if the facial feature of the face image to be processed matches the template facial feature, determining that the user to which the template facial feature belongs is the user to which the face image to be processed belongs; If the user to which the audio to be processed belongs and the user to which the facial image to be processed belongs are the same target user, interference audio other than the audio to be processed of the target user is suppressed to obtain target audio, and the target audio is sent to the second terminal.
2. The method according to claim 1, characterized in that For each template voiceprint feature, if the voiceprint feature of the audio to be processed matches the template voiceprint feature, before determining that the user to which the template voiceprint feature belongs is the user to which the audio to be processed belongs, the method further includes: For each audio to be processed, extract the features of the audio to be processed to obtain the voiceprint features of the audio to be processed; For each template voiceprint feature, if the difference between the voiceprint feature of the audio to be processed and the template voiceprint feature is less than a preset difference threshold, it is determined that the voiceprint feature of the audio to be processed matches the template voiceprint feature.
3. The method according to claim 2, characterized in that If the difference between the voiceprint feature of the audio to be processed and the voiceprint feature of the template is less than a preset difference threshold, after determining that the voiceprint feature of the audio to be processed matches the voiceprint feature of the template, the method further includes: If the user to which the audio to be processed belongs and the user to which the facial image to be processed belongs are not the same target user, increasing the preset difference threshold from the first value to the second value; For each template voiceprint feature, if the difference between the voiceprint feature of the audio to be processed and the template voiceprint feature is less than the adjusted difference threshold, it is determined that the voiceprint feature of the audio to be processed matches the template voiceprint feature, and the user to whom the template voiceprint feature that matches the voiceprint feature of the audio to be processed belongs is determined to be the user to whom the audio to be processed belongs; If the user to which the audio to be processed belongs and the user to which the facial image to be processed belongs are the same target user, interference audio other than the audio to be processed of the target user is suppressed to obtain target audio, and the target audio is sent to the second terminal.
4. The method according to claim 2, characterized in that After determining that the user to which the template voiceprint feature that matches the voiceprint feature of the audio to be processed belongs is the user to which the audio to be processed belongs, the method further includes: If the user to which the audio to be processed belongs and the user to which the facial image to be processed belongs are not the same user, it is determined that the audio to be processed is interference audio.
5. The method according to claim 3, characterized in that If the user to which the audio to be processed belongs and the user to which the facial image to be processed belong are not the same target user, after increasing the preset difference threshold from the first value to the second value, the method further includes: If the difference between the voiceprint feature of the audio to be processed and the voiceprint features of each template is greater than the adjusted preset difference threshold, it is determined that the audio to be processed is interference audio.
6. The method according to claim 1, characterized in that Before suppressing interference audio other than the target user's audio to be processed to obtain target audio if the user to which the audio to be processed belongs and the user to which the facial image to be processed belongs are the same target user, and sending the target audio to the second terminal, the method further includes: Performing lip movement detection on the face image to be processed; If lip movement of the user in the face image to be processed is detected, it is detected whether the user to whom the audio to be processed belongs and the user to whom the face image to be processed belongs are the same target user.
7. The method according to claim 1, characterized in that The step of obtaining the collected audio to be processed and the facial image to be processed includes: Obtaining audio captured by a microphone of the first terminal, performing audio preprocessing on the acquired audio to obtain original audio, extracting a vocal portion from the original audio to obtain audio to be processed; wherein the audio preprocessing includes acoustic echo cancellation, automatic gain control, and noise suppression; A video image containing a user image captured by a camera is obtained, and image preprocessing is performed on the obtained video image to obtain an original video image, and a face image in the original video image is extracted to obtain a face image to be processed; wherein the image preprocessing includes: automatic exposure, automatic white balance and automatic focus.
8. The method according to claim 1, characterized in that For each template voiceprint feature, if the voiceprint feature of the audio to be processed matches the template voiceprint feature, before determining that the user to which the template voiceprint feature belongs is the user to which the audio to be processed belongs, the method further includes: Obtain the collected audio and facial image of the user to be registered; Performing feature extraction on the audio to be registered to obtain a template voiceprint feature of the audio to be registered, and performing feature extraction on the facial image to be registered to obtain a template facial feature of the facial image to be registered; Performing lip movement detection on the face image to be registered; If lip movement is detected in the face image to be registered of the user to be registered, the template voiceprint features of the audio to be registered, the template facial features of the face image to be registered, and the corresponding user identification of the user to be registered are recorded in a preset storage area locally on the first terminal.
9. The method according to claim 8, characterized in that The preset storage area is a secure storage area of the first terminal.
10. A terminal, characterized in that: include: one or more processors and memory; The memory is coupled to the one or more processors, and is configured to store computer program codes, where the computer program codes include computer instructions. The one or more processors invoke the computer instructions to enable the terminal to execute the method according to any one of claims 1 to 9.
11. A computer-readable storage medium, characterized in that The method comprises a computer program, which, when executed on a terminal, causes the terminal to execute the method according to any one of claims 1 to 9.
12. A computer program product, characterized in that The computer program product comprises executable instructions, and when the executable instructions are executed on a terminal, the terminal is caused to execute the method according to any one of claims 1 to 9.
13. A chip system, characterized in that: The chip system is applied to a terminal, and the chip system includes one or more processors, which are used to call computer instructions to enable the terminal to input data into the chip system, and execute the method described in any one of claims 1-9 to process the data and output the processing results.
Citation Information
Patent Citations
User authentication method and device on basis of audios and videos
CN103973441A
Fusion type voice recognition method, device and system, equipment and storage medium
CN111883130A
Speech enhancement method, electronic equipment, storage medium and chip system
CN116072136A
Voice noise reduction method and device based on voiceprint recognition, equipment and medium
CN116312570A
Voiceprint information authentication method and device, electronic equipment and storage medium
CN117037845A