An audio processing method, a terminal and a storage medium

By combining audio and video preprocessing technologies in the terminal and using voiceprint and facial feature matching to identify and suppress interfering audio, the problem of noise interference in audio and video calls in public places is solved, achieving higher audio noise reduction accuracy and user experience.

CN120455583BActive Publication Date: 2026-04-28HONOR DEVICE CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HONOR DEVICE CO LTD
Filing Date
2024-09-25
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

When making audio and video calls in public places, the audio data collected by the terminal contains environmental noise and interfering human voices, which leads to a decrease in user experience. The accuracy of voiceprint noise reduction in existing technologies is low, and it is easy to cause recognition abnormalities due to changes in the target user's voiceprint or incorrect learning.

Method used

By acquiring the audio and facial images to be processed in the terminal, the system uses voiceprint and facial feature matching technology to identify the target user and suppress interfering audio. It combines audio and video preprocessing to improve the accuracy of feature extraction, adjusts the difference threshold to avoid recognition errors, and confirms user consistency through lip movement detection.

Benefits of technology

It improves the accuracy of suppressing interfering audio, enhances audio noise reduction, improves user experience, and ensures clear transmission of the target audio.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120455583B_ABST
    Figure CN120455583B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide an audio processing method, a terminal and a storage medium, and relate to the technical field of computers. The method comprises: in the process of audio communication between a first terminal and a second terminal, obtaining collected audio to be processed and a face image to be processed; if a voiceprint feature of the audio to be processed matches a template voiceprint feature, determining that a user to which the template voiceprint feature belongs is a user to which the audio to be processed belongs; if a face feature of the face image to be processed matches the template face feature, determining that a user to which the template face feature belongs is a user to which the face image to be processed belongs; if the user to which the audio to be processed belongs and the user to which the face image to be processed belong are a same target user, suppressing interference audio other than the audio to be processed of the target user to obtain target audio, and sending the target audio to the second terminal, which can improve the accuracy of suppressing the interference audio and improve the experience of the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to an audio processing method, terminal, and storage medium. Background Technology

[0002] Users can make audio and video calls using devices such as mobile phones, tablets, and PCs. However, if users are in public places, such as offices, cafes, restaurants, airports, and train stations, the audio data collected by the device during audio and video calls will contain environmental noise and interfering human voices, which will degrade the user experience.

[0003] Therefore, in order to ensure that the other party in a video call can hear the user's voice more clearly, the terminal needs to perform noise reduction processing on the collected audio data to remove environmental noise and interfering human voices, thereby improving the user experience. Summary of the Invention

[0004] The purpose of this application is to provide an audio processing method, terminal, and storage medium to improve the accuracy of suppressing interfering audio, achieve audio noise reduction, and enhance the user experience. The specific technical solution is as follows:

[0005] Firstly, in order to achieve the above objectives, embodiments of this application provide an audio processing method, which is applied to a first terminal, and the method includes:

[0006] During the audio call between the first terminal and the second terminal, the collected audio to be processed and the face image to be processed are acquired.

[0007] For each template voiceprint feature, if the voiceprint feature of the audio to be processed matches the template voiceprint feature, the user to which the template voiceprint feature belongs is determined to be the user to which the audio to be processed belongs.

[0008] For each template face feature, if the face feature of the face image to be processed matches the template face feature, the user to which the template face feature belongs is determined to be the user to which the face image to be processed belongs.

[0009] If the user to which the audio to be processed belongs and the user to which the face image to be processed belongs are the same target user, the interfering audio other than the audio to be processed of the target user is suppressed to obtain the target audio, and the target audio is sent to the second terminal.

[0010] As can be seen from the above, the technical solution provided in this embodiment, when the target user's audio to be processed and the template voiceprint features, as well as the target user's face image to be processed and the template face features, are all successfully matched, it indicates that the target user is the user conducting the audio call. That is, the target user's audio to be processed is not interfering audio. Therefore, suppressing interfering audio other than the target user's audio to be processed can improve the accuracy of suppressing interfering audio, achieve audio noise reduction, and improve the user experience.

[0011] In one embodiment of this application, before determining that the user to which the template voiceprint feature belongs is the user to which the audio to be processed belongs, for each template voiceprint feature, if the voiceprint feature of the audio to be processed matches the template voiceprint feature, the method further includes:

[0012] For each audio file to be processed, feature extraction is performed to obtain the voiceprint features of that audio file.

[0013] For each template voiceprint feature, if the difference between the voiceprint feature of the audio to be processed and the template voiceprint feature is less than a preset difference threshold, it is determined that the voiceprint feature of the audio to be processed matches the template voiceprint feature.

[0014] In one embodiment of this application, after determining that the voiceprint features of the audio to be processed match the voiceprint features of the template if the difference between the voiceprint features of the audio to be processed and the voiceprint features of the template are less than a preset difference threshold, the method further includes:

[0015] If the user to which the audio to be processed belongs and the user to which the face image to be processed belongs are not the same target user, the preset difference threshold is increased from the first value to the second value;

[0016] For each template voiceprint feature, if the difference between the voiceprint feature of the audio to be processed and the template voiceprint feature is less than the adjusted difference threshold, it is determined that the voiceprint feature of the audio to be processed matches the template voiceprint feature, and the user to which the template voiceprint feature that matches the voiceprint feature of the audio to be processed belongs is determined to be the user to which the audio to be processed belongs.

[0017] If the user to which the audio to be processed belongs and the user to which the face image to be processed belongs are the same target user, the interfering audio other than the audio to be processed of the target user is suppressed to obtain the target audio, and the target audio is sent to the second terminal.

[0018] As can be seen from the above, the technical solution provided in this embodiment, by adjusting the preset difference threshold, avoids the problem of recognition errors caused by changes in the voiceprint characteristics of the target user, further improves the accuracy of suppressing interfering audio, and enhances the user experience.

[0019] In one embodiment of this application, after determining that the user to whom the template voiceprint feature matching the voiceprint feature of the audio to be processed belongs is the user to whom the audio to be processed belongs, the method further includes:

[0020] If the user to which the audio to be processed belongs and the user to which the face image to be processed belongs are not the same user, the audio to be processed is determined to be interference audio.

[0021] As can be seen from the above, the technical solution provided in this embodiment determines that the audio to be processed where the facial image features and voiceprint features do not match as interference audio. It can accurately identify interference audio and suppress it, thereby improving the accuracy of interference audio suppression and enhancing the user experience.

[0022] In one embodiment of this application, after increasing the preset difference threshold from a first value to a second value if the user to which the audio to be processed belongs and the user to which the face image to be processed belongs are not the same target user, the method further includes:

[0023] If the difference between the voiceprint features of the audio to be processed and the voiceprint features of each template is greater than the adjusted preset difference threshold, the audio to be processed is determined to be interference audio.

[0024] As can be seen from the above, the technical solution provided in this embodiment determines that the audio to be processed that does not match the voiceprint features of each template is interference audio. It can accurately identify interference audio and suppress it, thereby improving the accuracy of interference audio suppression and enhancing the user experience.

[0025] In one embodiment of this application, before the step of suppressing interfering audio other than the target user's audio to be processed and the user of the face image to be processed being the same target user, obtaining the target audio, and sending the target audio to the second terminal, the method further includes:

[0026] Perform lip movement detection on the face image to be processed;

[0027] If lip movement is detected in the user in the face image to be processed, it is determined whether the user to whom the audio to be processed belongs and the user to whom the face image to be processed belongs are the same target user.

[0028] As can be seen from the above, the technical solution provided in this embodiment, if the user in the face image to be processed has lip movement, it indicates that the user within the field of view of the camera of the first terminal has made a sound. Then, it continues to detect whether the user who made the sound is the same user as the user within the field of view of the camera of the first terminal. This avoids the problem of misidentifying the sound made by other users as the sound made by the target user, which would lead to audio suppression errors. This further improves the accuracy of suppressing interfering audio and enhances the user experience.

[0029] In one embodiment of this application, acquiring the collected audio to be processed and the face image to be processed includes:

[0030] The audio collected by the microphone of the first terminal is acquired, and the acquired audio is preprocessed to obtain the original audio. The human voice part of the original audio is extracted to obtain the audio to be processed. The audio preprocessing includes: acoustic echo cancellation, automatic gain control and noise suppression.

[0031] The system acquires video images containing user images captured by a camera, performs image preprocessing on the acquired video images to obtain raw video images, and extracts face images from the raw video images to obtain face images to be processed; wherein, the image preprocessing includes: automatic exposure, automatic white balance and automatic focus.

[0032] As can be seen from the above, the technical solution provided in this embodiment can improve the accuracy of the subsequently extracted voiceprint features by performing audio preprocessing on the collected audio, and can improve the accuracy of the subsequently extracted facial features by performing video preprocessing on the collected video, further improving the accuracy of suppressing interfering audio and improving the user experience.

[0033] In one embodiment of this application, before determining that the user to which the template voiceprint feature belongs is the user to which the audio to be processed belongs, for each template voiceprint feature, if the voiceprint feature of the audio to be processed matches the template voiceprint feature, the method further includes:

[0034] Acquire the audio of the user to be registered, as well as the facial image of the user to be registered;

[0035] Feature extraction is performed on the audio to be registered to obtain the template voiceprint features of the audio to be registered, and feature extraction is performed on the face image to be registered to obtain the template face features of the face image to be registered;

[0036] Lip movement detection is performed on the face image to be registered;

[0037] If lip movement is detected in the face image to be registered, the template voiceprint features of the audio to be registered, the template face features of the face image to be registered, and the user identifier of the user to be registered are recorded in the preset storage area of ​​the first terminal.

[0038] As can be seen from the above, the technical solution provided in this embodiment can record the template voiceprint features of the audio to be registered, the template face features of the face image to be registered, and the user identifier of the user to be registered. Subsequently, if the target user's audio to be processed matches the template voiceprint features, and the target user's face image matches the template face features, it indicates that the target user is the user conducting the audio call, that is, the target user's audio to be processed is not interfering audio. Therefore, suppressing interfering audio other than the target user's audio to be processed can improve the accuracy of interfering audio suppression, achieve audio noise reduction, and improve the user experience.

[0039] In one embodiment of this application, the preset storage area is the secure storage area of ​​the first terminal.

[0040] As can be seen from the above, the technical solution provided in this embodiment, with the preset storage area being the secure storage area of ​​the first terminal, can improve the security of the registered template voiceprint features and template face features, and protect the user's privacy.

[0041] Secondly, embodiments of this application also provide a terminal, including:

[0042] One or more processors and memory;

[0043] The memory is coupled to the one or more processors, and the memory is used to store computer program code, the computer program code including computer instructions, which the one or more processors call to cause the terminal to perform any of the audio processing methods described above.

[0044] Thirdly, embodiments of this application also provide a computer-readable storage medium including a computer program that, when run on a terminal, causes the terminal to execute any of the audio processing methods described above.

[0045] Fourthly, embodiments of this application also provide a computer program product, the computer program product comprising executable instructions, which, when executed on a terminal, cause the terminal to perform any of the audio processing methods described above.

[0046] Fifthly, embodiments of this application also provide a chip system applied to a terminal. The chip system includes one or more processors, which are used to invoke computer instructions to cause the terminal to execute any of the audio processing methods described above.

[0047] The beneficial effects of the solutions provided in the embodiments of the second, third, fourth and fifth aspects above can be found in the beneficial effects of the solutions provided in the embodiments of the first aspect above. Attached Figure Description

[0048] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0049] Figure 1 A structural diagram of a terminal provided in an embodiment of this application;

[0050] Figure 2 A software structure block diagram of a terminal provided in an embodiment of this application;

[0051] Figure 3 A schematic diagram illustrating an application scenario of the first audio processing method provided in this application embodiment;

[0052] Figure 4 A schematic diagram illustrating an application scenario of the second audio processing method provided in this application embodiment;

[0053] Figure 5 A schematic diagram illustrating an application scenario of the third audio processing method provided in this application embodiment;

[0054] Figure 6 A flowchart illustrating the first audio processing method provided in this application embodiment;

[0055] Figure 7 A flowchart illustrating the second audio processing method provided in this application embodiment;

[0056] Figure 8 A flowchart illustrating the third audio processing method provided in this application embodiment;

[0057] Figure 9 A flowchart illustrating the fourth audio processing method provided in this application embodiment;

[0058] Figure 10 A flowchart for acquiring audio data is provided as an embodiment of this application;

[0059] Figure 11 A flowchart for acquiring image data is provided as an embodiment of this application;

[0060] Figure 12 A flowchart illustrating the first type of audio and image data transmission provided in this application embodiment;

[0061] Figure 13 A flowchart illustrating the registration template voiceprint features and template face features is provided in this application embodiment;

[0062] Figure 14 A flowchart for obtaining template voiceprint features is provided in an embodiment of this application;

[0063] Figure 15 A flowchart for obtaining template facial features is provided in an embodiment of this application;

[0064] Figure 16 A flowchart illustrating the second type of audio and image data transmission provided in this application embodiment;

[0065] Figure 17 This is a structural diagram of a chip system provided in an embodiment of this application. Detailed Implementation

[0066] To better understand the technical solution of this application, the embodiments of this application will be described in detail below with reference to the accompanying drawings.

[0067] To facilitate a clear description of the technical solutions in the embodiments of this application, the terms "first" and "second" are used in the embodiments of this application to distinguish identical or similar items with essentially the same function and effect. For example, "first instruction" and "second instruction" are used to distinguish different user instructions and do not limit their order. Those skilled in the art will understand that the terms "first" and "second" do not limit the quantity or execution order, and the terms "first" and "second" are not necessarily different.

[0068] It should be noted that, in this application, the words "exemplarily" or "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as "exemplarily" or "for example" in this application should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of words such as "exemplarily" or "for example" is intended to present the relevant concepts in a specific manner.

[0069] The audio processing method provided in this application is applied to a terminal. The terminal can be a mobile phone, tablet computer, personal digital assistant (PDA), smartwatch, wearable terminal, augmented reality (AR) device, virtual reality (VR) device, robot, smart glasses, or other device capable of processing audio and capturing images. Thus, the terminal can use the audio processing method provided in this application to suppress interfering audio during audio and video calls, improving the accuracy of interference suppression, achieving audio noise reduction, and enhancing the user experience.

[0070] For example, Figure 1 A structural diagram of terminal 100 is shown. Terminal 100 may include a processor 110, a display screen 120, a camera 130, internal memory 140, a Subscriber Identification Module (SIM) card interface 150, a Universal Serial Bus (USB) interface 160, a charging management module 170, a battery management module 171, a battery 172 with battery cells and battery protection devices, a sensor module 180, a mobile communication module 190, a wireless communication module 200, antenna 1, and antenna 2, etc. The sensor module 180 may include a pressure sensor 180A, a fingerprint sensor 180B, a touch sensor 180C, an ambient light sensor 180D, etc.

[0071] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the terminal 100. In other embodiments of this application, the terminal 100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0072] Processor 110 may include one or more processing units, such as a Central Processing Unit (CPU), an Application Processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural network processing unit (NPU). Different processing units may be independent components or integrated into one or more processors. In some embodiments, terminal 100 may also include one or more processors 110. The controller can generate operation control signals based on instruction opcodes and timing signals to control instruction fetching and execution. In other embodiments, processor 110 may also include a memory for storing instructions and data. For example, the memory in processor 110 may be a cache memory. This memory can store instructions or data that processor 110 has just used or is repeatedly used. If processor 110 needs to reuse the instruction or data, it can directly retrieve it from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the terminal 100 in processing data or executing instructions.

[0073] In some embodiments, the processor 110 may include one or more interfaces. These interfaces may include an Inter-Integrated Circuit (I2C) interface, an Inter-Integrated Circuit Sound (I2S) interface, a Pulse Code Modulation (PCM) interface, a Universal Asynchronous Receiver / Transmitter (UART) interface, a Mobile Industry Processor Interface (MIPI) interface, a General-Purpose Input / Output (GPIO) interface, a SIM card interface, and / or a USB interface, etc. The USB interface 160 is a USB standard-compliant interface, specifically a Mini USB interface, a Micro USB interface, a USB Type-C interface, etc. The USB interface 160 can be used to connect a charger to charge the terminal 100, and can also be used for data transfer between the terminal 100 and peripheral devices. The USB interface 160 can also be used to connect headphones for audio playback.

[0074] It is understood that the interface connection relationships between the modules illustrated in the embodiments of this application are for illustrative purposes only and do not constitute a structural limitation on the terminal 100. In other embodiments of this application, the terminal 100 may also adopt different interface connection methods or a combination of multiple interface connection methods as described in the above embodiments.

[0075] The wireless communication function of terminal 100 can be implemented through antenna 1, antenna 2, mobile communication module 190, wireless communication module 200, modem processor and baseband processor.

[0076] Antennas 1 and 2 are used to transmit and receive electromagnetic wave signals. Each antenna in terminal 100 can be used to cover one or more communication frequency bands. Different antennas can also be multiplexed to improve antenna utilization. For example, antenna 1 can be multiplexed as a diversity antenna for a wireless local area network. In some other embodiments, the antennas can be used in conjunction with tuning switches.

[0077] Terminal 100 implements display functions through a GPU, display screen 120, and application processor. The GPU is a microprocessor for image processing, connected to the display screen 120 and the application processor. The GPU is used to perform mathematical and geometric calculations and for graphics rendering. Processor 110 may include one or more GPUs, which execute program instructions to generate or modify display information.

[0078] The display screen 120 is used to display images, videos, etc. The display screen 120 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a miniature LED, a microLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the terminal 100 may include one or more display screens 120.

[0079] In some embodiments of this application, when the display panel uses materials such as OLED, AMOLED, and FLED, the above-mentioned Figure 1 The display screen 120 can be bent. Here, "the display screen 120 can be bent" means that the display screen can be bent to any angle at any part and can maintain that angle. For example, the display screen 120 can be folded from the middle left to right. It can also be folded from the middle up and down.

[0080] The display screen 120 of terminal 100 can be a flexible screen. Currently, flexible screens are attracting much attention due to their unique characteristics and enormous potential. Compared to traditional screens, flexible screens are highly flexible and bendable, providing users with new interaction methods based on their bendability and meeting more user needs for terminals. For terminals equipped with foldable displays, the foldable display can switch between a small screen in folded mode and a large screen in unfolded mode at any time. Therefore, users are increasingly using split-screen functionality on terminals equipped with foldable displays.

[0081] Terminal 100 can perform shooting functions through ISP, camera 130, video codec, GPU, display 120 and application processor, wherein camera 130 includes a front camera and a rear camera.

[0082] The ISP is used to process data fed back from the camera 130. For example, during shooting, when the shutter is opened, light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, and the camera's photosensitive element transmits the electrical signal to the ISP for processing, transforming it into an image visible to the naked eye. The ISP can perform algorithmic optimization of image noise, brightness, and color. The ISP can also optimize parameters such as exposure and color temperature of the shooting scene. In some embodiments, the ISP can be set in the camera 130.

[0083] Camera 130 is used to capture photos or videos. An object is projected onto a photosensitive element through a lens, generating an optical image. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then transmitted to an ISP (Internet Service Provider) for conversion into a digital image signal. The ISP outputs the digital image signal to a DSP (Digital Signal Processor) for processing. The DSP converts the digital image signal into standard red-green-blue (RGB), YUV, or other image signals. In some embodiments, terminal 100 may include one or N cameras 130, where N is a positive integer greater than 1.

[0084] A digital signal processor (DSP) is used to process digital signals. Besides digital image signals, it can also process other digital signals. For example, when terminal 100 selects a frequency point, the DSP can perform Fourier transforms on the frequency energy.

[0085] Video codecs are used to compress or decompress digital video. Terminal 100 may support one or more video codecs. Thus, terminal 100 can play or record video in various encoding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG 2, MPEG 3, and MPEG 4.

[0086] NPU stands for Neural Network (NN) computing processor. By borrowing the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, it can rapidly process input information and continuously learn on its own. NPUs can enable intelligent cognitive applications in terminals, such as image recognition, facial recognition, speech recognition, and text understanding.

[0087] The internal memory 140 can be used to store one or more computer programs, which include instructions. The processor 110 can execute the instructions stored in the internal memory 140, thereby causing the terminal 100 to perform the image generation method provided in some embodiments of this application, as well as various applications and data processing. The internal memory 140 may include a program storage area and a data storage area. The program storage area may store the operating system; it may also store one or more applications (such as a gallery, contacts, etc.). The data storage area may store data created by the terminal 100 during use (such as photos, contacts, etc.). Furthermore, the internal memory 140 may include high-speed random access memory and non-volatile memory, such as one or more disk storage components, flash memory components, Universal Flash Storage (UFS), etc. In some embodiments, the processor 110 can execute instructions stored in the internal memory 140 and / or instructions stored in memory disposed in the processor 110, thereby causing the terminal 100 to perform the audio processing method provided in the embodiments of this application, as well as other applications and data processing.

[0088] The internal memory 140 can be used to store the relevant programs of the audio processing method provided in the embodiments of this application. The processor 110 can be used to call the relevant programs of the audio processing method stored in the internal memory 140 when displaying information, and execute the audio processing method of the embodiments of this application.

[0089] The sensor module 180 may include a pressure sensor 180A, a fingerprint sensor 180B, a touch sensor 180C, an ambient light sensor 180D, etc.

[0090] Pressure sensor 180A is used to sense pressure signals and convert them into electrical signals. In some embodiments, pressure sensor 180A can be disposed on display screen 120. Pressure sensor 180A can be of many types, such as resistive pressure sensor, inductive pressure sensor, or capacitive pressure sensor. A capacitive pressure sensor can include at least two parallel plates with conductive material. When force is applied to pressure sensor 180A, the capacitance between the electrodes changes, and terminal 100 determines the pressure intensity based on the change in capacitance. When a touch operation is applied to display screen 120, terminal 100 detects the touch operation based on pressure sensor 180A. Terminal 100 can also calculate the touch position based on the detection signal from pressure sensor 180A. In some embodiments, touch operations applied to the same touch position but with different touch operation intensities can correspond to different operation commands. For example, when a touch operation with an intensity less than a first pressure threshold is applied to the SMS application icon, a command to view an SMS is executed; when a touch operation with an intensity greater than or equal to the first pressure threshold is applied to the SMS application icon, a command to create a new SMS is executed.

[0091] The fingerprint sensor 180B is used to collect fingerprints. The terminal 100 can use the collected fingerprint characteristics to perform functions such as unlocking, accessing the app lock, taking photos, and answering calls.

[0092] Touch sensor 180C, also known as a touch device, can be disposed on display screen 120. The touch sensor 180C and display screen 120 together form a touchscreen, also known as a touch display. Touch sensor 180C is used to detect touch operations applied to or near it. Touch sensor 180C can transmit the detected touch operation to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through display screen 120. In other embodiments, touch sensor 180C may also be disposed on the surface of terminal 100, and in a different location from display screen 120.

[0093] The ambient light sensor 180D is used to sense the ambient light intensity. The terminal 100 can adaptively adjust the brightness of the display screen 120 based on the sensed ambient light intensity. The ambient light sensor 180D can also be used to automatically adjust the white balance during shooting. The ambient light sensor 180D can also transmit environmental information about the device's location to the GPU.

[0094] The ambient light sensor 180D is also used to acquire the brightness, light ratio, color temperature, and other parameters of the environment in which the camera 130 captures images.

[0095] Figure 2This is a software architecture block diagram for a terminal to which this application's embodiments apply. The terminal's software system can adopt a layered architecture, event-driven architecture, microkernel architecture, microservice architecture, or cloud architecture. A layered architecture divides the terminal's software system into several layers, each with a clear role and division of labor. Layers communicate with each other through software interfaces. In some embodiments, the software system can be divided into three layers: the application layer, the application framework layer, and the hardware abstract layer (HAL).

[0096] The application layer can include a series of application packages. The application layer runs applications by calling the application programming interface (API) provided by the application framework layer. For example... Figure 2 As shown, the application package may include multiple applications, such as camera, clock, browser, and music programs. Understandably, the port of each of these applications can be used to receive data.

[0097] The application framework layer provides APIs and a programming framework for applications within the application layer. The application framework layer includes predefined functions. For example... Figure 2 As shown, the application framework layer may include a window manager, content provider, view system, resource manager, notification manager, and Dynamic Host Configuration Protocol (DHCP) module, etc.

[0098] The driver layer is the layer between hardware and software, used to drive the hardware and make it work. Multiple drivers can be installed in the driver layer to operate the hardware. Examples include camera drivers, display drivers, audio drivers, and sensor drivers.

[0099] In addition, the terminal also includes a hardware layer, which may include a camera, speaker, display screen and battery, etc. The hardware layer is connected to the driver layer.

[0100] The following describes the application scenarios of the audio processing method provided in the embodiments of this application.

[0101] When users make audio and video calls using terminals (such as mobile phones, tablets, and PCs), if the users are in public places such as offices, cafes, restaurants, airports, and train stations, the audio data collected by the terminal will contain environmental noise and interfering human voices, which will reduce the user experience.

[0102] The following uses a mobile phone as an example to illustrate the concept. Figure 3The application scenarios of the embodiments of this application will be described. Figure 3 This is merely an illustrative example and does not limit the audio processing method provided in this application to only [specific applications]. Figure 3 The application scenarios shown.

[0103] The target user is making a voice call with contact 1 using an instant messaging application on their mobile phone. Besides the target user's voice, there are also interfering users making sounds, and an alarm clock in the environment is also ringing. Therefore, the audio captured by the phone includes the target user's audio, the interfering users' audio, and environmental noise. Thus, to ensure the other party in the voice call can hear the target user clearly, the phone needs to suppress the interfering users' audio and environmental noise in the captured audio, i.e., perform noise reduction processing on the target user's audio, to improve the user experience.

[0104] Currently, terminals can perform voiceprint noise reduction on environmental noise and interfering human voices, but there are still many problems with current voiceprint noise reduction, such as recognition anomalies caused by changes in the target user's voiceprint, and errors in voiceprint learning, resulting in low accuracy of voiceprint noise reduction.

[0105] The following explanation, based on the principle of audio noise reduction on the terminal, explains the reasons for the low accuracy of voiceprint noise reduction.

[0106] Before voiceprint noise reduction, the target user's voiceprint features need to be registered with the terminal, meaning the terminal saves the target user's voiceprint features as template voiceprint features. Correspondingly, in the subsequent noise reduction stage, if the voiceprint features of the audio collected by the terminal match the saved template voiceprint features, it indicates that the collected audio is the target user's audio, and other audio besides the target user's audio can be suppressed.

[0107] However, as Figure 4 As shown, during the registration phase, in addition to the target user making a sound, an interfering user is also making a sound. The terminal may then incorrectly save the interfering user's voiceprint features as the template voiceprint features. Correspondingly, as... Figure 5 As shown, if the voiceprint features of the target user's audio do not match the template voiceprint features, the terminal will perform noise reduction on the target user's audio. Alternatively, if the target user's voiceprint features change, such as due to a cold, the voiceprint features of the target user's audio will also not match the template voiceprint features, and the terminal will again perform noise reduction on the target user's audio. This results in the terminal outputting silent audio to the other party in the voice call, degrading the user experience.

[0108] The audio processing method provided in this application, when the target user's audio to be processed matches the template voiceprint features and the target face image matches the template face features, indicates that the target user is the user conducting the audio call, that is, the target user's audio to be processed is not interfering audio. Therefore, interfering audio other than the target user's audio to be processed is suppressed, which can improve the accuracy of suppressing interfering audio, achieve audio noise reduction, and improve the user experience.

[0109] Next, the audio processing method provided in this application will be described in detail through specific embodiments.

[0110] See Figure 6 , Figure 6 A flowchart of an audio processing method provided in this application embodiment, the method being applied to a terminal, the method comprising the following steps:

[0111] S601: During an audio call between the first terminal and the second terminal, acquire the collected audio to be processed and the face image to be processed.

[0112] S602: For each template voiceprint feature, if the voiceprint feature of the audio to be processed matches the template voiceprint feature, determine that the user to which the template voiceprint feature belongs is the user to which the audio to be processed belongs.

[0113] S603: For each template face feature, if the face feature of the face image to be processed matches the template face feature, determine that the user to which the template face feature belongs is the user to which the face image to be processed belongs.

[0114] S604: If the user to which the audio to be processed belongs and the user to which the face image to be processed belongs are the same target user, suppress the interfering audio other than the target user's audio to be processed, obtain the target audio, and send the target audio to the second terminal.

[0115] As can be seen from the above, the technical solution provided in this embodiment, when the target user's audio to be processed and the template voiceprint features, as well as the target user's face image to be processed and the template face features, are all successfully matched, it indicates that the target user is the user conducting the audio call. That is, the target user's audio to be processed is not interfering audio. Therefore, suppressing interfering audio other than the target user's audio to be processed can improve the accuracy of suppressing interfering audio, achieve audio noise reduction, and improve the user experience.

[0116] In step S601, during the audio call between the user and the second terminal using the first terminal, the microphone of the first terminal collects the user's audio (i.e., the audio to be processed), and the sensor of the first terminal collects the user's image (i.e., the face image to be processed).

[0117] The sensor can be a camera on the first terminal, or an ultrasonic detector (such as radar) on the first terminal. The first terminal can acquire images of the user by taking pictures with a camera, or by using an ultrasonic detector to perform ultrasonic detection. This application does not limit the method of acquiring user images; the following embodiments use the method of acquiring user images by taking pictures with a camera as an example for illustration.

[0118] In one scenario, a user uses an instant messaging application installed on a first terminal to make a video call with a second terminal. During the call, the microphone of the first terminal captures the user's audio to be processed, and the camera of the first terminal captures the user's facial image to be processed.

[0119] In another scenario, a user uses an instant messaging application installed on a first terminal to make a voice call with a second terminal. Since the user is not using video calling, the microphone on the first terminal captures the user's audio during the call. If the user authorizes the camera to capture images during the voice call, the camera on the first terminal captures the user's facial image.

[0120] In some embodiments, step S601 may include the following steps: acquiring audio captured by the microphone of the first terminal, performing audio preprocessing on the acquired audio to obtain original audio, and extracting the human voice portion from the original audio to obtain audio to be processed. The audio preprocessing includes acoustic echo cancellation, automatic gain control, and noise suppression. Acquiring a video image containing the user's image captured by the camera, performing image preprocessing on the acquired video image to obtain original video image, and extracting the face image from the original video image to obtain the face image to be processed. The image preprocessing includes automatic exposure, automatic white balance, and automatic focus.

[0121] The audio collected by the first terminal is of low quality and may contain noise and other issues. Therefore, the first terminal uses the 3A call algorithm to preprocess the collected audio to obtain the original audio. Then, the human voice portion is extracted from the original audio to obtain the audio to be processed. The audio to be processed is the sound emitted by the user during the audio call between the first terminal and the second terminal.

[0122] When there is only one user making the sound, the audio to be processed for that user is extracted. When there are multiple users making the sound, the audio to be processed for each user can be extracted separately, and each audio to be processed is processed according to the method provided in the embodiments of this application.

[0123] The 3A call algorithm includes: Acoustic Echo Cancellation (AEC), Automatic Gain Control (AGC), and Noise Suppression (NS).

[0124] The video images captured by the first terminal may also be of low quality, such as low resolution. Therefore, the first terminal uses the 3A video algorithm to preprocess the captured video images to obtain the original video images. The 3A video algorithm includes: Auto Exposure (AE), Auto White Balance (AWB), and Auto Focus (AF). Then, facial images are extracted from the original video images to obtain the facial images to be processed. The facial images to be processed are the faces of users who are within the field of view of the first terminal's camera and are facing the first terminal's camera during the audio call between the first and second terminals.

[0125] As can be seen from the above, the technical solution provided in this embodiment can improve the accuracy of the subsequently extracted voiceprint features by performing audio preprocessing on the collected audio, and can improve the accuracy of the subsequently extracted facial features by performing video preprocessing on the collected video, further improving the accuracy of suppressing interfering audio and improving the user experience.

[0126] For steps S602 and S603, the template voiceprint features and template face features are saved locally on the first terminal during the registration phase. The registration process for the template face features and template voiceprint features is described in subsequent embodiments.

[0127] In some embodiments, Figure 6 Based on this, see Figure 7 Before step S602, the method may further include the following steps:

[0128] S605: For each audio to be processed, feature extraction is performed on the audio to be processed to obtain the voiceprint features of the audio to be processed.

[0129] S606: For each template voiceprint feature, if the difference between the voiceprint feature of the audio to be processed and the template voiceprint feature is less than a preset difference threshold, it is determined that the voiceprint feature of the audio to be processed matches the template voiceprint feature.

[0130] After obtaining the audio to be processed, the first terminal extracts the voiceprint features of the audio to be processed and compares the voiceprint features of the audio to be processed with the locally stored template voiceprint features. If the voiceprint features of the audio to be processed match any template voiceprint feature, the user to which the template voiceprint feature belongs is determined to be the user to which the audio to be processed belongs (which can be called the first user), that is, the audio to be processed is the sound emitted by the first user.

[0131] In some embodiments, voiceprint features may include: the fundamental frequency of the audio to be processed, the fundamental period of the audio to be processed, the harmonic components of the audio to be processed, and the Mel frequency cepstrum coefficient (MFCC) of the audio to be processed. The fundamental frequency (i.e., the fundamental tone frequency) refers to the frequency at which the vocal cords open and close. The fundamental period refers to the period of vocal cord vibration.

[0132] For each template voiceprint feature, if the difference between the voiceprint feature of the audio to be processed and the voiceprint feature of the template is less than a preset difference threshold, it indicates that the voiceprint feature of the audio to be processed is similar to the voiceprint feature of the template, and it can be determined that the voiceprint feature of the audio to be processed matches the voiceprint feature of the template.

[0133] If the difference between the voiceprint features of the audio to be processed and the voiceprint features of each template is greater than the preset difference threshold, it indicates that the audio to be processed is not the voice of a registered user. It may be that the voiceprint features of the target user have changed, causing the audio to be processed to be identified as the audio of another user. For example, the voiceprint features of the target user may have changed due to a cold. In this case, the preset difference threshold will be increased from the first value to the second value.

[0134] The voiceprint features of the audio to be processed are compared again with the voiceprint features of each template. If the difference between the voiceprint features of the audio to be processed and any template voiceprint feature is less than the adjusted preset difference threshold, it is determined that the voiceprint features of the audio to be processed match the template voiceprint features. The user to whom the template voiceprint features that match the voiceprint features of the audio to be processed belong is determined to be the user to whom the audio to be processed belongs (i.e., the first user). In other words, the user to whom the template voiceprint features matched this time belong is determined to be the user who made the sound. Furthermore, if the first user to whom the audio to be processed belongs and the second user to whom the face image to be processed belongs are the same target user, it indicates that the user who made the sound and the user who is facing the camera of the first terminal within the field of view of the first terminal are the same user. That is, the target user is using the first terminal to have an audio call with the user using the second terminal. In this case, all the audio in the original audio except for the target user's audio to be processed is interference audio. The first terminal then suppresses the interference audio except for the target user's audio to be processed to obtain the target audio. The target audio is noise-reduced and of higher quality. The first terminal then sends the target audio to the second terminal.

[0135] After acquiring the face image to be processed, the first terminal, since the face image to be processed is the face image of a user within the camera's field of view during an audio call between the first and second terminals, extracts the facial features of the face image to be processed and compares it with locally stored template facial features. If the facial features of the face image to be processed match any template facial feature, the user to whom the template facial feature belongs is determined to be the user to whom the face image to be processed belongs (which can be called the second user). In other words, the second user is the face image of a user within the camera's field of view during an audio call between the first and second terminals. Facial features can include key points in the face image, such as the 68 key points obtained from face detection.

[0136] Furthermore, it is determined whether the first user to whom the audio to be processed belongs and the second user to whom the image to be processed belong are the same user (i.e., the target user). In other words, it is detected whether the user who makes the sound during the audio call between the first terminal and the second terminal is the same user as the user who is in the field of view of the camera of the first terminal and is facing the camera of the first terminal.

[0137] In some embodiments, Figure 6 Based on this, see Figure 8 Before step S604, the method may further include the following steps:

[0138] S607: Perform lip movement detection on the face image to be processed.

[0139] S608: If lip movement is detected in the user in the face image to be processed, check whether the user to which the audio to be processed belongs and the user to which the face image to be processed belongs are the same target user.

[0140] The first terminal performs lip movement detection on the face image to be processed. For example, the face image to be processed is input into a pre-trained lip movement detection model to obtain the detection result of whether the user in the face image to be processed has lip movement. If the user in the face image to be processed does not have lip movement, it means that the user within the field of view of the first terminal's camera has not made a sound. Therefore, there is no need to detect whether the first user and the second user are the same target user, and the audio to be processed can be suppressed directly.

[0141] If the user in the face image to be processed has lip movement, it indicates that the user within the field of view of the first terminal's camera has made a sound. Then, it continues to detect whether the user who made the sound is the same user as the user within the field of view of the first terminal's camera. This avoids misidentifying the sound made by other users as the sound made by the target user, which would lead to audio suppression errors. This further improves the accuracy of suppressing interfering audio, achieves audio noise reduction, and improves the user experience.

[0142] Regarding step S604, if the first user to whom the audio to be processed belongs and the second user to whom the face image to be processed belongs are the same target user, it indicates that the user who made the sound and the user within the field of view of the camera of the first terminal are the same target user. That is, the target user is using the first terminal to conduct an audio call with the user using the second terminal. Therefore, all audio in the original audio except for the target user's audio is interference audio. The first terminal then suppresses the interference audio except for the target user's audio to obtain the target audio. The target audio is noise-reduced and of higher quality. The first terminal then sends the target audio to the second terminal. Interference audio includes environmental noise and sounds made by users other than the target user.

[0143] In some embodiments, Figure 7 Based on this, see Figure 9 After detecting whether the user to whom the audio to be processed belongs and the user to whom the face image to be processed belongs are the same user, the method may further include the following steps:

[0144] S609: If the user to which the audio to be processed belongs and the user to which the face image to be processed belongs are not the same target user, the preset difference threshold is increased from the first value to the second value.

[0145] S610: For each template voiceprint feature, if the difference between the voiceprint feature of the audio to be processed and the template voiceprint feature is less than the adjusted preset difference threshold, determine that the voiceprint feature of the audio to be processed matches the template voiceprint feature, and determine that the user to which the template voiceprint feature that matches the voiceprint feature of the audio to be processed belongs is the user to which the audio to be processed belongs.

[0146] S611: If the user to which the audio to be processed belongs and the user to which the face image to be processed belongs are the same target user, suppress the interfering audio other than the target user's audio to be processed, obtain the target audio, and send the target audio to the second terminal.

[0147] If the first user to which the audio to be processed belongs and the second user to which the face image to be processed belongs are not the same target user, it indicates that the user who made the sound and the user within the field of view of the camera of the first terminal are not the same target user. It may be that the change in the voiceprint characteristics of the target user caused the audio to be processed to be identified as the audio of another user, such as the change in voiceprint characteristics caused by the target user having a cold. In this case, the preset difference threshold is increased from the first value to the second value.

[0148] The voiceprint features of the audio to be processed are compared again with the voiceprint features of each template. If the difference between the voiceprint features of the audio to be processed and any template voiceprint feature is less than the adjusted preset difference threshold, it is determined that the voiceprint features of the audio to be processed match the template voiceprint features. The user to whom the template voiceprint features that match the voiceprint features of the audio to be processed belong is determined to be the user to whom the audio to be processed belongs (i.e., the first user). In other words, the user to whom the matched template voiceprint features belong is determined to be the user who made the sound. Furthermore, if the first user to whom the audio to be processed belongs and the second user to whom the face image to be processed belongs are the same target user, it indicates that the user who made the sound and the user within the field of view of the camera of the first terminal are the same user. That is, the target user is using the first terminal to have an audio call with the user using the second terminal. In this case, all the audio in the original audio except for the target user's audio to be processed is interference audio. The first terminal then suppresses the interference audio except for the target user's audio to be processed to obtain the target audio. The target audio is noise-reduced and of higher quality. The first terminal then sends the target audio to the second terminal.

[0149] The first and second values ​​of the preset difference threshold can be determined by technicians based on experimental results. For example, voiceprint features that have not changed and voiceprint features that have changed from the same user are collected, and the voiceprint features collected in the two cases are compared to determine the preset difference threshold that would allow the two voiceprint features to be identified as the same user.

[0150] As can be seen from the above, the technical solution provided in this embodiment avoids the problem of recognition errors caused by changes in the voiceprint characteristics of the target user by adjusting the preset difference threshold, further improving the accuracy of suppressing interfering audio, achieving audio noise reduction, and improving the user experience.

[0151] In some embodiments, if the difference between the voiceprint features of the audio to be processed and the voiceprint features of each template is greater than the adjusted preset difference threshold, it indicates that the audio to be processed is not the voice of a registered user, nor is it the target user's voiceprint feature change that causes the audio to be processed to be identified as the audio of another user. The first terminal then determines that the audio to be processed is interference audio.

[0152] As can be seen from the above, the technical solution provided in this embodiment determines that the audio to be processed that does not match the voiceprint features of each template is interference audio. It can accurately identify interference audio and suppress it, further improving the accuracy of suppressing interference audio, achieving audio noise reduction, and improving the user experience.

[0153] In some embodiments, if the first user to whom the audio to be processed belongs and the second user to whom the face image to be processed belongs are still not the same user, indicating that the voiceprint feature change of the target user caused the audio to be processed to be identified as the audio of another user, the first terminal determines that the audio to be processed is interference audio.

[0154] As can be seen from the above, the technical solution provided in this embodiment determines that the audio to be processed where the facial image features and voiceprint features do not match as interference audio. It can accurately identify interference audio and suppress it, further improving the accuracy of suppressing interference audio, achieving audio noise reduction, and improving the user experience.

[0155] The following combination Figure 10 , Figure 11 and Figure 12 The data stream during audio noise reduction is explained. Figure 10 and Figure 11 For the process executed simultaneously by the first terminal, in order to clearly explain the audio data stream and image data stream, the following is provided: Figure 10 and Figure 11 They will be introduced separately.

[0156] See Figure 10 , Figure 10 This is a flowchart of an embodiment of the present application for acquiring audio data.

[0157] S1001: The APP initiates an audio / video call.

[0158] In this step, the APP refers to an application at the application layer that provides audio call functionality to the user, such as an instant messaging application, conferencing application, or social application. The APP initiates an audio / video call after receiving instructions from the user.

[0159] S1002: The APP calls the audio service.

[0160] In this step, the application framework layer includes various services of the first terminal, such as audio services (i.e., audio FWR), camera services, sensor services, etc. When the APP initiates an audio or video call according to the user's instructions, it calls the audio FWR in the application framework layer.

[0161] S1003: Audio FWR calls audio HAL to start recording.

[0162] In this step, the driver layer includes drivers for various hardware components in the first terminal, such as the audio driver (i.e., audio HAL), camera driver, and sensor driver. When the app initiates an audio / video call according to the user's instructions, it calls the audio service in the application framework layer. After receiving the app's recording request, the audio FWR calls the audio HAL to start the recording function.

[0163] S1004: Audio HAL activates DSP.

[0164] In this step, the audio HAL activates the DSP after detecting the call to the audio FWR.

[0165] S1005: Audio HAL calls the microphone to record audio.

[0166] In this step, after the audio HAL detects the call to the audio FWR, it calls the microphone to start recording to begin the recording service.

[0167] S1006: DSP loads audio preprocessing algorithm.

[0168] In this step, after the audio HAL activates the DSP, the DSP loads the audio preprocessing algorithm (i.e., acoustic echo cancellation, automatic gain control, and noise suppression in the above embodiment).

[0169] In this embodiment of the application, the execution order of the above steps S1005 and S1006 is not limited, and steps S1005 and S1006 can be executed simultaneously.

[0170] S1007: The microphone transmits the captured audio to the DSP.

[0171] In this step, after the microphone is activated, recording begins, and the recorded audio is transmitted to the DSP. Figure 12 Step 1: The microphone transmits audio data to the DSP.

[0172] S1008: DSP performs audio preprocessing.

[0173] In this step, the DSP preprocesses the audio transmitted from the microphone according to the loaded audio preprocessing algorithm to obtain the preprocessed audio (i.e., the original audio in the aforementioned embodiment). Figure 12 In this process, the DSP uses the 3A call algorithm to process the audio data transmitted from the microphone.

[0174] S1009: The DSP transmits the pre-processed audio to the audio HAL.

[0175] In this step, the DSP transmits the preprocessed audio to the audio HAL, i.e. Figure 12In step ②, the DSP transmits the audio processed using the 3A call algorithm to the audio HAL.

[0176] S1010: Audio HAL performs voiceprint noise reduction on pre-processed audio.

[0177] In this step, the audio HAL extracts the human voice portion (i.e., the audio to be processed) from the preprocessed audio and extracts the voiceprint features of the audio to be processed. The voiceprint features of the audio to be processed are compared with the locally stored template voiceprint features, and the user to which the template voiceprint features that match the voiceprint features of the audio to be processed belong is determined, thus obtaining the user to whom the audio to be processed belongs.

[0178] Furthermore, when the user to which the audio to be processed belongs and the user to which the face image to be processed belongs are the same target user, the audio HAL suppresses the interfering audio in the preprocessed audio except for the target user's audio, and obtains the denoised audio (i.e. the target audio).

[0179] That is Figure 12 In step ③, the audio HAL compares the voiceprint features of the audio to be processed with the template voiceprint features stored in the secure OS (i.e., the preset storage area). The secure OS stores N voiceprint features, namely the voiceprints of User 1, User 2, ..., User N in the voiceprint management. If the voiceprint of User 1 matches the voiceprint features of the audio to be processed, then User 1 is determined to be the user to whom the audio to be processed belongs. Furthermore, when the user to whom the face image to be processed belongs is also User 1, the audio HAL suppresses the interfering audio in the preprocessed audio except for the audio of User 1, obtaining the denoised audio (i.e., the target audio).

[0180] S1011: Audio HAL transmits the noise-reduced audio to Audio FWR.

[0181] In this step, the audio HAL transmits the noise-reduced audio to the audio FWR, i.e. Figure 12 Step 4: The audio HAL transmits the noise-reduced audio to the audio FWR.

[0182] S1012: Audio FWR performs audio encoding on the noise-reduced audio.

[0183] In this step, the audio FWR performs audio encoding on the noise-reduced audio transmitted via the audio HAL, that is... Figure 12 In the process, the audio FWR performs audio encoding on the noise-reduced audio transmitted by the audio HAL to obtain the recording file (AudioRecord).

[0184] S1013: Audio FWR transmits the encoded audio to the APP.

[0185] In this step, the audio FWR transmits the encoded audio to the APP, which is... Figure 12 Step 5: The audio FWR transmits the AudioRecord to the APP, and the APP then uses the AudioRecord to make audio calls with the second terminal.

[0186] See Figure 11 , Figure 11 This is a flowchart illustrating the acquisition of image data, as provided in an embodiment of this application.

[0187] S1101: The APP initiates an audio / video call.

[0188] In this step, the APP refers to an application at the application layer that provides audio and video calling functionality to the user, such as an instant messaging application, conferencing application, or social application. The APP initiates an audio and video call after receiving instructions from the user.

[0189] S1102: The APP calls the camera service.

[0190] In this step, the application framework layer includes various services of the first terminal, such as audio services, camera services (i.e., camera FWR), sensor services, etc. When the APP initiates an audio or video call according to the user's instructions, it calls the audio FWR in the application framework layer.

[0191] S1103: The camera's FWR calls the camera's HAL to start shooting.

[0192] In this step, the driver layer includes drivers for various hardware components in the first terminal, such as audio drivers, camera drivers (i.e., camera HAL), and sensor drivers. When the app initiates an audio / video call according to the user's instructions, it calls the camera FWR in the application framework layer. After receiving the app's shooting request, the camera FWR calls the camera HAL to start the shooting function.

[0193] S1104: Camera HAL Activated Compute Digital Signal Processor (CDSP).

[0194] In this step, the camera HAL activates the CDSP after detecting the call to the camera FWR.

[0195] S1105: Camera HAL calls the camera to take a picture.

[0196] In this step, after the camera HAL detects the call to the camera FWR, it calls the camera to take a picture to start the video recording service.

[0197] S1106: CDSP loading video preprocessing algorithm.

[0198] In this step, after the camera HAL activates the CDSP, the CDSP loads the video preprocessing algorithm (i.e., automatic exposure, automatic white balance, and automatic focus in the above embodiment).

[0199] In this embodiment of the application, the execution order of the above steps S1105 and S1106 is not limited, and steps S1105 and S1106 can be executed simultaneously.

[0200] S1107: The camera transmits the captured video to the DSP.

[0201] In this step, after the camera is turned on, it begins to record and transmits the captured video to the CDSP, i.e. Figure 12 Step 1: The camera transmits image data to the CDSP.

[0202] S1108: CDSP performs video preprocessing.

[0203] In this step, the CDSP preprocesses the image transmitted from the camera according to the loaded video preprocessing algorithm to obtain the preprocessed video image (i.e., the original video image in the aforementioned embodiment). Figure 12 In this process, CDSP performs image processing.

[0204] S1109: The CDSP transmits the pre-processed video to the camera HAL.

[0205] In this step, the CDSP transmits the preprocessed video to the camera HAL, i.e. Figure 12 In step ②, the CDSP transmits the processed video to the camera HAL.

[0206] S1110: Camera HAL performs face recognition and lip movement detection on the pre-processed video.

[0207] In this step, the camera's HAL extracts face images (i.e., the face images to be processed) from the preprocessed video and performs face recognition on these images. Specifically, it extracts the facial features of the face images to be processed. These features are then compared with locally stored template facial features to determine the user to whom the template facial features matching the face features of the face images to be processed belong, thus obtaining the user to whom the face images to be processed belong.

[0208] Furthermore, the camera HAL performs lip movement detection on the face image to be processed. When lip movement is detected in the face image to be processed, it checks whether the user to whom the audio to be processed belongs and the user to whom the face image to be processed belongs are the same target user.

[0209] That is Figure 12In step ③, the camera HAL compares the facial features of the face image to be processed with the template facial features stored in the security OS. The security OS stores N facial features, namely the facial features of user 1, user 2, ... user N in the face recognition management. If the facial features of user 1 match the facial features of the face image to be processed, then user 1 is determined to be the user to whom the face image to be processed belongs. Furthermore, when both the user to whom the face image to be processed belongs and the user to whom the audio to be processed belongs are user 1, the audio HAL suppresses the interfering audio in the preprocessed audio except for the audio of user 1, obtaining the denoised audio (i.e., the target audio).

[0210] S1111: Camera HAL transmits the processed video to camera FWR.

[0211] In this step, after facial recognition identifies the user to whom the face image to be processed belongs, the camera HAL transmits the original video image to the camera FWR, i.e. Figure 12 Step 4: The camera HAL transmits the raw video image to the camera FWR.

[0212] S1112: Camera FWR performs video encoding on the raw video images.

[0213] In this step, the camera FWR performs video encoding on the raw video images transmitted by the camera HAL, that is... Figure 12 In this process, the camera's FWR performs video encoding on the raw video images transmitted by the camera's HAL to obtain a video file.

[0214] S1113: The camera's FWR transmits the encoded video to the app.

[0215] In this step, the camera's FWR transmits the encoded video to the app, which is... Figure 12 Step 5: The camera's FWR transmits the video file to the APP, and the APP then uses the video file to make video calls with the second terminal.

[0216] As can be seen from the above, the technical solution provided in this embodiment, when the target user's audio to be processed and the template voiceprint features, as well as the target user's face image to be processed and the template face features, are all successfully matched, it indicates that the target user is the user conducting the audio call. That is, the target user's audio to be processed is not interfering audio. Therefore, suppressing interfering audio other than the target user's audio to be processed can improve the accuracy of suppressing interfering audio, achieve audio noise reduction, and improve the user experience.

[0217] The following describes the methods for registering template voiceprint features and template facial features on the first terminal. See [link / reference] Figure 13 , Figure 13A flowchart illustrating the registration of template voiceprint features and template face features is provided for embodiments of this application. The method includes the following steps:

[0218] S1301: Acquire the audio of the user to be registered and the facial image of the user to be registered.

[0219] S1302: Extract features from the audio to be registered to obtain the template voiceprint features of the audio to be registered, and extract features from the face image to be registered to obtain the template face features of the face image to be registered.

[0220] S1303: Perform lip movement detection on the face image to be registered.

[0221] S1304: If lip movement is detected in the face image to be registered, the template voiceprint features of the audio to be registered, the template face features of the face image to be registered, and the user identifier of the user to be registered are recorded in the preset storage area of ​​the first terminal.

[0222] As can be seen from the above, the technical solution provided in this embodiment can record the template voiceprint features of the audio to be registered, the template face features of the face image to be registered, and the user identifier of the user to be registered. Subsequently, if the target user's audio to be processed matches the template voiceprint features, and the target user's face image matches the template face features, it indicates that the target user is the user conducting the audio call, that is, the target user's audio to be processed is not interfering audio. Therefore, suppressing interfering audio other than the target user's audio to be processed can improve the accuracy of interfering audio suppression, achieve audio noise reduction, and improve the user experience.

[0223] During the registration phase, the first terminal acquires audio captured by the microphone, performs audio preprocessing on the acquired audio, and extracts the human voice portion from the preprocessed audio to obtain the registration audio for the user to be registered. The audio preprocessing method is described in the foregoing embodiments. Additionally, the terminal acquires a user image of the user to be registered captured by the camera, performs video preprocessing on the acquired user image, and extracts the face image from the preprocessed video user image to obtain the registration face image. The video preprocessing method is described in the foregoing embodiments.

[0224] In one application scenario, when the first terminal is not engaged in audio or video calls with other terminals, the first terminal displays preset text according to the user's instructions, causing the user to read the preset text aloud. Correspondingly, the first terminal acquires the sound emitted by the user while reading the preset text, extracts the human voice portion to obtain the audio to be registered, and acquires the user image while the user reads the preset text, extracts the facial image to obtain the facial image to be registered.

[0225] In another application scenario, during an audio / video call between the first terminal and another terminal, the first terminal acquires audio collected within a preset time period, extracts the human voice portion to obtain the audio to be registered, and acquires user images collected within the same preset time period, extracts the facial images to obtain the facial images to be registered. The duration of the preset time period is set according to requirements.

[0226] Next, the first terminal extracts features from the audio to be registered, obtaining template voiceprint features, and extracts features from the face image to be registered, obtaining template face features. Furthermore, the first terminal performs lip movement detection on the face image to be processed. If lip movement is detected in the face image to be registered, it indicates that the user within the terminal's camera's field of view is making a sound. In other words, the user within the terminal's camera's field of view is performing voiceprint and face image registration. The first terminal then binds the template voiceprint features and template face features of the user to be registered. This means that the template voiceprint features of the audio to be registered, the template face features of the face image to be registered, and the user identifier of the user to be registered are recorded in a local preset storage area. The user identifier of the user to be registered can be a unique identifier such as a username or user number.

[0227] The default storage area is a secure storage area on the first terminal's local machine, such as a secure OS. The secure OS is a secure storage area within the first terminal's operating system. Developers can access this secure storage area through a specified encrypted interface, while other personnel cannot access it. This enhances the security of registered template voiceprint and facial features, protecting user privacy.

[0228] The following combination Figure 14 , Figure 15 and Figure 16 The data stream during audio registration is explained. Figure 14 and Figure 15 For the process executed simultaneously by the first terminal, and to clearly explain the audio data stream and image data stream, Figure 14 and Figure 15 They will be introduced separately.

[0229] See Figure 14 , Figure 14 This is a flowchart for obtaining template voiceprint features, provided as an embodiment of this application.

[0230] S1401: Receive user instruction.

[0231] In this step, the APP can be used for settings. When the first terminal is not engaged in audio or video calls with other terminals, the settings application on the first terminal displays preset text according to the user's instructions, which means the settings application has received instructions to capture audio.

[0232] Alternatively, an app can be an application within the application layer that provides audio call functionality to users, such as instant messaging, conferencing, or social networking applications. The app initiates an audio or video call after receiving instructions from the user.

[0233] S1402: The APP calls the audio service.

[0234] In this step, the application framework layer includes various services of the first terminal, such as audio service (i.e., audio FWR), camera service, sensor service, etc. After receiving the instruction to collect audio, the APP calls the audio FWR in the application framework layer.

[0235] S1403: Audio FWR calls audio HAL to start recording.

[0236] In this step, the driver layer includes drivers for various hardware components in the first terminal, such as the audio driver (i.e., audio HAL), camera driver, and sensor driver. When the app initiates an audio / video call according to the user's instructions, it calls the audio service in the application framework layer. After receiving the app's recording request, the audio FWR calls the audio HAL to start the recording function.

[0237] S1404: Audio HAL activates DSP.

[0238] In this step, the audio HAL activates the DSP after detecting the call to the audio FWR.

[0239] S1405: Audio HAL calls the microphone to record.

[0240] In this step, after the audio HAL detects the call to the audio FWR, it calls the microphone to start recording to begin the recording service.

[0241] S1406: DSP loads audio preprocessing algorithm.

[0242] In this step, after the audio HAL activates the DSP, the DSP loads the audio preprocessing algorithm (i.e., acoustic echo cancellation, automatic gain control, and noise suppression in the above embodiment).

[0243] In this embodiment of the application, the execution order of the above steps S1405 and S1406 is not limited, and steps S1405 and S1406 can be executed simultaneously.

[0244] S1407: The microphone transmits the captured audio to the DSP.

[0245] In this step, after the microphone is activated, recording begins, and the recorded audio is transmitted to the DSP. Figure 16 Step 1: The microphone transmits audio data to the DSP.

[0246] S1408: DSP performs audio preprocessing.

[0247] In this step, the DSP preprocesses the audio transmitted from the microphone according to the loaded audio preprocessing algorithm to obtain the preprocessed audio, i.e. Figure 16 In this process, the DSP uses the 3A call algorithm to process the audio data transmitted from the microphone.

[0248] S1409: The DSP transmits the pre-processed audio to the audio HAL.

[0249] In this step, the DSP transmits the preprocessed audio to the audio HAL, i.e. Figure 16 In step ②, the DSP transmits the audio processed using the 3A call algorithm to the audio HAL.

[0250] S1410: Audio HAL extracts voiceprint features from preprocessed audio.

[0251] In this step, the audio HAL extracts the human voice portion (i.e., the audio to be registered) from the preprocessed audio and extracts the voiceprint features of the audio to be registered.

[0252] S1411: Audio HAL saves voiceprint features as template voiceprint features.

[0253] In this step, when lip movement is detected in the face image of the user to be registered, the audio HAL saves the voiceprint features of the audio to be registered as template voiceprint features. That is, it records the template voiceprint features of the audio to be registered, the template face features of the face image to be registered, and the user identifier of the user to be registered in the local preset storage area.

[0254] That is Figure 16 In step ③, the audio HAL saves the voiceprint features of the audio to be registered as template voiceprint features in the security OS. The security OS stores N voiceprint features, namely the voiceprint of user 1, the voiceprint of user 2, ..., the voiceprint of user N in voiceprint management.

[0255] In one application scenario, when the APP is in the settings application, the registration process ends after the template voiceprint features of the audio to be registered, the template facial features of the face image to be registered, and the user identifier of the user to be registered are recorded in the local preset storage area.

[0256] In another application scenario, when the APP is an application providing audio call functionality to users within the application layer, it records the template voiceprint features of the audio to be registered, the template facial features of the face image to be registered, and the user identifier of the user to be registered in a local preset storage area. After these correspondences are matched, the first terminal executes the subsequent audio noise reduction process. (Audio noise reduction process reference) Figure 10 , Figure 11 and Figure 12 Related information.

[0257] Figure 16 The audio data stream processing procedures in steps ④ and ⑤ are similar to those in the middle. Figure 12 The processing of the audio data stream in steps ④ and ⑤ is similar, as described in the foregoing embodiments.

[0258] See Figure 15 , Figure 15 This is a flowchart for obtaining template facial features, provided as an embodiment of this application.

[0259] S1501: Receive user instruction.

[0260] In this step, the APP can be used for settings. When the first terminal is not engaged in audio or video calls with other terminals, the settings application on the first terminal displays preset text according to the user's instructions, which means the settings application has received instructions to capture audio.

[0261] Alternatively, an app can be an application within the application layer that provides audio call functionality to users, such as instant messaging, conferencing, or social networking applications. The app initiates an audio or video call after receiving instructions from the user.

[0262] S1502: The APP calls the camera service.

[0263] In this step, the application framework layer includes various services of the first terminal, such as audio services, camera services (i.e., camera FWR), sensor services, etc. When the APP initiates an audio or video call according to the user's instructions, it calls the audio FWR in the application framework layer.

[0264] S1503: The camera's FWR calls the camera's HAL and starts shooting.

[0265] In this step, the driver layer includes drivers for various hardware components in the first terminal, such as audio drivers, camera drivers (i.e., camera HAL), and sensor drivers. When the app initiates an audio / video call according to the user's instructions, it calls the camera FWR in the application framework layer. After receiving the app's shooting request, the camera FWR calls the camera HAL to start the shooting function.

[0266] S1504: Camera HAL activates CDSP.

[0267] In this step, the camera HAL activates the CDSP after detecting the call to the camera FWR.

[0268] S1505: Camera HAL calls the camera to take a picture.

[0269] In this step, after the camera HAL detects the call to the camera FWR, it calls the camera to take a picture to start the video recording service.

[0270] S1506: CDSP loading video preprocessing algorithm.

[0271] In this step, after the camera HAL activates the CDSP, the CDSP loads the video preprocessing algorithm (i.e., automatic exposure, automatic white balance, and automatic focus in the above embodiment).

[0272] In this embodiment of the application, the execution order of the above steps S1505 and S1506 is not limited, and steps S1505 and S1506 can be executed simultaneously.

[0273] S1507: The camera transmits the captured video to the DSP.

[0274] In this step, after the camera is turned on, it begins to record and transmits the captured video to the CDSP, i.e. Figure 16 Step 1: The camera transmits image data to the CDSP.

[0275] S1508: CDSP performs video preprocessing.

[0276] In this step, the CDSP preprocesses the images transmitted from the camera according to the loaded video preprocessing algorithm to obtain the preprocessed video image, i.e. Figure 16 In this process, CDSP performs image processing.

[0277] S1509: The CDSP transmits the pre-processed video to the camera HAL.

[0278] In this step, the CDSP transmits the preprocessed video to the camera HAL, i.e. Figure 16 In step ②, the CDSP transmits the processed video to the camera HAL.

[0279] S1510: Camera HAL performs face detection, feature extraction, and lip movement detection on the pre-processed video.

[0280] In this step, the camera's HAL performs face detection on the preprocessed video images and extracts the face images (i.e., the face images to be registered) from the preprocessed video images. Then, it extracts the facial features of the face images to be registered. Furthermore, the camera's HAL performs lip movement detection on the face images to be registered.

[0281] S1511: The camera HAL saves facial features as template facial features.

[0282] In this step, when the camera HAL detects lip movements in the face image of the user to be registered, it saves the facial features as template facial features.

[0283] That is Figure 16 In step ③, the camera HAL saves the facial features of the face image to be registered as template facial features in the secure OS. The secure OS stores N facial features, namely the facial features of User 1, User 2, ..., User N in the face recognition management.

[0284] In one application scenario, when the APP is in the settings application, the registration process ends after the template voiceprint features of the audio to be registered, the template facial features of the face image to be registered, and the user identifier of the user to be registered are recorded in the local preset storage area.

[0285] In another application scenario, when the APP is an application providing audio call functionality to users within the application layer, it records the template voiceprint features of the audio to be registered, the template facial features of the face image to be registered, and the user identifier of the user to be registered in a local preset storage area. After these correspondences are matched, the first terminal executes the subsequent audio noise reduction process. (Audio noise reduction process reference) Figure 10 , Figure 11 and Figure 12 Related information.

[0286] Figure 16 The audio data stream processing procedures in steps ④ and ⑤ are similar to those in the middle. Figure 12 The processing of the audio data stream in steps ④ and ⑤ is similar, as described in the foregoing embodiments.

[0287] As can be seen from the above, the technical solution provided in this embodiment can record the template voiceprint features of the audio to be registered, the template face features of the face image to be registered, and the user identifier of the user to be registered. Subsequently, if the target user's audio to be processed matches the template voiceprint features, and the target user's face image matches the template face features, it indicates that the target user is the user conducting the audio call, that is, the target user's audio to be processed is not interfering audio. Therefore, suppressing interfering audio other than the target user's audio to be processed can improve the accuracy of interfering audio suppression, achieve audio noise reduction, and improve the user experience.

[0288] In a specific implementation, this application also provides a terminal, which includes one or more processors and a memory; the memory is coupled to one or more processors, and the memory is used to store computer program code, which includes computer instructions, and one or more processors call the computer instructions to cause the terminal to perform some or all of the steps in the above method embodiments.

[0289] This application also provides a computer-readable storage medium including a computer program that, when run on a terminal, causes the terminal to perform some or all of the steps described in the method embodiments. The storage medium may be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.

[0290] In a specific implementation, this application also provides a computer program product, which includes executable instructions. When the executable instructions are executed on a terminal, the terminal performs some or all of the steps in the above method embodiments.

[0291] like Figure 17 As shown, this application also provides a chip system applied to a terminal. The chip system includes one or more processors 1701. The processors 1701 are used to call computer instructions to cause the terminal to input data to be processed into the chip system. The chip system processes the data based on the audio processing method provided in the embodiments of this application and outputs the processing result.

[0292] In one possible implementation, the chip system also includes input and output interfaces for inputting and outputting data.

[0293] The various embodiments of the mechanisms disclosed in this application can be implemented in hardware, software, firmware, or a combination of these implementation methods. Embodiments of this application can be implemented as computer programs or program code executable on a programmable system, the programmable system including at least one processor, a storage system (including volatile and non-volatile memory and / or storage elements), at least one input device, and at least one output device.

[0294] Program code can be applied to input instructions to execute the functions described in this application and generate output information. The output information can be applied to one or more output devices in a known manner. For the purposes of this application, the processing system includes any system having a processor such as, for example, a Digital Signal Processor (DSP), a microcontroller, an Application Specific Integrated Circuit (ASIC), or a microprocessor.

[0295] The program code can be implemented using a high-level procedural language or an object-oriented programming language to communicate with the processing system. Assembly language or machine language can also be used when needed. In fact, the mechanisms described in this application are not limited to any particular programming language. In either case, the language can be a compiled language or an interpreted language.

[0296] In some cases, the disclosed embodiments may be implemented in hardware, firmware, software, or any combination thereof. The disclosed embodiments may also be implemented as instructions carried or stored thereon on one or more temporary or non-temporary machine-readable (e.g., computer-readable) storage media, which may be read and executed by one or more processors. For example, the instructions may be distributed via a network or through other computer-readable media. Therefore, machine-readable media may include any mechanism for storing or transmitting information in a machine-readable (e.g., computer-readable) form, including but not limited to floppy disks, optical disks, CD-ROMs, compact disc read-only memory (CD-ROMs), magneto-optical disks, read-only memory, random access memory, erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), magnetic cards or optical cards, flash memory, or tangible machine-readable storage for transmitting information (e.g., carrier waves, infrared signals, digital signals, etc.) using the Internet in the form of electrical, optical, acoustic, or other forms of propagated signals. Therefore, machine-readable media includes any type of machine-readable medium suitable for storing or transmitting electronic instructions or information in a machine-readable (e.g., computer-readable) form.

[0297] In the accompanying drawings, some structural or methodological features may be shown in a specific arrangement and / or order. However, it should be understood that such a specific arrangement and / or order may not be necessary. Rather, in some embodiments, these features may be arranged in a manner and / or order different from that shown in the accompanying drawings. Furthermore, including structural or methodological features in a particular figure does not imply that such features are required in all embodiments, and in some embodiments, these features may be omitted or may be combined with other features.

[0298] It should be noted that all units / modules mentioned in the device embodiments of this application are logical units / modules. Physically, a logical unit / module can be a physical unit / module, a part of a physical unit / module, or a combination of multiple physical units / modules. The physical implementation of these logical units / modules themselves is not the most important factor; the combination of functions implemented by these logical units / modules is the key to solving the technical problems proposed in this application. Furthermore, to highlight the innovative aspects of this application, the above-described device embodiments of this application have not introduced units / modules that are not closely related to solving the technical problems proposed in this application. This does not mean that the above-described device embodiments do not contain other units / modules.

[0299] It should be noted that in the examples and description of this patent, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one" does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0300] Although this application has been illustrated and described with reference to certain preferred embodiments thereof, those skilled in the art should understand that various changes in form and detail may be made thereto without departing from the spirit and scope of this application.

Claims

1. An audio processing method, characterized in that, The method is applied to a first terminal, and the method includes: During the audio call between the first terminal and the second terminal, the collected audio to be processed and the face image to be processed are acquired. For each audio file to be processed, feature extraction is performed to obtain the voiceprint features of that audio file. For each template voiceprint feature, if the difference between the voiceprint feature of the audio to be processed and the template voiceprint feature is less than a preset difference threshold, it is determined that the voiceprint feature of the audio to be processed matches the template voiceprint feature, and the user to which the template voiceprint feature belongs is determined to be the user to which the audio to be processed belongs. For each template face feature, if the face feature of the face image to be processed matches the template face feature, the user to which the template face feature belongs is determined to be the user to which the face image to be processed belongs. If the user to which the audio to be processed belongs and the user to which the face image to be processed belongs are not the same target user, the preset difference threshold is increased from the first value to the second value; For each template voiceprint feature, if the difference between the voiceprint feature of the audio to be processed and the template voiceprint feature is less than the adjusted difference threshold, it is determined that the voiceprint feature of the audio to be processed matches the template voiceprint feature, and the user to which the template voiceprint feature that matches the voiceprint feature of the audio to be processed belongs is determined to be the user to which the audio to be processed belongs. If the user to which the audio to be processed belongs and the user to which the face image to be processed belongs are the same target user, the interfering audio other than the audio to be processed of the target user is suppressed to obtain the target audio, and the target audio is sent to the second terminal.

2. The method according to claim 1, characterized in that, After determining that the user to whom the template voiceprint features matching the voiceprint features of the audio to be processed belong is the user to whom the audio to be processed belongs, the method further includes: If the user to which the audio to be processed belongs and the user to which the face image to be processed belongs are not the same user, the audio to be processed is determined to be interference audio.

3. The method according to claim 2, characterized in that, After increasing the preset difference threshold from a first value to a second value if the user to which the audio to be processed belongs and the user to which the face image to be processed belongs are not the same target user, the method further includes: If the difference between the voiceprint features of the audio to be processed and the voiceprint features of each template is greater than the adjusted preset difference threshold, the audio to be processed is determined to be interference audio.

4. The method according to claim 1, characterized in that, Before the step of suppressing interfering audio (excluding the target user's audio) to obtain the target audio and sending the target audio to the second terminal, if the user to which the audio to be processed belongs and the user to which the face image to be processed belongs are the same target user, the method further includes: Perform lip movement detection on the face image to be processed; If lip movement is detected in the user in the face image to be processed, it is determined whether the user to whom the audio to be processed belongs and the user to whom the face image to be processed belongs are the same target user.

5. The method according to claim 1, characterized in that, The acquisition of the collected audio to be processed and the face image to be processed includes: The audio collected by the microphone of the first terminal is acquired, and the acquired audio is preprocessed to obtain the original audio. The human voice part of the original audio is extracted to obtain the audio to be processed. The audio preprocessing includes: acoustic echo cancellation, automatic gain control and noise suppression. The system acquires video images containing user images captured by a camera, performs image preprocessing on the acquired video images to obtain raw video images, and extracts face images from the raw video images to obtain face images to be processed; wherein, the image preprocessing includes: automatic exposure, automatic white balance and automatic focus.

6. The method according to claim 1, characterized in that, Before determining, for each template voiceprint feature, if the voiceprint feature of the audio to be processed matches the template voiceprint feature, and thus determining that the user to which the template voiceprint feature belongs is the user to which the audio to be processed belongs, the method further includes: Acquire the audio of the user to be registered, as well as the facial image of the user to be registered; Feature extraction is performed on the audio to be registered to obtain the template voiceprint features of the audio to be registered, and feature extraction is performed on the face image to be registered to obtain the template face features of the face image to be registered; Lip movement detection is performed on the face image to be registered; If lip movement is detected in the face image to be registered, the template voiceprint features of the audio to be registered, the template face features of the face image to be registered, and the user identifier of the user to be registered are recorded in the preset storage area of ​​the first terminal.

7. The method according to claim 6, characterized in that, The preset storage area is the secure storage area of ​​the first terminal.

8. An audio processing terminal, characterized in that, include: One or more processors and memory; The memory is coupled to the one or more processors, the memory being used to store computer program code, the computer program code including computer instructions, the one or more processors invoking the computer instructions to cause the terminal to perform the method as described in any one of claims 1-7.

9. A computer-readable storage medium, characterized in that, Includes a computer program that, when run on a terminal, causes the terminal to perform the method of any one of claims 1-7.

10. A computer program product, characterized in that, The computer program product includes executable instructions that, when executed on a terminal, cause the terminal to perform the method of any one of claims 1-7.

11. A chip system, characterized in that, The chip system is applied to a terminal. The chip system includes one or more processors. The processors are used to call computer instructions to cause the terminal to input data into the chip system and to execute the method of any one of claims 1-7 to process the data and output the processing result.

Citation Information

Patent Citations

  • User authentication method and device on basis of audios and videos

    CN103973441A

  • Fusion type voice recognition method, device and system, equipment and storage medium

    CN111883130A

  • Voice noise reduction method and device based on voiceprint recognition, equipment and medium

    CN116312570A