Sound signal processing method and electronic equipment

By processing the video and sound signals recorded by electronic devices through microphones, the sound energy and diffuse field noise in non-target directions are reduced, solving the problem of interference signals when recording videos and improving the quality of sound signals and user experience.

CN115914517BActive Publication Date: 2025-09-30BEIJING HONOR DEVICE CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202110927121.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-08-12
Publication Date
2025-09-30
Estimated Expiration
2041-08-12

AI Technical Summary

Technical Problem

When recording videos, existing electronic devices will collect sounds of non-target objects and environmental noise, resulting in low recorded sound quality and more interference signals.

Method used

The sound signal is collected and processed by the microphone to reduce the sound energy in non-target directions, suppress diffuse field noise, and improve the sound signal quality of the target object.

Benefits of technology

Reduces interfering sound signals in recorded videos, improves sound signal quality and signal-to-noise ratio, and provides a better user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115914517B_ABST
    Figure CN115914517B_ABST
Patent Text Reader

Abstract

An embodiment of the present application provides a sound signal processing method and an electronic device, which can reduce interfering sound signals in the sound of recorded videos and improve the quality of the sound signals of recorded videos. The method is applied to electronic devices. The method includes: the electronic device obtains a first sound signal. The first sound signal is the sound signal of the recorded video. The electronic device processes the first sound signal to obtain a second sound signal. When playing the recorded video file, the electronic device outputs the second sound signal. The energy of the sound signal in the non-target direction in the second sound signal is lower than the energy of the sound signal in the non-target direction in the first sound signal. The non-target direction is the direction outside the field of view of the camera when recording the video.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of electronic technology, and in particular to a sound signal processing method and electronic equipment. Background Art

[0002] Currently, the recording function of electronic devices has become a frequently used function. With the development of short videos and live social software (such as Kuaishou and Douyin), recording high-quality video files has become a demand.

[0003] Existing electronic devices collect sound signals from their surroundings when recording video, but some of these signals are interference signals and are not what the user wants. For example, when recording a selfie or live stream using a front-facing camera, the device will collect both the user's own voice and the surrounding sounds. This results in the recorded selfie sound being unclear, containing a lot of interference, and resulting in low sound quality. Summary of the Invention

[0004] The embodiments of the present application provide a sound signal processing method and an electronic device, which can reduce interfering sound signals in the sound of recorded videos and improve the quality of the sound signals of recorded videos.

[0005] To achieve the above objectives, this application provides the following technical solutions:

[0006] In a first aspect, embodiments of the present application provide a method for processing sound signals. The method is applied to an electronic device comprising a camera and a microphone. A first target object is within the camera's shooting range, while a second target object is not. The first target object being within the camera's shooting range may mean that the first target object is within the camera's field of view. The method comprises: the electronic device activates the camera. A preview interface is displayed, the preview interface including a first control. A first operation on the first control is detected. In response to the first operation, shooting begins. At a first moment, a shooting interface is displayed, the shooting interface including a first image, the first image being an image captured in real time by the camera, the first image including the first target object, and the first image not including the second target object. The first moment may be any moment during the shooting process. At the first moment, the microphone captures first audio, the first audio including a first audio signal and a second audio signal, the first audio signal corresponding to the first target object, and the second audio signal corresponding to the second target object. A second operation on the first control on the shooting interface is detected. In response to the second operation, shooting is stopped and the first video is saved, wherein the first moment of the first video includes a first image and a second audio, the second audio includes a first audio signal and a third audio signal, the third audio signal is obtained by the electronic device processing the second audio signal, and the energy of the third audio signal is less than the energy of the second audio signal.

[0007] Generally speaking, when a user uses an electronic device to record a video, the device's microphone will collect sound signals from the device's surroundings. For example, the device will collect sound signals within the camera's field of view when recording the video, as well as sound signals outside the camera's field of view when recording the video, and also collect ambient noise. In this case, the sound signals outside the camera's field of view and ambient noise during video recording become interference signals.

[0008] For example, when the electronic device records the sound signal (i.e., the second audio signal) of the second target object (such as non-target object 1 or non-target object), the energy of the second audio signal can be reduced to obtain a third audio signal. In this way, in an embodiment of the present application, the electronic device can process the sound signal of the recorded video (such as the sound signal collected by the microphone) and reduce the energy of the interference signal (such as the energy of the second audio signal), so that when the recorded video file is played, the energy of the output third audio signal is lower than the energy of the sound signal of the non-target orientation in the second audio signal, so as to reduce the interference sound signal in the sound signal of the recorded video and improve the quality of the sound signal of the recorded video.

[0009] In a possible implementation, the third audio signal is obtained by the electronic device processing the second audio signal, including: configuring the gain of the second audio signal to be less than 1. The third audio signal is obtained according to the second audio signal and the gain of the second audio signal.

[0010] In one possible implementation, the third audio signal is obtained by the electronic device processing the second audio signal, including: the electronic device calculates the probability that the second audio signal is within the target direction. The target direction is the direction within the field of view of the camera when recording the video. The first target object is within the target direction, and the second target object is not within the target direction. The electronic device determines the gain of the second audio signal based on the probability that the second audio signal is within the target direction. If the probability that the second audio signal is within the target direction is greater than a preset probability threshold, the gain of the second audio signal is equal to 1. If the probability that the second audio signal is within the target direction is less than or equal to the preset probability threshold, the gain of the second audio signal is less than 1. The electronic device obtains the third audio signal based on the energy of the second audio signal and the gain of the second audio signal.

[0011] In this solution, the electronic device may determine the gain of the second audio signal according to the probability that the second audio signal is within the target direction, so as to reduce the energy of the second audio signal and obtain the third audio signal.

[0012] In one possible implementation, the first audio signal further includes a fourth audio signal, which is a diffuse-field noise audio signal. The second audio signal further includes a fifth audio signal, which is a diffuse-field noise audio signal. The fifth audio signal is obtained by the electronic device processing the fourth audio signal, and the energy of the fifth audio signal is less than that of the fourth audio signal.

[0013] In a possible implementation, the fifth audio signal is obtained by the electronic device processing the fourth audio signal, including: configuring the gain of the fourth audio signal to be less than 1. The fifth audio signal is obtained according to the energy of the fourth audio signal and the gain of the fourth audio signal.

[0014] In one possible implementation, the fifth audio signal is obtained by the electronic device processing the fourth audio signal, including: suppressing the fourth audio signal to obtain a sixth audio signal; and compensating the sixth audio signal to obtain the fifth audio signal. The sixth audio signal is a diffuse-field noise audio signal, and has less energy than the fourth audio signal, and the sixth audio signal is less than the fifth audio signal.

[0015] It should be noted that during the processing of the fourth audio signal, the energy of the sixth audio signal obtained from the fourth audio signal may be very low, resulting in an uneven diffuse field noise. Therefore, by performing noise compensation on the sixth audio signal, the energy of the fifth audio signal obtained from the processing of the fourth audio signal can be made more stable, thereby improving the user's listening experience.

[0016] In one possible implementation, the method further includes: after the microphone collects the first audio at the first moment, the electronic device processes the first audio to obtain the second audio. In other words, the electronic device can process the audio signal in real time when the audio signal is collected.

[0017] In one possible implementation, the method further includes: after stopping recording in response to the second operation, the electronic device processes the first audio to obtain a second audio. In other words, the electronic device can obtain a sound signal from the video file when recording of the video file ends. Then, the sound signal is processed frame by frame in chronological order.

[0018] In a second aspect, an embodiment of the present application provides a method for processing a sound signal. The method is applied to an electronic device. The method includes: the electronic device obtains a first sound signal. The first sound signal is a sound signal of a recorded video. The electronic device processes the first sound signal to obtain a second sound signal. When playing the recorded video file, the electronic device outputs the second sound signal. The energy of the sound signal in the non-target direction in the second sound signal is lower than the energy of the sound signal in the non-target direction in the first sound signal. The non-target direction is the direction outside the field of view of the camera when recording the video.

[0019] Generally speaking, when a user uses an electronic device to record a video, the electronic device will collect sound signals around the electronic device through the microphone. For example, the electronic device will collect sound signals within the camera's field of view when recording the video, as well as sound signals outside the camera's field of view when recording the video, and the electronic device will also collect ambient noise.

[0020] In an embodiment of the present application, the electronic device can process the sound signal of the recorded video (such as the sound signal collected by the microphone) and suppress the sound signal of the non-target direction in the sound signal, so that when the recorded video file is played, the energy of the sound signal of the non-target direction in the output second sound signal is lower than the energy of the sound signal of the non-target direction in the first sound signal, so as to reduce the interfering sound signal in the sound signal of the recorded video and improve the quality of the sound signal of the recorded video.

[0021] In one possible implementation, the electronic device acquires the first sound signal, including: the electronic device collects the first sound signal in real time through a microphone in response to a first operation, wherein the first operation is used to trigger the electronic device to start recording or live broadcasting.

[0022] For example, the electronic device can collect the first sound signal in real time through the microphone when the camera's recording function is activated and recording begins. For another example, the electronic device can collect the sound signal in real time through the microphone when a live broadcast application (such as Tik Tok or Kuaishou) is activated to start a live video broadcast. During the recording or live broadcast process, the electronic device processes each frame of the sound signal it collects.

[0023] In one possible implementation, before the electronic device acquires the first sound signal, the method further includes: the electronic device recording a video file. The electronic device acquiring the first sound signal includes: the electronic device acquiring the first sound signal from the video file in response to completion of video file recording.

[0024] For example, the electronic device can obtain a sound signal from the video file when the video file recording is finished, and then process the sound signal frame by frame in chronological order.

[0025] In one possible implementation, the electronic device acquires the first sound signal, including: the electronic device acquires the first sound signal from a video file stored in the electronic device in response to a second operation, wherein the second operation is used to trigger the electronic device to process the video file to improve the sound quality of the video file.

[0026] For example, an electronic device processes the sound in a video file stored locally on the electronic device. When the electronic device detects a user instruction to process the video file (e.g., clicking a "denoising" option button in a video file operation interface), the electronic device begins to acquire the sound signal of the video file and processes the sound signal frame by frame in chronological order.

[0027] In one possible implementation, a first sound signal includes multiple time-frequency speech signals. The electronic device processes the first sound signal to obtain a second sound signal, including: the electronic device identifying the orientation of each time-frequency speech signal in the first sound signal. If the orientation of a first time-frequency speech signal in the first sound signal is a non-target orientation, the electronic device reduces the energy of the first time-frequency speech signal to obtain the second sound signal. The first time-frequency speech signal is any one of the multiple time-frequency speech signals in the first sound signal.

[0028] In one possible implementation, the first sound signal includes multiple time-frequency voice signals. The electronic device processes the first sound signal to obtain a second sound signal, including: the electronic device calculates the probability that each time-frequency voice signal in the first sound signal is within a target orientation. The target orientation is the orientation within the field of view of the camera when recording the video. The electronic device determines the gain of the second time-frequency voice signal based on the probability that the second time-frequency voice signal in the first sound signal is within the target orientation; the second time-frequency voice signal is any one of the multiple time-frequency voice signals in the first sound signal; if the probability that the second time-frequency voice signal is within the target orientation is greater than a preset probability threshold, the gain of the second time-frequency voice signal is equal to 1; if the probability that the second time-frequency voice signal is within the target orientation is less than or equal to the preset probability threshold, the gain of the second time-frequency voice signal is less than 1. The electronic device obtains the second sound signal based on each time-frequency voice signal in the first sound signal and the corresponding gain.

[0029] In one possible implementation, the energy of the diffuse-field noise in the second sound signal is lower than the energy of the diffuse-field noise in the first sound signal. It should be understood that reducing the energy of non-target sound signals in the first sound signal does not necessarily reduce all diffuse-field noise. To ensure the quality of the recorded video sound signal, it is also necessary to reduce the diffuse-field noise to improve the signal-to-noise ratio of the recorded video sound signal.

[0030] In one possible implementation, a first sound signal includes multiple time-frequency speech signals. An electronic device processes the first sound signal to obtain a second sound signal, including: identifying, by the electronic device, whether each time-frequency speech signal in the first sound signal is diffuse-field noise. If a third time-frequency speech signal in the first sound signal is diffuse-field noise, the electronic device reduces the energy of the third time-frequency speech signal to obtain the second sound signal. The third time-frequency speech signal is any one of the multiple time-frequency speech signals in the first sound signal.

[0031] In one possible implementation, the first sound signal includes multiple time-frequency voice signals. The electronic device processes the first sound signal to obtain the second sound signal, further comprising: the electronic device identifying whether each time-frequency voice signal in the first sound signal is diffuse field noise. The electronic device determines the gain of the fourth time-frequency voice signal based on whether the fourth time-frequency voice signal in the first sound signal is diffuse field noise; wherein the fourth time-frequency voice signal is any one of the multiple time-frequency voice signals in the first sound signal; if the fourth time-frequency voice signal is diffuse field noise, the gain of the fourth time-frequency voice signal is less than 1; if the fourth time-frequency voice signal is a coherent signal, the gain of the fourth time-frequency voice signal is equal to 1. The electronic device obtains the second sound signal based on each time-frequency voice signal in the first sound signal and the corresponding gain.

[0032] In one possible implementation, a first sound signal includes multiple time-frequency speech signals; the electronic device processes the first sound signal to obtain a second sound signal, further comprising: calculating a probability that each time-frequency speech signal in the first sound signal is within a target location; the target location is a location within the field of view of a camera during video recording; and identifying whether each time-frequency speech signal in the first sound signal is diffuse-field noise. The electronic device determines the gain of the fifth time-frequency voice signal in the first sound signal based on the probability that the fifth time-frequency voice signal is within the target direction and whether the fifth time-frequency voice signal is diffuse field noise. The fifth time-frequency voice signal is any one of the multiple time-frequency voice signals in the first sound signal. If the probability that the fifth time-frequency voice signal is within the target direction is greater than a preset probability threshold and the fifth time-frequency voice signal is a coherent signal, the gain of the fifth time-frequency voice signal is equal to 1. If the probability that the fifth time-frequency voice signal is within the target direction is greater than a preset probability threshold and the fifth time-frequency voice signal is diffuse field noise, the gain of the fifth time-frequency voice signal is less than 1. If the probability that the fifth time-frequency voice signal is within the target direction is less than or equal to the preset probability threshold, the gain of the fifth time-frequency voice signal is less than 1. The electronic device obtains a second sound signal based on each time-frequency voice signal in the first sound signal and the corresponding gain.

[0033] In one possible implementation, the electronic device determines the gain of the fifth time-frequency voice signal based on the probability that the fifth time-frequency voice signal in the first sound signal is within the target location and whether the fifth time-frequency voice signal is diffuse field noise. The method includes: determining a first gain of the fifth time-frequency voice signal based on the probability that the fifth time-frequency voice signal is within the target location; if the probability that the fifth time-frequency voice signal is within the target location is greater than a preset probability threshold, the first gain of the fifth time-frequency voice signal is equal to 1; if the probability that the fifth time-frequency voice signal is within the target location is less than or equal to the preset probability threshold, the first gain of the fifth time-frequency voice signal is less than 1. The electronic device determines a second gain of the fifth time-frequency voice signal based on whether the fifth time-frequency voice signal is diffuse field noise; if the fifth time-frequency voice signal is diffuse field noise, the second gain of the fifth time-frequency voice signal is less than 1; if the fifth time-frequency voice signal is a coherent signal, the second gain of the fifth time-frequency voice signal is equal to 1. The electronic device determines a gain of the fifth time-frequency voice signal based on the first gain and the second gain of the fifth time-frequency voice signal; the gain of the fifth time-frequency voice signal is the product of the first gain and the second gain of the fifth time-frequency voice signal.

[0034] In one possible implementation, if the fifth time-frequency speech signal is diffuse field noise and the product of the first gain and the second gain of the fifth time-frequency speech signal is less than a preset gain value, the gain of the fifth time-frequency speech signal is equal to the preset gain value.

[0035] In a third aspect, an embodiment of the present application provides an electronic device. The electronic device includes: a microphone; a camera; one or more processors; a memory; and a communication module. The microphone is used to capture sound signals during video recording or live broadcast; and the camera is used to capture image signals during video recording or live broadcast. The communication module is used to communicate with external devices. One or more computer programs are stored in the memory, and the one or more computer programs include instructions. When the instructions are executed by the processor, the electronic device executes the method described in the first aspect and any possible implementation thereof.

[0036] In a fourth aspect, an embodiment of the present application provides a chip system, which is applied to an electronic device. The chip system includes one or more interface circuits and one or more processors. The interface circuits and processors are interconnected via a circuit. The interface circuit is configured to receive a signal from a memory of the electronic device and send the signal to the processor, the signal including computer instructions stored in the memory. When the processor executes the computer instructions, the electronic device performs the method described in the first aspect and any possible implementation thereof.

[0037] In a fifth aspect, an embodiment of the present application provides a computer storage medium, which includes computer instructions. When the computer instructions are executed on a foldable electronic device, the electronic device executes the method described in the first aspect and any possible implementation thereof.

[0038] In a sixth aspect, an embodiment of the present application provides a computer program product, which, when executed on a computer, enables the computer to execute the method described in the first aspect and any possible design thereof.

[0039] It can be understood that the beneficial effects that can be achieved by the electronic device described in the third aspect, the chip system described in the fourth aspect, the computer storage medium described in the fifth aspect, and the computer program product described in the sixth aspect provided above can be referred to the beneficial effects in the first aspect and any possible implementation thereof, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 A diagram illustrating an application scenario of the sound signal processing method provided in an embodiment of the present application;

[0041] Figure 2 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application;

[0042] Figure 3 A schematic diagram of the microphone position of an electronic device provided in an embodiment of the present application;

[0043] Figure 4The process of the sound signal processing method provided in the embodiment of the present application Figure 1 ;

[0044] Figure 5 A comparative schematic diagram of converting a time-domain sound signal collected by a microphone of an electronic device provided in an embodiment of the present application into a frequency-domain sound signal;

[0045] Figure 6 A diagram showing the correspondence between speech frames and frequency points of a frequency domain sound signal involved in an embodiment of the present application;

[0046] Figure 7 The process of the sound signal processing method provided in the embodiment of the present application Figure 2 ;

[0047] Figure 8 The sound signal scenario involved in the sound signal processing method provided in the embodiment of the present application Figure 1 ;

[0048] Figure 9 The probability distribution diagram of the nth frame of speech in 36 directions provided in the embodiment of the present application;

[0049] Figure 10 The sound signal scenario involved in the sound signal processing method provided in the embodiment of the present application Figure 2 ;

[0050] Figure 11 Execution of the embodiment of this application Figure 7 Speech spectrograms before and after the method shown;

[0051] Figure 12 The process of the sound signal processing method provided in the embodiment of the present application Figure 3 ;

[0052] Figure 13 The process of the sound signal processing method provided in the embodiment of the present application Figure 4 ;

[0053] Figure 14 The sound signal scenario involved in the sound signal processing method provided in the embodiment of the present application Figure 3 ;

[0054] Figure 15 The sound signal scenario involved in the sound signal processing method provided in the embodiment of the present application Figure 4 ;

[0055] Figure 16 Execution of the embodiment of this application Figure 7 and Figure 13 Comparative speech spectrograms of the methods shown;

[0056] Figure 17 The process of the sound signal processing method provided in the embodiment of the present application Figure 5 ;

[0057] Figure 18 The sound signal scenario involved in the sound signal processing method provided in the embodiment of the present application Figure 5 ;

[0058] Figure 19 Execution of the embodiment of this application Figure 13 and Figure 17 Comparative speech spectrograms of the methods shown;

[0059] Figure 20 The process of the sound signal processing method provided in the embodiment of the present application Figure 5 ;

[0060] Figure 21A The interface involved in the sound signal processing method provided in the embodiment of the present application Figure 1 ;

[0061] Figure 21B Scenario of the sound signal processing method provided in the embodiment of the present application Figure 1 ;

[0062] Figure 22 The process of the sound signal processing method provided in the embodiment of the present application Figure 6 ;

[0063] Figure 23 The process of the sound signal processing method provided in the embodiment of the present application Figure 7 ;

[0064] Figure 24 The interface involved in the sound signal processing method provided in the embodiment of the present application Figure 2 ;

[0065] Figure 25 Scenario of the sound signal processing method provided in the embodiment of the present application Figure 2 ;

[0066] Figure 26 The process of the sound signal processing method provided in the embodiment of the present application Figure 8 ;

[0067] Figure 27 The interface involved in the sound signal processing method provided in the embodiment of the present application Figure 3 ;

[0068] Figure 28 The interface involved in the sound signal processing method provided in the embodiment of the present application Figure 4 ;

[0069] Figure 29 The interface involved in the sound signal processing method provided in the embodiment of the present application Figure 5 ;

[0070] Figure 30 The interface involved in the sound signal processing method provided in the embodiment of the present application Figure 6 ;

[0071] Figure 31 A schematic structural diagram of a chip system provided in an embodiment of the present application. DETAILED DESCRIPTION

[0072] For ease of understanding, some concepts related to the embodiments of the present application are exemplarily described for reference, as shown below:

[0073] Target object: An object within the field of view of a camera (e.g., a front-facing camera), such as a person or animal. A camera's field of view is determined by its field of view (FOV). The larger the camera's FOV, the wider its field of view.

[0074] Non-target objects: Objects that are not within the camera's field of view. For example, objects on the back of the phone are non-target objects.

[0075] Diffuse field noise: During video or audio recording, the sound emitted by the target object or non-target object is reflected by walls, floors, or ceilings.

[0076] The technical solutions in the embodiments of the present application will be described below in conjunction with the accompanying drawings in the embodiments of the present application. In the description of the embodiments of the present application, unless otherwise specified, " / " means or, for example, A / B can mean A or B; "and / or" in this article is merely a description of the association relationship of associated objects, indicating that three relationships can exist, for example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, in the description of the embodiments of the present application, "multiple" means two or more than two.

[0077] In the following, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of the technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of this embodiment, unless otherwise specified, "plurality" means two or more.

[0078] At present, with the development of short videos and live social software (such as Kuaishou, Douyin and other applications), the recording function of electronic devices has become a frequently used function, and electronic devices that can record high-quality video files have become a demand.

[0079] Existing electronic devices collect sound signals around them when recording videos. However, some of these sound signals are interference signals and are not what users want. For example, when an electronic device uses a camera (such as a front-facing camera or a rear-facing camera) to record a video, the electronic device will collect the sound of the target object within the camera's field of view (FOV), the sound of non-target objects outside the camera's field of view (FOV), and some ambient noise. In this case, the sound of non-target objects may become interference objects, affecting the sound quality of the video recorded by the electronic device.

[0080] Taking the front camera recording as an example, usually, the front camera of an electronic device is used to facilitate users to record short videos or small videos. Figure 1 As shown, when a user uses the front camera of an electronic device to record a short video, there may be a child playing on the back of the electronic device (i.e. Figure 1 Non-target object 1 in the user (i.e. Figure 1 On the same side of the target object in the image, there may be other objects, such as a puppy barking or a little girl singing and dancing (i.e. Figure 1 Therefore, during the recording process, the electronic device will inevitably record the sound of non-target object 1 or non-target object 2. However, for users, they prefer to record short videos that highlight their own voice (i.e., the voice of the target object) and suppress non-own voices, such as Figure 1 The sounds produced by non-target object 1 and non-target object 2.

[0081] In addition, due to the influence of the shooting environment, during the shooting process of the electronic device, the short video recorded by the electronic device may contain a lot of noise caused by the environment, such as Figure 1 The diffuse field noise 1 and diffuse field noise 2 shown in the figure cause the short video recorded by the electronic device to have large and harsh noise, which interferes with the sound of the target object and affects the sound quality of the recorded short video.

[0082] To solve the above problems, the present invention provides a method for processing sound signals, which can be applied to electronic devices and can suppress non-camera voices during selfie videos and improve the signal-to-noise ratio of selfie voices. Figure 1Taking the video shooting scene shown as an example, the sound signal processing method provided in the embodiment of the present application can remove the sounds of non-target objects 1 and non-target objects 2, retain the sound of the target object, and can also reduce the impact of diffuse field noise 1 and diffuse field noise 2 on the sound of the target object, thereby improving the signal-to-noise ratio of the audio signal of the target object after recording, smoothing the background noise, and making the user's listening experience better.

[0083] The sound signal processing method provided in the embodiments of the present application can be used for recording video with a front camera of an electronic device, and can also be used for recording video with a rear camera of an electronic device. The electronic device can be a mobile terminal such as a mobile phone, a tablet computer, a wearable device (such as a smart watch), an in-vehicle device, an augmented reality (AR) / virtual reality (VR) device, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), or a professional camera. The embodiments of the present application do not impose any restrictions on the specific type of electronic device.

[0084] For example, Figure 2 1 shows a schematic structural diagram of an electronic device 100. The electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, an earphone interface 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display 194, and a subscriber identification module (SIM) card interface 195. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, an air pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.

[0085] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU). The different processing units may be independent devices or integrated into one or more processors.

[0086] The controller may be the nerve center and command center of the electronic device 100. The controller may generate an operation control signal according to the instruction operation code and the timing signal to complete the control of fetching and executing instructions.

[0087] Processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in processor 110 is a cache memory. This memory can store instructions or data that have just been used or are being recycled by processor 110. If processor 110 needs to use the same instruction or data again, it can directly retrieve it from the memory. This avoids duplicate accesses, reduces processor 110 latency, and thus improves system efficiency.

[0088] Electronic device 100 implements display functionality through a GPU, display screen 194, and an application processor. A GPU is a microprocessor for image processing that connects display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. Processor 110 may include one or more GPUs that execute program instructions to generate or modify display information.

[0089] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode or an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), Miniled, MicroLed, Micro-oLed, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the electronic device 100 may include 1 or N display screens 194, where N is a positive integer greater than 1. In an embodiment of the present application, the display screen 194 can be used to display a preview interface and a shooting interface in a shooting mode, etc.

[0090] The electronic device 100 can implement a shooting function through an ISP, a camera 193, a video codec, a GPU, a display screen 194, and an application processor.

[0091] The ISP processes data fed back by camera 193. For example, when taking a photo, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, which is then passed to the ISP for processing and transformed into a visible image. The ISP can also perform algorithmic optimization on image noise, brightness, and skin tone. It can also optimize parameters such as exposure and color temperature of the captured scene. In some embodiments, the ISP can be located within camera 193.

[0092] The camera 193 is used to capture still images or videos. The object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, and then passes the electrical signal to the ISP for conversion into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in a standard RGB, YUV or other format. In some embodiments, the electronic device 100 may include 1 or N cameras 193, where N is a positive integer greater than 1.

[0093] In addition, the camera 193 may also include a depth camera for measuring the distance to the object to be photographed, as well as other cameras. For example, the depth camera may include a three-dimensional (3D) depth camera, a time-of-flight (TOF) depth camera, or a binocular depth camera.

[0094] The digital signal processor is used to process digital signals. In addition to processing digital image signals, it can also process other digital signals. For example, when the electronic device 100 selects a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy.

[0095] Video codecs are used to compress or decompress digital video. Electronic device 100 may support one or more video codecs. This allows electronic device 100 to play or record videos in various encoding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, and MPEG4.

[0096] The internal memory 121 can be used to store computer executable program code, and the executable program code includes instructions. The processor 110 executes various functional applications and data processing of the electronic device 100 by running the instructions stored in the internal memory 121. The internal memory 121 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc. The data storage area can store data created during the use of the electronic device 100 (such as audio data, a phone book, etc.), etc. In addition, the internal memory 121 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc. In other embodiments, the processor 110 enables the electronic device 100 to execute the method provided in the embodiment of the present application, as well as various functional applications and data processing by running the instructions stored in the internal memory 121, and / or the instructions stored in the memory provided in the processor.

[0097] The electronic device 100 can implement audio functions such as music playback and recording through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the headphone jack 170D, and the application processor.

[0098] The audio module 170 is used to convert digital audio information into analog audio signal output, and is also used to convert analog audio input into digital audio signals. The audio module 170 can also be used to encode and decode audio signals. In some embodiments, the audio module 170 can be provided in the processor 110, or some functional modules of the audio module 170 can be provided in the processor 110.

[0099] The speaker 170A, also called a "speaker", is used to convert audio electrical signals into sound signals. The electronic device 100 can listen to music or listen to hands-free calls through the speaker 170A.

[0100] The receiver 170B, also called a "handset", is used to convert audio electrical signals into sound signals. When the electronic device 100 receives a call or a voice message, the user can place the receiver 170B close to the ear to hear the voice.

[0101] Microphone 170C, also known as a "microphone" or "microphone," is used to convert sound signals into electrical signals. When making a call, sending a voice message, or recording an audio or video file, a user can speak by placing their mouth close to microphone 170C, inputting the sound signal into microphone 170C. The electronic device 100 can be equipped with one or more microphones 170C. For example, the electronic device 100 can be equipped with three, four, or more microphones 170C to collect sound signals, reduce noise, identify the direction of the sound source, implement directional recording, and suppress sounds from non-target directions.

[0102] The headphone jack 170D is used to connect a wired headphone and can be the USB interface 130 or a 3.5mm open mobile terminal platform (OMTP) standard interface or a cellular telecommunications industry association of the USA (CTIA) standard interface.

[0103] The touch sensor 180K is also called a "touch panel." The touch sensor 180K can be disposed on the display screen 194. The touch sensor 180K and the display screen 194 form a touch screen, also called a touch screen. The touch sensor 180K is used to detect touch operations applied thereto or in the vicinity thereof. The touch sensor can transmit the detected touch operations to the application processor to determine the type of touch event. Visual output related to the touch operations can be provided via the display screen 194. In other embodiments, the touch sensor 180K can also be disposed on the surface of the electronic device 100, in a location different from that of the display screen 194.

[0104] It should be understood that the structures illustrated in the embodiments of the present application do not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may include more or fewer components than shown, or may combine or separate certain components, or arrange the components differently. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0105] The following describes in detail the sound signal processing method provided in the embodiment of the present application, taking the electronic device as a mobile phone 300 and the recording and shooting by the front camera of the electronic device as an example.

[0106] For example, Figure 3 As shown, the mobile phone 300 is provided with three microphones, namely the top microphone 301, the bottom microphone 302 and the back microphone 303, which are used to collect the user's voice signals when the user makes a call or records audio or video files. During the process of recording or recording video on the mobile phone, the top microphone 301, the bottom microphone 302 and the back microphone 303 respectively collect the sound signals of the recording environment where the user is located (including the sound of the target object, the sound of non-target objects, and the noise generated by the environment). For example, when the mobile phone 300 is Figure 3 When recording video and collecting audio in the direction shown, the top microphone 301 can collect the left channel sound signal, and the bottom microphone 302 can collect the right channel sound signal. Figure 3 When recording video and collecting audio after rotating 180 degrees in the direction shown, the top microphone 301 can collect the right channel voice signal, and the bottom microphone 302 can collect the left channel voice signal. In this regard, the three microphones in the mobile phone 300 can collect the voice signal of the channel, which can be different depending on the usage scenario. The above description in the embodiment of the present application is only illustrative and does not constitute a limitation.

[0107] In addition, the sound signal collected by the back microphone 303 can be combined with the sound signals collected by the top microphone 301 and the bottom microphone 302 to determine the direction of the sound signal collected by the mobile phone.

[0108] by Figure 1 As an example of the video shooting scene shown in FIG, the mobile phone 300 Figure 3 The three microphones shown (i.e., top microphone 301, bottom microphone 302, and back microphone 303) can collect the sound of the target object, the sound of non-target object 1, the sound of non-target object 2, and the diffuse field noise 1 and diffuse field noise 2 formed by the reflection of the above-mentioned target object and non-target object sounds through the environment.

[0109] It should be understood that Figure 1 In the video shooting scene shown, the user's main shooting subject is himself (for example, Figure 1 Therefore, users do not want to capture non-target objects (such as Figure 1 The sound signal processing method in the embodiment of the present application can enable the mobile phone to process the collected sound signals, suppress the sound signals of non-target objects, highlight the sound of the target object, and improve the sound quality of the captured video.

[0110] In some embodiments, as Figure 4 As shown, an embodiment of the present application provides a sound signal processing method including:

[0111] 400. The mobile phone obtains the sound signal.

[0112] For example, when a user is recording a video, the mobile phone can Figure 3 The three microphones shown (i.e., top microphone 301, bottom microphone 302, and back microphone 303) collect sound signals. In the following embodiments, sound signals may also be referred to as audio signals.

[0113] Typically, if Figure 5 As shown in (a) in the figure, the sound signal collected by the microphone is a time domain signal, which is used to characterize the change of the amplitude of the sound signal over time. In order to facilitate the analysis and processing of the sound signal collected by the microphone, the sound signal collected by the microphone can be converted into a frequency domain signal through Fourier transform, such as fast Fourier transform (FFT) or discrete Fourier transform (DFT), for example Figure 5 As shown in (b) in Figure 5 In (a), the time domain signal is represented by time / amplitude, where the horizontal axis is the sampling time and the vertical axis is the amplitude of the sound signal. Figure 5 The sound signal shown in (a) can be converted into the following after FFT or DFT transformation: Figure 5 The frequency domain signal corresponding to the speech spectrum shown in (b) in FIG. In the speech spectrum, the horizontal axis is time, the vertical axis is frequency, and the coordinate value of the horizontal and vertical axes is the sound signal energy. For example, Figure 5 In (b), the energy of the sound data at the time-frequency signal position 511 is higher, and the energy of the sound data at the time-frequency signal position 512 (the white box above, which cannot be clearly shown in the figure, is explained here) is lower.

[0114] It should be understood that the greater the amplitude of the sound signal, the greater the energy of the sound signal and the higher the decibel of the sound.

[0115] It should be noted that when the sound signal collected by the microphone is transformed into a frequency domain signal through FFT or DFT, the sound signal will be framed and each frame of the sound signal will be processed separately. A frame of sound signal can be transformed into a frequency domain signal including multiple (such as 1024, 512) frequency domain sampling points (i.e., frequency points) through FFT or DFT. Figure 6 As shown, after the sound signal collected by the microphone is converted into a frequency domain signal with 1024 frequency points, multiple time-frequency points can be used to represent the sound signal collected by the microphone. Figure 6 A box in represents a time-frequency point. Figure 6 The horizontal axis represents the number of frames of the sound signal (which can be called voice frames), and the vertical axis represents the frequency of the sound signal.

[0116] After the sound signals collected by the three microphones, namely the top microphone 301, the bottom microphone 302 and the back microphone 303, are converted into frequency domain signals, X L (t, f) represents the left channel time-frequency speech signal, that is, the sound signal corresponding to different time-frequency points in the left channel sound signal collected by the top microphone 301. Similarly, X R (t, f) represents the right channel time-frequency speech signal, that is, the sound signal energy corresponding to different time-frequency points in the right channel sound signal collected by the bottom microphone 302; it can be expressed by X 背 (t, f) represents the time-frequency speech signals of the left and right channels, i.e., the sound signals corresponding to different time-frequency points in the left and right channel surround sound signals collected by the back microphone 303. Here, t represents the number of frames of the sound signal (which can be called a speech frame), and f represents the frequency of the sound signal.

[0117] 401. The mobile phone processes the collected sound signals and suppresses sound signals from non-target directions.

[0118] As can be seen from the above description, the sound signals collected by the mobile phone include the sound of the target object and the sound of the non-target object. However, the sound of the non-target object is not what the user wants and needs to be suppressed. Normally, the target object is within the field of view of the camera (such as the front camera), and the field of view of the camera can be used as the target orientation of the embodiment of the present application. Therefore, the above-mentioned non-target orientation refers to the orientation that is not within the field of view of the camera (such as the front camera).

[0119] For example, to suppress the sound signal in the non-target direction, the time-frequency gain calculation of the sound signal can be realized by calculating the probability of the sound source direction and the probability of the target direction in sequence. The sound signal in the target direction and the sound signal in the non-target direction can be distinguished by the difference in the size of the time-frequency gain of the sound signal. Figure 7 As shown, suppressing the sound signal of non-target direction may include sound source direction probability calculation, target direction probability calculation and time-frequency point gain g mask(t,f) is calculated in three steps, as follows:

[0120] (1) Calculation of sound source direction probability

[0121] For example, in an embodiment of the present application, assuming that when the mobile phone is recording a video, the direction directly in front of the mobile phone screen is the 0° direction, the direction directly behind the mobile phone screen is the 180° direction, the direction directly to the right of the mobile phone screen is the 90° direction, and the direction directly to the left of the mobile phone screen is the 270° direction.

[0122] In the embodiment of the present application, the 360° spatial orientation formed by the front, back, left, and right sides of the mobile phone screen can be divided into multiple spatial orientations. For example, the 360° spatial orientation can be divided into 36 spatial orientations with a spatial orientation interval of 10°, as shown in Table 1 below.

[0123] Table 1 Comparison table of 360° spatial orientations divided into 36 spatial orientations

[0124]

[0125] Taking the example of a mobile phone using the front camera to record video, assuming that the field of view FOV of the front camera of the mobile phone is [310°, 50°], the target direction is the [310°, 50°] direction in the 360° spatial direction formed by the front, back, left and right of the mobile phone screen. The target object photographed by the mobile phone is usually within the angle range of the field of view FOV of the front camera, that is, within the target direction. Suppressing sound signals in non-target directions means suppressing the sound of objects outside the angle range of the field of view FOV of the front camera, for example Figure 1 Non-target object 1 and non-target object 2 are shown.

[0126] For environmental noise (such as diffuse field noise 1 and diffuse field noise 2), the environmental noise may be in the target direction or in the non-target direction. It should be noted that diffuse field noise 1 and diffuse field noise 2 can be essentially the same noise. In the embodiments of the present application, for the sake of distinction, diffuse field noise 1 is used as the environmental noise in the non-target direction, and diffuse field noise 2 is used as the environmental noise in the target direction.

[0127] For example, Figure 1 Take the shooting scene shown as an example, Figure 8 The figure shows the spatial orientation of the target object, non-target object 1, non-target object 2, diffuse field noise 1, and diffuse field noise 2. For example, the spatial orientation of the target object is approximately 340°, the spatial orientation of non-target object 1 is approximately 150°, the spatial orientation of non-target object 2 is approximately 60°, the spatial orientation of diffuse field noise 1 is approximately 230°, and the spatial orientation of diffuse field noise 2 is approximately 30°.

[0128] When the phone is recording the target object, the microphone (such as Figure 3 The top microphone 301, bottom microphone 302 and back microphone 303 in the left channel will collect the sound of the target object, non-target objects (such as non-target object 1 and non-target object 2) and environmental noise (such as diffuse field noise 1 and diffuse field noise 2). The collected sound signal can be converted into a frequency domain signal (hereinafter referred to as a time-frequency speech signal) through FFT or DFT, which are the left channel time-frequency speech signal X L (t, f), right channel time-frequency speech signal X R (t, f) and the back mixed channel time-frequency speech signal X 背 (t, f).

[0129] The time-frequency speech signal X collected by three microphones L (t, f), X R (t, f) and X 背 (t, f), can be synthesized into a time-frequency speech signal X(t, f). The time-frequency speech signal X(t, f) can be input into the sound source orientation probability calculation model to calculate the probability P of the input time-frequency speech signal existing in each orientation. k (t, f). Where t represents the number of sound frames (i.e., speech frames), f represents the frequency point, and k is the spatial orientation number. The sound source orientation probability calculation model is used to calculate the orientation probability of the sound source. For example, the sound source orientation probability calculation model can be a Complex Angular Central Gaussian Mixture Model (cACGMM). Taking the number of spatial orientations K as 36 as an example, 1≤k≤36, and That is to say, the sum of the probabilities of the sound signal existing in 36 directions in the same frame and at the same frequency is 1. Figure 9 As shown in FIG, it is a schematic diagram of the probability of the sound signal corresponding to 1024 frequency points in the n-th frame sound signal existing in 36 spatial directions. Figure 9 In the figure, each small box represents the probability of the sound signal corresponding to a certain frequency point in the n-th frame speech signal at a certain spatial orientation. Figure 9 The dotted box in indicates that the sum of the probabilities of the sound signal corresponding to frequency point 3 in the n-th frame speech signal in 36 spatial directions is 1.

[0130] For example, assume the front camera's field of view (FOV) is [310°, 50°]. Since the target object is the subject the phone wants to record, it is typically located within the front camera's FOV. Therefore, the probability that the target object's sound signal originates from the direction [310°, 50°] is highest. The specific probability distribution can be as shown in Table 2.

[0131] Table 2 Probability distribution of the target object’s sound source in 36 spatial locations

[0132]

[0133] For non-target objects (such as non-target object 1 or non-target object 2), generally, the probability of the non-target object appearing in the field of view FOV of the front camera is small, which may be less than 0.5 or even 0.

[0134] For ambient noise (such as diffuse field noise 1 and diffuse field noise 2), since diffuse field noise 1 is ambient noise in the non-target direction, the probability of diffuse field noise 1 appearing in the front camera's field of view (FOV) is low, possibly less than 0.5 or even 0. Since diffuse field noise 2 is ambient noise in the target direction, the probability of diffuse field noise 2 appearing in the front camera's field of view (FOV) is high, possibly greater than 0.8 or even 1.

[0135] It should be understood that the above-mentioned probabilities of the target object, non-target object and environmental noise appearing in the target direction are examples and do not constitute a limitation to the embodiments of the present application.

[0136] (2) Target direction probability calculation

[0137] The target location probability calculation refers to the sum of the probabilities of the time-frequency speech signal existing in each location within the target location, which can also be called the spatial clustering probability of the target location. Therefore, the spatial clustering probability P(t,f) of the time-frequency speech signal in the target location can be calculated using the following formula (1):

[0138]

[0139] Among them, k1~k2 are the angle indexes of the target orientation, and can also be the spatial orientation sequence numbers of the target orientation. k (t,f) is the probability that the current time-frequency speech signal exists in the k direction; P(t,f) is the sum of the probabilities that the current time-frequency speech signal exists in the target direction.

[0140] For example, still taking the direction directly in front of the mobile phone screen as 0°, the field of view FOV of the front camera of the mobile phone as [310°, 50°], that is, the target orientation as [310°, 50°] as an example.

[0141] For the target object, taking the probability distribution of the target object's sound in 36 spatial directions shown in Table 2 above as an example, k1~k2 are the probabilities corresponding to serial numbers 32, 33, 34, 35, 36, 1, 2, 3, 4, 5, and 6 respectively. Therefore, the total probability P(t,f) of the target object's time-frequency speech signal existing in the target direction is 0.4+0.3+0.3=1.

[0142] A similar calculation method can be used to calculate the sum of the probabilities P(t,f) of the time-frequency speech signals of non-target objects existing at the target location. Alternatively, the sum of the probabilities P(t,f) of the video speech signals of environmental noise (such as diffuse field noise 1 and diffuse field noise 2) existing at the target location can be calculated.

[0143] For non-target objects, the sum of the probabilities P(t,f) of the time-frequency speech signals of the non-target objects existing in the target location may be less than 0.5 or even 0.

[0144] For environmental noise, such as diffuse field noise 1, diffuse field noise 1 is environmental noise in the non-target direction. The probability of the time-frequency speech signal of diffuse field noise 1 existing in the target direction is small. Therefore, the total probability P(t,f) of the time-frequency speech signal of diffuse field noise 1 existing in the target direction may be less than 0.5 or even 0.

[0145] For environmental noise, such as diffuse field noise 2, diffuse field noise 2 is the environmental noise within the target direction. The probability of the time-frequency speech signal of diffuse field noise 2 existing in the target direction is relatively high. Therefore, the total probability P(t,f) of the time-frequency speech signal of diffuse field noise 2 existing in the target direction may be greater than 0.8 or even 1.

[0146] (3) Time-frequency gain g mask (t,f) calculation

[0147] As can be seen from the above description, the main purpose of suppressing sound signals from non-target directions is to retain the sound signals of the target object and suppress the sound signals of non-target objects. Normally, the target objects are within the field of view FOV of the front camera, so the sound signals of the target objects mostly come from the target direction, that is, the probability of the sound of the target object appearing in the target direction is usually higher. On the contrary, for non-target objects, non-target objects are usually not within the field of view FOV of the front camera, so the sound signals of non-target objects mostly come from non-target directions, that is, the probability of the sound of non-target objects appearing in the target direction is usually lower.

[0148] Based on this, the current time-frequency point gain g can be achieved through the above target orientation clustering probability P(t,f) mask (t,f) calculation, please refer to the following formula (2):

[0149]

[0150] Among them, P th To preset the probability threshold, it can be configured through parameters, such as P th Set to 0.8; g mask-minIt is the time-frequency point gain when the current time-frequency speech signal is located at the non-target direction, which can be configured by parameters, such as g mask-min Set to 0.2.

[0151] When the total probability P(t,f) of the current time-frequency speech signal existing in the target direction is greater than the probability threshold P th When the current time-frequency speech signal is in the target direction, the time-frequency point gain g of the current time-frequency speech signal is mask (t,f)=1. Accordingly, when the sum of the probabilities of the time-frequency speech signal existing in the target direction P(t,f) is less than or equal to the probability threshold P th When , it can be considered that the current time-frequency speech signal is not within the target direction. In this case, the parameter g can be set mask-min , as the time-frequency point gain g when the current time-frequency speech signal is not in the target direction mask (t,f), for example g mask (t,f)=g mask-min =0.2.

[0152] In this way, if the current time-frequency speech signal is within the target direction, the current time-frequency speech signal is most likely to come from the target object, so the time-frequency point gain g when the current time-frequency speech signal is within the target direction is mask (t,f) is configured to 1, which can retain the sound of the target object to the greatest extent. If the current time-frequency speech signal is not within the target direction, the current time-frequency speech signal is most likely from a non-target object (such as non-target object 1 or non-target object 2), so the time-frequency point gain g when the current time-frequency speech signal is no longer within the target direction is set. mask The (t,f) is configured as 0.2, which can effectively suppress the sound of non-target objects (such as non-target object 1 or non-target object 2).

[0153] It should be understood that the time-frequency speech signal of the environmental noise may exist in the target direction, such as diffuse field noise 2; it may also exist in the non-target direction, such as diffuse field noise 1. Therefore, the time-frequency point gain g of the time-frequency speech signal of the environmental noise, such as diffuse field noise 2, is mask (t,f) is likely to be 1; environmental noise, such as the time-frequency gain g of the time-frequency speech signal of diffuse field noise 1 mask (t,f) is more likely to be g mask-min , such as 0.2. That is, suppressing the sound of non-target objects in the above manner cannot suppress the energy of the ambient noise.

[0154] 402. Output the processed sound signal.

[0155] Normally, a mobile phone has two speakers, namely, a speaker at the top of the mobile phone screen (hereinafter referred to as speaker 1) and a speaker at the bottom of the mobile phone (hereinafter referred to as speaker 2). Among them, when the mobile phone outputs audio (i.e., sound signal), speaker 1 can be used to output the left channel audio signal, and speaker 2 can be used to output the right channel audio signal. Of course, when the mobile phone outputs audio, speaker 1 can also be used to output the right channel audio signal, and speaker 2 can be used to output the left channel audio signal. In this regard, the embodiments of the present application do not make special limitations. It is understandable that when an electronic device has only one speaker, it can output (left channel audio signal + right channel audio signal) / 2, it can also output left channel audio signal + right channel audio signal), and it can also be output after fusion, which is not limited in this application.

[0156] In order to make the audio signal recorded by the mobile phone be output by speaker 1 and speaker 2. After the sound signal is processed by the above method, the output sound signal can be divided into the left channel audio output signal Y L (t,f) and right channel audio output signal Y R (t,f).

[0157] For example, the time-frequency gain g of each type of sound signal obtained by the above calculation is mask (t,f), and combined with the sound input signal collected by the microphone, such as the left channel time-frequency speech signal X L (t, f) or right channel time-frequency speech signal X R (t, f), we can get the sound signal after suppressing the sound of non-target objects, such as Y L (t,f) and Y R (t,f). Specifically, the sound signal Y output after processing L (t,f) and Y R (t,f) can be calculated by the following formula (3) and formula (4) respectively:

[0158] Y L (t,f)=X L (t, f)*g mask (t,f); Formula (3)

[0159] Y R (t,f)=X R (t, f)*g mask (t,f). Formula (4)

[0160] For example, for the target object, the left channel audio output signal Y L The energy of (t,f) is equal to the left channel time-frequency speech signal X L (t, f) energy; right channel audio output signal Y RThe energy of (t,f) is equal to the right channel time-frequency speech signal X R The energy of (t, f), that is, Figure 10 The acoustic signal of the target object is fully preserved.

[0161] For non-target objects, the left channel audio output signal Y L The energy of (t,f) is equal to the left channel time-frequency speech signal X L 0.2 times the energy of (t, f); the right channel audio output signal Y R The energy of (t,f) is equal to the right channel time-frequency speech signal X R 0.2 times the energy of (t, f), that is, Figure 10 The acoustic signals of non-target objects are shown to be effectively suppressed.

[0162] For environmental noise, such as diffuse field noise 2 in the target direction, the left channel audio output signal Y L The energy of (t,f) is equal to the left channel time-frequency speech signal X L (t, f) energy; right channel audio output signal Y R The energy of (t,f) is equal to the right channel time-frequency speech signal X R The energy of (t, f), that is, Figure 10 The acoustic signal of the diffuse field noise 2 is shown to be unsuppressed.

[0163] For environmental noise, such as diffuse field noise 1 in a non-target direction, the left channel audio output signal Y L The energy of (t,f) is equal to the left channel time-frequency speech signal X L 0.2 times the energy of (t, f); the right channel audio output signal Y R (t,f) is equal to the right channel time-frequency speech signal X R 0.2 times the energy of (t, f), that is, Figure 10 The acoustic signal of the diffuse field noise 1 is shown to be effectively suppressed.

[0164] In summary, if Figure 10 As shown, after suppressing the time-frequency gain g in the sound signal of the non-target direction mask (t,f) calculation, the time-frequency speech signals within the target direction (such as the time-frequency speech signals of the target object and the time-frequency speech signals of the diffuse field noise 2) are completely preserved, and the time-frequency speech signals outside the target direction (such as the time-frequency speech signals of the non-target object 1, the time-frequency speech signals of the non-target object 2, and the time-frequency speech signals of the diffuse field noise 1) are effectively suppressed. For example, Figure 11 (a) is the time-frequency speech signal before the above 401 processing. Figure 11(b) is the time-frequency speech signal after the above 401 processing. Figure 11 (a) and Figure 11 The box in (b) is the time-frequency speech signal of the non-target direction. Figure 11 (a) and Figure 11 As can be seen from (b) in FIG. 4 , after the above-mentioned 401 processing of suppressing the sound signal in the non-target direction, the time-frequency speech signal in the non-target direction is suppressed, that is, the energy of the time-frequency speech signal in the box is significantly reduced.

[0165] It should be understood that after the above time-frequency gain g mask The time-frequency speech signal output by (t, f) calculation only suppresses the time-frequency speech signal in the non-target direction. However, there may still be environmental noise (such as diffuse field noise 2) in the target direction, so the environmental noise of the output time-frequency speech signal is still large, the signal-to-noise ratio of the output sound signal is small, and the speech signal quality is low.

[0166] Based on this, in other embodiments, the sound signal processing method provided in the embodiments of the present application can also improve the signal-to-noise ratio of the output voice signal and improve the clarity of the voice signal by suppressing the diffuse field noise. Figure 12 As shown, the sound signal processing method provided in the embodiment of the present application may include 400-401 and 1201-1203.

[0167] 1201. Process the collected sound signal to suppress diffuse field noise.

[0168] After the sound signal is collected in step 400, the collected sound signal can be processed in step 1201 to suppress the diffuse field noise. For example, the diffuse field noise can be suppressed by sequentially calculating the coherent-to-diffuse power ratio (CDR) to achieve the time-frequency gain g when suppressing the diffuse field noise. cdr (t,f) calculation. Through the time-frequency point gain g cdr The size of (t,f) is different, which can distinguish the coherent signals (such as the sound signals of the target object and the sound signals of non-target objects) and the diffuse field noise in the sound signal. Figure 13 As shown, the suppression of diffuse field noise can include the calculation of the coherent diffusion ratio CDR and the time-frequency gain g cdr (t,f) is calculated in two steps, as follows:

[0169] (1) Coherence diffusion ratio (CDR) calculation

[0170] The coherent-to-diffuse power ratio (CDR) refers to the power ratio of the coherent signal (i.e., the speech signal of the target object or non-target object) to the diffuse field noise. L (t, f), right channel time-frequency speech signal X R (t, f) and the back mixed channel time-frequency speech signal X 背 (t, f) Using existing technology, such as coherent-to-diffuse power ratio estimation for dereverberation, the coherent-to-diffuse power ratio is realized. calculate.

[0171] For example, in the above Figure 1 In the shooting scene shown, for the sound signal of the target object, the coherent diffusion ratio of the sound signal of the target object is is infinite∞. For the sound signal of a non-target object (such as non-target object 1 or non-target object 2), the coherent diffusion ratio of the sound signal of the non-target object is also infinite∞. For diffuse field noise (such as diffuse field noise 1 or diffuse field noise 2), the coherent diffusion ratio of the diffuse field noise is 0.

[0172] (2) Time-frequency gain g cdr (t,f) calculation

[0173] As can be seen from the above description, the main purpose of suppressing diffuse field noise is to retain the sound of the coherent signal (such as the target object) and reduce the energy of the diffuse field noise.

[0174] For example, the coherent diffusion ratio Determine the time-frequency gain g of the coherent signal (i.e. the sound signal of the target object and the sound signal of the non-target object) cdr (t,f), determine the time-frequency gain g of the incoherent signal (i.e., diffuse field noise) cdr (t,f), that is, through the time-frequency point gain g cdr (t,f), distinguishing coherent or incoherent signals.

[0175] For example, the time-frequency gain g of the coherent signal can be cdr (t,f) is kept as 1, which can make the time-frequency gain g of the incoherent signal cdr (t,f) is reduced, such as setting it to 0.3. In this way, the sound signal of the target object can be retained and the diffuse field noise can be suppressed to reduce the ability of the diffuse field noise.

[0176] For example, the time-frequency gain g can be calculated using the following formula (5):cdr (t,f):

[0177]

[0178] Among them, g cdr-min To suppress the minimum gain after diffuse field noise, you can configure the parameters, for example, you can set g cdr-min 0.3 g cdr (t,f) is the time-frequency gain after suppressing diffuse field noise. μ is the overestimation factor, which can be configured through parameters, for example, setting μ to 1.

[0179] In this way, for the target object, due to the coherent diffusion ratio of the sound signal of the target object is infinite∞, and is substituted into the above formula (5), then the time-frequency gain g of the target object's sound signal after suppressing the diffuse field noise is cdr (t,f)=1.

[0180] For non-target objects (such as non-target object 1 and non-target object 2), the coherent diffusion ratio of the sound signal of the non-target object is is also infinite∞, and when substituted into the above formula (5), the time-frequency gain g of the sound signal of the non-target object after suppressing the diffuse field noise is obtained. cdr (t,f)=1.

[0181] For diffuse field noise (such as diffuse field noise 1 and diffuse field noise 2), due to the coherent diffusion ratio of the diffuse field noise is 0, and substituting it into the above formula (5), the time-frequency gain g of the diffuse field noise is cdr (t,f)=0.3.

[0182] It can be seen from this that Figure 14 As shown, the time-frequency gain g cdr (t,f) calculation, the time-frequency gain g of the coherent signal (such as the sound signal of the target object and the sound signal of the non-target object) cdr (t,f) is 1, the coherent signal can be completely preserved. However, after the time-frequency gain g cdr (t,f) calculation, the time-frequency gain g of the diffuse field noise (such as diffuse field noise 1 and diffuse field noise 2) cdr (t,f) is 0.3, so the diffuse field noise is effectively suppressed, that is, the energy of the diffuse field noise is significantly reduced compared to before processing.

[0183] 1202. Fusing the suppression of non-target direction sounds and the suppression of diffuse field noise.

[0184] It should be understood that the main purpose of suppressing the sound signal of the non-target direction in the above 401 is to retain the sound signal of the target object and suppress the sound signal of the non-target object. The main purpose of suppressing the diffuse field noise in the above 1201 is to suppress the diffuse field noise and protect the coherent signal (i.e. the sound signal of the target object or the non-target object). Figure 12 or Figure 13 In the sound signal processing method shown in FIG, after executing 1201 to suppress the diffuse field noise, the time-frequency gain g obtained by suppressing the sound signal in the non-target direction in 401 can be mask (t,f), and the time-frequency gain g obtained by suppressing the diffuse field noise by 1201 above cdr (t,f) is fused and calculated to obtain the fusion gain g mix (t,f). According to the fusion gain g mix (t,f) processes the input time-frequency speech signal to obtain a clear time-frequency speech signal of the target object, reduces the energy of the diffuse field noise, and improves the signal-to-noise ratio of the audio (sound) signal.

[0185] For example, the following formula (VI) can be used to calculate the fusion gain g: mix (t,f) calculation:

[0186] g mix (t,f)=g mask (t,f)*g cdr (t,f); Formula (6)

[0187] Among them, g mix (t,f) is the mixed gain after gain fusion calculation.

[0188] For example, still with the above Figure 1 As an example of the shooting scene shown in the figure, the time-frequency point signal of the target object, the time-frequency speech signal of the non-target object (such as non-target object 1 and non-target object 2), the time-frequency speech signal of the diffuse field noise 1, and the time-frequency speech signal of the diffuse field noise 2 are respectively fused. The fusion gain g mix (t,f) is calculated.

[0189] For example:

[0190] Time-frequency speech signal of the target object: g mask (t,f)=1,g cdr (t,f)=1, then g mix (t,f)=1;

[0191] Time-frequency speech signal of non-target object: g mask (t,f)=0.2,g cdr (t,f)=1, then gmix (t,f)=0.2;

[0192] Time-frequency speech signal of diffuse field noise 1: g mask (t,f)=0.2,g cdr (t,f)=0.3, then g mix (t,f)=0.06;

[0193] Time-frequency speech signal of diffuse field noise 2: g mask (t,f)=1,g cdr (t,f)=0.3, then g mix (t,f)=0.3.

[0194] From the above fusion gain g mix (t,f) calculation shows that the time-frequency gain g obtained by suppressing the sound signal of the non-target direction is mask (t,f), and the time-frequency gain g obtained by suppressing the diffuse field noise cdr (t,f) is fused and calculated to suppress the diffuse field noise with large energy, such as diffuse field noise 2.

[0195] 1203, output the processed sound signal.

[0196] For example, the fusion gain g in the above 1202 is performed mix After (t,f) is calculated, the fusion gain g of various sound signals obtained by the above calculation can be obtained mix (t,f), and combined with the sound input signal collected by the microphone, such as the left channel time-frequency speech signal X L (t, f) or right channel time-frequency speech signal X R (t, f), we can get the sound signal after suppressing the sound of non-target objects and suppressing the fusion of diffuse field noise, such as Y L (t,f) and Y R (t,f). Specifically, the processed sound signal Y L (t,f) and Y R (t,f) can be calculated by the following formula (7) and formula (8) respectively:

[0197] Y L (t,f)=X L (t, f)*g mix (t,f); Formula (7)

[0198] Y R (t,f)=X R (t, f)*g mix (t,f). Formula (8)

[0199] For example, for the target object, the left channel audio output signal Y L The energy of (t,f) is equal to the left channel time-frequency speech signal X L (t, f) energy; right channel audio output signal Y R The energy of (t,f) is equal to the right channel time-frequency speech signal X R The energy of (t, f), that is, Figure 15 The acoustic signal of the target object is fully preserved.

[0200] For non-target objects, the left channel audio output signal Y L The energy of (t,f) is equal to the left channel time-frequency speech signal X L 0.2 times the energy of (t, f); the right channel audio output signal Y R The energy of (t,f) is equal to the right channel time-frequency speech signal X R 0.2 times the energy of (t, f), that is, Figure 15 The acoustic signals of non-target objects are shown to be effectively suppressed.

[0201] For environmental noise, such as diffuse field noise 2 in the target direction, the left channel audio output signal Y L The energy of (t,f) is equal to the left channel time-frequency speech signal X L 0.3 times the energy of (t, f); the right channel audio output signal Y R The energy of (t,f) is equal to the right channel time-frequency speech signal X R 0.3 times the energy of (t, f), that is, Figure 15 The acoustic signal of the diffuse field noise 2 is shown to be effectively suppressed.

[0202] For environmental noise, such as diffuse field noise 1 in a non-target direction, the left channel audio output signal Y L The energy of (t,f) is equal to the left channel time-frequency speech signal X L 0.06 times the energy of (t, f); the right channel audio output signal Y R The energy of (t,f) is equal to the right channel time-frequency speech signal X R 0.06 times the energy of (t, f), that is, Figure 15 The acoustic signal of the diffuse field noise 1 is shown to be effectively suppressed.

[0203] In summary, if Figure 15 As shown, after suppressing the time-frequency gain g in the sound signal of the non-target direction mask(t,f) calculation, the time-frequency speech signal of the target object in the target direction is completely preserved, and the time-frequency speech signal outside the target direction (such as the time-frequency speech signal of non-target object 1, the time-frequency speech signal of non-target object 2, and the time-frequency speech signal of diffuse field noise 1) is effectively suppressed. In addition, the time-frequency speech signal of diffuse field noise 2 located in the target direction is also effectively suppressed, thereby improving the signal-to-noise ratio of the speech signal and improving the clarity of the selfie sound. For example, Figure 16 (a) is the time-frequency speech signal that only suppresses the sound signal in the non-target direction. Figure 16 (b) is the time-frequency speech signal obtained by fusion processing after suppressing the sound signal of non-target direction and suppressing the diffuse field noise. Figure 16 (a) and Figure 16 In (b), we can see Figure 16 The background color of the speech spectrum shown in (a) is slightly lighter. Figure 16 The background noise energy of the time-frequency speech signal in the speech spectrum diagram (a) is relatively large. Figure 16 The background color of the speech spectrogram shown in (b) is darker. Figure 16 In the speech spectrogram (b) of Figure 1, the background color of the corresponding time-frequency speech signal is darker, and the energy of the background noise (i.e., diffuse noise) is lower. Therefore, it can be determined that after the fusion processing of suppressing the sound signal in the non-target direction and suppressing the diffuse noise, the noise energy of the output sound signal is reduced, and the signal-to-noise ratio of the sound signal is significantly improved.

[0204] It should be noted that in a noisy environment, only relying on the fusion gain g in 1202 above mix (t,f) calculation to separate the sound of the target object from the sound of the non-target object, there will be a problem of unstable background noise (ie, environmental noise). For example, after the fusion gain g in the above 1202 mix After (t,f) is calculated, the fusion gain g of the time-frequency speech signal of diffuse field noise 1 and diffuse field noise 2 is mix (t,f) has a large gap, which makes the background noise of the output audio signal unstable.

[0205] In order to solve the problem of unstable background noise of the audio signal after the above processing, in other embodiments, noise compensation can be performed on the diffuse field noise, and then a secondary noise reduction process can be performed to make the background noise of the audio signal more stable. Figure 17 As shown, the above-mentioned sound signal processing method may include 400-401, 1201-1202, and 1701-1702.

[0206] 1701. Compensate for diffuse field noise.

[0207] For example, in the diffuse field noise compensation stage, the diffuse field noise can be compensated by the following formula (IX):

[0208] g out (t,f)=MAX(g mix (t,f),MIN(1–g cdr (t,f),g min Formula (9)

[0209] Among them, g min is the minimum gain of the diffuse field noise (i.e., the preset gain value), which can be configured through parameters. For example, g min Set to 0.3. g out (t,f) is the time-frequency point gain of the time-frequency speech signal after diffuse field noise compensation.

[0210] For example, still with the above Figure 1 Taking the shooting scene shown in the figure as an example, the time-frequency speech signal of the target object, the time-frequency speech signal of the non-target object (such as non-target object 1 and non-target object 2), the time-frequency speech signal of the diffuse field noise 1, and the time-frequency speech signal of the diffuse field noise 2 are respectively calculated to obtain the time-frequency point gain g out (t,f).

[0211] For example:

[0212] Time-frequency speech signal of the target object: g mask (t,f)=1,g cdr (t,f)=1,g mix (t,f)=1; then g out (t,f)=1;

[0213] Time-frequency speech signal of non-target object: g mask (t,f)=0.2,g cdr (t,f)=1, then g mix (t,f)=0.2; then g out (t,f)=0.2;

[0214] Time-frequency speech signal of diffuse field noise 1: g mask (t,f)=0.2,g cdr (t,f)=0.3, then g mix (t,f)=0.06; then g out (t,f)=0.3;

[0215] Time-frequency speech signal of diffuse field noise 2: g mask (t,f)=1,g cdr (t,f)=0.3, then gmix (t,f)=0.3; then g out (t,f)=0.3.

[0216] It can be seen that after diffuse field noise compensation, the time-frequency point gain of diffuse field noise 1 increases from 0.06 to 0.3, and the time-frequency point gain of diffuse field noise 2 remains at 0.3, thereby making the background noise of the output sound signal (such as diffuse field noise 1 and diffuse field noise 2) more stable, so that the user's listening experience is better.

[0217] 1702. Output the processed sound signal.

[0218] The difference from the above 1203 is that the processed sound signal, such as the left channel audio output signal Y L (t,f) and right channel audio output signal Y R (t,f) is the time-frequency gain g of the time-frequency speech signal after diffuse field noise compensation out (t,f). Specifically, it can be calculated using the following formula (10) and formula (11):

[0219] Y R (t,f)=X R (t,f)*g out (t,f); Formula (10)

[0220] Y L (t,f)=X L (t,f)*g out (t,f); Formula (XI)

[0221] Among them, X R (t,f) is the right channel time-frequency speech signal collected by the microphone, X L (t,f) is the left channel time-frequency speech signal collected by the microphone.

[0222] For example, for the target object, the left channel audio output signal Y L The energy of (t,f) is equal to the left channel time-frequency speech signal X L (t, f) energy; right channel audio output signal Y R The energy of (t,f) is equal to the right channel time-frequency speech signal X R The energy of (t, f), that is, Figure 18 The acoustic signal of the target object is fully preserved.

[0223] For non-target objects, the left channel audio output signal Y L The energy of (t,f) is equal to the left channel time-frequency speech signal X L 0.2 times the energy of (t, f); the right channel audio output signal YR The energy of (t,f) is equal to the right channel time-frequency speech signal X R 0.2 times the energy of (t, f), that is, Figure 18 The acoustic signals of non-target objects are shown to be effectively suppressed.

[0224] For environmental noise, such as diffuse field noise 2 in the target direction, the left channel audio output signal Y L The energy of (t,f) is equal to the left channel time-frequency speech signal X L 0.3 times the energy of (t, f); the right channel audio output signal Y R The energy of (t,f) is equal to the right channel time-frequency speech signal X R 0.3 times the energy of (t, f), that is, Figure 18 The acoustic signal of the diffuse field noise 2 is shown to be effectively suppressed.

[0225] For environmental noise, such as diffuse field noise 1 in a non-target direction, the left channel audio output signal Y L The energy of (t,f) is equal to the left channel time-frequency speech signal X L 0.3 times the energy of (t, f); the right channel audio output signal Y R The energy of (t,f) is equal to the right channel time-frequency speech signal X R 0.3 times the energy of (t, f), that is, Figure 18 The acoustic signal of the diffuse field noise 1 is shown to be effectively suppressed.

[0226] In summary, if Figure 18 As shown, the time-frequency gain g after diffuse field noise compensation out (t,f) calculation, reducing the time-frequency gain g of diffuse field noise 1 and diffuse field noise 2 out The difference between (t,f) makes the background noise of the output sound signal (such as diffuse field noise 1 and diffuse field noise 2) more stable, making the user's listening experience more comfortable. Figure 19 (a) in the figure shows the time-frequency speech signal output without diffuse field noise compensation processing. Figure 19 (b) shows the time-frequency speech signal output after diffuse field noise compensation. Figure 19 (a) and Figure 19 In (b), we can see Figure 19 The background color of the speech spectrum shown in (b) is more uniform, indicating that the energy of the background noise is more uniform. Therefore, it can be determined that after the diffuse field noise compensation processing, the background noise of the output sound signal (such as the diffuse field noise) is more uniform and smooth, thereby making the user's listening experience better.

[0227] It should be noted that after the above Figure 7 、 Figure 13 and Figure 17 After the sound signal processing method shown in the figure is used, all left channel audio output signals Y are obtained. L (t,f) and right channel audio output signal Y R (t,f), and then through the inverse fast Fourier transform (IFFT) or inverse discrete Fourier transform (IDFT), the output time-frequency speech signal is transformed into a time-domain amplitude signal, which can be output through the mobile phone's speakers 1 and 2.

[0228] The above is an introduction to the specific process of the sound signal processing method provided in the embodiment of the present application. The following describes how to use the above sound signal processing method in combination with different application scenarios.

[0229] Scenario 1: The user uses the front camera to record a video.

[0230] In some embodiments, for the above scenario 1, the mobile phone processes each frame of image data it collects, or processes each audio data corresponding to each frame of image data it collects. For example, Figure 20 As shown, a sound signal processing method provided in an embodiment of the present application may include:

[0231] 2001. The mobile phone activates the front camera recording function and starts recording.

[0232] In an embodiment of the present application, when a user wants to use a mobile phone to record a video, the user can activate the recording function of the mobile phone. For example, the mobile phone can activate the camera application, or activate other applications with a recording function (such as AR applications such as Douyin or Kuaishou), thereby activating the recording function of the application.

[0233] For example, the mobile phone detects that the user clicks Figure 21A After operating the camera icon 2101 shown in (a) in FIG, the video recording function of the camera application is started and the following is displayed: Figure 21A In another example, the mobile phone displays the desktop or non-camera application interface, detects the user's voice command to open the camera application, starts the recording function, and displays the following Figure 21A (b) shows the preview interface of the front camera video.

[0234] It should be noted that the mobile phone can also start the recording function in response to other user touch operations, voice commands or quick gestures. The embodiment of the present application does not limit the operation of triggering the mobile phone to start the recording function.

[0235] When the phone displays Figure 21A When the front camera video preview interface shown in (b) is displayed, the phone detects that the user clicks Figure 21A After the video button 2102 shown in (b) is operated, the front camera video is started and the following is displayed: Figure 21A The recording interface of the front camera as shown in (c) above is displayed, and the recording timer is started.

[0236] 2002. The mobile phone captures the Nth frame image and processes the Nth frame image.

[0237] For example, during video recording, a mobile phone can generate an image stream and an audio stream. The image stream is used to capture image data and perform image processing on each frame. The audio stream is used to capture audio data and perform sound pickup and denoising on each frame.

[0238] For example, taking the first frame of an image as an example, after the mobile phone captures the first frame of the image, the mobile phone can process the first frame of the image, such as image denoising and tone mapping. After the mobile phone captures the second frame of the image, the mobile phone can process the second frame of the image. Similarly, after the mobile phone captures the Nth frame of the image, the mobile phone can process the Nth frame of the image. Where N is a positive integer.

[0239] 2003. The mobile phone collects the audio corresponding to the Nth frame image and processes the audio corresponding to the Nth frame image.

[0240] With the above Figure 1 Take the shooting environment and subject shown in the figure as an example, the mobile phone Figure 21B In the video recording process shown in the video recording interface shown in (a), the image data collected by the mobile phone is the image of the target object, and the audio data collected by the mobile phone includes not only the sound of the target object, but also the sound of non-target objects (such as non-target object 1 and non-target object 2), as well as environmental noise (such as diffuse field noise 1 and diffuse field noise 2).

[0241] Assuming one image frame is 30 milliseconds (ms) and one audio frame is 10 milliseconds, counting frames from the start of recording, the audio corresponding to the Nth image frame is the audio of frames 3N-2, 3N-1, and 3N. For example, the audio corresponding to the 1st image frame is the audio of frames 1, 2, and 3. For another example, the audio corresponding to the 2nd image frame is the audio of frames 4, 5, and 6. For another example, the audio corresponding to the 10th image frame is the audio of frames 28, 29, and 30.

[0242] Taking the audio corresponding to the first frame of the image as an example, the mobile phone processes the audio corresponding to the first frame of the image, and needs to process the first frame of audio, the second frame of audio, and the third frame of audio respectively.

[0243] For example, after collecting the first frame of audio, the mobile phone can execute the above Figure 7 The sound signal processing method shown in the figure processes the first frame of audio, suppresses the sound of non-target objects, and highlights the sound of the target object; or the mobile phone can execute the above Figure 13 The sound signal processing method shown in the figure suppresses the sound of non-target objects to highlight the sound of the target object, and suppresses the diffuse field noise to reduce the noise energy in the audio signal and improve the signal-to-noise ratio of the audio signal; or the mobile phone can perform the above Figure 17 The sound signal processing method shown suppresses the sound of non-target objects to highlight the sound of the target object, and suppresses diffuse field noise to reduce the noise energy in the audio signal, improve the signal-to-noise ratio of the audio signal, and smoothes the background noise (i.e., environmental noise) to provide users with a better listening experience.

[0244] Similarly, after collecting the second or third frame of audio, the mobile phone can also perform the above Figure 7 or Figure 13 or Figure 17 The sound signal processing method shown in the figure processes the audio signal. It should be understood that the above 2003 may correspond to Figure 7 or Figure 13 or Figure 17 400 steps in.

[0245] It should be understood that the audio corresponding to any subsequent frame of image can be processed according to the audio processing process corresponding to the first frame of image, and no further details will be given here. Of course, in the embodiment of the present application, when processing the audio corresponding to the Nth frame of image, the above method can be executed for each frame of audio collected, or the three frames of audio corresponding to the Nth frame of image can be collected and then each of the three frames of audio can be processed separately. The embodiment of the present application does not impose any special restrictions.

[0246] 2004. The mobile phone synthesizes the processed N-th frame image and the audio corresponding to the processed N-th frame image to obtain the N-th frame video data.

[0247] For example, after the Nth frame image is processed and the audio corresponding to the Nth frame image, such as the 3N-2th frame audio, the 3N-1th frame and the 3Nth frame audio, is also processed, the mobile phone can obtain the Nth frame image from the image stream and obtain the 3N-2th frame audio, the 3N-1th frame and the 3Nth frame audio from the audio stream, and then synthesize the 3N-2th frame audio, the 3N-1th frame and the 3Nth frame audio with the Nth frame image in timestamp order to form the Nth frame video data.

[0248] 2005. When the recording ends, after the last frame image and the audio corresponding to the last frame image are processed and the last frame video data is obtained, the first frame video data to the last frame video data are synthesized and saved as video file A.

[0249] For example, when the phone detects that the user clicks Figure 21B After the end recording button 2103 shown in (a) is operated, the mobile phone responds to the user's click operation, ends the recording, and displays the following Figure 21B (b) shows the preview interface of the front camera video.

[0250] During this process, the mobile phone responds to the user clicking the end recording button 2103, stops collecting images and audio, and completes processing of the last frame of image and the audio corresponding to the last frame of image, obtains the last frame of time-frequency data, synthesizes the first frame of video data to the last frame of video data in the order of timestamps, and saves it as video file A. At this time, Figure 21B The preview file displayed in the preview window 2104 in the preview interface of the front camera video shown in (b) is video file A.

[0251] When the phone detects that the user Figure 21B After the preview window 2104 shown in (b) is clicked, the mobile phone responds to the user's click operation on the preview window 2104 and can display the following Figure 21B The video file playback interface shown in (c) is used to play video file A.

[0252] When the mobile phone plays video file A, the sound signal of the played video file A no longer has the sound signals of non-target objects (such as non-target object 1 and non-target object 2), and the background noise (such as diffuse field noise 1 and diffuse field noise 2) in the sound signal of the played video file A is small and smooth, which can give the user a good listening experience.

[0253] It should be understood that Figure 21B The scenario shown in the following example is to execute the above Figure 17 The scenario of the sound signal processing method shown is not shown. Figure 7or Figure 13 Scenario of the sound signal processing method shown.

[0254] It should be noted that in the above Figure 20 In the method shown, after the mobile phone completes the acquisition and processing of the Nth frame image, and the acquisition and processing of the audio corresponding to the Nth frame image, the Nth frame image and the audio corresponding to the Nth frame image are not synthesized into the Nth frame video data. After the recording is completed and the last frame image and the audio corresponding to the last frame image are processed, all images and all audios can be synthesized into video file A in the order of timestamps.

[0255] In other embodiments, for the above scenario 1, the sound pickup and denoising operation on the audio data in the video file may be performed after the video file is recorded.

[0256] For example, Figure 22 As shown, the video shooting method provided by the embodiment of the present application is the same as the above Figure 20 The difference between the video recording method shown is that after the mobile phone starts recording, it first collects all the image data and audio data, then processes and synthesizes the image data and audio data, and finally saves the synthesized video file. Specifically, the method includes:

[0257] 2201. The mobile phone activates the front camera recording function and starts recording.

[0258] For example, the method of activating the front camera recording function of the mobile phone and activating the recording can refer to the description in 2001 above, which will not be repeated here.

[0259] 2202. The mobile phone collects image data and audio data respectively.

[0260] For example, during video recording, the mobile phone can be divided into an image stream and an audio stream. The image stream is used to collect multiple frames of image data during the video recording process. The audio stream is used to collect multiple frames of audio data during the video recording process.

[0261] For example, in an image stream, the phone sequentially captures the first image frame, the second image frame, and so on, and the last image frame. In an audio stream, the phone sequentially captures the first audio frame, the second audio frame, and so on, and so on, and the last image frame.

[0262] With the above Figure 1 Take the shooting environment and subject shown in the figure as an example, the mobile phone Figure 21BIn the video recording process shown in the video recording interface shown in (a), the image data collected by the mobile phone is the image of the target object, and the audio data collected by the mobile phone includes not only the sound of the target object, but also the sound of non-target objects (such as non-target object 1 and non-target object 2), as well as environmental noise (such as diffuse field noise 1 and diffuse field noise 2).

[0263] 2203. When recording ends, the mobile phone processes the collected image data and audio data respectively.

[0264] For example, when the phone detects that the user clicks Figure 21B After the end recording button 2103 shown in (a) is operated, the mobile phone responds to the user's click operation, ends the recording, and displays the following Figure 21B (b) shows the preview interface of the front camera video.

[0265] During this process, in response to the user clicking the end recording button 2103 , the mobile phone processes the collected image data and the collected audio data respectively.

[0266] For example, the mobile phone may process each frame of the collected image data separately, such as image denoising, tone mapping, etc., to obtain processed image data.

[0267] For example, the mobile phone can also process each frame of audio in the collected audio data. Figure 7 The sound signal processing method shown processes each frame of audio in the first audio data, suppresses the sound of non-target objects, and highlights the sound of the target object. Figure 13 The sound signal processing method shown processes each frame of audio in the collected audio data, suppresses the sound of non-target objects to highlight the sound of the target object, and suppresses diffuse field noise to reduce the noise energy in the audio signal and improve the signal-to-noise ratio of the audio signal. For another example, executing the above Figure 17 The sound signal processing method shown processes each frame of audio in the collected audio data, suppresses the sound of non-target objects to highlight the sound of the target object, and suppresses diffuse field noise to reduce the noise energy in the audio signal, thereby improving the signal-to-noise ratio of the audio signal. It can also smooth background noise (i.e., ambient noise) to provide users with a better listening experience.

[0268] After the mobile phone processes each frame of audio in the collected audio data, the processed audio data can be obtained.

[0269] It should be understood that the above 2202 and 2203 may correspond to Figure 7 or Figure 13 or Figure 17 400 steps in.

[0270] 2204. The mobile phone synthesizes the processed image data and the processed audio data to obtain video file A.

[0271] It should be understood that the processed image data and the processed audio data need to be synthesized into a video file before they can be shared or played by users. Therefore, after the mobile phone executes the above 2203 to obtain the processed image data and the processed audio data, the processed image data and the processed audio data can be synthesized to form a video file A.

[0272] 2205. The mobile phone saves video file A.

[0273] At this point, the mobile phone can save the video file A. For example, when the mobile phone detects that the user Figure 21B After the preview window 2104 shown in (b) is clicked, the mobile phone responds to the user's click operation on the preview window 2104 and can display the following Figure 21B The video file playback interface shown in (c) is used to play video file A.

[0274] When the mobile phone plays video file A, the sound signal of the played video file A no longer has the sound signals of non-target objects (such as non-target object 1 and non-target object 2), and the background noise (such as diffuse field noise 1 and diffuse field noise 2) in the sound signal of the played video file A is small and smooth, which can give the user a good listening experience.

[0275] Scenario 2: The user uses the front camera for live streaming.

[0276] In this scenario, the data collected during a live broadcast is displayed to the user in real time. Therefore, the images and audio captured during the live broadcast are processed in real time and displayed promptly to the user. This scenario includes at least mobile phone A, a server, and mobile phone B. Both mobile phone A and mobile phone B communicate with the server. Mobile phone A can be a live broadcast recording device, recording audio and video files and transmitting them to the server. Mobile phone B can be a live broadcast display device, retrieving the audio and video files from the server and displaying the content on the live broadcast interface for the user to watch.

[0277] For example, for the above scenario 2, if Figure 23 As shown, a video recording method provided in an embodiment of the present application is applied to mobile phone A. The method may include:

[0278] 2301. Mobile phone A activates the front camera live recording and starts live broadcast.

[0279] In an embodiment of the present application, when a user wants to use a mobile phone to record a live broadcast, he or she can start a live broadcast application in the mobile phone, such as Tik Tok or Kuaishou, and start the live broadcast recording.

[0280] For example, taking the TikTok app as an example, the phone detects that the user clicks Figure 24 After the operation of the video live broadcast button 2401 shown in (a) in FIG, the video live broadcast function of the TikTok application is started, and the following is displayed: Figure 24 The video live streaming acquisition interface shown in (b) of Figure 1. At this time, the Douyin application will collect image data and sound data.

[0281] 2302. Mobile phone A captures the Nth frame image and processes the Nth frame image.

[0282] This process is the same as above Figure 20 The 2202 shown is similar and will not be repeated here.

[0283] 2303. Mobile phone A collects the audio corresponding to the Nth frame image and processes the audio corresponding to the Nth frame image.

[0284] With the above Figure 1 Take the shooting environment and subject shown in the figure as an example. Figure 25 As shown, in the live video acquisition interface, the Nth frame image captured by the mobile phone is the image of the target object. The audio corresponding to the Nth frame image captured by the mobile phone will not only include the sound of the target object, but may also include the sound of non-target objects (such as non-target object 1 and non-target object 2), as well as environmental noise (such as diffuse field noise 1 and diffuse field noise 2).

[0285] This process is the same as above Figure 20 The 2203 shown is similar and will not be described again here.

[0286] 2304. Mobile phone A synthesizes the processed N-th frame image and the audio corresponding to the processed N-th frame image to obtain the N-th frame video data.

[0287] This process is the same as above Figure 20 It should be understood that when the Nth frame image and the processed audio corresponding to the Nth frame image are synthesized into the Nth frame video data, it can be displayed to the user.

[0288] 2305. Mobile phone A sends the Nth frame of video data to the server, so that mobile phone B displays the Nth frame of video data.

[0289] For example, after obtaining the Nth frame of video data, mobile phone A can send the Nth frame of video data to a server. It should be understood that the server is typically a live broadcast application server, such as the Douyin application server. When a user watching the live broadcast opens the live broadcast application on mobile phone B, such as the Douyin application, mobile phone B can display the Nth frame of video on the live broadcast display interface for the user to watch.

[0290] It should be noted that after the above 2303 pairs of audio corresponding to the N-th frame image, such as the 3N-2 frame audio, the 3N-1 frame and the 3N frame audio are processed, Figure 25 As shown, the audio signal in the Nth frame of video data output by mobile phone B is a processed audio signal, which only retains the sound of the target object (ie, the live broadcast object).

[0291] It should be understood that Figure 25 The scenario shown in the following example is to execute the above Figure 17 The scenario of the sound signal processing method shown is not shown. Figure 7 or Figure 13 Scenario of the sound signal processing method shown.

[0292] Scenario 3: Audio capture and processing of video files in a mobile phone album.

[0293] In some cases, electronic devices (such as mobile phones) do not support processing of recorded sound data during video file recording. In order to improve the clarity of the sound data of the video file, suppress the noise signal, and improve the signal ratio, the electronic device can perform the above processing on the video file saved in the mobile phone album and the original sound data saved. Figure 7 、 Figure 13 or Figure 17 The sound signal processing method removes the sound of non-target objects or diffuse field noise, making the noise smoother and improving the user's listening experience. The saved original sound data refers to the sound data collected by the mobile phone microphone when recording the video file in the mobile phone album.

[0294] For example, the embodiment of the present application further provides a sound pickup processing method for processing the sound pickup of video files in a mobile phone album. Figure 26 As shown, the sound pickup processing method includes:

[0295] 2601. The mobile phone obtains the first video file in the album.

[0296] In an embodiment of the present application, a user wants to perform sound pickup processing on the audio of a video file in a mobile phone album to remove the sound of a non-target object or remove diffuse field noise. The video file that the user wants to process can be selected from the mobile phone album to perform sound pickup processing.

[0297] For example, the mobile phone detects that the user clicks Figure 27 After the preview box 2701 of the first video file shown in (a) is operated, the mobile phone responds to the user's click operation and displays the following Figure 27 The operation interface for the first video file shown in (b) of FIG. In this operation interface, the user can play, share, collect, edit or delete the first video file.

[0298] For example, in Figure 27 In the operation interface of the first video file shown in (b), the user can also perform the sound pickup and noise reduction operation on the first video file. For example, the mobile phone detects that the user clicks Figure 28 After the operation of the "More" option 2801 shown in (a) of FIG. 1 is performed, the mobile phone responds to the user's click operation and displays the following Figure 28 The operation selection box 2802 shown in (b) of FIG. A "denoising process" option button 2803 is provided in the operation selection box 2802. The user can click the "denoising process" option button 2803 to perform sound denoising on the first video file.

[0299] For example, the mobile phone detects that the user clicks Figure 29 After the "denoising processing" option button 2803 shown in (a) is operated, the mobile phone can obtain the first video file in response to the user's click operation to perform sound denoising processing on the first video file. At this time, the mobile phone can display the following Figure 29 The waiting interface for executing the sound pickup and denoising process shown in (b) is displayed, and the following 2602 to 2604 are executed in the background to perform sound pickup and denoising on the first video file.

[0300] 2602. The mobile phone separates the first video file into first image data and first audio data.

[0301] It should be understood that the goal of performing sound pickup and denoising on the first video file is to remove the sound of non-target objects and suppress background noise (i.e., diffuse field noise). Therefore, the mobile phone needs to separate the audio data in the first video file in order to perform sound pickup and denoising on the audio data in the first video file.

[0302] For example, after the mobile phone obtains the first video file, the image data and audio data in the first video file can be separated into first image data and first audio data. The first image data can be the set of image frames from the first to the last frame in the image stream of the first video file. The first audio data can be the set of audio frames from the first to the last frame in the audio stream of the first video file.

[0303] 2603. The mobile phone performs image processing on the first image data to obtain second image data; and processes the first audio data to obtain second audio data.

[0304] For example, the mobile phone may process each frame of the first image data, such as performing image denoising and tone mapping, to obtain the second image data, wherein the second image data is a collection of images obtained after processing each frame of the first image data.

[0305] Exemplarily, the mobile phone can also process each frame of audio in the first audio data. For example, executing the above Figure 7 The sound signal processing method shown processes each frame of audio in the first audio data, suppresses the sound of non-target objects, and highlights the sound of the target object. Figure 13 The sound signal processing method shown processes each frame of audio in the first audio data, suppresses the sound of non-target objects to highlight the sound of the target object, and suppresses diffuse field noise to reduce the noise energy in the audio signal and improve the signal-to-noise ratio of the audio signal. For another example, executing the above Figure 17 The sound signal processing method shown processes each frame of audio in the first audio data, suppresses the sound of non-target objects to highlight the sound of the target object, and suppresses diffuse field noise to reduce the noise energy in the audio signal, thereby improving the signal-to-noise ratio of the audio signal. It can also smooth background noise (i.e., ambient noise) to provide users with a better listening experience.

[0306] After the mobile phone processes each frame of audio in the first audio data, the second audio data can be obtained, wherein the second audio data is a collection of audio data obtained after processing each frame of audio in the first audio data.

[0307] It should be understood that the above 2601 to 2603 can all correspond to Figure 7 or Figure 13 or Figure 17 400 steps in.

[0308] 2604. The mobile phone combines the second image data and the second audio data into a second video file.

[0309] It should be understood that the processed second image data and second audio data need to be synthesized into a video file before they can be shared or played by the user. Therefore, after the mobile phone executes the above 2603 to obtain the second image data and second audio data, the second image data and second audio data can be synthesized to form a second video file. In this case, the second video file is the first video file after the sound is picked up and denoised.

[0310] 2605. The mobile phone saves the second video file.

[0311] For example, when the mobile phone executes 2604 and synthesizes the second image data and the second audio data into a second video file, the sound pickup and denoising process is completed, and the mobile phone can display the following: Figure 30 The file save tab 3001 shown in FIG. In the file save tab 3001, the user may be prompted "Noise removal complete, do you want to replace the original file?" and provided with a first option button 3002 and a second option button 3003 for the user to select. The first option button 3002 may display "Yes" to instruct the user to replace the original first video file with the processed second video file. The second option button 3003 may display "No" to instruct the user to save the processed second video file separately without replacing the first video file, i.e., retaining both the first and second video files.

[0312] For example, if the mobile phone detects that the user Figure 30 In response to the user's click operation on the first option button 3002, the mobile phone replaces the original first video file with the processed second video file. Figure 30 In response to the user's click operation on the second option button 3003 , the mobile phone saves the first video file and the second video file respectively.

[0313] It is understandable that in the above embodiment, one frame of image does not correspond to one frame of audio. However, in some embodiments, one frame of image corresponds to multiple frames of audio, or multiple frames of image corresponds to one frame of audio. For example, one frame of image may correspond to three frames of audio. Figure 23 During real-time synthesis, the mobile phone may synthesize the processed Nth frame image and the processed audio of the 3N-2th frame, the 3N-1th frame, and the 3Nth frame to obtain the Nth frame video data.

[0314] The present application also provides another method for processing sound signals. This method is applied to an electronic device comprising a camera and a microphone. A first target object is within the camera's shooting range, and a second target object is not within the camera's shooting range. The first target object being within the camera's shooting range may mean that the first target object is within the camera's field of view. For example, the first target object may be the target object in the aforementioned embodiment. The second target object may be non-target object 1 or non-target object 2 in the aforementioned embodiment.

[0315] The method includes:

[0316] The electronic device activates the camera.

[0317] Display the preview interface, which includes a first control. The first control can be Figure 21A The recording button 2102 shown in (b) can also be Figure 24 The start video live broadcast button 2401 shown in (a) in FIG.

[0318] A first operation on a first control is detected. In response to the first operation, shooting is started. The first operation may be a click operation on the first control by a user.

[0319] At the first moment, the shooting interface is displayed, and the shooting interface includes a first image. The first image is an image captured in real time by the camera. The first image includes the first target object, and the first image does not include the second target object. The first moment can be any moment in the shooting process. The first image can be Figure 20 、 Figure 22 、 Figure 23 Each frame image in the method shown. The first target object can be the target object in the above embodiment. The second target object can be the non-target object 1 or the non-target object 2 in the above embodiment.

[0320] At the first moment, the microphone collects the first audio, which includes a first audio signal and a second audio signal. The first audio signal corresponds to the first target object, and the second audio signal corresponds to the second target object. Figure 1 Taking the illustrated scenario as an example, the first audio signal may be a sound signal of a target object, and the second audio signal may be a sound signal of a non-target object 1 or a sound signal of a non-target object 2.

[0321] A second operation on the first control of the shooting interface is detected. The first control of the shooting interface can be Figure 21B The end recording button 2103 shown in (a) in FIG. The second operation may be a click operation of the user on the first control of the shooting interface.

[0322] In response to the second operation, filming is stopped and the first video is saved. The first moment of the first video includes the first image and the second audio. The second audio includes the first audio signal and a third audio signal. The third audio signal is obtained by processing the second audio signal by the electronic device. The energy of the third audio signal is less than that of the second audio signal. For example, the third audio signal may be the processed sound signal of non-target object 1 or the sound signal of non-target object 2.

[0323] Optionally, the first audio may further include a fourth audio signal, which is a diffuse field noise audio signal; the second audio may further include a fifth audio signal, which is a diffuse field noise audio signal. The fifth audio signal is obtained by the electronic device processing the fourth audio signal. The energy of the fifth audio signal is less than the energy of the fourth audio signal. For example, the fourth audio signal may be the sound signal of the diffuse field noise 1 in the above embodiment, or the sound signal of the diffuse field noise 2 in the above embodiment. The fifth audio signal may be the sound signal of the electronic device performing the above Figure 13 or Figure 17 The processed sound signal of diffuse field noise 1 or the sound signal of diffuse field noise 2 is obtained.

[0324] Optionally, the fifth audio signal is obtained by the electronic device processing the fourth audio signal, including: suppressing the fourth audio signal to obtain the sixth audio signal. For example, the sixth audio signal may be obtained by performing Figure 13 The processed sound signal of diffuse field noise 1 is obtained.

[0325] The sixth audio signal is compensated to obtain a fifth audio signal. The sixth audio signal is a diffuse field noise audio signal, the energy of the sixth audio signal is less than that of the fourth audio signal, and the sixth audio signal is less than that of the fifth audio signal. For example, the fifth audio signal at this time may be Figure 17 The processed sound signal of diffuse field noise 1 is obtained.

[0326] Other embodiments of the present application provide an electronic device comprising: a microphone; a camera; one or more processors; a memory; and a communication module. The microphone is used to capture sound signals during video recording or live broadcasting; and the camera is used to capture image signals during video recording or live broadcasting. The communication module is used to communicate with external devices. The memory stores one or more computer programs, each of which includes instructions. When the processor executes the computer instructions, the electronic device can perform the various functions or steps performed by the mobile phone in the above-described method embodiments.

[0327] The embodiment of the present application also provides a chip system. The chip system can be applied to foldable electronic devices. Figure 31As shown, the chip system includes at least one processor 3101 and at least one interface circuit 3102. The processor 3101 and the interface circuit 3102 can be interconnected via lines. For example, the interface circuit 3102 can be used to receive signals from other devices (such as a memory of an electronic device). For another example, the interface circuit 3102 can be used to send signals to other devices (such as the processor 3101). Exemplarily, the interface circuit 3102 can read instructions stored in the memory and send the instructions to the processor 3101. When the instructions are executed by the processor 3101, the electronic device can execute the various steps in the above embodiments. Of course, the chip system can also include other discrete devices, which are not specifically limited in the embodiments of the present application.

[0328] An embodiment of the present application also provides a computer storage medium, which includes computer instructions. When the computer instructions are executed on the above-mentioned foldable electronic device, the electronic device executes the various functions or steps executed by the mobile phone in the above-mentioned method embodiment.

[0329] The embodiment of the present application further provides a computer program product, which, when executed on a computer, enables the computer to execute the functions or steps executed by the mobile phone in the above method embodiment.

[0330] Through the description of the above embodiments, those skilled in the art will clearly understand that for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. The specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0331] The functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0332] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as flash memory, mobile hard disk, read-only memory, random access memory, magnetic disk or optical disk.

[0333] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions within the technical scope disclosed in the present invention should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.

Claims

1. A sound signal processing method, characterized in that: Applicable to electronic devices, wherein the electronic devices include a camera and a microphone; The first target object is within the field of view of the camera, and the second target object is not within the field of view of the camera. The method includes: The electronic device activates a camera; Display a preview interface, wherein the preview interface includes a first control; detecting a first operation on the first control; In response to the first operation, starting shooting; At a first moment, a shooting interface is displayed, where the shooting interface includes a first image, the first image is an image captured in real time by the camera, the first image includes the first target object, and the first image does not include the second target object; At the first moment, the microphone collects first audio, the first audio including a first audio signal and a second audio signal, the first audio signal corresponding to the first target object, the second audio signal corresponding to the second target object, the first audio including multiple time-frequency speech signals, the multiple time-frequency speech signals being frequency domain representations of the first audio; detecting a second operation on the first control of the shooting interface; In response to a second operation, filming is stopped and a first video is saved, where the first video at a first moment includes a first image and a second audio signal, the second audio signal being obtained by the electronic device processing the first audio signal, the second audio signal including a first audio signal and a third audio signal, and energy of the third audio signal being less than energy of the second audio signal; The electronic device processes the first audio to obtain the second audio, including: Calculating the probability of each time-frequency speech signal in multiple spatial orientations; wherein the 360° space formed by the front, back, left, and right sides of the screen of the electronic device is divided into multiple spatial orientations; Summing the probabilities of the time-frequency speech signal being in a target orientation in the multiple spatial orientations to obtain a probability of each time-frequency speech signal being in the target orientation, where the target orientation is an orientation within the field of view of the camera when recording the video; Configuring the gain of the time-frequency speech signal according to the probability that the time-frequency speech signal is within the target direction; wherein, if the probability that the time-frequency speech signal is within the target direction is greater than a preset probability threshold, the gain of the time-frequency speech signal is equal to 1; if the probability that the time-frequency speech signal is within the target direction is less than or equal to the preset probability threshold, the gain of the time-frequency speech signal is less than 1; The second audio is obtained by processing the multiple time-frequency speech signals and the gain of each of the time-frequency speech signals.

2. The method according to claim 1, characterized in that The first audio further includes a fourth audio signal, and the fourth audio signal is a diffuse field noise audio signal; The second audio further includes a fifth audio signal, which is a diffuse field noise audio signal. The fifth audio signal is obtained by the electronic device processing the fourth audio signal, and energy of the fifth audio signal is less than energy of the fourth audio signal.

3. The method according to claim 2, characterized in that The fifth audio signal is obtained by the electronic device processing the fourth audio signal, including: performing suppression processing on the fourth audio signal to obtain a sixth audio signal; The fifth audio signal is obtained by performing compensation processing on the sixth audio signal, wherein the sixth audio signal is a diffuse field noise audio signal, has less energy than the fourth audio signal, and has less energy than the fifth audio signal.

4. The method according to claim 2, characterized in that The fifth audio signal is obtained by the electronic device processing the fourth audio signal, including: configuring the gain of the fourth audio signal to be less than 1; The fifth audio signal is obtained according to the fourth audio signal and the gain of the fourth audio signal.

5. The method according to any one of claims 1 to 4, characterized in that The method further comprises: At the first moment, after the microphone collects the first audio, the electronic device processes the first audio to obtain the second audio.

6. The method according to any one of claims 1 to 4, characterized in that The method further comprises: After stopping shooting in response to a second operation, the electronic device processes the first audio to obtain the second audio.

7. An electronic device, characterized in that: The electronic device comprises: microphone; Camera; one or more processors; Memory; Communication module; The microphone is used to collect sound signals during video recording or live broadcast; the camera is used to collect image signals during video recording or live broadcast; the communication module is used to communicate with an external device; the memory stores one or more computer programs, and the one or more computer programs include instructions. When the instructions are executed by the processor, the electronic device executes the method as described in any one of claims 1 to 6.

8. A computer-readable storage medium storing instructions, characterized in that: When the instructions are executed on an electronic device, the electronic device is caused to execute the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Voice recognition method and device based on audio and video recording, equipment and storage medium

    CN111816183A

  • Post-processing including median filtering of noise suppression gains

    US20130322640A1

  • Sound processing method and apparatus

    US20200118580A1