A recording processing method and related apparatus

By using sound source localization and separation technology, the pickup quality can be monitored and displayed in real time, solving the problem of insufficient pickup quality detection during the recording process and improving audio quality and user experience.

CN115691555BActive Publication Date: 2026-06-05BEIJING HONOR DEVICE CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING HONOR DEVICE CO LTD
Filing Date
2021-07-31
Publication Date
2026-06-05

Smart Images

  • Figure CN115691555B_ABST
    Figure CN115691555B_ABST
Patent Text Reader

Abstract

The application provides a recording processing method and related devices. The method can include: an electronic device can perform sound source positioning based on the sound collected by a microphone, obtain the position of a target sound source and the number of sound sources in the recording environment, and then perform sound source separation on the sound collected by the microphone according to the position of the target sound source and the number of sound sources in the recording environment to obtain the sound corresponding to the target sound source, i.e., a target audio signal. The electronic device can also determine the signal-to-noise ratio and display the current sound pickup quality to the user. This method can monitor and display the sound pickup quality to the user in real time, so that the user can adjust in time when the sound pickup quality is poor, thereby obtaining high-quality audio and improving the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of audio processing, and more particularly to a recording processing method and related apparatus. Background Technology

[0002] Sound pickup is a crucial step in recording video and audio. The quality of sound pickup directly impacts the quality of the user's recorded video and audio. Sound pickup, as the name suggests, refers to the process of collecting sound (audio). Ideally, the audio collected by the user should only include the target speech, i.e., the sound the user needs. However, there are often many interfering sounds during audio collection, which can potentially lead to low-quality audio finalized by the user.

[0003] Currently, audio quality can be improved by performing noise reduction processing on the collected audio (e.g., acoustic signal processing). However, in situations with a lot of interference, this processing method often cannot suppress the interference while ensuring the target speech remains undamaged. Furthermore, during video and audio recording, the ambient noise perceived by the user differs from the sound collected by the microphone, which may lead to the user only discovering the poor audio quality after recording is complete.

[0004] Therefore, how to detect the sound pickup quality in a timely manner and remind the user is an urgent problem that needs to be solved. Summary of the Invention

[0005] This application provides a recording processing method and related apparatus. It can locate sound sources based on sound collected by a microphone, obtain the location of the target sound source and the number of sound sources in the current recording environment, and then perform sound source separation on the sound collected by the microphone according to the location of the target sound source and the number of sound sources in the recording environment to obtain the audio signal generated by the target sound source, i.e., the target audio signal. The electronic device can also determine the signal-to-noise ratio and display the current sound pickup quality to the user. This method can monitor sound pickup quality in real time and display it to the user, allowing the user to adjust it promptly when the sound pickup quality is poor, thereby obtaining high-quality audio and improving the user experience.

[0006] Firstly, this application provides a recording processing method. This method can be applied to an electronic device. The method may include: receiving a recording start command; responding to the recording start command by acquiring a first audio signal; the first audio signal including audio signals generated by at least one sound source; processing the first audio signal to obtain the location of a target sound source; processing the first audio signal based on the location of the target sound source to obtain a second audio signal; and during recording, outputting prompt information based on the first and second audio signals. The second audio signal is the audio signal generated by the target sound source. The prompt information is used to characterize the sound pickup quality in the current recording environment.

[0007] In the solution provided in this application, the electronic device can receive a recording start command triggered by a user. In response to this command, the electronic device can acquire a first audio signal. It is understood that the first audio signal may include audio signals generated by one or more sound sources in the current recording environment. The electronic device can perform sound source separation on the first audio signal to obtain an audio signal generated by a target sound source, i.e., a second audio signal. The electronic device can also output prompt information based on the first and second audio signals. It is understood that this prompt information can characterize the sound pickup quality in the current recording environment. This method can monitor the sound pickup quality in real time and display it to the user, allowing the user to adjust promptly when the sound pickup quality is poor, thereby obtaining high-quality audio and improving the user experience.

[0008] In some embodiments of this application, the prompt information may be presented in the form of text, images, etc., and this application does not limit this.

[0009] In some embodiments of this application, the electronic device can determine the controllable response power at different locations in the current recording environment, and determine the location with the largest controllable response power as the location of the target sound source.

[0010] In conjunction with the first aspect, in one possible implementation of the first aspect, the first audio signal is processed based on the location of the target sound source to obtain a second audio signal. Specifically, this includes: amplifying the audio signal located in the direction of the target sound source based on its location to obtain a third audio signal; separating audio signals generated by different sound sources from the first audio signal using a first separation method based on the number of sound sources in the current recording environment to obtain a first separated audio signal set; separating audio signals generated by different sound sources from the first audio signal using a second separation method to obtain a second separated audio signal set; determining a fourth audio signal based on the third audio signal and the first separated audio signal set; determining a fifth audio signal based on the third audio signal and the second separated audio signal set; and determining the audio signal with the smallest absolute value among the third, fourth, and fifth audio signals as the second audio signal. The fourth audio signal is the audio signal in the first separated audio signal set that has the highest correlation with the third audio signal. The fifth audio signal is the audio signal in the second separated audio signal set that has the highest correlation with the third audio signal.

[0011] In the solution provided in this application, the electronic device can acquire the audio signal generated by the target sound source in various ways. For example, the electronic device can enhance the audio signal in the direction of the target sound source based on its location. Furthermore, the electronic device can separate audio signals generated by different sound sources from a first audio signal using different separation methods (e.g., a first separation method, a second separation method, etc.). The electronic device can also combine the above methods to acquire the audio signal generated by the target sound source to achieve a better separation effect, making the separated audio signal generated by the target sound source closer to the source signal (the audio signal generated by the target sound source that is unaffected by interference signals). This approach not only improves the accuracy of sound source separation but also improves the accuracy of the output prompt information.

[0012] In some embodiments of this application, electronic devices can enhance audio signals located in the direction of a target sound source using beamforming methods.

[0013] In some embodiments of this application, the first separation method can be a deep learning method. It is understood that the first separation method can also be a fixed-point algorithm, a support vector machine method, a Gaussian mixture model-based method, a non-negative matrix factorization-based method, a multi-repetition structure separation method, etc., and this application does not impose any limitations on this.

[0014] In some embodiments of this application, the second separation method can be a blind source separation method. It is understood that the second separation method can also be a fixed-point algorithm, a support vector machine method, a Gaussian mixture model-based method, a non-negative matrix factorization-based method, a multi-repeating structure separation method, etc., and this application does not impose any limitations on this.

[0015] In conjunction with the first aspect, in one possible implementation of the first aspect, the first audio signal is processed based on the location of the target sound source to obtain a second audio signal. Specifically, this includes: amplifying the audio signal located in the direction of the target sound source based on its location to obtain a third audio signal; separating audio signals generated by different sound sources from the first audio signal using a first separation method based on the number of sound sources in the current recording environment to obtain a first set of separated audio signals; determining a fourth audio signal based on the third audio signal and the first set of separated audio signals; and determining the audio signal with the smallest absolute value among the third and fourth audio signals as the second audio signal. The fourth audio signal is the audio signal in the first set of separated audio signals that has the strongest correlation with the third audio signal.

[0016] In the solution provided in this application, the electronic device can acquire the audio signal generated by the target sound source in various ways. For example, the electronic device can enhance the audio signal in the direction of the target sound source based on its location. Alternatively, the electronic device can separate audio signals generated by different sound sources from a first audio signal using different separation methods (e.g., a first separation method). The electronic device can also combine the above two methods to acquire the audio signal generated by the target sound source to achieve a better separation effect, making the separated audio signal generated by the target sound source closer to the source signal (the audio signal generated by the target sound source that is unaffected by interference signals). This approach not only improves the accuracy of sound source separation but also improves the accuracy of the output prompt information.

[0017] In conjunction with the first aspect, in one possible implementation of the first aspect, the first audio signal is processed based on the location of the target sound source to obtain a second audio signal. Specifically, this includes: enhancing the audio signal located in the direction of the target sound source based on its location to obtain a third audio signal; separating audio signals generated by different sound sources from the first audio signal using a second separation method to obtain a second set of separated audio signals; determining a fifth audio signal based on the third audio signal and the second set of separated audio signals; and identifying the audio signal with the smallest absolute value between the third and fifth audio signals as the second audio signal. The fifth audio signal is the audio signal in the second set of separated audio signals that has the strongest correlation with the third audio signal.

[0018] In the solution provided in this application, the electronic device can obtain the audio signal generated by the target sound source in various ways. For example, the electronic device can enhance the audio signal in the direction of the target sound source based on its location. Alternatively, the electronic device can separate audio signals generated by different sound sources from the first audio signal using different separation methods (e.g., a second separation method). The electronic device can also combine the above two methods to obtain the audio signal generated by the target sound source to achieve a better separation effect, making the separated audio signal generated by the target sound source closer to the source signal (the audio signal generated by the target sound source that is not affected by interference signals). This method not only improves the accuracy of sound source separation but also improves the accuracy of the output prompt information.

[0019] In conjunction with the first aspect, in one possible implementation of the first aspect, outputting prompt information based on the first audio signal and the second audio signal specifically includes: obtaining the signal-to-noise ratio (SNR) in the current recording environment based on the first audio signal and the second audio signal; determining the quality level of the sound pickup quality in the current recording environment by comparing the SNR with a preset threshold; and displaying the quality level of the sound pickup quality in the current recording environment. Different quality levels represent different sound pickup qualities.

[0020] In the solution provided in this application, the electronic device can obtain the signal-to-noise ratio (SNR) of the current recording environment based on the first audio signal and the second audio signal. It is understood that the second audio signal is the useful signal, and the difference between the first and second audio signals is the interference signal. The electronic device can also compare the SNR with a preset threshold to determine the quality level of the sound pickup in the current recording environment and display this quality level on the screen. In other words, the electronic device can monitor the sound pickup quality in real time and display it to the user, allowing the user to adjust the settings promptly when the sound pickup quality is poor, thereby obtaining high-quality audio and improving the user experience.

[0021] As is understandable, quality ratings can characterize the quality of sound pickup. For example, the highest quality rating can indicate good sound pickup quality.

[0022] It should be noted that the quality level displayed on the screen of an electronic device is not limited to text; it can also be in the form of images or other formats.

[0023] In some embodiments of this application, the quality level can directly represent the level of sound pickup quality in the current recording environment. For example, if the second quality level is "poor sound pickup quality," the electronic device can directly display the text "poor sound pickup quality" on the display screen.

[0024] It is understood that a preset threshold may include one or more thresholds. The preset threshold can be set according to actual needs, and this application does not impose any restrictions on it.

[0025] In conjunction with the first aspect, in one possible implementation of the first aspect, outputting prompt information based on the first audio signal and the second audio signal includes: obtaining the signal-to-noise ratio in the current recording environment based on the first audio signal and the second audio signal; and displaying the signal-to-noise ratio.

[0026] In the solution provided in this application, the electronic device can obtain the signal-to-noise ratio (SNR) of the current recording environment based on the first audio signal and the second audio signal, and directly display it on the display screen. It is understood that a higher SNR indicates better sound pickup quality in the current recording environment. In other words, the electronic device can monitor the sound pickup quality in real time and display it to the user, allowing the user to adjust the settings promptly when the sound pickup quality is poor, thereby obtaining high-quality audio and improving the user experience.

[0027] Secondly, this application provides an electronic device, including a display screen, one or more memories, and one or more processors. The one or more processors are coupled to the one or more memories, which store computer program code, including computer instructions. The processor can be used to: receive a recording start command; in response to the recording start command, acquire a first audio signal; the first audio signal includes audio signals generated by at least one sound source; process the first audio signal to obtain the location of a target sound source; based on the location of the target sound source, process the first audio signal to obtain a second audio signal; and during recording, output prompt information based on the first and second audio signals. The second audio signal is the audio signal generated by the target sound source. The prompt information is used to characterize the sound pickup quality in the current recording environment.

[0028] In conjunction with the second aspect, in one possible implementation of the second aspect, when the processor processes the first audio signal based on the location of the target sound source to obtain the second audio signal, it can specifically perform the following: based on the location of the target sound source, amplify the audio signal located in the direction of the target sound source to obtain a third audio signal; based on the number of sound sources in the current recording environment, separate audio signals generated by different sound sources from the first audio signal using a first separation method to obtain a first separated audio signal set; separate audio signals generated by different sound sources from the first audio signal using a second separation method to obtain a second separated audio signal set; determine a fourth audio signal based on the third audio signal and the first separated audio signal set; determine a fifth audio signal based on the third audio signal and the second separated audio signal set; and determine the audio signal with the smallest absolute value among the third, fourth, and fifth audio signals as the second audio signal. Wherein, the fourth audio signal is the audio signal in the first separated audio signal set that has the highest correlation with the third audio signal. The fifth audio signal is the audio signal in the second separated audio signal set that has the highest correlation with the third audio signal.

[0029] In conjunction with the second aspect, in one possible implementation of the second aspect, when the processor processes the first audio signal based on the location of the target sound source to obtain the second audio signal, it can specifically perform the following: based on the location of the target sound source, amplify the audio signal located in the direction of the target sound source to obtain a third audio signal; based on the number of sound sources in the current recording environment, separate the audio signals generated by different sound sources from the first audio signal using a first separation method to obtain a first set of separated audio signals; determine a fourth audio signal based on the third audio signal and the first set of separated audio signals; and determine the audio signal with the smallest absolute value among the third and fourth audio signals as the second audio signal. The fourth audio signal is the audio signal in the first set of separated audio signals that has the strongest correlation with the third audio signal.

[0030] In conjunction with the second aspect, in one possible implementation of the second aspect, when the processor processes the first audio signal based on the location of the target sound source to obtain the second audio signal, it can specifically perform the following: based on the location of the target sound source, amplify the audio signal located in the direction of the target sound source to obtain a third audio signal; separate audio signals generated by different sound sources from the first audio signal using a second separation method to obtain a second set of separated audio signals; determine a fifth audio signal based on the third audio signal and the second set of separated audio signals; and determine the audio signal with the smallest absolute value among the third and fifth audio signals as the second audio signal. The fifth audio signal is the audio signal in the second set of separated audio signals that has the strongest correlation with the third audio signal.

[0031] In conjunction with the second aspect, in one possible implementation of the second aspect, the processor, when outputting prompt information based on the first audio signal and the second audio signal, may specifically be used to: obtain the signal-to-noise ratio (SNR) in the current recording environment based on the first audio signal and the second audio signal; and determine the quality level of the sound pickup quality in the current recording environment by comparing the SNR with a preset threshold. It is understood that the electronic device may also include a display screen. The display screen can be used to display the quality level of the sound pickup quality in the current recording environment. Different quality levels represent different sound pickup qualities.

[0032] In conjunction with the second aspect, in one possible implementation of the second aspect, the processor, when outputting prompt information based on the first audio signal and the second audio signal, may specifically be used to: obtain the signal-to-noise ratio (SNR) of the current recording environment based on the first audio signal and the second audio signal. It is understood that the electronic device may also include a display screen. The display screen can be used to display the SNR.

[0033] Thirdly, this application provides a computer storage medium including instructions that, when executed on an electronic device, cause the electronic device to perform any of the possible implementations of the first aspect.

[0034] Fourthly, embodiments of this application provide a chip applied to an electronic device. The chip includes one or more processors, which are used to invoke computer instructions to cause the electronic device to execute any of the possible implementations in the first aspect described above.

[0035] Fifthly, embodiments of this application provide a computer program product containing instructions that, when the computer program product is run on a device, cause the electronic device to execute any of the possible implementations of the first aspect.

[0036] Understandably, the electronic device provided in the second aspect, the computer storage medium provided in the third aspect, the chip provided in the fourth aspect, and the computer program product provided in the fifth aspect are all used to execute any possible implementation of the first aspect. Therefore, the beneficial effects they can achieve can be referenced to the beneficial effects of any possible implementation of the first aspect, and will not be repeated here. Attached Figure Description

[0037] Figure 1 A schematic diagram of the hardware structure of an electronic device 100 provided in an embodiment of this application;

[0038] Figure 2 A schematic diagram of the software structure of an electronic device 100 provided in an embodiment of this application;

[0039] Figures 3A-3I A set of user interface diagrams provided for embodiments of this application;

[0040] Figure 4 This application provides a sound source localization method.

[0041] Figure 5 This application provides a method for separating sound sources.

[0042] Figure 6 This application provides yet another method for sound source separation.

[0043] Figure 7 This application provides yet another method for sound source separation.

[0044] Figure 8 This application provides a recording processing method for its embodiments. Detailed Implementation

[0045] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. In the description of the embodiments of this application, unless otherwise stated, " / " means "or," for example, A / B can mean A or B; "and / or" in the text is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Furthermore, in the description of the embodiments of this application, "multiple" refers to two or more than two.

[0046] It should be understood that the terms "first," "second," etc., in the specification, claims, and drawings of this application are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or apparatuses.

[0047] In this application, the reference to "embodiment" means that a specific feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a mutually exclusive, independent, or alternative embodiment. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described in this application can be combined with other embodiments.

[0048] Throughout its long history, media has undergone numerous transformations alongside technological innovation. Today, new media—a form of communication that utilizes digital technology to deliver information and services to users through computer networks, wireless communication networks, satellites, and terminals such as computers, mobile phones, and digital televisions—has become the mainstream media. New media accelerates information dissemination, enabling users to share information in a shorter time.

[0049] Video and audio are both important media channels in new media. Especially with the rise of short videos and audiobooks, more and more users are using their devices to record videos and audio and share them with others through new media. Therefore, improving the quality of recorded videos and audio has become a major demand for users.

[0050] Sound pickup is a crucial step in recording video and audio. The quality of sound pickup directly impacts the quality of the user's recorded video and audio. Sound pickup, as the name suggests, refers to the process of collecting sound (audio). Ideally, the audio collected by the user should only include the target speech, i.e., the sound the user needs. However, there are often many interfering sounds during audio collection, which can potentially lead to low-quality audio finalized by the user.

[0051] Currently, audio quality can be improved by performing noise reduction processing on the collected audio (e.g., acoustic signal processing). However, in situations with a lot of interference, this processing method often cannot suppress the interference while ensuring the target speech remains undamaged. Furthermore, during video and audio recording, the ambient noise perceived by the user differs from the sound collected by the microphone, which may lead to the user only discovering the poor audio quality after recording is complete.

[0052] This application provides a recording processing method and related apparatus, which can perform sound source localization on the sound collected by a microphone to obtain the location of the target sound source and the number of sound sources in the recording environment. Then, based on the location of the target sound source and the number of sound sources in the recording environment, the sound collected by the microphone is separated to obtain the target sound source, i.e., the target audio signal. The signal-to-noise ratio is then determined to indicate the current sound pickup quality to the user. This method can monitor the sound pickup quality in real time and display it to the user, allowing the user to adjust the settings promptly when the sound pickup quality is poor, thereby obtaining high-quality audio and improving the user experience.

[0053] The apparatus involved in the embodiments of this application is described below.

[0054] Figure 1 This is a schematic diagram of the hardware structure of an electronic device 100 provided in an embodiment of this application.

[0055] Electronic device 100 may include processor 110, external memory interface 120, internal memory 121, Universal Serial Bus (USB) interface 130, charging management module 140, power management module 141, battery 142, antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, sensor module 180, button 190, motor 191, indicator 192, camera 193, display screen 194, and Subscriber Identification Module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an accelerometer sensor 180E, a distance sensor 180F, a proximity sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.

[0056] It is understood that the structures illustrated in the embodiments of the present invention do not constitute a specific limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0057] Processor 110 may include one or more processing units, such as an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural network processing unit (NPU). These different processing units may be independent devices or integrated into one or more processors.

[0058] The controller can be the nerve center and command center of the electronic device 100. The controller can generate operation control signals according to the instruction opcode and timing signals to complete the control of fetching and executing instructions.

[0059] In the embodiments provided in this application, the electronic device 100 can use the processor 110 to perform sound source localization and sound source separation on the acquired audio signal to obtain the audio signal corresponding to the target sound source. The electronic device 110 can also determine the signal-to-noise ratio and determine the current sound pickup quality based on the signal-to-noise ratio.

[0060] The processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store instructions or data that the processor 110 has just used or that are used repeatedly. If the processor 110 needs to use the instruction or data again, it can retrieve it directly from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.

[0061] In some embodiments, the processor 110 may include one or more interfaces. The USB interface 130 is an interface compliant with the USB standard specification, specifically a Mini USB interface, a Micro USB interface, a USB Type-C interface, etc. The USB interface 130 can be used to connect a charger to charge the electronic device 100, and can also be used for data transfer between the electronic device 100 and peripheral devices. It can also be used to connect headphones for audio playback. This interface can also be used to connect other electronic devices 100, such as AR devices.

[0062] The charging management module 140 receives charging input from the charger. While charging the battery 142, the charging management module 140 can also supply power to the electronic device 100 through the power management module 141.

[0063] The power management module 141 is used to connect the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140 to power the processor 110, internal memory 121, external memory, display 194, camera 193, and wireless communication module 160, etc.

[0064] The wireless communication function of electronic device 100 can be realized through antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modem processor and baseband processor, etc.

[0065] Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in electronic device 100 can be used to cover one or more communication frequency bands. Different antennas can also be reused to improve antenna utilization.

[0066] The mobile communication module 150 can provide solutions for wireless communication, including 2G / 3G / 4G / 5G, applied to the electronic device 100. The mobile communication module 150 may include at least one filter, switch, power amplifier, low-noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves via antenna 1, and perform filtering, amplification, and other processing on the received electromagnetic waves before transmitting them to a modem processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modem processor and convert it into electromagnetic waves for radiation via antenna 1.

[0067] The modem processor may include a modulator and a demodulator. The modulator modulates the low-frequency baseband signal to be transmitted into a mid-to-high frequency signal. The demodulator demodulates the received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After processing by the baseband processor, the low-frequency baseband signal is transmitted to the application processor. The application processor outputs sound signals through audio devices (not limited to speaker 170A, receiver 170B, etc.) or displays images or videos through the display screen 194.

[0068] The wireless communication module 160 can provide solutions for wireless communication applications on the electronic device 100, including Wireless Local Area Networks (WLANs) (such as Wireless Fidelity (Wi-Fi) networks), Bluetooth (BT), Global Navigation Satellite System (GNSS), Frequency Modulation (FM), Near Field Communication (NFC), and Infrared (IR) technologies. The wireless communication module 160 can be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via antenna 2, performs frequency modulation and filtering of the electromagnetic wave signals, and sends the processed signal to processor 110. The wireless communication module 160 can also receive signals to be transmitted from processor 110, perform frequency modulation and amplification, and convert them into electromagnetic waves for radiation via antenna 2.

[0069] In some embodiments, antenna 1 of electronic device 100 is coupled to mobile communication module 150, and antenna 2 is coupled to wireless communication module 160, so that electronic device 100 can communicate with networks and other devices through wireless communication technology.

[0070] Electronic device 100 implements display functions through a GPU, a display screen 194, and an application processor. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations and for graphics rendering. Processor 110 may include one or more GPUs, which execute program instructions to generate or modify display information.

[0071] Display screen 194 is used to display images, videos, etc. Display screen 194 includes a display panel. The display panel may be a Liquid Crystal Display (LCD), an Organic Light-Emitting Diode (OLED), an Active-Matrix Organic Light-Emitting Diode (AMOLED), a Flexible Light-Emitting Diode (FLED), Mini LED, Micro LED, Micro-OLED, Quantum Dot Light-Emitting Diodes (QLED), etc. In some embodiments, electronic device 100 may include one or N displays 194, where N is a positive integer greater than 1.

[0072] In some embodiments of this application, the display screen 194 may display the signal-to-noise ratio calculated by the processor 110.

[0073] In some other embodiments of this application, the display screen 194 may display the current pickup quality determined by the processor.

[0074] Electronic device 100 can acquire data through ISP, camera 193, video codec, GPU, display 194, and application processor.

[0075] The ISP (Image Signal Processor) is used to process data fed back from the camera 193. For example, when taking a picture, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, and the camera's photosensitive element transmits the electrical signal to the ISP for processing, converting it into an image or video visible to the naked eye. The ISP can also perform algorithmic optimization of image noise, brightness, and skin tone. The ISP can also optimize parameters such as exposure and color temperature of the shooting scene. In some embodiments, the ISP can be set in the camera 193.

[0076] Camera 193 is used to capture still images or videos. An object passes through the lens, generating an optical image that is projected onto a photosensitive element. This photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then passed to an ISP (Internet Service Provider) for conversion into a digital image or video signal. The ISP outputs the digital image or video signal to a DSP (Digital Signal Processor) for further processing. The DSP converts the digital image or video signal into standard RGB, YUV, or other image or video signal formats.

[0077] In some embodiments, the electronic device 100 may include one or N cameras 193, where N is a positive integer greater than 1. For example, in some embodiments, the electronic device 100 may use the N cameras 193 to acquire images with multiple exposure intensities, and then, in video post-processing, the electronic device 100 may synthesize an HDR image based on the multiple images with multiple exposure intensities using HDR technology.

[0078] In this embodiment, the electronic device 100 may include two cameras 193. These two cameras are a main camera and an auxiliary camera, respectively. If the user triggers a shooting action, the main camera and the auxiliary camera can simultaneously acquire images, which are then merged into a single image by the electronic device 100. This single image is then displayed to the user.

[0079] In this embodiment, the light-sensing capability of the camera 193 can be characterized by a sensitivity coefficient. It can also be understood that the light-sensing capability of the photosensitive element in the camera 193 can be represented by the sensitivity coefficient. Under the same exposure intensity, the larger the sensitivity coefficient, the stronger the light-sensing capability of the photosensitive element, and the brighter the image acquired by the camera 193.

[0080] In this embodiment, the main camera and the auxiliary camera can acquire Raw images and transmit them to the ISP for processing. Based on the brightness of the Raw image and the exposure intensity of the two cameras, the ISP can adjust the exposure intensity of the main camera or the auxiliary camera to ensure that the brightness of the images acquired by the two cameras is consistent.

[0081] A digital signal processor (DSP) is used to process digital signals. Besides processing digital image or video signals, it can also process other digital signals. For example, when electronic device 100 selects a frequency, the DSP is used to perform Fourier transforms on the frequency energy.

[0082] Video codecs are used to compress or decompress digital video. Electronic device 100 may support one or more video codecs. Thus, electronic device 100 can play or record video in various encoding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.

[0083] NPU stands for Neural Network (NN) computing processor. By borrowing the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, it can rapidly process input information and continuously learn on its own. NPUs enable intelligent cognitive applications in electronic devices, such as image recognition, facial recognition, speech recognition, and text understanding.

[0084] The external storage interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 through the external storage interface 120 to perform data storage functions. For example, music, video, and other files can be saved on the external memory card.

[0085] Internal memory 121 can be used to store computer executable program code, which includes instructions. Processor 110 executes various functional applications and data processing of electronic device 100 by running the instructions stored in internal memory 121. Internal memory 121 may include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback, image and video playback, etc.). The data storage area may store data created during the use of electronic device 100 (such as audio data, phonebook, etc.).

[0086] Electronic device 100 can implement audio functions, such as music playback and recording, through audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, and application processor.

[0087] The audio module 170 is used to convert digital audio information into analog audio signal output, and also to convert analog audio input into digital audio signal.

[0088] The loudspeaker 170A, also known as a "loudspeaker", is used to convert audio electrical signals into sound signals.

[0089] The receiver 170B, also known as the "earpiece", is used to convert audio electrical signals into sound signals.

[0090] Microphone 170C, also known as a "microphone" or "voice transducer," is used to convert sound signals into electrical signals. Electronic device 100 may include a microphone array. The microphone array may include at least one microphone 170C.

[0091] In some embodiments of this application, the electronic device 100 may acquire audio signals via a microphone array.

[0092] The 170D headphone jack is used to connect wired headphones.

[0093] The sensor module 180 may include one or more sensors, which may be of the same or different types. Understandably, Figure 1 The sensor module 180 shown is only an exemplary division method, and there may be other division methods, which are not limited in this application.

[0094] Pressure sensor 180A is used to sense pressure signals and convert them into electrical signals. In some embodiments, pressure sensor 180A may be disposed on display screen 194. When a touch operation is applied to display screen 194, electronic device 100 detects the intensity of the touch operation based on pressure sensor 180A. Electronic device 100 may also calculate the touch position based on the detection signal from pressure sensor 180A. In some embodiments, touch operations applied to the same touch position but with different touch operation intensities may correspond to different operation commands.

[0095] The gyroscope sensor 180B can be used to determine the motion attitude of the electronic device 100. In some embodiments, the angular velocity of the electronic device 100 about three axes (i.e., the x, y, and z axes) can be determined by the gyroscope sensor 180B. The gyroscope sensor 180B can be used for image stabilization.

[0096] The barometric pressure sensor 180C is used to measure air pressure. In some embodiments, the electronic device 100 calculates altitude using the air pressure value measured by the barometric pressure sensor 180C to assist in positioning and navigation.

[0097] The magnetic sensor 180D includes a Hall sensor. The electronic device 100 can use the magnetic sensor 180D to detect the opening and closing of the flip cover.

[0098] The accelerometer 180E can detect the magnitude of acceleration of electronic device 100 in various directions (typically three axes). When electronic device 100 is stationary, it can detect the magnitude and direction of gravity. It can also be used to identify the posture of electronic device 100, and can be applied to applications such as screen orientation switching and pedometers.

[0099] A distance sensor 180F is used to measure distance. Electronic device 100 can measure distance via infrared or laser. In some embodiments, during a shooting scene, electronic device 100 can utilize the distance sensor 180F to measure distance for rapid focusing.

[0100] The proximity sensor 180G may include, for example, a light-emitting diode (LED) and a light detector, such as a photodiode. The LED may be an infrared LED. The electronic device 100 emits infrared light outward through the LED. The electronic device 100 uses the photodiode to detect infrared reflected light from a nearby object. When sufficient reflected light is detected, it can be determined that an object is near the electronic device 100. When insufficient reflected light is detected, the electronic device 100 can determine that no object is near the electronic device 100.

[0101] The 180L ambient light sensor is used to detect ambient light intensity.

[0102] The fingerprint sensor 180H is used to acquire fingerprints.

[0103] The 180J temperature sensor is used to detect temperature.

[0104] Touch sensor 180K, also known as a "touch panel," can be located on display screen 194. The touch sensor 180K and display screen 194 together form a touchscreen, also known as a "touch screen." Touch sensor 180K detects touch operations applied to or near it. The touch sensor can transmit the detected touch operation to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through display screen 194. In other embodiments, touch sensor 180K may also be located on the surface of electronic device 100, in a different position than display screen 194.

[0105] In one embodiment of this application, a user uses electronic device 100 to perform time-lapse photography or continuous shooting, requiring the acquisition of a series of images. In the time-lapse or continuous shooting scenario, electronic device 100 can adopt AE mode. That is, electronic device 100 automatically adjusts the AE value. During the preview of this series of images, if the user performs a touch operation on display screen 194, touch AE mode may be triggered. In touch AE mode, electronic device 100 can adjust the brightness of the corresponding location on the display screen touched by the user and perform high-weighted metering. This makes the weight of the user-touched area significantly higher than other areas when calculating the average brightness of the image, resulting in an average brightness that is closer to the average brightness of the user-touched area.

[0106] The bone conduction sensor 180M can acquire vibration signals.

[0107] Buttons 190 include a power button, volume buttons, etc. Buttons 190 can be mechanical buttons or touch-sensitive buttons. Electronic device 100 can receive button input and generate key signal inputs related to user settings and function control of electronic device 100.

[0108] Motor 191 can generate vibration alerts. Motor 191 can be used for incoming call vibration alerts or for touch vibration feedback. For example, different vibration feedback effects can correspond to touch operations performed on different applications (such as taking photos, playing audio, etc.). Motor 191 can also correspond to different vibration feedback effects for touch operations performed on different areas of the display screen 194. Different application scenarios (such as time reminders, receiving messages, alarm clocks, games, etc.) can also correspond to different vibration feedback effects. The touch vibration feedback effect can also be customized.

[0109] Indicator 192 can be an indicator light, used to indicate charging status, power changes, or to indicate messages, missed calls, notifications, etc.

[0110] The SIM card interface 195 is used to connect a SIM card. The SIM card can be inserted into or removed from the SIM card interface 195 to achieve contact and separation with the electronic device 100. The electronic device 100 can support one or N SIM card interfaces, where N is a positive integer greater than 1. The electronic device 100 interacts with the network through the SIM card to realize functions such as calls and data communication. In some embodiments, the electronic device 100 uses an eSIM, i.e., an embedded SIM card. The eSIM card can be embedded in the electronic device 100 and cannot be separated from the electronic device 100.

[0111] Figure 2 This is a schematic diagram of the software structure of an electronic device 100 provided in an embodiment of this application.

[0112] A layered architecture divides software into several layers, each with a clear role and function. Layers communicate with each other through software interfaces. In some embodiments, the system is divided into four layers, from top to bottom: the application layer, the application framework layer, the runtime and system libraries, and the kernel layer.

[0113] The application layer can include a series of application packages.

[0114] like Figure 2 As shown, the application package may include applications (also known as apps) such as camera, gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, video, and SMS.

[0115] The application framework layer provides an Application Programming Interface (API) and programming framework for applications in the application layer. The application framework layer includes some predefined functions.

[0116] like Figure 2 As shown, the application framework layer may include a window manager, content provider, view system, phone manager, resource manager, notification manager, etc.

[0117] The window manager is used to manage windowed applications. It can retrieve screen size, determine the presence of a status bar, lock the screen, and capture screenshots, among other things.

[0118] Content providers store and retrieve data, making that data accessible to applications. This data may include videos, images, audio, made and received phone calls, browsing history and bookmarks, phone books, etc.

[0119] A view system includes visual controls, such as controls for displaying text and controls for displaying images. View systems can be used to build applications. A display interface can consist of one or more views. For example, a display interface including a text notification icon could include views for displaying text and views for displaying images.

[0120] The phone manager is used to provide communication functions for electronic device 100. For example, it manages call status (including connection and disconnection).

[0121] The file explorer provides applications with various resources, such as localized strings, icons, images, layout files, video files, and more.

[0122] The notification manager allows applications to display notifications in the status bar. These notifications can be used to deliver informational messages and can disappear automatically after a short pause, requiring no user interaction. For example, the notification manager can be used to notify users of completed downloads or message alerts. The notification manager can also display notifications as icons or scrolling text in the top status bar, such as notifications from background applications, or as dialog-style notifications on the screen. Examples include displaying text messages in the status bar, emitting sounds, vibrating electronic devices, and flashing indicator lights.

[0123] The runtime consists of the core libraries and the virtual machine. The runtime is responsible for system scheduling and management.

[0124] The core library consists of two parts: one part is the functionalities that the programming language (e.g., Java) needs to call, and the other part is the system's core library.

[0125] The application layer and application framework layer run in a virtual machine. The virtual machine executes the programming files (e.g., Java files) of the application layer and application framework layer as binary files. The virtual machine is used to perform functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection.

[0126] System libraries can include multiple functional modules. For example: Surface Manager, Media Libraries, 3D graphics processing libraries (e.g., OpenGL ES), 2D graphics engines (e.g., SGL), etc.

[0127] The Surface Manager is used to manage the display subsystem and provides the fusion of two-dimensional (2D) and three-dimensional (3D) layers for multiple applications.

[0128] The media library supports playback and recording of various common audio and video formats, as well as still image files. It supports multiple audio and video encoding formats, such as MPEG4, H.264, MP3, AAC, AMR, JPG, and PNG.

[0129] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, compositing, and layer processing.

[0130] A 2D graphics engine is a graphics engine for 2D drawing.

[0131] The kernel layer is the layer between hardware and software. The kernel layer includes at least the display driver, camera driver, audio driver, sensor driver, and virtual card driver.

[0132] The following describes some recording scenarios provided in the embodiments of this application.

[0133] It is understood that the term "user interface" in the specification, claims, and drawings of this application refers to the medium interface through which an application or operating system interacts and exchanges information with the user. It realizes the conversion between the internal form of information and the form acceptable to the user. The most common form of user interface is the graphical user interface (GUI), which refers to a user interface related to computer operation displayed graphically. It can be an interface element such as an icon, window, or control displayed on the screen of an electronic device. The control can include visual interface elements such as icons, buttons, menus, tabs, text boxes, dialog boxes, status bars, navigation bars, and widgets.

[0134] 1. Recording scene ( Figures 3A-3F )

[0135] (1) Start recording and begin sound pickup. Figures 3A-3C )

[0136] Figure 3A An exemplary user interface 310 on an electronic device 100 is shown for displaying applications installed on the electronic device 100.

[0137] User interface 310 displays a page with application icons, which may include multiple application icons (e.g., weather app icon, calendar app icon, email app icon, settings app icon, app store app icon, notepad app icon, photo album app icon, voice recorder app icon 311, etc.). Below these multiple application icons, a page indicator may also be displayed to indicate the positional relationship between the currently displayed page and other pages. Below the page indicator are multiple tray icons (e.g., camera app icon 312, browser app icon, messaging app icon, dialer app icon). The tray icons remain displayed when switching pages. This embodiment does not limit the content displayed on user interface 310.

[0138] Electronic device 100 can detect user actions on camera application icon 312, and in response to the action, electronic device 100 can display... Figure 3B The user interface 320 shown.

[0139] The user interface 320 may include a parameter adjustment area 321, a preview area 322, a camera mode option 323, an album shortcut control 324, a shutter control 325, and a camera flip control 326.

[0140] Preview area 322 can be used to display a preview image. This preview image is an image captured in real time by the electronic device 100 through a camera. The electronic device can refresh the display content in preview area 322 in real time so that the user can preview the image currently captured by the camera.

[0141] Camera mode option 323 may display one or more shooting mode options. These one or more shooting mode options may include: Night mode option 3231, Smart portrait mode option 3232, Photo mode option 3233, Video mode option 3234, and more options 3235.

[0142] The photo album shortcut control 324 can be used to open the photo album application. In response to a user operation, such as a touch operation, applied to the photo album shortcut control 324, the electronic device 100 can open the photo album application.

[0143] The shutter control 325 can be used to listen for user actions that trigger taking a picture. The electronic device 100 can detect the user actions performed on the shutter control 325, and in response to the actions, the electronic device 100 can save the preview image in the preview area 322 as a picture in the photo album application.

[0144] The camera flip control 326 can be used to listen for user actions that trigger the camera flip.

[0145] The controls included in the user interface 320 may be buttons or other forms of controls, and this application does not limit this. In addition, the user interface 320 may also include more or fewer controls, and this application embodiment does not limit this.

[0146] It is understood that the user operations mentioned in this application include, but are not limited to, touch, voice control, and gestures.

[0147] Electronic device 100 can detect user operation on recording mode option 3234, and in response to the operation, electronic device 100 can display as follows: Figure 3C The user interface 330 shown is understood to be a video preview interface in a camera application.

[0148] The controls included in user interface 330 are essentially the same as those included in user interface 320. Additionally, user interface 330 may include a video preview area 331 and a start recording control 332. The video preview area 331 can be used to display a preview image.

[0149] The electronic device 100 can detect user operations on the start recording control 332, and in response to the operation, the electronic device 100 can start recording video. At the same time, the microphone array in the electronic device 100 can also collect audio signals.

[0150] (2) Electronic equipment 100 test sound pickup quality ( Figures 3D-3F )

[0151] Electronic device 100 can locate sound sources based on the acquired audio signals, obtaining the location and number of target sound sources. Then, based on the location and number of target sound sources, it performs sound source separation, extracting the audio signals corresponding to the target sound sources. Electronic device 100 can calculate the signal-to-noise ratio (SNR) based on the audio signals corresponding to the target sound sources and display the SNR to the user. It is understood that the SNR mentioned here refers to the ratio of the audio signal corresponding to the target sound source to other interference signals in the environment.

[0152] like Figure 3D As shown, the electronic device 100 can display a user interface 340. It can be understood that the user interface 340 can be a recording interface in a camera application.

[0153] The controls included in user interface 340 are basically the same as those included in user interface 330. The difference is that user interface 340 does not have a camera mode option 323. User interface 340 may include a shutter control 341, an end recording control 342, a pause recording control 343, and a recording time control 344. The shutter control 341 can be used to trigger a photo, meaning the user can trigger the shutter control 341 to take a photo during recording. The end recording control 342 can be used to end video recording. The pause recording control 343 can be used to temporarily stop video recording. The recording time control 344 can indicate the current recording time. Figure 3D As shown, the recording time control 344 displays 00:00:02, which means that the current video has been recorded for 2 seconds (s).

[0154] It is understood that the user interface 340 may also include a video display area 345. The video display area 345 may include a display area 3451. The display area 3451 may be used to display the signal-to-noise ratio calculated by the electronic device 100. For example... Figure 3D As shown, display area 3451 displays the current signal-to-noise ratio as 5 dB. Users can determine whether the current sound pickup quality is good based on the content displayed in display area 3451, and thus decide whether to stop recording.

[0155] It is understood that the display area 3451 may include pop-up windows, etc., and this application does not limit the form of the display area 3451.

[0156] In some embodiments of this application, the electronic device 100 can determine the current sound pickup quality based on the calculated signal-to-noise ratio, and then display the current sound pickup quality to the user. It is understood that the electronic device 100 can display, for example... Figure 3E The user interface shown. (As shown) Figure 3E As shown, display area 3451 can show that the current sound pickup quality is good.

[0157] Furthermore, the content displayed in display area 3451 can change according to the sound pickup quality detected by electronic device 100. For example... Figure 3F As shown, the electronic device 100 can display a user interface 350. The controls included in the user interface 350 are identical to those included in the user interface 340. The recording time control 344 included in the user interface 350 displays 00:00:05, indicating that the current video has been recorded for 5 seconds. At this time, the signal-to-noise ratio displayed in the display area 3451 is 33dB.

[0158] 2. Recording scenarios ( Figures 3G-3I )

[0159] (1) Start picking up sound ( Figure 3G )

[0160] Electronic device 100 can detect user actions on the recorder application icon 311, and in response to the action, electronic device 100 can display... Figure 3G The user interface shown is 360°.

[0161] The user interface 360 ​​may include a display area 361 and a start recording control 362. The display area 361 may display a list of recording files, including already recorded audio files. The start recording control 362 can be used to begin recording audio.

[0162] Electronic device 100 can detect user operation on start recording control 362, and in response to the user operation, electronic device 100 can start recording audio.

[0163] (2) Electronic equipment 100 test sound pickup quality

[0164] Electronic device 100 can locate sound sources based on the acquired audio signals, obtaining the location and number of target sound sources. Then, based on the location and number of target sound sources, it performs sound source separation, extracting the audio signals corresponding to the target sound sources. Electronic device 100 can calculate the signal-to-noise ratio (SNR) based on the audio signals corresponding to the target sound sources and display the SNR to the user. It is understood that the SNR mentioned here refers to the ratio of the audio signal corresponding to the target sound source to other interference signals in the environment.

[0165] Electronic device 100 can display such as Figure 3H The user interface shown is 370.

[0166] The user interface 379 may include a recording display area 371, a recording file list shortcut control 372, a pause recording control 373, a stop recording control 374, and a recording time control 375. The recording display area 371 displays the current recording volume. The recording file list shortcut control 372 displays a list of recording files. The pause recording control 373 pauses audio recording. The stop recording control 374 stops audio recording. The recording time control 375 displays the duration of the currently recorded audio. Figure 3H As shown, the recording time control 375 displays 00:00:03, which means that 3 seconds of audio has been recorded.

[0167] It is understood that the recording display area 371 may include display area 3712. Display area 3721 may include display area 3712. Display area 3712 may be used to display the signal-to-noise ratio calculated by electronic device 100. For example... Figure 3HAs shown, display area 3712 displays the current signal-to-noise ratio as 5 dB. Users can judge whether the current sound pickup quality is good based on the content displayed in display area 3712, and thus decide whether to stop recording.

[0168] It is understood that the display area 3451 may include pop-up windows, etc., and this application does not limit the form of the display area 3712.

[0169] In some embodiments of this application, the electronic device 100 can determine the current sound pickup quality based on the calculated signal-to-noise ratio, and then display the current sound pickup quality to the user. It is understood that the electronic device 100 can display, for example... Figure 3I The user interface shown. (As shown) Figure 3I As shown, display area 3712 can display the current poor sound pickup quality.

[0170] Furthermore, the content displayed in display area 3712 can change depending on the sound pickup quality detected by electronic device 100. This is understandable; please refer to the above description for further details.

[0171] The following is combined Figure 4 This application provides a method for locating a sound source.

[0172] S401: Electronic device 100 acquires audio signals through a microphone array.

[0173] During sound recording, the vibrations of the sound can propagate to the diaphragm of the microphone in the electronic device 100, forming a changing electric field and producing an audio signal. In other words, the microphone can convert sound signals into electrical signals.

[0174] It is understood that the electronic device 100 includes one or more microphones, which together form a microphone array. The electronic device 100 can acquire audio signals through this microphone array.

[0175] In some embodiments of this application, the microphone array in the electronic device 100 may include three microphones. For example, the electronic device 100 may include a top microphone, a rear microphone, and a bottom microphone.

[0176] It is understandable that the audio signal acquired by the electronic device 100 through the microphone array is an audio signal in the time domain. The time domain can describe the relationship between mathematical functions or physical signals and time.

[0177] S402: Electronic device 100 performs frame-by-frame processing and Fourier transform on the acquired audio signal to obtain the audio signal in the frequency domain.

[0178] Overall, audio signals exhibit non-stationary and time-varying characteristics. However, within a short timeframe, these characteristics remain largely unchanged, indicating relative stability. In other words, audio signals possess short-term stationarity. Therefore, audio signal processing often involves frame segmentation to reduce the impact of overall non-stationary and time-varying characteristics. This means that audio signals are typically analyzed in segments. Each segment is called a frame, and the length of a frame is called the frame length. Frame lengths are typically 10ms-30ms.

[0179] In some embodiments of this application, the frame length can be 10ms.

[0180] Electronic device 100 can perform a Fourier transform on the framed audio signal to obtain the audio signal in the frequency domain. In other words, electronic device 100 can convert an audio signal in the time domain into an audio signal in the frequency domain using a Fourier transform. The frequency domain is a coordinate system used to describe the frequency characteristics of a signal.

[0181] Specifically, the audio signal of the m-th microphone in the i-th frame can be denoted as: x m (i, k). Understandable, x m (i, k) is a discrete-time domain signal. Electronic device 100 can process x. m Perform a discrete Fourier transform on (i, k) and convert x to x m Transforming (i, k) into the frequency domain yields the audio signal in the frequency domain, denoted as: X m (i, k). Understandable. Where k represents a discrete frequency point. k = 0, 1, 2, ..., N-1.

[0182] It should be noted that before performing a Fourier transform on the framed audio signal, the electronic device 100 can also perform other processing on the audio signal. For example, the electronic device 100 can perform windowing processing on the audio signal, which can improve the resolution of the Fourier transform result (i.e., the spectrum).

[0183] S403: Electronic device 100 divides the space into grids, calculates the output power (i.e., controllable response power) of the microphone array guided to each grid based on the audio signal in the frequency domain, and determines the position of the target sound source as the position of the position with the maximum controllable response power.

[0184] Electronic device 100 can adopt a sound source localization method based on phase transformation weighted controllable response power (SRP-PHAT).

[0185] Electronic device 100 can divide the space containing the sound source into a grid, with each grid containing a hypothetical sound source. Electronic device 100 can calculate the time delay difference between each hypothetical sound source and a pair of microphones at specified locations (i.e., two microphones in a microphone array), and calculate the value of the Generalized Cross-Correlation (GCC) function based on this time delay difference. Electronic device 100 can also sum multiple GCC function values ​​to obtain the controllable response power. The hypothetical sound source location corresponding to the controllable response power is the estimated sound source location.

[0186] Understandably, a controllable response can be expressed as: The controllable response power can be expressed as: The estimated location of the sound source can be expressed as:

[0187] Where q represents the coordinate vector of the imaginary sound source. τ m (q) represents the time delay difference between the imaginary sound source reaching the m-th microphone and reaching the reference microphone. R represents the space where the sound source is located.

[0188] It is understandable that when traversing a hypothetical sound source in a grid, rectangular coordinates, polar coordinates, cylindrical coordinates, or spherical coordinates can be used. Similarly, the coordinates of the hypothetical sound source can also be represented using rectangular coordinates, polar coordinates, cylindrical coordinates, or spherical coordinates.

[0189] It should be noted that the electronic device 100 can determine the number of sound sources in the current environment based on the number of maximum values ​​of the controllable response power. Specifically, the electronic device 100 determines the number of maximum values ​​of the controllable response power as the number of sound sources in the current environment. Of course, the electronic device 100 can also determine the number of sound sources in the current environment through other methods, and this application does not limit this.

[0190] In some embodiments of this application, the electronic device 100 can further determine the location of the target sound source by combining images. For example, in scenarios where a user uses the electronic device 100 to record video, the microphone of the electronic device 100 will pick up sound, and at the same time, its camera will capture images and display them on the screen. In these scenarios, the user can select the target sound source on the image by touch, voice control, gestures, etc.

[0191] Understandably, users can select one or more target sound sources. Additionally, the camera capturing the image can be either a front-facing camera or a rear-facing camera.

[0192] Specifically, the electronic device 100 can display images captured by a camera on a screen and can also detect user actions (touch, voice control, gestures, etc.) applied to the image. The electronic device 100 can determine the target sound source based on the user action, and thus determine the location of the target sound source. For example, if a user selects a target sound source (e.g., a person) in an image displayed on the screen and clicks on the corresponding area, the electronic device 100 can detect the user action, record the location of the click, and use it as the location of the target sound source.

[0193] In some embodiments of this application, the location of the target sound source selected by the user and recorded by the electronic device 100 can be represented by rectangular coordinates, polar coordinates, cylindrical coordinates, or spherical coordinates.

[0194] In some embodiments of this application, the location of the target sound source selected by the user, recorded by the electronic device 100, can also be represented by a relative position. For example, if the user clicks on a location in the upper left area of ​​an image, the electronic device 100 can record that the target sound source is located in the upper left area of ​​the image.

[0195] It is understood that the electronic device 100 may also record the location of the target sound source selected by the user in other ways, and this application does not limit this.

[0196] In some embodiments of this application, the electronic device 100 can display the image captured by the camera on the screen while simultaneously picking up sound, but the user has not selected a target sound source. In this case, the electronic device 100 can combine stored user images (e.g., facial data recorded for unlocking the electronic device 100) to determine whether the user of the device is present in the image. If the image displayed by the electronic device 100 includes the user of the device (e.g., the electronic device 100 detects a face identical to the recorded facial data), the electronic device 100 can default to using the user of the device as the target sound source and record the user's position in the image. If the image displayed by the electronic device 100 does not include the user of the device, all people in the image can be used as target sound sources. It is understood that the electronic device 100 can determine the number and position of people in the image by detecting faces. The electronic device 100 can record the position of the face and use it as the position of the target sound source.

[0197] In some embodiments of this application, the electronic device 100 can compare the pre-recorded location of a user-selected target sound source with the location of a target sound source after... Figure 4 The location of the target sound source determined by the method shown is compared to the location of the target sound source, and the location of the target sound source is finally determined.

[0198] In some embodiments of this application, if the user selects himself as the target sound source, or if the electronic device 100 defaults to the user of the device as the target sound source, the electronic device 100 can use the stored voiceprint information of the user of the device to compare with the audio signal collected in step S401, and finally determine the location of the target sound source.

[0199] The following is combined Figure 5 This application provides a specific embodiment of a sound source separation method.

[0200] S501: Electronic device 100 enhances the audio signal located in the direction of the target sound source in audio signal X using beamforming (BF) method, and the enhanced audio signal is denoted as S. BF (i, k). Where i represents the frame number and k represents the discrete frequency point. The audio signal X is the audio signal in the frequency domain after processing the audio signal acquired by the microphone array.

[0201] The basic principle of BF (Browser-Focused) is as follows: For a microphone array, due to the different distribution positions of the microphones, there will be a certain time difference in the audio signals they collect. This can be used to determine the direction and location of the target sound source. By aligning the audio signals collected by each microphone, interference signals can be canceled, thereby enhancing the audio signal corresponding to the target sound source. It can be understood that alignment here refers to eliminating the time difference in audio signal collection between the microphones. For example, the relative time delay of the audio signals collected by each microphone can be compensated so that when the signals arrive at the microphone array, they can be equivalent to arriving at each microphone simultaneously on the same wavefront.

[0202] In the time domain, the audio signals collected by each microphone in the microphone array will have a time difference, which means that in the frequency domain, the corresponding audio signals will have a phase difference. It can be understood that the corresponding audio signals referred to here are the audio signals in the frequency domain obtained after processing (e.g., step S402) the audio signals in the time domain. These audio signals in the frequency domain constitute the audio signal X. The electronic device 100 can adjust the phase of these audio signals in the frequency domain so that the audio signals in some directions achieve constructive interference, while the audio signals in other directions achieve destructive interference, thereby enhancing the audio signals in some directions.

[0203] Understandably, according to the principle of wave superposition, if the crests (or troughs) of two waves arrive at the same point simultaneously, the two waves are said to be in phase at that point. At this time, the interfering waves will produce their maximum amplitude. This phenomenon is called constructive interference. Conversely, according to the principle of wave superposition, if the crest of one wave and the trough of another wave arrive at the same point simultaneously, the two waves are said to be out of phase at that point. At this time, the interfering waves will produce their minimum amplitude. This phenomenon is called destructive interference.

[0204] Specifically, the electronic device 100 can determine the direction vectors of the audio signals collected by different microphones based on the location of the target sound source, and then perform weighted coherent superposition of these audio signals based on these direction vectors and the weight vector of the beamformer to finally obtain the audio signal in the direction of the target sound source. This audio signal in the direction of the target sound source is the enhanced audio signal. Let it be denoted as S. BF (i, k). Where i represents the frame number and k represents the discrete frequency point.

[0205] S502: The electronic device 100 uses the acquired number of sound sources M as prior information, and performs sound source separation on the audio signal X using a deep learning method, denoting the separated audio signal as S. NN (n NN (i, k). Where n NN This refers to a single sound source after the sound source has been separated. NN ≤M. i represents the number of frames. k represents the discrete frequency points.

[0206] It is understandable that sound source separation of audio signal X using deep learning methods can include, but is not limited to, sound source separation methods based on deep neural networks (DNN) and sound source separation methods based on convolutional neural networks (CNN).

[0207] The following explanation uses a DNN-based sound source separation method as an example.

[0208] Electronic device 100 performs sound source separation using a DNN-based sound source separation method, primarily through a DNN sound source separation model. Electronic device 100 may include a DNN sound source separation model. Electronic device 100 can perform sound source separation on the audio signal X input to this model and output the separated audio signal S. NN (n NN (i, k). Where n NN This refers to a single sound source after the sound source has been separated. NN ≤M. i represents the number of frames. k represents the discrete frequency points.

[0209] The following is a brief introduction to DNN.

[0210] A DNN can be understood as a neural network with many hidden layers. The internal neural network of a DNN can be divided into three categories: input layer, hidden layer, and output layer. Generally, the first layer is the input layer, the last layer is the output layer, and the layers in between are hidden layers. All layers are fully connected. That is, any neuron in the i-th layer is connected to all neurons in the adjacent layers. The operation of each layer in a DNN can be understood as the following linear relationship expression: y = α(Wx + b). Where x is the input vector, y is the output vector, b is the offset vector, W is the weight matrix, and α() is the activation function. It can be understood that the input and output of each layer in a DNN are related to the input and output of the adjacent layers. That is, given the output of the i-th layer in a DNN, the output of the (i-1)-th layer and the output of the (i+1)-th layer can be calculated.

[0211] The following is a brief introduction to the DNN sound source separation model.

[0212] It is understandable that electronic device 100 can train DNN to obtain DNN sound source separation model.

[0213] Specifically, the electronic device 100 mixes audio signals from multiple sound sources in a certain proportion to obtain a mixed audio signal. The electronic device 100 can extract features from the mixed audio signal, and can also use the extracted audio features as input to a DNN, with the output of the DNN being the separated audio signal.

[0214] Understandably, the mixed audio signal serves as the input to the DNN. Ideally, the output of the DNN should be the original audio signal before mixing.

[0215] It should be noted that during the training of the DNN, the electronic device 100 needs to ensure that the DNN's output is equal to or very close to the ideal DNN output. Specifically, the electronic device 100 needs to find suitable offset vector b and weight matrix W so that the DNN's output is equal to or very close to the ideal DNN output.

[0216] To achieve the above objectives, the electronic device 100 can select a suitable loss function to measure the output loss during the training process and optimize the loss function to find its minimum extreme value. The corresponding series of offset vectors b and weight matrix W are the suitable offset vectors b and weight matrix W. It can be understood that the DNN corresponding to the suitable offset vectors b and weight matrix W is the DNN sound source separation model.

[0217] It is understandable that the features extracted by the electronic device 100 can be audio features at the frame level or time-frequency unit level.

[0218] It is understood that the audio signals used to train the DNN are training samples. The audio signals included in the training samples can be determined according to actual needs, and this application does not impose any restrictions on this.

[0219] It should be noted that the above content is only an example provided in this application, and DNN sound source separation models can also be obtained through other means, which are not limited in this application.

[0220] S503: Electronic device 100 performs sound source separation on audio signal X using the Blind Source Separation (BSS) method, and denotes the separated audio signal as S. BSS (n BSS (i,k). Where n BSS This refers to a single sound source after the sound source has been separated. BSS ≤M. i represents the number of frames. k represents the discrete frequency points.

[0221] The following is a brief introduction to BSS.

[0222] BSS (Breakpoint Separation Matrix) refers to the process of recovering the individual components of a source signal from the observed signal based solely on the statistical properties of the input source signal, without knowing the parameters of the source signal and the transmission channel. The ultimate goal of BSS is to find a separation matrix that makes the output signal approximate the true source signal as closely as possible.

[0223] Understandingly, Independent Component Analysis (ICA) is a solution for blind source separation. The basic idea of ​​ICA is to assume that the source signals generating the observed signals are independent of each other, and the goal of finding a separation matrix is ​​to make the components in the output result as independent as possible. Therefore, ICA can be viewed as an optimization problem. Its objective function is a function that measures the independence of the separation results, and the relationship is: ICA algorithm = objective function + optimization algorithm. The objective function can employ separation criteria based on independence measures (e.g., non-Gaussian maximization criterion, mutual information minimization criterion, information maximization criterion, and maximum likelihood criterion). The optimization algorithm can employ batch processing algorithms, adaptive algorithms, successive extraction methods, etc.

[0224] It is understandable that the audio signal collected by the microphone array is actually a mixture of audio signals from different sound sources. That is, the audio signal X is formed by mixing the audio signals corresponding to multiple sound sources. In other words, the observed signal mentioned above can be the audio signal X, and the source signal can be the audio signal corresponding to the sound source.

[0225] The following is a brief introduction to the BSS-based sound source separation method.

[0226] The basic idea of ​​using the BSS method to separate audio signal X is to find a suitable separation matrix that can separate audio signals from different sound sources in audio signal X, and make the separated audio signals as consistent as possible with the real source signals.

[0227] In some embodiments of this application, the electronic device 100 can use ICA to find a suitable separation matrix to achieve the best separation effect, and then use the separation matrix to perform sound source separation on the audio signal X to obtain the separated audio signal S. BSS (n BSS (i, k). Where n BSS This refers to a single sound source after the sound source has been separated. BSS ≤M. i represents the number of frames. k represents the discrete frequency points.

[0228] S504: Electronic device 100 determines a first correlation and a second correlation. The first correlation is S. BF (i, k) and S NN (n NN The correlation between (i, k) is shown. The second correlation is S. BF (i, k) and S BSS (n BSS The correlation of (i, k).

[0229] Electronic device 100 can determine the first correlation, i.e., S. BF (i, k) and S NN (n NN The correlation between (i, k). Let the first correlation be denoted as ρ. BF,NN (n NN ). ρ BF,NN (n NN The calculation method for ) is as follows:

[0230]

[0231] Among them, cov(S BF (i, k), S NN (n NN i, k)) is the audio signal S BF (i, k) and audio signal S NN (n NN The covariance of (i, k). abs(S) BF (i, k) represents the audio signal S. BF The absolute value of (i, k). abs(S NN (n NN i, k)) is the audio signal S NN (n NN The absolute value of (i, k).

[0232] Electronic device 100 can also determine a second correlation, namely S. BF (i, k) and S BSS (n BSS The correlation between (i, k). Let the first correlation be denoted as ρ. BF,BSS (n BSS ).

[0233]

[0234] Among them, cov(S BF (i, k), S BSS (n BSS i, k)) is the audio signal S BF (i, k) and audio signal S BSS (n BSS The covariance of (i, k). abs(S) BF (i, k) represents the audio signal S. BF The absolute value of (i, k). abs(S BSS (n BSS ,i,k)) is the audio signal SB SS (n BSS The absolute value of (i, k).

[0235] S505: Electronic device 100 determined S BF The audio signal with the smallest absolute value among (i, k), audio signal S, and audio signal T is the audio signal corresponding to the target sound source. Here, audio signal S is the audio signal with the highest first correlation among the audio signals obtained by the electronic device 100 through deep learning to separate audio signal X. Audio signal T is the audio signal with the highest second correlation among the audio signals obtained by the electronic device 100 through the BSS method to separate audio signal X.

[0236] Electronic device 100 can determine audio signal S. Audio signal S is the audio signal with the highest first correlation among the audio signals obtained by electronic device 100 through separation of audio signal X using a deep learning method. Audio signal S is as follows:

[0237] S NN (n1, i, k), ifρ BF,NN (n1)=max(ρ BF,NN (n NN )), n1≤M

[0238] Electronic device 100 can determine audio signal T. Audio signal T is the audio signal with the highest second correlation among the audio signals obtained by electronic device 100 through the BSS method of separating audio signal X. Audio signal T is expressed as follows:

[0239] S BSS (n2, i, k), ifρ BF,BSS (n2)=max(ρ BF,BSS (n BSS n2≤M

[0240] Electronic device 100 can determine S BF The audio signal with the smallest absolute value among (i, k), audio signal S, and audio signal T is the audio signal corresponding to the target sound source. The audio signal corresponding to the finally determined target sound source is denoted as S. target (i, k). Where i represents the frame number and k represents the discrete frequency point. S target The calculation method for (i, k) is as follows:

[0241]

[0242] Where Smin(i,k) represents abs(S BF (i, k)), abs(S) NN (n1, i, k)) and abs(S) BSS The minimum value among the three (n2, i, k) is Smin(i, k) = min(abs(S BF (i, k)), abs(S) NN (n1, i, k)), abs(S) BSS (n2, i, k))).

[0243] It is understood that this application does not restrict the order in which steps S501, S502 and S503 are executed.

[0244] Understandably, the above method combines the BF method, deep learning method and BSS method for sound source separation. The audio signal corresponding to the separated target sound source is closer to the original audio signal, which also makes the subsequent calculation of signal-to-noise ratio more accurate.

[0245] The following is combined Figure 6 This application provides another method for sound source separation.

[0246] S601: Electronic device 100 enhances the audio signal located in the direction of the target sound source in audio signal X using the BF method, and the enhanced audio signal is denoted as S. BF (i, k). Where i represents the frame number and k represents the discrete frequency point. The audio signal X is the audio signal in the frequency domain after processing the audio signal acquired by the microphone array.

[0247] Specifically, the electronic device 100 can determine the direction vectors of the audio signals collected by different microphones based on the location of the target sound source, and then perform weighted coherent superposition of these audio signals based on these direction vectors and the weight vector of the beamformer to finally obtain the audio signal in the direction of the target sound source. This audio signal in the direction of the target sound source is the enhanced audio signal. Let it be denoted as S. BF (i, k). Where i represents the frame number and k represents the discrete frequency point.

[0248] It is understood that the relevant description of BF can be found in step S501, and this application will not repeat it here.

[0249] S602: The electronic device 100 uses the acquired number of sound sources M as prior information, and performs sound source separation on the audio signal X using a deep learning method, denoting the separated audio signal as S. NN (n NN (i, k). Where n NN This refers to a single sound source after the sound source has been separated. NN ≤M. i represents the number of frames. k represents the discrete frequency points.

[0250] It is understandable that the specific implementation of sound source separation of audio signal X using deep learning methods can be found in step S502, and will not be repeated here.

[0251] S603: Electronic device 100 determines the first correlation. The first correlation is S. BF (i, k) and S NN (n NN The correlation of (i, k).

[0252] Electronic device 100 can determine S BF (i, k) and S NN (n NN The correlation between (i, k). This is understandable; for details, please refer to step S504, which will not be elaborated here.

[0253] S604: Electronic device 100 determined S BF The audio signal with the smallest absolute value among (i, k) and audio signal S is the audio signal corresponding to the target sound source. Here, audio signal S is the audio signal with the highest first correlation among the audio signals obtained by the electronic device 100 through deep learning to separate audio signal X.

[0254] It is understandable that the relevant description of the audio signal S can be found in step S505, and will not be repeated here.

[0255] The electronic device 100 can record the audio signal corresponding to the finally determined target sound source as S. target(i, k). Where i represents the frame number and k represents the discrete frequency point. S target The calculation method for (i, k) is as follows:

[0256]

[0257] Where Smin(i,k) represents abs(S BF (i, k)) and abs(S) NN The minimum value in (n1, i, k)). That is, Smin(i, k) = min(abs(S BF (i, k)), abs(S) NN (n1, i, k))).

[0258] It is understood that this application does not restrict the order in which steps S601 and S602 are executed.

[0259] The following is combined Figure 7 This application provides another method for sound source separation.

[0260] S701: Electronic device 100 enhances the audio signal located in the direction of the target sound source in audio signal X using the BF method, and the enhanced audio signal is denoted as S. BF (i, k). Where i represents the frame number and k represents the discrete frequency point. The audio signal X is the audio signal in the frequency domain after processing the audio signal acquired by the microphone array.

[0261] Specifically, the electronic device 100 can determine the direction vectors of the audio signals collected by different microphones based on the location of the target sound source, and then perform weighted coherent superposition of these audio signals based on these direction vectors and the weight vector of the beamformer to finally obtain the audio signal in the direction of the target sound source. This audio signal in the direction of the target sound source is the enhanced audio signal. Let it be denoted as S. BF (i, k). Where i represents the frame number and k represents the discrete frequency point.

[0262] It is understood that the relevant description of BF can be found in step S501, and this application will not repeat it here.

[0263] S702: Electronic device 100 performs source separation on audio signal X using the BSS method, and the separated audio signal is denoted as S. BSS (n BSS (i, k). Where n BSS This refers to a single sound source after the sound source has been separated. BSS ≤M. i represents the number of frames. k represents the discrete frequency points.

[0264] It is understandable that the specific implementation of the BSS method for source separation of audio signal X can be referred to step S503, and will not be repeated here.

[0265] S703: Electronic device 100 determines the second correlation. The second correlation is S. BF (i, k) and S BSS (n BSS The correlation of (i, k).

[0266] Electronic device 100 can determine S BF (i, k) and S BSS (n BSS The correlation between (i, k). This is understandable; for details, please refer to step S504, which will not be elaborated here.

[0267] S704: Electronic device 100 confirmed S BF The audio signal with the smallest absolute value among (i, k) and audio signal T is the audio signal corresponding to the target sound source. Here, audio signal T is the audio signal with the highest second correlation among the audio signals obtained by electronic device 100 through the BSS method separating audio signal X.

[0268] It is understandable that the relevant description of the audio signal T can be found in step S505, and will not be repeated here.

[0269] The electronic device 100 can record the audio signal corresponding to the finally determined target sound source as S. target (i, k). Where i represents the frame number and k represents the discrete frequency point. S target The calculation method for (i, k) is as follows:

[0270]

[0271] Where Smin(i,k) represents abs(S BF (i, k)) and abs(S) BSS The minimum value in (n2, i, k)). That is, Smin(i, k) = min(abs(S BF (i, k)), abs(S) BSS (n2, i, k))).

[0272] It is understood that this application does not restrict the order in which steps S701 and S702 are executed.

[0273] The following is combined Figure 8 This application provides a specific recording processing method based on an embodiment.

[0274] S801: Electronic device 100 acquires audio signals through a microphone array.

[0275] It is understood that the electronic device 100 can acquire audio signals through the microphone array included in the device. It is understood that the specific content of step S801 can be referred to step S401, and will not be repeated here.

[0276] S802: Electronic device 100 performs sound source localization based on the acquired audio signal, that is, determines the location and number of target sound sources based on the acquired audio signal.

[0277] Specifically, the electronic device 100 can process the acquired audio signal, converting it into an audio signal in the frequency domain for subsequent operations. The electronic device 100 can use a sound source localization method based on SRP-PHAT to determine the location and number of target sound sources. Furthermore, in scenarios such as video recording, the electronic device 100 can further determine the location of the target sound source by combining images and user operations. It is understood that the specific content of step S802 can be referred to in steps S402 and S403, and will not be repeated here.

[0278] S803: Electronic device 100 performs sound source separation on the acquired audio signal based on the location and number of the target sound source to obtain the audio signal of the target sound source.

[0279] Specifically, the electronic device 100 can perform sound source separation on the acquired audio signal using the BF method, deep learning method, and BSS method, based on the location and number of the target sound source. The electronic device 100 can then determine the audio signal S of the target sound source based on the results obtained from the sound source separation using these three methods. target (i, k). This is understandable; for details of step S803, please refer to [link / reference]. Figure 5 , Figure 6 and Figure 7 The embodiments shown will not be described in detail here.

[0280] S804: Electronic device 100 determines the interference signal based on the audio signal of the target sound source.

[0281] Based on the above, the audio signal of the m-th microphone in the i-th frame is denoted as: x m (i, k). The corresponding audio signal in the frequency domain is: X m (i, k).

[0282] It is understandable that the electronic device 100 can determine the audio signal S from the target sound source. target (i, k), X m (i, k) is used to determine the interference signal. The interference signal is denoted as S. other (i, k). Then we have: S other(i, k) = X m (i, k)-S target (i, k).

[0283] S805: Electronic device 100 determines the signal-to-noise ratio based on the audio signal and interference signal of the target sound source.

[0284] Electronic device 100 can determine the signal-to-noise ratio (SNR) based on the audio signal and interference signal of the target sound source. The SNR is denoted as SNR(i), where i represents the frame number. The SNR(i) is calculated as follows:

[0285]

[0286] Among them, RMS target (i) represents S target The root mean square of (i, k). RMS other (i) represents S other The root mean square of (i, k).

[0287] It should be noted that the electronic device 100 can display the signal-to-noise ratio on the screen to remind the user to pay attention to the current sound pickup quality.

[0288] In some embodiments of this application, the electronic device 100 may display on a screen the quality level to which the sound pickup quality in the current recording environment belongs.

[0289] For example, if the signal-to-noise ratio (SNR) is greater than a first threshold, the electronic device 100 determines the quality level of the sound pickup in the current recording environment as a first quality level. If the SNR is less than a second threshold, the electronic device 100 determines the quality level of the sound pickup in the current recording environment as a second quality level. If the SNR is not less than the second threshold and not greater than the first threshold, the electronic device 100 determines the quality level of the sound pickup in the current recording environment as a third quality level. It can be understood that a first quality level can indicate good sound pickup quality in the current recording environment. A second quality level can indicate poor sound pickup quality in the current recording environment. A third quality level can indicate average sound pickup quality in the current recording environment.

[0290] It is understood that the first and second thresholds can be set according to actual needs, and this application does not impose any restrictions on them.

[0291] For example, the first threshold can be set to 20dB and the second threshold can be set to 0dB.

[0292] It should be noted that the electronic device mentioned in the claims can be the electronic device 100 in the embodiments of this application.

[0293] In some embodiments of this application, the first audio signal may be the audio signal X in the foregoing embodiments. The second audio signal may be S in the foregoing embodiments. target (i, k).

[0294] In some embodiments of this application, the third audio signal may be the S signal in the foregoing embodiments. BF (i, k).

[0295] In some embodiments of this application, the first set of separated audio signals can be S in the foregoing embodiments. NN (n NN (i, k).

[0296] In some embodiments of this application, the second set of separated audio signals can be S in the foregoing embodiments. BSS (n BSS (i, k). In some embodiments of this application, the fourth audio signal may be the audio signal S in the foregoing embodiments.

[0297] In some embodiments of this application, the fifth audio signal may be the audio signal T in the foregoing embodiments.

[0298] In some embodiments of this application, the preset threshold mentioned in the claims may include the first threshold and the second threshold in the foregoing embodiments.

[0299] Additionally, it should be noted that the audio signal corresponding to the sound source mentioned in this application has the same meaning as the audio signal generated by the sound source and the audio signal from the sound source.

[0300] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit it. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.

Claims

1. A recording processing method, characterized in that, Applied to electronic devices, the method includes: Receive recording start command; In response to the recording start command, a first audio signal is acquired; the first audio signal includes an audio signal generated by at least one sound source. The first audio signal is processed to obtain the location of the target sound source; The audio signal with the smallest absolute value among the third audio signal, the fourth audio signal, and / or the fifth audio signal is determined as the second audio signal; the second audio signal is the audio signal generated by the target sound source; During the recording process, prompt information is output based on the first audio signal and the second audio signal; the prompt information is used to characterize the sound pickup quality in the current recording environment; The third audio signal is an audio signal obtained by enhancing the audio signal located in the direction of the target sound source based on the location of the target sound source; the fourth audio signal is the audio signal with the highest correlation to the third audio signal in the first set of separated audio signals, which is an audio signal set obtained by separating audio signals generated by different sound sources from the first audio signal using a first separation method based on the number of sound sources in the current recording environment; the fifth audio signal is the audio signal with the highest correlation to the third audio signal in the second set of separated audio signals, which is an audio signal set obtained by separating audio signals generated by different sound sources from the first audio signal using a second separation method; the pickup quality is determined based on the signal-to-noise ratio in the current recording environment, which is obtained based on the first audio signal and the second audio signal.

2. The method as described in claim 1, characterized in that, When determining the audio signal with the smallest absolute value among the third, fourth, and fifth audio signals as the second audio signal, the method further includes the following steps before determining the audio signal with the smallest absolute value among the third, fourth, and fifth audio signals as the second audio signal: Based on the location of the target sound source, the audio signal located in the direction of the target sound source is amplified to obtain the third audio signal; Based on the number of sound sources in the current recording environment, the first separation method is used to separate audio signals generated by different sound sources from the first audio signal to obtain the first separated audio signal set; The second separation method is used to separate audio signals generated by different sound sources from the first audio signal to obtain the second separated audio signal set; The fourth audio signal is determined based on the third audio signal and the first set of separated audio signals; The fifth audio signal is determined based on the third audio signal and the second set of separated audio signals.

3. The method as described in claim 1, characterized in that, When determining the audio signal with the smallest absolute value among the third and fourth audio signals as the second audio signal, the method further includes, before determining the audio signal with the smallest absolute value among the third and fourth audio signals as the second audio signal: Based on the location of the target sound source, the audio signal located in the direction of the target sound source is amplified to obtain the third audio signal; Based on the number of sound sources in the current recording environment, the first separation method is used to separate audio signals generated by different sound sources from the first audio signal to obtain the first separated audio signal set; The fourth audio signal is determined based on the third audio signal and the first set of separated audio signals.

4. The method as described in claim 1, characterized in that, When determining the audio signal with the smallest absolute value among the third and fifth audio signals as the second audio signal, the method further includes, before determining the audio signal with the smallest absolute value among the third and fifth audio signals as the second audio signal: Based on the location of the target sound source, the audio signal located in the direction of the target sound source is amplified to obtain the third audio signal; The second separation method is used to separate audio signals generated by different sound sources from the first audio signal to obtain the second separated audio signal set; The fifth audio signal is determined based on the third audio signal and the second set of separated audio signals.

5. The method according to any one of claims 1-4, characterized in that, Based on the first audio signal and the second audio signal, a prompt message is output, specifically including: The signal-to-noise ratio in the current recording environment is obtained based on the first audio signal and the second audio signal; By comparing the signal-to-noise ratio with a preset threshold, the quality level of the sound pickup in the current recording environment is determined. Displays the quality level of the sound pickup in the current recording environment; Different quality levels represent different pickup qualities.

6. The method according to any one of claims 1-4, characterized in that, Based on the first audio signal and the second audio signal, output prompt information, including: The signal-to-noise ratio in the current recording environment is obtained based on the first audio signal and the second audio signal; The signal-to-noise ratio is displayed.

7. An electronic device comprising one or more memories and one or more processors, characterized in that, The one or more processors are coupled to the one or more memories, the one or more memories being used to store computer program code, the computer program code including computer instructions; The processor is used to receive the recording start command; The processor is further configured to, in response to the recording start command, acquire a first audio signal; the first audio signal includes an audio signal generated by at least one sound source; The processor is further configured to process the first audio signal to obtain the location of the target sound source; The processor is further configured to determine the audio signal with the smallest absolute value among the third audio signal, the fourth audio signal, and / or the fifth audio signal as the second audio signal; the second audio signal is the audio signal generated by the target sound source; The processor is further configured to output prompt information based on the first audio signal and the second audio signal during the recording process; the prompt information is used to characterize the sound pickup quality in the current recording environment. The third audio signal is an audio signal obtained by enhancing the audio signal located in the direction of the target sound source based on the location of the target sound source; the fourth audio signal is the audio signal with the highest correlation to the third audio signal in the first set of separated audio signals, which is an audio signal set obtained by separating audio signals generated by different sound sources from the first audio signal using a first separation method based on the number of sound sources in the current recording environment; the fifth audio signal is the audio signal with the highest correlation to the third audio signal in the second set of separated audio signals, which is an audio signal set obtained by separating audio signals generated by different sound sources from the first audio signal using a second separation method; the pickup quality is determined based on the signal-to-noise ratio in the current recording environment, which is obtained based on the first audio signal and the second audio signal.

8. The electronic device as claimed in claim 7, characterized in that, The processor, when determining the audio signal with the smallest absolute value among the third, fourth, and fifth audio signals as the second audio signal, further configures itself, before determining the audio signal with the smallest absolute value among the third, fourth, and fifth audio signals as the second audio signal: Based on the location of the target sound source, the audio signal located in the direction of the target sound source is amplified to obtain the third audio signal; Based on the number of sound sources in the current recording environment, the first separation method is used to separate audio signals generated by different sound sources from the first audio signal to obtain the first separated audio signal set; The second separation method is used to separate audio signals generated by different sound sources from the first audio signal to obtain the second separated audio signal set; The fourth audio signal is determined based on the third audio signal and the first set of separated audio signals; The fifth audio signal is determined based on the third audio signal and the second set of separated audio signals.

9. The electronic device as claimed in claim 7, characterized in that, The processor, when determining the audio signal with the smallest absolute value among the third and fourth audio signals as the second audio signal, further configures itself, before determining the audio signal with the smallest absolute value among the third and fourth audio signals as the second audio signal: Based on the location of the target sound source, the audio signal located in the direction of the target sound source is amplified to obtain the third audio signal; Based on the number of sound sources in the current recording environment, the first separation method is used to separate audio signals generated by different sound sources from the first audio signal to obtain the first separated audio signal set; The fourth audio signal is determined based on the third audio signal and the first set of separated audio signals.

10. The electronic device as claimed in claim 7, characterized in that, The processor, when determining the audio signal with the smallest absolute value among the third and fifth audio signals as the second audio signal, further configures itself, before determining the audio signal with the smallest absolute value among the third and fifth audio signals as the second audio signal: Based on the location of the target sound source, the audio signal located in the direction of the target sound source is amplified to obtain the third audio signal; The second separation method is used to separate audio signals generated by different sound sources from the first audio signal to obtain the second separated audio signal set; The fifth audio signal is determined based on the third audio signal and the second set of separated audio signals.

11. The electronic device according to any one of claims 7-10, characterized in that, When the processor is used to output prompt information based on the first audio signal and the second audio signal, it is specifically used for: The signal-to-noise ratio in the current recording environment is obtained based on the first audio signal and the second audio signal; By comparing the signal-to-noise ratio with a preset threshold, the quality level of the sound pickup in the current recording environment is determined. The electronic device further includes a display screen; the display screen is used to display the quality level of the sound pickup in the current recording environment; Different quality levels represent different pickup qualities.

12. The electronic device according to any one of claims 7-10, characterized in that, When the processor is used to output prompt information based on the first audio signal and the second audio signal, it is specifically used for: The signal-to-noise ratio in the current recording environment is obtained based on the first audio signal and the second audio signal; The electronic device further includes a display screen for displaying the signal-to-noise ratio.

13. A computer storage medium, characterized in that, include: Computer instructions; when the computer instructions are executed on an electronic device, causing the electronic device to perform the method of any one of claims 1-6.