Method and electronic device for audio enhancement

By displaying audio enhancement controls and performing audio enhancement processing when switching focus, the problem of poor sound pickup during long-distance shooting was solved, improving video and audio quality and user experience.

CN122437997APending Publication Date: 2026-07-21HUAWEI DEVICE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HUAWEI DEVICE CO LTD
Filing Date
2025-11-14
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

When shooting from a distance, electronic devices have poor sound pickup performance at long distances, resulting in weak and unclear voices in the recorded video, which affects the viewing experience.

Method used

By displaying audio enhancement controls during focus switching and performing audio enhancement processing while recording video, including enhancing target sounds and suppressing non-target sounds, the timing and direction of audio enhancement are determined by combining speech and visual detection.

Benefits of technology

It improves the audio quality of long-distance video recording, enhances target sound, reduces ambient noise, and improves user experience and the usage rate of audio enhancement features.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122437997A_ABST
    Figure CN122437997A_ABST
Patent Text Reader

Abstract

The application provides a method and an electronic device for audio enhancement. The method is applied to a first device, and the method comprises: displaying a first interface, the first interface being an interface for recording a first video; in response to detecting an operation of switching a focal length from a first focal length to a second focal length by a user, performing audio enhancement processing on sound corresponding to the first video while recording the first video, wherein the second focal length is greater than the first focal length. Through the method and the electronic device, the electronic device can perform enhancement processing on the recorded sound while shooting a long-focus video, which helps to solve the problem of poor long-distance sound pickup effect of the electronic device and can improve the sound effect of the recorded long-focus video.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of electronic devices, and more specifically, to a method and electronic device for audio enhancement. Background Technology

[0002] Currently, most electronic devices with camera functions support high-magnification optical zoom and digital zoom, allowing users to capture distant subjects as needed. However, in scenarios where users record video from a distance, long-distance audio pickup remains a weakness in the industry.

[0003] Typically, the microphones built into electronic devices can accurately pick up sound in close-range scenarios. However, when shooting distant subjects, the sound signal of the subject is significantly attenuated, while environmental noise is amplified. This often results in the recorded video having weak or unclear voices, severely impacting the final viewing experience of the audio recording.

[0004] Therefore, how to improve the long-distance sound pickup effect of electronic devices has become an urgent problem to be solved. Summary of the Invention

[0005] This application provides an audio enhancement method and an electronic device. Through this method and electronic device, the electronic device can enhance the recorded sound while shooting telephoto video, helping to solve the problem of poor sound pickup at long distances and improving the sound quality of the recorded telephoto video.

[0006] In a first aspect, an audio enhancement method is provided, which is applied to a first device. The method includes: displaying a first interface, the first interface being an interface for recording a first video; in response to detecting an operation by a user switching the focal length from a first focal length to a second focal length, performing audio enhancement processing on the sound corresponding to the first video while recording the first video, wherein the second focal length is greater than the first focal length.

[0007] In some embodiments, the camera application of the first device can be opened and switched to recording mode. In recording mode, the first device can be controlled to start recording a first video. At this time, the first device displays a first interface.

[0008] In some embodiments, the operation of a user switching the focal length from a first focal length to a second focal length can be, for example, the operation of a user switching the zoom level of a first video recording from a first zoom level to a second zoom level, wherein the first zoom level is less than the second zoom level. For example, the first zoom level is less than a zoom level threshold, and the second zoom level is greater than a zoom level threshold.

[0009] In one example, the zoom threshold could be, for instance, 5x zoom.

[0010] In some embodiments, while enhancing the audio of the first video, the first device may also retain the original audio corresponding to the first video, i.e., the audio that has not undergone audio enhancement processing.

[0011] The original audio format can be stereo, spatial audio, or other formats.

[0012] In this embodiment of the application, the electronic device can enhance the recorded sound while shooting telephoto video, which helps to solve the problem of poor sound pickup effect of electronic devices at long distances and can improve the sound effect corresponding to the recorded telephoto video.

[0013] In conjunction with the first aspect, in one possible implementation, in response to detecting that the user switches the focal length from a first focal length to a second focal length, audio enhancement processing is performed on the audio corresponding to the first video while recording the first video, including: in response to detecting that the user switches the focal length from the first focal length to the second focal length, displaying a first control on the first interface, the first control being used by the user to enable audio enhancement processing on the audio corresponding to the first video; and in response to the user's operation on the first control, performing audio enhancement processing on the audio corresponding to the first video while recording the first video.

[0014] The user's action on the first control can be, for example, clicking the first control.

[0015] In this embodiment, the electronic device can display a control on the video recording interface for the user to enable the audio enhancement function while recording telephoto video. This makes it easier for the user to perceive that the electronic device has an audio enhancement function, thereby increasing the usage rate of the audio enhancement function. Guiding the user to use the audio enhancement function of the electronic device at an appropriate time helps to solve the problem of poor sound pickup at long distances, and can improve the sound quality of the recorded telephoto video. Furthermore, the user can choose whether to enable the audio enhancement function, which can further enhance the user experience of using the electronic device.

[0016] In conjunction with the first aspect, in one possible implementation, the first control is also used to prompt the user to enable audio enhancement for the sound corresponding to the first video.

[0017] In some embodiments, the first control may provide prompts through icons, text, or other means. Text prompts may include, for example, "Super Hearing," or "Enable Super Hearing for clearer target sounds," and so on.

[0018] In conjunction with the first aspect, in one possible implementation, in response to detecting that the user has switched the focus from a first focal length to a second focal length, displaying a first control on the first interface includes: in response to detecting that the user has switched the focus from the first focal length to the second focal length, determining whether there is human voice in the audio corresponding to the first video; and if it is determined that there is human voice in the audio corresponding to the first video, displaying the first control on the first interface.

[0019] In some embodiments, the first device may include a voice detection module, which can detect human voices.

[0020] In this embodiment, during the recording of telephoto video, the electronic device only displays a control for enabling the audio enhancement function on the interface after further detecting that the recorded sound includes human voices. This makes the timing of displaying the control for enabling the audio enhancement function more accurate, reduces unnecessary display of the control on the interface, and avoids affecting the user's video recording experience.

[0021] In conjunction with the first aspect, in one possible implementation, in response to detecting that the user has switched the focus from a first focal length to a second focal length, displaying a first control on the first interface includes: in response to detecting that the user has switched the focus from the first focal length to the second focal length, determining whether there is a subject emitting sound in the frame of the first video; and if it is determined that there is a subject emitting sound in the frame of the first video, displaying the first control on the first interface.

[0022] In some embodiments, the first device may include a sound detection module and / or a visual detection module, which can detect whether there is a subject emitting sound in the frame of the first video.

[0023] The entity that makes the sound can include, for example, people, animals, plants, etc.

[0024] In this embodiment, during the recording of telephoto video, the electronic device only displays the control for enabling the audio enhancement function on the interface after further detecting the presence of a sound-emitting entity in the captured image. This makes the display timing of the audio enhancement function control more accurate, reduces unnecessary display of the control on the interface, and avoids affecting the user's video recording experience.

[0025] In conjunction with the first aspect, in one possible implementation, the first focal length is less than a focal length threshold, and the second focal length is greater than or equal to the focal length threshold.

[0026] In this embodiment, the first device has a preset focal length threshold. When the focal length is greater than the focal length threshold during the video recording process, the first device provides audio enhancement services to the user. This can effectively solve the problem of poor sound pickup effect of electronic devices at long distances and improve the sound effect of the recorded video.

[0027] In conjunction with the first aspect, in one possible implementation, the method further includes: displaying a second interface after the first video recording ends, the second interface being a playback interface for the first video, the second interface including a second control, the second control being in an on state, used to indicate that the current state of the first video is an audio-enhanced state.

[0028] In some embodiments, after the first video recording is completed, the first device saves the first video to the album. The user can view the first video through the album application of the first device, that is, the user can enter the second interface by using the album application of the first device.

[0029] It is understandable that when the first video is currently in an audio-enhanced state, the sound emitted by the first device from the first video is the audio-enhanced sound.

[0030] In this embodiment of the application, a control can be displayed on the video preview interface to indicate whether the currently previewed video has undergone audio enhancement processing, which can improve the user's perception of the application of audio enhancement function.

[0031] In conjunction with the first aspect, in one possible implementation, the second interface also includes a third control for the user to adjust the audio enhancement level corresponding to the first video.

[0032] In this embodiment of the application, when previewing a video, the user can preview videos with different levels of audio enhancement, and the user can keep the video with the most desired level of audio enhancement.

[0033] In conjunction with the first aspect, in one possible implementation, the method further includes: in response to detecting a user clicking the second control, the state of the second control is switched from an on state to a off state, for indicating that the current state of the first video is switched from an audio-enhanced state to an unenhanced state.

[0034] In other words, in response to detecting a user clicking the second control, the first device cancels the audio enhancement effect on the first video. This allows the user to experience both the enhanced and original audio versions of the first video through the second control, enabling them to compare the two sound effects and retain the desired one based on their needs.

[0035] In other words, when the second control is "on", the first video played by the user is a video with enhanced audio; when the second control is "off", the first video played by the user is the original video without audio enhancement.

[0036] It is understandable that when the first video is in a state without audio enhancement, the sound emitted by the first device from the first video is the sound without audio enhancement when the user plays the first video.

[0037] In this embodiment of the application, after the video recording is completed, the video preview interface displays a switch control for audio enhancement effects. Users can preview the original audio effect and the audio enhancement effect through this switch control, which allows for convenient comparison of the two audio enhancement effects. This helps users quickly and clearly perceive the effect after audio processing and also helps to improve users' recognition of the product.

[0038] In conjunction with the first aspect, in one possible implementation, audio enhancement processing is performed on the sound corresponding to the first video, including: enhancing the sound emitted by a first subject in the frame of the first video; and suppressing sounds other than the sound emitted by the first subject.

[0039] In conjunction with the first aspect, in one possible implementation, the sound emitted by the first subject in the first video frame is enhanced, including enhancing the sound from the location of the first subject.

[0040] In some embodiments, the first device may detect the location of the first subject through voice detection and / or visual detection.

[0041] In this embodiment of the application, the first device can enhance the sound from the main direction and suppress the sound from other directions, thereby further improving the quality of the effective sound in the recorded video.

[0042] In conjunction with the first aspect, in one possible implementation, the method further includes: updating the location of the first subject through speech detection and / or visual detection; wherein, enhancing the sound emitted by the first subject in the frame of the first video further includes: enhancing the sound from the updated location of the first subject.

[0043] In this embodiment, the location of the sound subject can be updated in a timely manner, so that the first device can also adjust the direction of audio enhancement in a timely manner, so that even if the location of the sound subject changes, the stability of the audio enhancement effect will not be affected.

[0044] Secondly, an audio enhancement method is provided, applied to a first device. The method includes: displaying a third interface, which is a playback interface for a second video, wherein the recording focal length of the second video is greater than a focal length threshold, and the second video has not undergone audio enhancement processing during recording; the third interface includes a fourth control, which is in a closed state to indicate that the current state of the second video is that it has not undergone audio enhancement; in response to detecting a user clicking the fourth control, performing audio enhancement processing on the sound corresponding to the second video, and switching the state of the fourth control to an open state.

[0045] In this embodiment of the application, a function switch for enhancing the audio of the video is provided to the user in the video preview interface. This can guide the user to use the audio enhancement function of the first device to enhance the audio of the recorded video, thereby improving the quality of the video and the user's device experience.

[0046] In conjunction with the second aspect, in one possible implementation, audio enhancement processing is performed on the sound corresponding to the second video, including: enhancing the sound emitted by the second subject in the frame of the second video; and suppressing sounds other than the sound emitted by the second subject.

[0047] In this embodiment of the application, the first device can enhance the sound from the main direction and suppress the sound from other directions, thereby further improving the quality of the effective sound in the recorded video.

[0048] Thirdly, an electronic device is provided, comprising a memory and a processor, wherein the memory is used to store computer program code, and the processor is used to execute the computer program code stored in the memory to implement the method in the first aspect or any possible implementation thereof, or to implement the method in the second aspect or any possible implementation thereof.

[0049] Fourthly, a computer-readable storage medium is provided, which stores a computer program or instructions that, when executed by a processor, implement the method of the first aspect or any possible implementation thereof, or implement the method of the second aspect or any possible implementation thereof.

[0050] Fifthly, a chip is provided, including a circuit for performing the method in the first aspect or any possible implementation thereof, or performing the method in the second aspect or any possible implementation thereof.

[0051] In a sixth aspect, a computer program product is provided, which stores a computer program or instructions that, when executed by a processor, implement the method in the first aspect or any possible implementation thereof, or implement the method in the second aspect or any possible implementation thereof. Attached Figure Description

[0052] Figure 1 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application; Figure 2 This is a software structure block diagram of the electronic device provided in the embodiments of this application; Figure 3 This is a schematic diagram of a user interface for enabling audio enhancement function provided in an embodiment of this application; Figure 4 This is a schematic diagram of a user interface using audio enhancement functionality provided in an embodiment of this application; Figure 5 This is a schematic diagram of another user interface using audio enhancement functionality provided in an embodiment of this application; Figure 6 This is a schematic diagram of another user interface using audio enhancement functionality provided in an embodiment of this application; Figure 7 This is a schematic diagram of another user interface using audio enhancement functionality provided in an embodiment of this application; Figure 8 This is another schematic diagram of audio enhancement provided in the embodiments of this application; Figure 9 This is a schematic diagram of another user interface for enabling audio enhancement function provided in the embodiments of this application; Figure 10 These are two more user interface diagrams illustrating the use of audio enhancement functions provided in the embodiments of this application; Figure 11 This is a schematic flowchart illustrating an audio enhancement method provided in an embodiment of this application; Figure 12 This is a schematic flowchart illustrating another audio enhancement method provided in the embodiments of this application. Detailed Implementation

[0053] The technical solutions of this application will now be described with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this application, and not all embodiments.

[0054] The technical solutions of the embodiments of this application will be described below with reference to the accompanying drawings. In the description of the embodiments of this application, unless otherwise stated, " / " means "or," for example, A / B can mean A or B; "and / or" in this text is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Furthermore, in the description of the embodiments of this application, "plural" or "multiple" refers to two or more than two.

[0055] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this embodiment, unless otherwise stated, "a plurality of" means two or more.

[0056] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "one embodiment," "some embodiments," "another embodiment," "other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0057] For example, Figure 1A schematic diagram of the structure of electronic device 100 is shown. Electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, antenna 1, antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, a sensor module 180, buttons 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an accelerometer sensor 180E, a distance sensor 180F, a proximity sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.

[0058] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0059] In some embodiments, the electronic device 100 may include at least one of the following: mobile phone, foldable electronic device, tablet computer, desktop computer, laptop computer, handheld computer, notebook computer, super mobile personal computer, netbook, cellular phone, personal digital assistant (PDA), augmented reality (AR) device, virtual reality (VR) device, artificial intelligence device, wearable device, in-vehicle device, smart home device, and smart city device. This application embodiment does not impose any special limitation on the type of electronic device 100.

[0060] Processor 110 may include one or more processing units, such as: application processor (AP), modem processor, graphics processing unit (GPU), coprocessor, image signal processor (ISP), controller, memory, video codec, digital signal processor (DSP), baseband processor, and / or neural network processing unit (NPU), etc. Different processing units may be independent devices or integrated into one or more processors.

[0061] The controller can be the nerve center and command center of the electronic device 100. The controller can generate operation control signals according to the instruction opcode and timing signals to complete the control of fetching and executing instructions.

[0062] The processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store instructions or data that the processor 110 has just used or that are used repeatedly. If the processor 110 needs to use the instruction or data again, it can retrieve it directly from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.

[0063] In some embodiments, the processor 110 may include one or more interfaces. Interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a SIM card interface, and / or a universal serial bus (USB) interface, etc.

[0064] It is understood that the interface connection relationships between the modules illustrated in the embodiments of this application are merely illustrative and do not constitute a structural limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may also employ different interface connection methods or combinations of multiple interface connection methods as described in the above embodiments.

[0065] The charging management module 140 receives charging input from a charger. The charger can be a wireless charger or a wired charger. In some wired charging embodiments, the charging management module 140 receives charging input from the wired charger via the USB interface 130. In some wireless charging embodiments, the charging management module 140 receives wireless charging input via the wireless charging coil of the electronic device 100. While charging the battery 142, the charging management module 140 can also supply power to the electronic device via the power management module 141.

[0066] The power management module 141 connects the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140, providing power to the processor 110, internal memory 121, external memory, display screen 194, camera 193, and wireless communication module 160, etc. The power management module 141 can also monitor parameters such as battery capacity, battery cycle count, and battery health status (leakage current, impedance). In some other embodiments, the power management module 141 may also be located within the processor 110. In other embodiments, the power management module 141 and the charging management module 140 may be located in the same device.

[0067] The wireless communication function of electronic device 100 can be realized through antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modem processor and baseband processor, etc.

[0068] The mobile communication module 150 can provide wireless communication solutions, including 2G / 3G / 4G / 5G, for use on the electronic device 100. In some embodiments, at least some functional modules of the mobile communication module 150 may be housed in the processor 110. In some embodiments, at least some functional modules of the mobile communication module 150 and at least some modules of the processor 110 may be housed in the same device.

[0069] The modem processor may include a modulator and a demodulator. The modulator modulates the low-frequency baseband signal to be transmitted into a mid-to-high frequency signal. The demodulator demodulates the received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After processing by the baseband processor, the low-frequency baseband signal is transmitted to the application processor. The application processor outputs sound signals through audio devices (not limited to speaker 170A, receiver 170B, etc.) or displays images or videos through the display screen 194. In some embodiments, the modem processor may be a separate device. In other embodiments, the modem processor may be independent of the processor 110 and may be housed in the same device as the mobile communication module 150 or other functional modules.

[0070] The wireless communication module 160 can provide solutions for wireless communication applications on the electronic device 100, including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared (IR) technology, etc.

[0071] In some embodiments, antenna 1 of electronic device 100 is coupled to mobile communication module 150, and antenna 2 is coupled to wireless communication module 160, so that electronic device 100 can communicate with networks and other devices through wireless communication technology.

[0072] Electronic device 100 implements display functions through a GPU, a display screen 194, and an application processor. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations and for graphics rendering. Processor 110 may include one or more GPUs, which execute program instructions to generate or modify display information.

[0073] Display screen 194 is used to display images, videos, etc. Display screen 194 includes a display panel. The display panel may be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a MiniLED, a MicroLED, a Micro-OLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, electronic device 100 may include one or N displays 194, where N is a positive integer greater than 1.

[0074] Electronic device 100 can perform shooting functions through ISP, camera 193, video codec, GPU, display 194 and application processor.

[0075] The ISP (Image Signal Processor) is used to process data fed back from the camera 193. For example, when taking a picture, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, and the camera's photosensitive element transmits the electrical signal to the ISP for processing, transforming it into an image visible to the naked eye. The ISP can also perform algorithmic optimization of image noise, brightness, and color. The ISP can also optimize parameters such as exposure and color temperature of the shooting scene. In some embodiments, the ISP can be set in the camera 193.

[0076] Camera 193 is used to capture still images or videos. An object is projected onto a photosensitive element by generating an optical image through the lens. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then passed to an ISP for conversion into a digital image signal. The ISP outputs the digital image signal to a DSP for processing. The DSP converts the digital image signal into image signals in standard RGB, YUV, or other formats. In some embodiments, the electronic device 100 may include one or N cameras 193, where N is a positive integer greater than 1.

[0077] The external storage interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 through the external storage interface 120 to perform data storage functions. For example, music, video, and other files can be saved on the external memory card.

[0078] Internal memory 121 can be used to store computer executable program code, which includes instructions. Processor 110 executes various functional applications and data processing of electronic device 100 by running the instructions stored in internal memory 121. Internal memory 121 may include a program storage area and a data storage area. The program storage area may store the operating system, at least one application required for a function (such as sound playback, image playback, etc.), etc. The data storage area may store data created during the use of electronic device 100 (such as audio data, phonebook, etc.). Furthermore, internal memory 121 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc.

[0079] Electronic device 100 can implement audio functions, such as music playback and recording, through audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, and application processor.

[0080] Pressure sensor 180A is used to sense pressure signals and convert them into electrical signals. In some embodiments, pressure sensor 180A can be disposed on display screen 194. There are many types of pressure sensors 180A, such as resistive pressure sensors, inductive pressure sensors, and capacitive pressure sensors. A capacitive pressure sensor may include at least two parallel plates with conductive material. When force is applied to pressure sensor 180A, the capacitance between the electrodes changes. Electronic device 100 determines the pressure intensity based on the change in capacitance. When a touch operation is applied to display screen 194, electronic device 100 detects the intensity of the touch operation based on pressure sensor 180A. Electronic device 100 can also calculate the touch position based on the detection signal from pressure sensor 180A. In some embodiments, touch operations applied to the same touch position but with different touch operation intensities can correspond to different operation commands. For example, when a touch operation with an intensity less than a first pressure threshold is applied to the SMS application icon, a command to view an SMS is executed. When a touch operation with an intensity greater than or equal to the first pressure threshold is applied to the SMS application icon, a command to create a new SMS is executed.

[0081] The proximity sensor 180G may include, for example, a light-emitting diode (LED) and a light detector, such as a photodiode. The LED may be an infrared LED. The electronic device 100 emits infrared light outward through the LED. The electronic device 100 uses the photodiode to detect infrared reflected light from nearby objects. When sufficient reflected light is detected, it can be determined that there is an object near the electronic device 100. When insufficient reflected light is detected, the electronic device 100 can determine that there is no object near the electronic device 100. The electronic device 100 may use the proximity sensor 180G to detect when a user holds the electronic device 100 close to their ear for a call, so as to automatically turn off the screen to save power. The proximity sensor 180G can also be used in holster mode and pocket mode for automatic unlocking and locking of the screen.

[0082] The fingerprint sensor 180H is used to collect fingerprints, and the collected fingerprint image includes the fingerprint's characteristic information. For example, the fingerprint image includes ridges (lines with a certain width and direction in the fingerprint image) and valleys (the recessed parts between the ridges). Fingerprint images from different fingers of different users are different. The fingerprint image is used for fingerprint recognition, fingerprint gesture determination, etc. The electronic device 100 can utilize the collected fingerprint characteristics to achieve fingerprint unlocking, app access lock, fingerprint photography, fingerprint answering of calls, etc. The electronic device 100 can have a fingerprint recognition function, which can identify whether the collected image is a fingerprint image and can also identify the user's identity based on the fingerprint image. The electronic device 100 can also have a fingerprint gesture recognition function. The fingerprint sensor 180H can be located on the back of the electronic device (e.g., below the rear camera, or in the lower middle part of the electronic device), the front (e.g., below the display screen 194), or the side. For example, the fingerprint sensor 180H can be disposed in the display screen 194 and integrated with the display screen 194 to realize the fingerprint recognition function of the electronic device 100. In this case, the fingerprint sensor 180H is configured as part of the display screen 194, or it can be disposed in the display screen 194 in other ways. The fingerprint sensor 180H in the embodiments of this application can adopt any type of sensing technology, including but not limited to optical, capacitive, piezoelectric or ultrasonic sensing technologies.

[0083] The fingerprint sensor 180H can be an optical sensor. Optical sensors utilize the principles of light refraction and reflection to detect the refracted and reflected light from an object touching the fingerprint sensor. The angles and brightness of light refracted and reflected on the uneven surface of the object vary. The optical sensor projects the received light onto its image acquisition device through an optical structure, forming a touch image. The electronic device detects whether the touch image is a fingerprint image. Here, the fingerprint image is a touch image that includes the fingerprint.

[0084] The fingerprint sensor 180H may also include a capacitive sensor. A capacitive sensor detects changes in capacitance between the surface of an object touching the fingerprint sensor and the sensor electrodes. When the object touches the surface of the capacitive sensor, the sensor detects the electrical signal generated by the capacitance change. Based on the electric field distribution of the electrical signal, it converts the signal into a digital signal, thereby forming a touch image.

[0085] The fingerprint sensor 180H may also include an ultrasonic sensor that detects ultrasonic pulses reflected and scattered by an object touching the fingerprint device, thereby determining the distance difference between the unevenness of the object's surface based on the difference in the ultrasonic pulse signals, and forming a touch image based on the distance difference.

[0086] The fingerprint sensor 180H can also use sensors based on other principles to acquire touch images, and this application does not limit this.

[0087] The electronic device provided in this application embodiment can run an operating system (OS). This operating system can be various operating systems currently used in the industry, such as HarmonyOS, an operating system developed based on OpenHarmony; or other operating systems such as Android. TM An operating system can refer to the iOS mobile operating system; it can also refer to various open-source operating systems or their derivatives, such as Linux OS and other embedded operating systems; or it can refer to future new operating systems, such as AI operating systems based on artificial intelligence. An operating system is a set of interconnected system software programs that manage and control the operation of electronic devices, utilize and run hardware and software resources, and provide public services to organize user interactions. In electronic devices, the operating system occupies a pivotal position, connecting to the physical hardware layer below and providing a runtime environment for application software above.

[0088] An operating system typically includes a kernel layer, a middleware layer, and an application layer. The application layer includes applications, which can include system applications and third-party applications. The middleware layer is a suite of software, or frameworks, that provides various services to application developers, such as databases, multimedia, and graphics, or capabilities like distributed scheduling and system expansion. For example, the middleware layer can also be broadly divided into a framework layer and / or a system service layer. The framework layer provides application programming interfaces (APIs) and programming frameworks for applications in the application layer. The system service layer includes the system's core capabilities, providing services to applications through the framework layer. The kernel layer is the layer between hardware and software. The kernel layer can include hardware drivers and the operating system kernel. In addition to providing hardware drivers, the kernel layer also supports functions such as memory management and system process management.

[0089] The electronic devices we use in our daily lives come in various types and forms, and are applied in a wide range of scenarios. Therefore, based on the different forms and functions of electronic devices, different application scenarios, and different user needs, the operating systems used in these devices may also differ. These operating systems share commonalities but also have their own unique characteristics. Different operating systems affect user experience, application ecosystem, and system performance. The basic functions implemented by the electronic device provided in this application can be achieved using a general-purpose operating system or a dedicated operating system.

[0090] To more clearly illustrate the implementation of the embodiments of this application under a specific operating system, exemplarily, Figure 2 The architecture of HarmonyOS is illustrated, and those skilled in the art can deduce the implementation of the embodiments of this application under other specific operating systems, such as Android. TM Implementation under the operating system.

[0091] like Figure 2 As shown, the software architecture of an electronic device can be divided into several layers. In some embodiments, from bottom to top, these layers are: kernel layer, system service layer, framework layer, and application layer. The layers communicate with each other through software interfaces. System functions can be tailored, added, or combined at the subsystem granularity in different device deployment scenarios, and each subsystem can also be tailored, added, or combined at the functional granularity.

[0092] (1) Kernel layer The Kernel Abstraction Layer (KAL) provides basic kernel capabilities to upper layers by shielding the differences between multiple kernels, including but not limited to process / thread management, memory management, file system, network management, and peripheral device management.

[0093] Kernel Subsystem: Supports the selection of a suitable OS kernel for different resource-constrained devices, including but not limited to Linux kernel, HarmonyOS kernel, LiteOS, etc.

[0094] Driver Subsystem: The driver framework is the foundation for the open system hardware ecosystem, providing unified peripheral access capabilities and a framework for driver development and management. The driver framework includes: display drivers, camera drivers, audio drivers, Bluetooth drivers, sensor drivers, etc.

[0095] (2) System service layer The system service layer comprises the core capabilities of the system, providing services to applications through the framework layer. This layer includes, but is not limited to, the following subsystems: The system's basic capability subsystems provide fundamental capabilities for the operation, scheduling, and migration of distributed applications across multiple devices. For example, they may include distributed soft bus, distributed data management, distributed task scheduling, and the Ark multi-language runtime. They may also include multi-modal input subsystems, graphics subsystems, security subsystems, and AI subsystems.

[0096] Basic software service subsystems: provide common and general software services; for example, event notification subsystem, telephone service subsystem, multimedia subsystem, etc.

[0097] Enhanced software service subsystem suite: Provides differentiated capability-enhancing software services for different devices; for example, it may include proprietary business subsystems for smart screens, wearable devices, and IoT devices.

[0098] Hardware service subsystem set: provides hardware services; for example, it may include location service subsystem, user IAM (Identity and Access Management) subsystem, wearable proprietary hardware service subsystem, biometric identification, IoT proprietary hardware service subsystem, etc.

[0099] Distributed task scheduling enables distributed service management (discovery, synchronization, registration, and invocation), supporting remote startup, remote invocation, remote connection, and migration of applications across devices.

[0100] Distributed data management enables data synchronization, data storage, data sharing, and data access across all scenarios and devices.

[0101] The distributed soft bus provides communication-related capabilities for seamless interconnection between multiple devices, including: WLAN service capabilities, Bluetooth service capabilities, soft bus, inter-process communication RPC (Remote Procedure Call) and other communication capabilities.

[0102] Ark Multilingual Runtime is a unified compilation runtime platform designed to support the joint compilation and execution of multiple programming languages ​​and multiple chip platforms.

[0103] (3) Framework layer The framework layer provides application programming interfaces (APIs) and programming frameworks for applications in the application layer. Examples include the ArkUI framework (which provides a complete infrastructure for UI development of system applications, including UI functionalities such as components, layouts, animations, and interactive events, as well as a real-time interface preview tool), the user application framework, and the Ability framework (an Ability is a lightweight application; the Ability framework schedules and manages the operation and lifecycle of Abilities). Different devices may run different operating systems, and therefore support different APIs.

[0104] The HarmonyOS API is a series of open capabilities provided to support HarmonyOS application development. The HarmonyOS API can be set at the framework layer or independently of the framework layer. Examples include: Audio API (audio service), Push API (push service), and Account API (account service).

[0105] (4) Application layer Applications can include system apps and extended / third-party apps. System apps can include the desktop, control bar, settings, contacts, phone, camera, etc., while extended / third-party apps can include social apps, travel apps, etc.

[0106] It should be understood that the technical solutions in the embodiments of this application can be used in Android. TM iOS TM In systems such as HarmonyOS.

[0107] The technical solutions of this application embodiment can be applied, for example, to electronic devices with audio and video recording functions. Exemplarily, they can be applied to televisions, desktop computers, laptops, in-vehicle screens, and portable electronic devices such as mobile phones, foldable screens, tablets, cameras, camcorders, video recorders, watches, bracelets, and smart glasses. They can also be applied to other electronic devices with audio and video recording functions, and further to electronic devices in future networks or in future evolved public land mobile networks (PLMNs).

[0108] Currently, when portable electronic devices such as smartphones record telephoto videos, although the devices can clearly capture the details of distant subjects through zoom, the built-in microphones of the devices cannot effectively collect clear audio corresponding to the distant subjects. This often results in unclear sound of the subject when recording telephoto videos.

[0109] In view of this, this application provides an audio enhancement method and an electronic device, which helps to solve the problem of poor sound pickup performance of electronic devices at long distances, especially in scenarios where electronic devices record telephoto videos, and can improve the sound quality of the audio corresponding to the recorded telephoto video. The audio enhancement method provided by the embodiments of this application is described below.

[0110] In some embodiments, the audio enhancement method of this application can be implemented through a first function of the electronic device. When the first function of the electronic device is enabled, during audio recording, the electronic device can enhance the sound of the subject, and perform noise reduction or other processing. Additionally, during video recording, the electronic device can enhance the sound of the subject from the direction of the camera, and suppress ambient noise.

[0111] In some implementations, during video recording, when the recorded image is magnified, the electronic device can simultaneously enhance the sound from the first direction and effectively suppress background noise from other directions through a first function.

[0112] For example, in some electronic devices in the embodiments of this application, the first function may be referred to as the "super hearing" function. Through the "super hearing" function, the electronic device can intelligently analyze and process audio signals, thereby presenting clearer and higher quality audio to the user.

[0113] For example, taking a video recording scenario, Figure 3 A schematic diagram of a user interface for enabling audio enhancement function provided in an embodiment of this application is shown.

[0114] like Figure 3 As shown in (a), electronic device 300 displays a video recording interface. The user is using the camera application of electronic device 300 to record a meeting scene. The video recording interface includes a viewfinder that displays a group of people sitting around a round table having a meeting. Among the multiple participants in the recorded meeting scene, participant 1 and participant 2 may be included, with participant 1 speaking.

[0115] The top area of ​​the video recording interface of electronic device 300 displays a settings icon, which users can use to adjust shooting settings. The video recording interface also includes multiple shooting modes, which users can switch between. These shooting modes include night scene, portrait, photo, video, professional, and flash capture; the currently selected shooting mode is video.

[0116] In addition, the video recording interface also includes controls 301, 302, and 303. Users can use control 301 to start, pause, and end recording; control 302 to switch the electronic device 300 to the photo album interface to view recorded videos or captured pictures; and control 303 to switch the electronic device 300's working camera between the front and rear cameras.

[0117] The video recording interface may also include a zoom control bar. This bar offers various selectable zoom levels, such as 1x, 2x, 5x, and 10x. The currently selected zoom level is 1x. Users can zoom in or out of the subject within the viewfinder by clicking on different zoom levels.

[0118] For example, users can... Figure 3 In the interface shown in (a), selecting the zoom level "5×" triggers the electronic device 300 to display the following: Figure 3 The interface shown in (b).

[0119] like Figure 3 As shown in (b), in response to the detection that the user has triggered the electronic device 300 to switch the video zoom from "1×" to "5×", the viewfinder in the video recording interface of the electronic device 300 zooms in to focus on the speaker 1.

[0120] Furthermore, in response to the detection that the user triggers the electronic device 300 to switch the video zoom from "1×" to "5×", the video recording interface of the electronic device 300 displays a first prompt to prompt the user to enable the "Super Hearing" function. This first prompt may be, for example, a "Super Hearing" bubble reminder.

[0121] In some embodiments, the electronic device 300 has a preset focal length multiplier threshold. If the focal length multiplier switched by the user is greater than the focal length multiplier threshold, a first prompt is displayed on the video recording interface of the electronic device 300.

[0122] Furthermore, after the user clicks the "Super Hearing" bubble, the electronic device 300 activates the "Super Hearing" function. In this way, the electronic device 300 can amplify the sound emitted by participant 1 and also reduce or suppress other sounds. Based on this, in the video recorded by the user, compared to the video recorded without using the "Super Hearing" function, participant 1's sound volume is higher and the clarity is also higher.

[0123] In some embodiments, after the user clicks the "Super Hearing" bubble, the video recording interface of the electronic device 300 can display the working status of the "Super Hearing" function. For example, the video recording interface can display a prompt "Super Hearing is working", or it can display an icon to indicate that Super Hearing is working, etc.

[0124] For example, when participant 1 is speaking, electronic device 300 can amplify participant 1's voice and suppress the voices of other participants, noise in the room, noise outside the window, etc.

[0125] In some embodiments, in response to detecting that a user triggers the electronic device 300 to switch the video zoom to a value greater than or equal to a focal length magnification threshold, and the electronic device 300 detects a sound emitted by the subject through voice detection, the electronic device 300 displays as shown in the image. Figure 3 The interface shown in (b) displays the "Super Hearing" bubble.

[0126] The sound emitted by the subject may include, for example, the sound emitted by a human, the sound emitted by an animal, the sound emitted by an audio playback device, and so on.

[0127] In some embodiments, when the video frame recorded by the electronic device 300 includes multiple subjects, in response to detecting that the user triggers the electronic device 300 to switch the recording zoom to a value greater than or equal to the focal length multiplier threshold, and the electronic device 300 detects the sound emitted by subject 1 among the multiple subjects through voice detection, the electronic device 300 can focus the zoom on subject 1, and the electronic device 300 can display a "super hearing" bubble. After the user triggers the "super hearing" function to be activated, the electronic device 300 enhances the sound emitted from the direction of subject 1 and reduces or suppresses the sound emitted from other directions.

[0128] In some cases, when "Super Hearing" is enabled, the electronic device 300 can also detect the position of the subject in the recording screen through the visual detection module. When the electronic device 300 detects that the subject 1 moves while emitting sound, the electronic device 300 can adjust the direction of sound enhancement according to the position information of the subject 1, so that the electronic device 300 can always enhance the sound emitted from the direction of the subject 1 and reduce or suppress the sound emitted from other directions, making the audio enhancement effect more stable.

[0129] Based on this, when "super hearing" is enabled, and when electronic device 300 detects through voice detection that the sound emitted by subject 1 among multiple subjects has switched to the sound emitted by subject 2 among multiple subjects, electronic device 300 can enhance the sound emitted from the direction of subject 2 and reduce or suppress the sound emitted from other directions.

[0130] For example, taking a video recording scenario, Figure 4 A schematic diagram of a user interface using audio enhancement functionality provided in an embodiment of this application is shown.

[0131] by Figure 3 Taking the illustrated embodiment as an example, the electronic device 300 uses... Figure 3 After capturing a video of the meeting scene in the illustrated embodiment, the electronic device 300 stores the recorded video of the meeting scene in its gallery after recording ends.

[0132] like Figure 4 As shown in (a), the electronic device 300 displays a playback interface for a meeting scene video. This playback interface includes a control 401, which is used by the user to enable or disable the "Super Hearing" effect of the meeting scene video. Since the meeting scene video used the "Super Hearing" function during recording, the control 701 displayed in the playback interface is in an enabled state, indicating that the "Super Hearing" function was used during the recording of the meeting scene video.

[0133] The playback interface may also include the date "September 3, 2025" when the video recording of the meeting scene was completed; it may also include a control for returning to the previous interface.

[0134] exist Figure 4 In the interface shown in (a), control 401 is in the "on" state, indicating that the "super hearing" effect of the meeting scene video is enabled. In this state, when the user clicks the play button for the meeting scene video, the electronic device 300 plays the meeting scene video, where the audio corresponding to the played video is enhanced audio. Furthermore, the interface may also include a control bar 402, which displays the current level of the "super hearing" effect in the meeting scene video. The control bar 402 includes a control ball 403, which the user can adjust by sliding the control ball 403 across the control bar 402. For example, sliding the control ball 403 to the right can enhance the current level of the "super hearing" effect in the meeting scene video, that is, increase the degree of audio enhancement.

[0135] Users can, for example, in such Figure 4 Clicking control 401 in the interface shown in (a) triggers the switch of control 401's state from the on state to the off state. At this time, the display interface of electronic device 300 is as follows: Figure 4 As shown in (b) above. In this case, when the user clicks the play button for the meeting scene video, the electronic device 300 plays the meeting scene video, wherein the audio corresponding to the played meeting scene video is the audio without audio enhancement.

[0136] For example, Figure 5 This illustration shows another user interface diagram using audio enhancement functionality provided in an embodiment of this application.

[0137] like Figure 5 As shown in (a), electronic device 500 displays a video recording interface. The user is using the camera application of electronic device 500 to record a meeting scene. The video recording interface includes a viewfinder showing a group of people sitting around a round table having a meeting. Among the participants in the recorded meeting scene, participant 1 may be present, and participant 1 is speaking. The video recording interface also includes a notification icon 501, indicating that the "Super Hearing" function of electronic device 500 is currently enabled, and is performing real-time amplification processing on the recorded voice of participant 1.

[0138] like Figure 5 As shown in (b), during the video recording process, while speaking, participant 1 stands up and moves to the right from the position of sitting around the table. In response to the electronic device 500 detecting that participant 1 has moved while making a sound, the electronic device 500 can adjust the direction of sound enhancement according to the position information of participant 1, so that the electronic device 500 can always enhance the sound emitted from the direction of participant 1 and reduce or suppress the sound emitted from other directions, making the audio enhancement effect more stable.

[0139] This embodiment can solve the problem of unstable sound pickup caused by target movement.

[0140] In some embodiments, during video recording, the electronic device 500 can detect the position of the subject emitting the sound in real time, for example, through a visual detection module.

[0141] For example, Figure 6 This illustration shows another user interface diagram using audio enhancement functionality provided in an embodiment of this application.

[0142] like Figure 6As shown in (a), electronic device 600 displays a video recording interface. The user is using the camera application of electronic device 600 to record a meeting scene. The video recording interface includes a viewfinder showing a group of people sitting around a round table having a meeting. The participants in the recorded meeting scene may include participant 1 and participant 2, with participant 1 speaking. The video recording interface also includes a prompt icon 601, indicating that the "Super Hearing" function of electronic device 600 is currently enabled, and is performing real-time amplification processing on the recorded voice of participant 1.

[0143] like Figure 6 As shown in (b), during the video recording process, the speaker is switched from participant 1 to participant 2. In response to the electronic device 600 detecting the switch from participant 1 to participant 2, the electronic device 600 can adjust the direction of sound enhancement according to the location information of participant 2, so that the electronic device 600 can switch to enhance the sound emitted from the direction of participant 2 and reduce or suppress the sound emitted from other directions, making the audio enhancement effect more stable.

[0144] In some embodiments, during video recording, the electronic device 600 can perform voice detection and / or location detection on the subject emitting the sound in real time. For example, location detection can be performed by a visual detection module, and voice detection can be performed by a voice detection module.

[0145] For example, Figure 7 This illustration shows another user interface diagram using audio enhancement functionality provided in an embodiment of this application.

[0146] After capturing a video of the meeting scene without using the "Super Hearing" function, the electronic device 700 stores the recorded video of the meeting scene in its gallery after recording ends.

[0147] like Figure 7 As shown in (a), the electronic device 700 displays a playback interface for a meeting scene video. This playback interface includes a control 701, which is used by the user to enable or disable the "Super Hearing" effect for the meeting scene video. Since the "Super Hearing" function was not used during the recording of the meeting scene video, the control 701 displayed on the playback interface is in a closed state, indicating that the "Super Hearing" function was not used during the recording of the meeting scene video.

[0148] In this situation, if users feel that the audio in the recorded meeting video is unclear, they can... Figure 4Clicking on control 701 in the interface shown in (a) will switch the state of control 701 from closed to open.

[0149] In response to the user switching the state of control 701 from closed to open, electronic device 700 can perform audio enhancement processing on the recorded meeting scene video. Specifically, it can enhance the main sound and reduce and / or suppress other sounds. At this time, electronic device 700 displays as shown below. Figure 7 The interface shown in (b) has control 701 in the "on" state, and can also display a prompt message to the user indicating that audio enhancement processing is being performed on the video of the meeting scene. This prompt message could be, for example, the text "Super Listening in Progress".

[0150] After the electronic device 700 finishes audio enhancement processing of the video of the meeting scene, the electronic device 700 displays as follows: Figure 7 The interface shown in (c) has control 701 in the "on" state and displays a play button for the meeting scene video. The electronic device 700 plays the meeting scene video. The audio corresponding to the played meeting scene video is enhanced audio. Furthermore, the interface may include a control bar 702, which displays the current level of the "super hearing" effect in the meeting scene video. The control bar 702 includes a control ball 703, which the user can adjust by sliding the control ball 703 across the control bar 702. For example, sliding the control ball 703 to the right can enhance the current level of the "super hearing" effect in the meeting scene video, that is, increase the degree of audio enhancement.

[0151] Users can also, for example, in such Figure 7 Clicking control 701 in the interface shown in (c) triggers the switch of control 701's state from the on state to the off state. At this time, the display interface of electronic device 700 is as follows: Figure 7 As shown in (a) above. In this case, when the user clicks the play button for the meeting scene video, the electronic device 700 plays the meeting scene video, wherein the audio corresponding to the played meeting scene video is the audio without audio enhancement.

[0152] In some embodiments, the electronic device can also identify the scene of the video recording, determine the subject of the sound based on the scene, and then enhance the audio of the sound from the subject based on the location of the subject.

[0153] For example, if the current scene is identified as a meeting scene and the main speaker is determined to be a person, audio enhancement can be performed on the person's voice.

[0154] For example, if the current scene is identified as a puppy playing on the lawn, and the main sound is determined to be a puppy, then audio enhancement can be applied to the sound made by the puppy.

[0155] For example, Figure 8 This illustration shows another audio enhancement operation diagram provided in an embodiment of this application.

[0156] like Figure 8 As shown in (a), the user is recording video in landscape mode on an electronic device, with a focal length of 1x. In the viewfinder, two elderly people are playing chess under a large tree. Not far away, a little boy is walking a small dog, which is running ahead and making happy noises. Two birds are perched on a branch of the tree, chirping happily.

[0157] Based on this, the user adjusts the electronic device's shooting focal length from 1x to 5x and points the lens at the bird on the branch. At this point, the electronic device can display as follows: Figure 8 The interface shown in (b) is an electronic device that can amplify the sounds of birds while suppressing the sounds of people under the tree and dogs.

[0158] Furthermore, the user points the electronic device's camera at the puppy under the tree. At this point, the electronic device can display something like... Figure 8 The interface shown in (c) is an electronic device that can amplify the sound of a puppy while suppressing the sounds of humans and birds in the trees.

[0159] In some embodiments, after the user adjusts the focus of the electronic device, the electronic device may first prompt the user to activate the "Super Hearing" function. If the user confirms that the "Super Hearing" function is activated, the electronic device will amplify the sound corresponding to the main subject in the picture.

[0160] In some embodiments, the electronic device can perform image analysis on the screen in the viewfinder to identify the subject of the sound.

[0161] For example, taking a video recording scenario, Figure 9 This illustration shows another user interface diagram for enabling audio enhancement functions provided in an embodiment of this application.

[0162] like Figure 9As shown in (a), electronic device 900 displays a video recording interface. The user is using the camera application of electronic device 900 to record a meeting scene. The video recording interface includes a viewfinder that displays a group of people sitting around a round table having a meeting. Among the multiple participants in the recorded meeting scene, participant 1 and participant 2 may be included, with participant 1 speaking.

[0163] The video recording interface also includes control 901, which is used by the user to activate the "super hearing" function of the electronic device.

[0164] In response to detecting user in Figure 9 Clicking control 901 in the interface shown in (a) activates the "Super Hearing" function of electronic device 900. This allows electronic device 900 to amplify the sound emitted by participant 1 and reduce or suppress other sounds. As a result, in the video recorded by the user, participant 1's voice has a higher volume and greater clarity compared to a video recorded without the "Super Hearing" function.

[0165] In some embodiments, after the user clicks the control 901, the electronic device 900 displays as shown below. Figure 9 The video recording interface shown in (b) has the control 901 switching from the off state to the on state. It can also display the working status of the "Super Hearing" function, such as displaying the prompt "Super Hearing in progress" in the video recording interface.

[0166] For example, taking a video recording scenario, Figure 10 Two more user interface diagrams using audio enhancement functions provided in embodiments of this application are shown.

[0167] like Figure 10 As shown in (a), the electronic device 1000 displays a playback interface for a meeting scene video. This playback interface includes a control 1001, which is used by the user to enable or disable the "Super Hearing" effect of the meeting scene video. Since the meeting scene video used the "Super Hearing" function during recording, the control 1001 displayed on the playback interface is in an "on" state, indicating that the "Super Hearing" function was used during the recording of the meeting scene video.

[0168] The playback interface may also include the date "September 3, 2025" when the video recording of the meeting scene was completed; it may also include a control for returning to the previous interface.

[0169] exist Figure 10In the interface shown in (a), control 1001 is in the on state, indicating that the "super hearing" effect of the meeting scene video is on.

[0170] Furthermore, the interface can also include audio track 1 and audio track 2. Audio track 1 is the audio track corresponding to the original audio; audio track 2 is the audio track corresponding to the super hearing aid.

[0171] In some embodiments, users can adjust or play the original audio through audio track 1; and can adjust or play the audio track corresponding to Super Hearing through audio track 2.

[0172] like Figure 10 As shown in (b), control 1001 is in the "on" state, indicating that the "Super Hearing" effect of the meeting scene video is enabled. Furthermore, this interface can also include audio track 1, audio track 2, and audio track 3. Audio track 1 is the audio track corresponding to the original audio; audio tracks 2 and 3 are the audio tracks corresponding to the Super Hearing effect.

[0173] In one example, the main sound subject of audio track 2 is a participant in the video frame, that is, the audio track that enhances the sound of that participant; the main sound subject of audio track 3 is another participant in the video frame, that is, the audio track that enhances the sound of that other participant.

[0174] For example, Figure 11 A schematic flowchart of an audio enhancement method 1100 provided in an embodiment of this application is shown.

[0175] like Figure 11 As shown, the method 1100 includes: S1101: The first device displays the first interface, which is the interface for recording the first video.

[0176] In some embodiments, the camera application of the first device can be opened and switched to recording mode. In recording mode, the first device can be controlled to start recording a first video. At this time, the first device displays a first interface.

[0177] S1102: In response to detecting that the user has switched the focal length from the first focal length to the second focal length, the first device performs audio enhancement processing on the audio corresponding to the first video while recording the first video, wherein the second focal length is greater than the first focal length.

[0178] In some embodiments, the first focal length is less than a focal length threshold, and the second focal length is greater than or equal to the focal length threshold. The focal length threshold is a preset value in the first device.

[0179] In some embodiments, the operation of a user switching the focal length from a first focal length to a second focal length can be, for example, the operation of a user switching the zoom level of a first video recording from a first zoom level to a second zoom level, wherein the first zoom level is less than the second zoom level. For example, the first zoom level is less than a zoom level threshold, and the second zoom level is greater than a zoom level threshold.

[0180] In one example, the zoom threshold could be, for instance, 5x zoom.

[0181] In some embodiments, in response to detecting a user switching the focal length from a first focal length to a second focal length, the first device may first display a first control on a first interface. This first control is used by the user to enable audio enhancement for the sound corresponding to the first video. In response to the user's operation on the first control, the first device performs audio enhancement processing on the sound corresponding to the first video while recording the first video. Figure 3 Taking the illustrated embodiment as an example, the first control could be, for example, Figure 3 The "Super Hearing" bubble reminder shown in (b) is shown in the image.

[0182] The user's action on the first control can be, for example, clicking the first control.

[0183] The first control can also be used to prompt the user to enable audio enhancement for the sound corresponding to the first video. This can be done through icons, text, or other means. Text prompts could include phrases like "Super Hearing" or "Enable Super Hearing for clearer target sound," and so on.

[0184] In some embodiments, while enhancing the audio of the first video, the first device may also retain the original audio corresponding to the first video, i.e., the audio that has not undergone audio enhancement processing.

[0185] In some embodiments, in response to detecting a user switching the focus from a first focal length to a second focal length, the first device may first determine whether human voice is detected; if human voice is detected, the first device displays a first control on the first interface. This reduces unnecessary display of the first control, effectively avoiding disruption to the user's use of the electronic device and contributing to a better user experience.

[0186] For example, the first device can detect human voices through a voice detection module.

[0187] In some embodiments, in response to detecting that a user switches the focus from a first focal length to a second focal length, the first device may first determine whether there is a subject emitting sound in the frame of the first video; if it is determined that there is a subject emitting sound in the frame of the first video, the first device displays a first control on the first interface. This reduces unnecessary display of the first control, effectively avoiding interference with the user's use of the electronic device and helping to improve the user experience.

[0188] The entity emitting the sound can include, for example, people, animals, plants, etc. Figure 3 Taking the illustrated embodiment as an example, in Figure 3 In the interface shown in (a), the user adjusts the zoom from 1x to 5x. The electronic device identifies the subject emitting sound in the recorded video. If it confirms the presence of "Attendee 1" emitting sound in the video, the first device displays the following: Figure 3 The interface shown in (b) is shown in the image.

[0189] In some embodiments, the first device may include a sound detection module and / or a visual detection module, which can detect whether there is a subject emitting sound in the frame of the first video.

[0190] In some embodiments, method 1100 may further include: after the first video recording ends, the first device displays a second interface. The second interface is a playback interface for the first video, and includes a second control. The second control is in an "on" state, and the "on" state of the second control indicates that the current state of the first video is an audio-enhanced state. Figure 4 Taking the illustrated embodiment as an example, the second interface may be, for example, Figure 4 In the interface shown in (a), the second control can be, for example, control 401.

[0191] In some embodiments, after the first video recording is completed, the first device saves the first video to the album. The user can view the first video through the album application of the first device, that is, the user can enter the second interface by using the album application of the first device.

[0192] It is understandable that when the first video is currently in an audio-enhanced state, the sound emitted by the first device from the first video is the audio-enhanced sound.

[0193] In some embodiments, the second interface may further include a third control for the user to adjust the audio enhancement level corresponding to the first video. Figure 4In the illustrated embodiment, the third control may be, for example, a control bar 402 and a control ball 403.

[0194] In some embodiments, method 1100 may further include: in response to detecting a user clicking the second control, switching the state of the second control from an on state to a off state, wherein the off state of the second control is used to indicate that the current state of the first video changes from an audio-enhanced state to an unenhanced state. That is, in response to detecting a user clicking the second control, the first device cancels the audio enhancement effect on the first video. In this way, the user can use the second control to experience both the audio-enhanced sound of the first video and the original audio of the first video, allowing for comparison of the two sound effects and thus retaining the desired sound effect according to actual needs.

[0195] In other words, when the second control is "on", the first video played by the user is a video with enhanced audio; when the second control is "off", the first video played by the user is the original video without audio enhancement.

[0196] It is understandable that when the first video is in a state without audio enhancement, the sound emitted by the first device from the first video is the sound without audio enhancement when the user plays the first video.

[0197] In some embodiments, audio enhancement processing of the sound corresponding to the first video may include: enhancing the sound emitted by a first subject in the frame of the first video; and suppressing sounds other than the sound emitted by the first subject. Figure 3 Taking the illustrated embodiment as an example, the first device enhances the audio of the sound emitted by participant 1, such as increasing the volume or increasing the sound frequency; and suppresses the sound emitted by other participants, such as noise reduction or reducing the volume.

[0198] In one implementation, the first device may enhance the sound from the location of a first subject in the frame of the first video.

[0199] In some embodiments, the first device may detect the location of the first subject through voice detection and / or visual detection.

[0200] In some embodiments, the first device may update the location of the first subject through voice detection and / or visual detection; and enhance the sound at the updated location of the first subject based on the updated location of the first subject.

[0201] In some examples, taking a person as the first subject, the first device can detect the face in the shooting scene through the visual detection module, thereby recognizing the movement of the person in the scene or the movement of the person in the frame, and updating the position of the person according to the movement of the person, and enhancing the sound of the updated position of the person, which can improve the stability of the audio enhancement effect.

[0202] For example, Figure 12 A schematic flowchart of another audio enhancement method 1200 provided in an embodiment of this application is shown.

[0203] like Figure 12 As shown, the method 1200 includes: S1201: The first device displays a third interface, which is the playback interface for the second video.

[0204] Among them, the recording focal length corresponding to the second video is greater than the focal length threshold.

[0205] The second video was not audio-enhanced during recording. The third interface includes a fourth control, which is in a closed state. The closed fourth control is used to indicate that the current state of the second video is that it has not been audio-enhanced.

[0206] S1202: In response to detecting a user clicking the fourth control, the first device performs audio enhancement processing on the sound corresponding to the second video, and the state of the fourth control is switched to the on state. Figure 7 Taking the illustrated embodiment as an example, the third interface is... Figure 7 The interface shown in (a) has a fourth control, which is the Super Hearing switch control 701.

[0207] In some embodiments, audio enhancement processing of the sound corresponding to the second video may include: enhancing the sound emitted by a second subject in the frame of the second video; and suppressing sounds other than those emitted by the second subject.

[0208] After the first device performs audio enhancement processing on the audio corresponding to the second video, the user can play the processed second video, which will then play the enhanced audio. The user can also adjust the degree of audio enhancement for the second video.

[0209] In this embodiment, a function switch for audio enhancement of the video is provided to the user in the video preview interface. This guides the user to use the audio enhancement function of the first device to enhance the audio of the recorded video, thereby improving video quality and enhancing the user's device experience. One or more modules or units described herein can be implemented in software, hardware, or a combination of both. When any of the above modules or units are implemented in software, the software exists as computer program instructions and is stored in memory. A processor can be used to execute the program instructions and implement the above method flow. The processor can include, but is not limited to, at least one of the following: a central processing unit (CPU), microprocessor, DSP, microcontroller unit (MCU), or artificial intelligence processor, etc., and various computing devices that run software. Each computing device may include one or more cores for executing software instructions to perform calculations or processing. The processor can be built into a system-on-a-chip (SoC) or application-specific integrated circuit (ASIC), or it can be a separate semiconductor chip. The processor's internal processing unit is used to execute software instructions for computation or processing. It may further include necessary hardware accelerators, such as field-programmable gate arrays (FPGAs), programmable logic devices (PLDs), or logic circuits that implement dedicated logic operations.

[0210] When the modules or units described herein are implemented in hardware, the hardware may be any one or any combination of a CPU, microprocessor, DSP, MCU, artificial intelligence processor, ASIC, SoC, FPGA, PLD, application-specific digital circuit, hardware accelerator, or non-integrated discrete device, which may run the necessary software or perform the above method flow independently of software.

[0211] When the modules or units described herein are implemented using software, they can be implemented in whole or in part as a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk (SSD)).

[0212] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0213] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0214] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0215] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0216] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0217] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0218] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for audio enhancement, characterized in that, The method is applied to a first device, and the method includes: The first interface is displayed, which is the interface for recording the first video; In response to detecting that the user switches the focal length from the first focal length to the second focal length, audio enhancement processing is performed on the audio corresponding to the first video while recording the first video, wherein the second focal length is greater than the first focal length.

2. The method according to claim 1, characterized in that, In response to detecting that the user has switched the focal length from a first focal length to a second focal length, the audio enhancement processing for the corresponding audio of the first video is performed simultaneously with recording the first video, including: In response to detecting that the user has switched the focal length from the first focal length to the second focal length, a first control is displayed on the first interface. The first control is used by the user to enable audio enhancement for the sound corresponding to the first video. In response to the user's operation on the first control, while recording the first video, audio enhancement processing is performed on the corresponding sound of the first video.

3. The method according to claim 2, characterized in that, The first control is also used to prompt the user to enable audio enhancement for the sound corresponding to the first video.

4. The method according to claim 2 or 3, characterized in that, In response to detecting that the user has switched the focal length from a first focal length to a second focal length, a first control is displayed on the first interface, including: In response to detecting that the user switches the focus from the first focus to the second focus, determine whether there is a human voice in the audio corresponding to the first video; If it is determined that there is human voice in the audio corresponding to the first video, the first control is displayed on the first interface.

5. The method according to claim 2 or 3, characterized in that, In response to detecting that the user has switched the focal length from a first focal length to a second focal length, a first control is displayed on the first interface, including: In response to detecting that the user switches the focus from the first focus to the second focus, it is determined whether there is a subject emitting sound in the frame of the first video; If it is determined that there is a subject emitting sound in the first video frame, the first control is displayed on the first interface.

6. The method according to any one of claims 1 to 5, characterized in that, The first focal length is less than the focal length threshold, and the second focal length is greater than or equal to the focal length threshold.

7. The method according to any one of claims 1 to 6, characterized in that, The method further includes: After the first video recording ends, a second interface is displayed. The second interface is the playback interface of the first video. The second interface includes a second control. The state of the second control is on, which is used to indicate that the current state of the first video is an audio-enhanced state.

8. The method according to claim 7, characterized in that, The second interface also includes a third control, which is used by the user to adjust the audio enhancement level corresponding to the first video.

9. The method according to claim 7 or 8, characterized in that, The method further includes: In response to detecting a user clicking the second control, the state of the second control changes from an on state to a off state, indicating that the current state of the first video changes from an audio-enhanced state to an unenhanced state.

10. The method according to any one of claims 1 to 9, characterized in that, The audio enhancement processing of the sound corresponding to the first video includes: The sound emitted by the first subject in the first video frame is enhanced. Suppress sounds other than those emitted by the first subject.

11. The method according to claim 10, characterized in that, The enhancement processing of the sound emitted by the first subject in the first video frame includes: The sound from the location of the first subject is enhanced.

12. The method according to claim 11, characterized in that, The method further includes: The location of the first subject is updated through voice detection and / or visual detection; The enhancement processing of the sound emitted by the first subject in the first video frame further includes: The sound from the location of the updated first subject is enhanced.

13. A method for audio enhancement, characterized in that, The method is applied to a first device, and the method includes: The third interface is displayed. The third interface is the playback interface of the second video. The recording focal length of the second video is greater than the focal length threshold. The second video has not undergone audio enhancement processing during the recording process. The third interface includes a fourth control. The state of the fourth control is closed, which is used to indicate that the current state of the second video is that it has not undergone audio enhancement. In response to detecting a user clicking the fourth control, audio enhancement processing is performed on the sound corresponding to the second video, and the state of the fourth control is switched to the on state.

14. The method according to claim 13, characterized in that, The audio enhancement processing of the sound corresponding to the second video includes: The sound emitted by the second subject in the second video frame is enhanced. Suppress sounds other than those emitted by the second subject.

15. An electronic device, characterized in that, include: One or more processors; One or more memory units; And one or more computer programs, wherein the one or more computer programs are stored in the one or more memories, the one or more computer programs including instructions that, when executed by the one or more processors, cause the electronic device to perform the method as claimed in any one of claims 1 to 12, or to perform the method as claimed in claim 13 or 14.

16. A computer-readable storage medium, characterized in that, The storage medium stores a program or instructions that, when executed by a processor, implement the method as described in any one of claims 1 to 12, or the method as described in claim 13 or 14.

17. A chip, characterized in that, The chip includes circuitry for performing the method as described in any one of claims 1 to 12, or for performing the method as described in claim 13 or 14.

18. A computer program product, characterized in that, The computer program product stores a program or instructions that, when executed by a processor, implement the method as described in any one of claims 1 to 12, or implement the method as described in claim 13 or 14.