An audio processing method and device

By recording multiple audio based on the shooting angle corresponding to the video picture in the multi-channel recording mode, the problem that audio cannot match the video picture and shooting angle in the prior art is solved, and the user's audio experience is improved.

CN113365013BActive Publication Date: 2025-06-17HUAWEI TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202010153655.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-03-06
Publication Date
2025-06-17
Estimated Expiration
2040-03-06

AI Technical Summary

Technical Problem

In the existing multi-channel recording mode, electronic devices can only record one audio channel and cannot match different video images and shooting angles, resulting in poor user audio experience.

Method used

In multi-channel recording mode, the electronic device records the corresponding audio of each video screen according to the corresponding shooting angle of each video screen to ensure that the recorded audio matches the video screen and shooting angle.

Benefits of technology

During video playback, users can choose to play audio that matches the video picture and shooting angle of their attention, improving the user's multi-channel recording audio experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113365013B_ABST
    Figure CN113365013B_ABST
Patent Text Reader

Abstract

An embodiment of the present application provides an audio processing method and device, which relate to the field of electronic technologies, and can record multiple video images and multiple audio channels simultaneously in the multi-channel video recording mode, and can play different audio during video playback, improving the audio experience of users in multi-channel video recording. The specific solution is as follows: After the electronic device detects the operation of the user turning on the camera, it displays a shooting preview interface; then it enters the multi-channel video recording mode; after the electronic device detects the shooting operation of the user, it displays a shooting interface, and the shooting interface includes multiple video images; then the electronic device records multiple video images; and respectively records the audio corresponding to each video image according to the shooting angle corresponding to each video image in the multiple video images. The embodiment of the present application is used in the multi-channel video recording process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the field of electronic technologies, and in particular, to an audio processing method and device. Background Art

[0002] With the improvement of the computing power and hardware capabilities of electronic devices such as mobile phones or tablet computers, the video recording function of electronic devices has become more and more powerful. For example, some electronic devices support multi-channel video recording, also known as multi-scene video recording.

[0003] In the existing multi-channel video recording mode, an electronic device can record audio and multi-channel video images. For example, the electronic device can record a panoramic video image and a close-up video image respectively. When playing back the video, the electronic device can play the audio and multi-channel video images. Summary of the Invention

[0004] Embodiments of the present application provide an audio processing method and device, which can record multi-channel video images and multi-channel audio simultaneously in the multi-channel video recording mode, and can play different audio during video playback, improving the user's audio experience of multi-channel video recording.

[0005] To achieve the above object, the embodiments of the present application adopt the following technical solutions:

[0006] On the one hand, embodiments of the present application provide an audio processing method, which includes: after the electronic device detects that the user opens the camera, it displays a shooting preview interface. Then, the electronic device enters the multi-channel video recording mode. After the electronic device detects the user's shooting operation, it displays a shooting interface, and the shooting interface includes multi-channel video images. Then, the electronic device records multi-channel video. The electronic device records the audio corresponding to each video image in the multi-channel video images according to the shooting angle corresponding to each video image in the multi-channel video images. Wherein, when the multi-channel video images include a first video image and a second video image, the electronic device records the audio corresponding to the first video image according to the shooting angle corresponding to the first video image; and records the audio corresponding to the second video image according to the shooting angle corresponding to the second video image.

[0007] In this solution, in the multi-channel video recording mode, the electronic device can record multi-channel audio corresponding to multi-channel video images according to the shooting angle corresponding to each video image while recording multi-channel video images, so that the recorded audio can be matched with different video images and shooting angles, so that when playing back the video, the user can select to play the audio matched with the video image and shooting angle that the user is concerned about, thereby improving the user's audio experience.

[0008] In a possible design, each video image corresponds to a shooting angle.

[0009] In this way, the audio corresponding to each video frame also corresponds to a shooting perspective.

[0010] In another possible design, the shooting perspective corresponding to each video frame is variable.

[0011] In this way, the video corresponding to each video frame is also dynamic audio that matches the real-time changing shooting perspective.

[0012] In another possible design, after the electronic device records the audio corresponding to each video frame respectively, the method further includes: after the electronic device detects an operation of the user to stop shooting, generating multiple recorded videos, and the multiple recorded videos further include the audio corresponding to each video frame. After the electronic device detects an operation of the user to play the multiple recorded videos, a video playback interface is displayed, and the video playback interface includes multiple video frames.

[0013] In this way, during video playback, the electronic device can default to playing each video frame.

[0014] In another possible design, after the electronic device detects an operation of the user to play the multiple recorded videos, the method further includes: after the electronic device detects an operation of the user to play the audio corresponding to the first video frame, playing the audio corresponding to the first video frame, and the first video frame is one of the multiple video frames. If the electronic device plays other audio corresponding to the multiple video frames before playing the audio corresponding to the first video frame, the electronic device stops playing the other audio.

[0015] In this solution, the electronic device can play one piece of audio indicated by the user and stop playing the audio of other paths.

[0016] In another possible design, the video playback interface includes audio playback controls corresponding to each video frame respectively, and the electronic device detecting an operation of the user to play the audio corresponding to the first video frame includes: the electronic device detecting an operation of the user to click the audio playback control corresponding to the first video frame.

[0017] In this solution, the user can indicate to play the audio corresponding to the audio playback control by clicking the audio playback control.

[0018] In another possible design, before the electronic device detects an operation of the user to play the audio corresponding to the first video frame, the method further includes: the electronic device defaulting to playing the audio corresponding to the second video frame among the multiple video frames.

[0019] In this solution, during video playback, the electronic device can default to playing a certain piece of audio, so that sound can be played for the user in a timely manner during video playback.

[0020] In another possible design, the electronic device defaults to playing the shooting perspective corresponding to the second video picture as a preset shooting perspective.

[0021] That is to say, the electronic device can default to playing the audio corresponding to the preset shooting perspective. For example, the preset shooting perspective can be a wide-angle perspective.

[0022] In another possible design, after the electronic device displays the video playback interface, the method further includes: after the electronic device detects an operation by the user to play the first video picture among multiple video pictures, the electronic device displays the first video picture on the video playback interface and stops displaying other video pictures except the first video picture. The electronic device automatically plays the audio corresponding to the first video picture.

[0023] In this solution, the electronic device can play only one of the video pictures according to the user's instruction and automatically play the audio corresponding to that video picture.

[0024] For example, the electronic device displays the first video picture on the video playback interface, including: magnifying or full-screen displaying the first video picture on the video playback interface. In this way, the single video picture instructed by the user to play can be highlighted and prominently displayed.

[0025] In another possible design, before the electronic device displays the shooting interface, the method further includes: the electronic device determines a target shooting mode, and the target shooting mode is used to represent the number of video picture lines to be recorded. The electronic device displays the shooting interface, including: the electronic device displays the shooting interface according to the number of video picture lines to be recorded corresponding to the target shooting mode.

[0026] In this solution, the electronic device can first determine the target shooting mode and then display the shooting interface according to the target representation mode.

[0027] In another possible design, the target shooting mode is further used to represent the correspondence between each video picture and the shooting perspective. The electronic device records the audio corresponding to each video picture according to the shooting perspective corresponding to each video picture among multiple video pictures, including: the electronic device determines the shooting perspective corresponding to each video picture according to the correspondence between each video picture and the shooting perspective corresponding to the target shooting mode. The electronic device records the audio corresponding to each video picture according to the shooting perspective corresponding to each video picture.

[0028] In this solution, the electronic device can determine the shooting perspective according to the target shooting mode, and thus record the audio corresponding to the video picture according to the shooting perspective.

[0029] For example, the target shooting mode includes: a wide-angle view and zoom view combination mode, a wide-angle view and front view combination mode, a zoom view and front view combination mode, or a wide-angle view, zoom view, and front view combination mode. The zoom ratio corresponding to the wide-angle view is less than or equal to a preset value, and the zoom ratio corresponding to the zoom view is greater than the preset value.

[0030] In another possible design, the target shooting mode is a preset shooting mode.

[0031] For example, the preset shooting mode can be a wide-angle view and zoom view combination mode.

[0032] In another possible design, the electronic device records the audio corresponding to each video frame in the multi-channel video frames according to the shooting view corresponding to each video frame in the multi-channel video frames, including: The electronic device determines the shooting view corresponding to each video frame according to the size relationship between the zoom ratio corresponding to each video frame and the preset value and the front / back feature. The electronic device records the audio corresponding to each video frame according to the shooting view corresponding to each video frame. When the multi-channel video frames include a first video frame and a second video frame, the electronic device determines the shooting view corresponding to the first video frame according to the size relationship between the zoom ratio corresponding to the first video frame and the preset value and the front / back feature; and determines the shooting view corresponding to the second video frame according to the size relationship between the zoom ratio corresponding to the second video frame and the preset value and the front / back feature.

[0033] In this solution, the electronic device can determine the shooting view according to the size relationship between the zoom ratio corresponding to the video frame and the preset value and the front / back feature, and thus record the audio corresponding to the video frame according to the shooting view.

[0034] In another possible design, after the electronic device determines the target shooting mode, the method further includes: The electronic device detects an operation of the user to switch the target shooting mode. The electronic device switches the target shooting mode.

[0035] That is to say, during the shooting process of multi-channel video recording, the electronic device can also switch the target shooting mode.

[0036] In another possible design, the shooting view corresponding to the first video frame in the multi-channel video frames is the first shooting view. The electronic device records the audio corresponding to the first video frame according to the shooting view corresponding to the first video frame, including: The electronic device obtains the audio data to be processed corresponding to the first video frame according to the collected sound signal. The electronic device performs timbre correction processing, forms a stereo sound beam, and performs gain control processing on the audio data to be processed according to the first shooting view. The electronic device records the audio corresponding to the first video frame according to the audio data after the gain control processing.

[0037] For example, the first shooting perspective can be a wide-angle perspective. After the electronic device performs audio processing such as tone correction processing, forming a stereo sound beam, and gain control processing on the audio data to be processed in sequence according to the wide-angle perspective, it can record the corresponding audio.

[0038] In another possible design, the shooting perspective corresponding to the first video frame in the multi-channel video frames is the second shooting perspective. The electronic device records the audio corresponding to the first video frame according to the shooting perspective corresponding to the first video frame, including: the electronic device obtains the audio data to be processed corresponding to the first video frame according to the collected sound signal. The electronic device performs tone correction processing, forms a stereo / mono sound beam, ambient noise control processing, and gain control processing on the audio data to be processed according to the second shooting perspective. The electronic device records the audio corresponding to the first video frame according to the audio data after the gain control processing.

[0039] For example, the first shooting perspective can be a zoom perspective. After the electronic device performs audio processing such as tone correction processing, forming a stereo / mono sound beam, ambient noise control processing, and gain control processing on the audio data to be processed in sequence according to the wide-angle perspective, it can record the corresponding audio.

[0040] In another possible design, the shooting perspective corresponding to the first video frame in the multi-channel video frames is the third shooting perspective. The electronic device records the audio corresponding to the first video frame according to the shooting perspective corresponding to the first video frame, including: the electronic device obtains the audio data to be processed corresponding to the first video frame according to the collected sound signal. The electronic device performs tone correction processing, forms a stereo / mono sound beam, human voice enhancement processing, and gain control processing on the audio data to be processed according to the third shooting perspective. The electronic device records the audio corresponding to the first video frame according to the audio data after the gain control processing.

[0041] For example, the first shooting perspective can be a front-facing perspective. After the electronic device performs audio processing such as tone correction processing, forming a stereo / mono sound beam, human voice enhancement processing, and gain control processing on the audio data to be processed in sequence according to the front-facing perspective, it can record the corresponding audio.

[0042] In another possible design, the electronic device includes a directional microphone that points in the direction of the rear camera. The shooting angle corresponding to the first video frame in the multiple video frames is the second shooting angle. The electronic device records the audio corresponding to the first video frame according to the shooting angle corresponding to the first video frame, including: the electronic device obtains the audio data to be processed corresponding to the first video frame according to the collected sound signal; the electronic device performs timbre correction processing, ambient noise control processing, and gain control processing on the audio data to be processed according to the second shooting angle; the electronic device records the audio corresponding to the first video frame according to the audio data after the gain control processing.

[0043] In another possible design, the first shooting angle can be a zoom angle. The electronic device can perform audio processing such as timbre correction processing, ambient noise control processing, and gain control processing on the audio data to be processed in sequence according to the wide-angle view, and then record the corresponding audio.

[0044] In another possible design, the electronic device includes a directional microphone that points in the direction of the front camera. The shooting angle corresponding to the first video frame in the multiple video frames is the third shooting angle. The electronic device records the audio corresponding to the first video frame according to the shooting angle corresponding to the first video frame, including: the electronic device obtains the audio data to be processed corresponding to the first video frame according to the collected sound signal; the electronic device performs timbre correction processing, human voice enhancement processing, and gain control processing on the audio data to be processed according to the third shooting angle; the electronic device records the audio corresponding to the first video frame according to the audio data after the gain control processing.

[0045] For example, the first shooting angle can be a front view angle. The electronic device can perform audio processing such as timbre correction processing, human voice enhancement processing, and gain control processing on the audio data to be processed in sequence according to the front view angle, and then record the corresponding audio.

[0046] On the other hand, an embodiment of the present application provides an electronic device, including: multiple microphones for collecting sound signals; a screen for displaying an interface; an audio playback component for playing audio; one or more processors; a memory; and one or more computer programs, where one or more computer programs are stored in the memory, and one or more computer programs include instructions that, when executed by the electronic device, cause the electronic device to perform the following steps: after detecting an operation by the user to open the camera, display a shooting preview interface; enter a multi-channel video recording mode; after detecting a shooting operation by the user, display a shooting interface, where the shooting interface includes multiple video frames; record multiple videos; and record the audio corresponding to each video frame in the multiple video frames according to the shooting angle corresponding to each video frame.

[0047] In this solution, in the multi-channel video recording mode, the electronic device can collect sound signals through a microphone, capture video images through a camera, display them on the screen interface, and play audio through an audio playback component. Moreover, while the electronic device is recording multi-channel video images, it records multi-channel audio corresponding to each multi-channel video image according to the shooting angle corresponding to each video image, so that the recorded audio can be matched with different video images and shooting angles. In this way, when the video is played back, the user can select to play the audio that matches the video image and shooting angle they are interested in, thereby improving the user's audio experience.

[0048] In a possible design, there are at least three microphones, and at least one microphone is respectively arranged on the top, bottom, and back of the electronic device.

[0049] In this way, the electronic device can collect sound signals from all directions through these at least three microphones, so as to subsequently obtain the sound signals within each sound pickup range.

[0050] For example, the microphone is an internal component or an external accessory.

[0051] Among them, when the microphone is an external accessory, the microphone can be a directional microphone. At this time, in the zoom view or the front view, it is not necessary to form a mono / stereo sound beam during the audio processing.

[0052] In another possible design, each video image corresponds to a shooting angle.

[0053] In another possible design, the shooting angle corresponding to each video image is variable.

[0054] In another possible design, when the instruction is executed by the electronic device, it also causes the electronic device to perform the following steps: when respectively recording the audio corresponding to each video image and detecting the user's operation of stopping shooting, generate multi-channel video recordings, and the multi-channel video recordings also include the audio corresponding to each video image. After detecting the user's operation of playing the multi-channel video recordings, display a video playback interface, and the video playback interface includes multi-channel video images.

[0055] In another possible design, when the instruction is executed by the electronic device, it also causes the electronic device to perform the following steps: after detecting the user's operation of playing the multi-channel video recordings and detecting the user's operation of playing the audio corresponding to the first video image, play the audio corresponding to the first video image, and the first video image is one of the multi-channel video images. If other audio in the multi-channel audio is played before playing the audio corresponding to the first video image, stop playing the other audio.

[0056] In another possible design, the video playback interface includes audio playback controls corresponding to each video frame respectively. Detecting an operation by the user to play the audio corresponding to the first video frame includes: detecting an operation by the user to click the audio playback control corresponding to the first video frame.

[0057] In another possible design, when the instruction is executed by the electronic device, it also causes the electronic device to perform the following steps: Before detecting an operation by the user to play the audio corresponding to the first video frame, by default, play the audio corresponding to the second video frame among the multiple video frames.

[0058] In another possible design, the shooting angle corresponding to the second video frame is a preset shooting angle.

[0059] In another possible design, when the instruction is executed by the electronic device, it also causes the electronic device to perform the following steps: When the video playback interface is displayed and an operation by the user to play the first video frame among the multiple video frames is detected, display the first video frame on the video playback interface and stop displaying other video frames other than the first video frame. Automatically play the audio corresponding to the first video frame.

[0060] In another possible design, when the instruction is executed by the electronic device, it also causes the electronic device to perform the following steps: Before the shooting interface is displayed, determine the target shooting mode, where the target shooting mode is used to represent the number of video frames to be recorded. Display the shooting interface, including: Display the shooting interface according to the number of video frames to be recorded corresponding to the target shooting mode.

[0061] In another possible design, the target shooting mode is also used to represent the correspondence between each video frame and the shooting angle; Recording the audio corresponding to each video frame respectively according to the shooting angle corresponding to each video frame among the multiple video frames includes: Determining the shooting angle corresponding to each video frame according to the correspondence between each video frame and the shooting angle corresponding to the target shooting mode; Recording the audio corresponding to each video frame respectively according to the shooting angle corresponding to each video frame.

[0062] In another possible design, the target shooting mode includes: a wide-angle view zoom view combination mode, a wide-angle view front view combination mode, a zoom view front view combination mode, or a wide-angle view zoom view front view combination mode, where the zoom ratio corresponding to the wide-angle view is less than or equal to a preset value, and the zoom ratio corresponding to the zoom view is greater than the preset value.

[0063] In another possible design, the target shooting mode is a preset shooting mode.

[0064] In another possible design, according to the shooting angles corresponding to each video frame in the multi-channel video frames, the audio corresponding to each video frame is recorded respectively, including: determining the shooting angle corresponding to each video frame according to the relationship between the zoom ratio corresponding to each video frame and a preset value and the front / rear characteristics. According to the shooting angles corresponding to each video frame, the audio corresponding to each video frame is recorded respectively.

[0065] In another possible design, when the instruction is executed by the electronic device, the electronic device is further caused to execute the following steps: after determining the target shooting mode, detecting an operation by the user to switch the target shooting mode; switching the target shooting mode.

[0066] In another possible design, the shooting angle corresponding to the first video frame in the multi-channel video frames is the first shooting angle. Recording the audio corresponding to the first video frame according to the shooting angle corresponding to the first video frame, including: obtaining the audio data to be processed corresponding to the first video frame according to the collected sound signal; performing timbre correction processing, forming a stereo sound beam and gain control processing on the audio data to be processed according to the first shooting angle; recording the audio corresponding to the first video frame according to the audio data after the gain control processing.

[0067] In another possible design, the shooting angle corresponding to the first video frame in the multi-channel video frames is the second shooting angle. Recording the audio corresponding to the first video frame according to the shooting angle corresponding to the first video frame, including: obtaining the audio data to be processed corresponding to the first video frame according to the collected sound signal; performing timbre correction processing, forming a stereo / mono sound beam, ambient noise control processing and gain control processing on the audio data to be processed according to the second shooting angle; recording the audio corresponding to the first video frame according to the audio data after the gain control processing.

[0068] In another possible design, the shooting angle corresponding to the first video frame in the multi-channel video frames is the third shooting angle. Recording the audio corresponding to the first video frame according to the shooting angle corresponding to the first video frame, including: obtaining the audio data to be processed corresponding to the first video frame according to the collected sound signal; performing timbre correction processing, forming a stereo / mono sound beam, human voice enhancement processing and gain control processing on the audio data to be processed according to the third shooting angle; recording the audio corresponding to the first video frame according to the audio data after the gain control processing.

[0069] In another possible design, the electronic device includes a directional microphone that points in the direction of the rear camera. The shooting angle corresponding to the first video frame among the multiple video frames is the second shooting angle. Recording the audio corresponding to the first video frame according to the shooting angle corresponding to the first video frame includes: obtaining the audio data to be processed corresponding to the first video frame according to the collected sound signal; performing timbre correction processing, ambient noise control processing, and gain control processing on the audio data to be processed according to the second shooting angle; and recording the audio corresponding to the first video frame according to the audio data after the gain control processing.

[0070] In another possible design, the electronic device includes a directional microphone that points in the direction of the front camera. The shooting angle corresponding to the first video frame among the multiple video frames is the third shooting angle. Recording the audio corresponding to the first video frame according to the shooting angle corresponding to the first video frame includes: obtaining the audio data to be processed corresponding to the first video frame according to the collected sound signal; performing timbre correction processing, human voice enhancement processing, and gain control processing on the audio data to be processed according to the third shooting angle; and recording the audio corresponding to the first video frame according to the audio data after the gain control processing.

[0071] On the other hand, an embodiment of the present application provides an audio processing device, which is included in the electronic device. The device has the function of implementing the behavior of the electronic device in any of the methods in the above aspects and possible designs, so that the electronic device executes the audio processing method executed by the electronic device in any of the possible designs in the above aspects. This function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes at least one module or unit corresponding to the above function. For example, the device may include an audio processing module, a microphone, or a camera, etc.

[0072] In yet another aspect, an embodiment of the present application provides an electronic device, including: one or more processors; and a memory in which code is stored. When the code is executed by the electronic device, the electronic device executes the audio processing method executed by the electronic device in any of the possible designs in the above aspects.

[0073] On the other hand, an embodiment of the present application provides a computer-readable storage medium, including computer instructions, which, when running on the electronic device, cause the electronic device to execute the audio processing method in any of the possible designs in the above aspects.

[0074] In yet another aspect, an embodiment of the present application provides a computer program product, which, when running on a computer, causes the computer to execute the audio processing method executed by the electronic device in any of the possible designs in the above aspects.

[0075] On the other hand, an embodiment of the present application provides a chip system, which is applied to an electronic device. The chip system includes one or more interface circuits and one or more processors; the interface circuits and the processors are interconnected by lines; the interface circuits are configured to receive signals from a memory of the electronic device and send the signals to the processors, and the signals include computer instructions stored in the memory; when the processors execute the computer instructions, the electronic device is caused to execute the audio processing method in any possible design of the above aspects.

[0076] For the beneficial effects corresponding to the above other aspects, reference may be made to the description of the beneficial effects of the method aspect, which will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0077] Figure 1A It is a schematic structural diagram of an electronic device provided by an embodiment of the present application;

[0078] Figure 1B It is a schematic layout diagram of a microphone provided by an embodiment of the present application;

[0079] Figure 1C It is another schematic layout diagram of a microphone provided by an embodiment of the present application;

[0080] Figure 2 It is a schematic software architecture diagram of an electronic device provided by an embodiment of the present application;

[0081] Figure 3 It is a schematic audio processing flowchart provided by an embodiment of the present application;

[0082] Figure 4 It is a set of schematic interface diagrams provided by an embodiment of the present application;

[0083] Figure 5 It is another set of schematic interface diagrams provided by an embodiment of the present application;

[0084] Figure 6 It is another set of schematic interface diagrams provided by an embodiment of the present application;

[0085] Figure 7 It is another set of schematic interface diagrams provided by an embodiment of the present application;

[0086] Figure 8 It is a schematic audio recording flowchart provided by an embodiment of the present application;

[0087] Figure 9 It is a schematic diagram of an audio processing scheme and an audio beam provided by an embodiment of the present application;

[0088] Figure 10 It is another schematic diagram of an audio processing scheme and an audio beam provided by an embodiment of the present application;

[0089] Figure 11 Another audio processing solution and an audio beam schematic diagram provided by an embodiment of the present application;

[0090] Figure 12 Another set of interface schematic diagrams provided by an embodiment of the present application;

[0091] Figure 13 Another set of interface schematic diagrams provided by an embodiment of the present application;

[0092] Figure 14 Another set of interface schematic diagrams provided by an embodiment of the present application;

[0093] Figure 15 Another audio beam schematic diagram provided by an embodiment of the present application;

[0094] Figure 16 Another audio processing solution schematic diagram provided by an embodiment of the present application;

[0095] Figure 17 Another audio processing solution schematic diagram provided by an embodiment of the present application;

[0096] Figure 18 Another audio processing solution schematic diagram provided by an embodiment of the present application. Detailed implementation manners

[0097] Next, the technical solutions in the embodiments of the present application will be described with reference to the accompanying drawings in the embodiments of the present application. Among them, in the description of the embodiments of the present application, unless otherwise specified, " / " means "or". For example, A / B may represent A or B; herein, "and / or" is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, in the description of the embodiments of the present application, "a plurality of" means two or more than two.

[0098] Hereinafter, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of this embodiment, unless otherwise specified, the meaning of "a plurality of" is two or more than two.

[0099] In the multi-channel video recording (or multi-scene video recording) mode, the electronic device can record multiple video frames during the video recording process, that is, record video frames of multiple lines.

[0100] In the embodiments of the present application, in some cases, the video images of different lines correspond to different shooting perspectives, and the shooting perspectives corresponding to each line of video images are fixed and unchanged during the current video recording process. The multi-channel video recording in this case can also be referred to as multi-perspective video recording. Hereinafter, this case is referred to as Case 1.

[0101] Among them, the shooting perspective can be divided according to whether the object to be shot is a front object or a rear object, and / or the magnitude of the zoom ratio. For example, in the embodiments of the present application, the shooting perspective can include a wide-angle perspective, a zoom perspective, a front perspective, etc. Among them, the wide-angle perspective and the zoom perspective belong to the rear perspective, and can also be respectively referred to as the rear wide-angle perspective and the rear zoom perspective. The wide-angle perspective can be the shooting perspective corresponding to a scene where the zoom ratio is less than or equal to a preset value K. For example, the preset value K can be 2, 1.5, or 1, etc. The zoom perspective can be the shooting perspective corresponding to a scene where the zoom ratio is greater than the preset value K. The front perspective is the shooting perspective corresponding to a front shooting scene such as a self-portrait.

[0102] Exemplarily, the shooting modes in the multi-channel video recording mode can be referred to Table 1. The shooting perspective corresponding to this shooting mode can be any combination of multiple shooting perspectives among the wide-angle perspective, the zoom perspective, or the front perspective. As shown in Table 1, each shooting mode can include multiple lines, and each line can correspond to a line of video image and a shooting perspective. The video images recorded by using the shooting modes in the multi-channel video recording mode can include any combination of multiple lines of video images among the video images under the wide-angle perspective, the video images under the zoom perspective, or the video images under the front perspective. The shooting modes in Table 1 represent the number of lines of the video images to be recorded and the corresponding relationship between each line of video image and the shooting perspective, and each line of video image corresponds to a shooting perspective.

[0103] Table 1

[0104]

[0105] In other cases, the type of the shooting perspective corresponding to the video images of different lines is variable. Hereinafter, this case is referred to as Case 2. Exemplarily, in this case, the shooting modes in the multi-channel video recording mode can be referred to Table 2. For example, the shooting mode a in Table 2 includes Line 1 and Line 2. The shooting perspective corresponding to Line 1 can be switched between the wide-angle perspective and the zoom perspective. Line 2 corresponds to the front perspective. The shooting modes in Table 2 represent the number of lines of the video images to be recorded, and the shooting perspective corresponding to each line of video image is variable.

[0106] Table 2

[0107]

[0108] The embodiment of the present application provides an audio processing method in a multi-channel video recording mode. In the multi-channel video recording mode, an electronic device can record multiple video images and multiple audio channels simultaneously. When playing back the multi-channel video recordings (hereinafter referred to as video playback), the electronic device can play different audio, thereby improving the user's audio experience in multi-channel video recording.

[0109] In some embodiments of the present application, in the multi-channel video recording mode, while the electronic device records the video images corresponding to multiple shooting perspectives, it can also record the audio corresponding to different shooting perspectives and video images. During video playback, the electronic device can play the audio that matches different shooting perspectives and video images, so that the played audio content corresponds to the shooting perspective and video image that the user is concerned about, improving the user's audio experience in multi-channel video recording.

[0110] For example, the audio content corresponding to the wide-angle perspective may include panoramic sounds in all directions around (i.e., sounds in a 360-degree range around), and the audio content corresponding to the zoom perspective mainly includes the sounds within the zoom range. Here, the zoom range refers to the shooting range corresponding to the current zoom ratio under the zoom perspective. The audio content corresponding to the front view mainly includes the human voices within the front view range.

[0111] For example, in Case 1, while the electronic device records each video image, it can also record the audio according to the shooting perspective adopted by each video image. During video playback, the electronic device can play the audio corresponding to the shooting perspective and video image that the user is concerned about.

[0112] Exemplarily, taking the video recording mode 4 shown in Table 1 as an example, the electronic device can record the video image under the wide-angle perspective corresponding to Line 1 and record the audio corresponding to Line 1 according to the wide-angle perspective; the electronic device can record the video image under the zoom perspective corresponding to Line 2 and record the audio corresponding to Line 2 according to the zoom perspective; the electronic device can record the video image under the front view corresponding to Line 3 and record the audio corresponding to Line 3 according to the front view.

[0113] In this way, during video playback, if the user is concerned about the video image of the wide-angle perspective and Line 1, the audio played by the electronic device can be the panoramic sound corresponding to the wide-angle perspective; if the user is concerned about the video image of the zoom perspective and Line 2, the audio played by the electronic device can be the sound within the zoom range; if the user is concerned about the video image of the front view and Line 3, the audio played by the electronic device can be the human voices within the front view range. Therefore, the audio played by the electronic device can be matched in real time with the shooting perspective and video image that the user is concerned about, thereby improving the user's audio experience.

[0114] For another example, in Case 2, while the electronic device records the video images corresponding to each line, it can also record the corresponding audio according to the changing shooting perspectives in each line. During video playback, the electronic device can play the audio that is real-time matched with the shooting perspective and the video image in Line 1. Exemplarily, taking the shooting mode a shown in Table 2 as an example, the electronic device can record the video image under the wide-angle perspective corresponding to Line 1 and record the audio corresponding to the wide-angle perspective of Line 1; then, the electronic device switches to record the video image under the zoom perspective corresponding to Line 1 and records the audio corresponding to the zoom perspective of Line 1; after that, the electronic device switches to record the video image under the wide-angle perspective corresponding to Line 1 and records the audio corresponding to the wide-angle perspective of Line 1. Moreover, while the electronic device records the video image and audio under the wide-angle perspective corresponding to Line 1, it can also record the video image and audio under the front-facing perspective corresponding to Line 2.

[0115] In this way, during video playback, when the user focuses on the video image of Line 1, the electronic device can play the panoramic sound corresponding to the wide-angle perspective, the sound within the zoom range corresponding to the zoom perspective, and then the panoramic sound corresponding to the wide-angle perspective as the shooting perspective of Line 1 changes, so that the played audio is real-time matched with the shooting perspective and the video image, improving the user's audio experience. When the user focuses on the video image of Line 2, the electronic device can play the human voice within the front-facing range corresponding to the front-facing perspective.

[0116] In the existing multi-channel video recording mode, the electronic device only records one channel of audio and can only play this one channel of audio during video playback. The audio content cannot be matched with different shooting perspectives and video images, nor can it be matched with the shooting perspective and video image that the user focuses on, resulting in a poor user audio experience.

[0117] The audio processing method provided by the embodiments of the present application can be applied to an electronic device. For example, the electronic device can specifically be a mobile phone, a tablet computer, a wearable device, a vehicle-mounted device, an augmented reality (AR) / virtual reality (VR) device, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), or a dedicated camera (such as a single-lens reflex camera, a compact camera), etc. The embodiments of the present application do not impose any restrictions on the specific type of the electronic device.

[0118] Exemplarily, Figure 1AShows a schematic structural diagram of the electronic device 100. The electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone interface 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. Among them, the sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.

[0119] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.

[0120] Among them, the controller may be the nerve center and command center of the electronic device 100. The controller may generate operation control signals according to the instruction operation code and timing signal to complete the control of fetching instructions and executing instructions.

[0121] A memory may also be provided in the processor 110 for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can hold the instructions or data that the processor 110 has just used or recycled. If the processor 110 needs to use the instruction or data again, it can be directly called from the memory. This avoids repeated accesses and reduces the waiting time of the processor 110, thus improving the efficiency of the system.

[0122] The electronic device 100 implements the display function through a GPU, a display screen 194, an application processor, etc. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 110 may include one or more GPUs, which execute program instructions to generate or change display information.

[0123] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a MiniLED, a MicroLED, a Micro-OLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the electronic device 100 may include 1 or N display screens 194, where N is a positive integer greater than 1.

[0124] In the embodiments of the present application, the display screen 194 can display the shooting preview interface, the video recording preview interface, and the shooting interface in the multi-channel video recording mode, and can also display the video playback interface, etc. during video playback.

[0125] The electronic device 100 can implement the shooting function through an ISP, a camera 193, a video codec, a GPU, a display screen 194, and an application processor, etc.

[0126] The ISP is used to process the data fed back by the camera 193. For example, when taking a photo, the shutter is opened, and light passes through the lens and is transmitted to the camera's photosensitive element. The light signal is converted into an electrical signal, and the camera's photosensitive element transmits the electrical signal to the ISP for processing and converts it into an image visible to the naked eye. The ISP can also perform algorithm optimization on the noise, brightness, and skin color of the image. The ISP can also optimize parameters such as exposure and color temperature of the shooting scene. In some embodiments, the ISP can be provided in the camera 193. For example, in the embodiments of the present application, the ISP can control the photosensitive element to perform exposure and take pictures according to the shooting parameters.

[0127] The camera 193 is used to capture static images or videos. An object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal and then transmits the electrical signal to the ISP to be converted into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in a standard RGB, YUV, etc. format. In some embodiments, the electronic device 100 may include one or N cameras 193, where N is a positive integer greater than 1. Among them, the camera 193 can be located in the edge area of the electronic device, and can be an under-screen camera or a retractable camera. The camera 193 can include a rear camera and can also include a front camera. The specific position and form of the camera 193 in the embodiments of the present application are not limited. The electronic device 100 can include cameras with one or more focal lengths. For example, cameras with different focal lengths can include a telephoto camera, a wide-angle camera, an ultra-wide-angle camera, or a panoramic camera, etc.

[0128] In the embodiments of the present application, in the multi-channel video recording mode, different cameras can be used to collect video images from different perspectives. For example, a wide-angle camera, an ultra-wide-angle camera, or a panoramic camera can collect video images from a wide-angle perspective, a telephoto camera or a wide-angle camera can collect video images from a zoom perspective, and a front camera can collect video images from a front perspective.

[0129] The digital signal processor is used to process digital signals. In addition to processing digital image signals, it can also process other digital signals. For example, when the electronic device 100 selects a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy, etc.

[0130] A video codec is used to compress or decompress digital video. The electronic device 100 can support one or more video codecs. In this way, the electronic device 100 can play or record videos in multiple coding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.

[0131] The NPU is a neural-network (NN) computing processor. By drawing on the structure of a biological neural network, such as the transmission pattern between human brain neurons, it can quickly process input information and can also continuously self-learn. Through the NPU, applications such as intelligent cognition of the electronic device 100 can be realized, such as image recognition, face recognition, speech recognition, text understanding, etc.

[0132] The internal memory 121 can be used to store computer-executable program code, and the executable program code includes instructions. The processor 110 executes various functional applications and data processing of the electronic device 100 by running the instructions stored in the internal memory 121. The internal memory 121 can include a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.). The data storage area can store data created during the use of the electronic device 100 (such as images, audio data, phone books, etc. collected by the electronic device 100).

[0133] In an embodiment of the present application, the processor 110 can record video images from multiple shooting perspectives in a multi-channel recording mode and record the audio corresponding to different shooting perspectives by running the instructions stored in the internal memory 121, so that when the video is played back, the audio corresponding to different shooting perspectives and video images can be played, making the played audio match the shooting perspective and video image that the user is concerned about.

[0134] The electronic device 100 can implement audio functions through the audio module 170, the speaker 170A, the receiver 170B, multiple microphones 170C, the headphone jack 170D, and the application processor, etc. Such as music playback, recording, etc.

[0135] The audio module 170 is used to convert digital audio data into an analog audio electrical signal for output, and is also used to convert the input analog audio electrical signal into digital audio data. For example, the audio module 170 is used to convert the analog audio electrical signal output by the microphone 170C into digital audio data.

[0136] Among them, the audio module 170 may further include an audio processing module. The audio processing module is used to perform audio processing on digital audio data in the multi-channel video recording mode, so as to generate audio corresponding to different shooting perspectives. For example, for the wide-angle perspective, the audio processing module may include a timbre correction module, a stereo beamforming module, and a gain control module, etc. For the zoom perspective, the audio processing module may include a timbre correction module, a stereo / mono beamforming module, an ambient noise control module, and a gain control module, etc. For the front perspective, the audio processing module may include a timbre correction module, a stereo / mono beamforming module, a voice enhancement module, and a gain control module, etc.

[0137] The audio module 170 may also be used to encode and decode audio data.

[0138] In some embodiments, the audio module 170 may be disposed in the processor 110, or some functional modules of the audio module 170 may be disposed in the processor 110.

[0139] The speaker 170A, also known as the "loudspeaker", is used to convert the analog audio electrical signal into a sound signal. The electronic device 100 can listen to music or make hands-free calls through the speaker 170A. In the embodiments of the present application, when playing back the video of multi-channel video recording, the speaker 170A can be used to play the audio corresponding to different shooting perspectives and video pictures.

[0140] The receiver 170B, also known as the "earpiece", is used to convert the analog audio electrical signal into a sound signal. When the electronic device 100 answers a call or a voice message, the voice can be received by bringing the receiver 170B close to the human ear.

[0141] The microphone 170C, also known as the "microphone" or "transmitter", is used to convert the sound signal into an analog audio electrical signal. When making a call or sending a voice message, the user can speak by bringing the mouth close to the microphone 170C to input the sound signal into the microphone 170C. In the embodiments of the present application, the electronic device 100 may include at least three microphones 170C, which can realize the function of collecting sound signals in all directions, converting the collected sound signals into analog audio electrical signals, and can also realize functions such as noise reduction, sound source identification, or directional recording.

[0142] Exemplarily, the layout of the microphones 170C on the electronic device 100 may be as Figure 1B shown. The electronic device 100 may include a microphone 1 disposed at the bottom, a microphone 2 disposed at the top, and a microphone 3 disposed on the back. The combination of microphones 1-3 can collect sound signals in all directions around the electronic device 100.

[0143] Exemplarily again, refer toFigure 1C , the electronic device 100 may also include a greater number of microphones 170C. As Figure 1C shown, the electronic device 100 may include a microphone 1 and a microphone 4 provided at the bottom, a microphone 2 provided at the top, a microphone 3 provided at the back, and a microphone 5 provided in front of the screen. These microphones can collect sound signals from all directions around the electronic device 100. The screen is a display screen 194 or a touch screen.

[0144] It should be noted that the microphone 170C may be an internal component of the electronic device 100 or an external accessory of the electronic device 100. For example, the electronic device 100 may include a microphone 1 provided at the bottom, a microphone 2 provided at the top, and an external accessory. Exemplarily, the external accessory may be a miniature microphone connected to the electronic device 100 (wired connection or wireless connection), or a headset with a microphone (such as a wired headset or a TWS headset, etc.).

[0145] In some embodiments, the microphone 170C may be a directional microphone that can collect sound signals in a specific direction.

[0146] A distance sensor 180F is used to measure distance. The electronic device 100 can measure distance through infrared or laser. In some embodiments, in a shooting scene, the electronic device 100 can use the distance sensor 180F to measure distance to achieve rapid focusing.

[0147] A touch sensor 180K, also known as a "touch panel". The touch sensor 180K may be disposed on the display screen 194, and the touch sensor 180K and the display screen 194 form a touch screen, also known as a "touch screen". The touch sensor 180K is used to detect touch operations acting thereon or nearby. The touch sensor can transmit the detected touch operation to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through the display screen 194. In some other embodiments, the touch sensor 180K may also be disposed on the surface of the electronic device 100, at a different position from the display screen 194.

[0148] For example, in the embodiments of the present application, the electronic device 100 can detect an operation for the user to indicate the start and / or stop of shooting through the touch sensor 180K.

[0149] It can be understood that the structure illustrated in the embodiments of the present application does not constitute a specific limitation on the electronic device 100. In some other embodiments of the present application, the electronic device 100 may include more or fewer components than shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The illustrated components can be implemented in hardware, software, or a combination of software and hardware.

[0150] In an embodiment of the present application, in the multi-channel video recording mode, the display screen 194 can display the shooting preview interface, video recording preview interface, and shooting interface during video recording. The camera 193 can be used to collect multi-channel video images. Multiple microphones 170C can be used to collect sound signals and generate analog audio electrical signals. The audio module 170 can generate digital audio data from the analog audio electrical signals and generate audio corresponding to different shooting perspectives and video images. During video playback, the display screen 194 can display the video playback interface. By running the instructions stored in the internal memory 121, the processor 110 can, according to the user's selection, control the speaker 170A to play the audio corresponding to the shooting perspective and video image that the user is concerned about, thereby enhancing the user's audio experience of multi-channel video recording.

[0151] The software system of the electronic device 100 can adopt a layered architecture, event-driven architecture, microkernel architecture, microservices architecture, or cloud architecture. In this embodiment of the present application, taking the Android system with a layered architecture as an example, the software structure of the electronic device 100 is exemplarily described.

[0152] Figure 2 It is the software structure block diagram of the electronic device 100 in this embodiment of the present application. The layered architecture divides the software into several layers, and each layer has a clear role and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers, from top to bottom are the application layer, application framework layer, Android runtime and system libraries, hardware abstraction layer (HAL), and kernel layer. The application layer can include a series of application packages.

[0153] As Figure 2 shown, the application packages can include applications such as camera, gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, video, short message, etc.

[0154] The application framework layer provides application programming interfaces (APIs) and programming frameworks for the applications in the application layer. The application framework layer includes some predefined functions.

[0155] As Figure 2 shown, the application framework layer can include a window manager, content provider, view system, telephone manager, resource manager, notification manager, etc.

[0156] The window manager is used to manage window programs. The window manager can obtain the display screen size, determine whether there is a status bar, lock the screen, capture the screen, etc.

[0157] The content provider is used to store and retrieve data and make this data accessible to applications. The data can include videos, images, audio, incoming and outgoing calls, browsing history and bookmarks, phone books, etc.

[0158] The view system includes visual controls, such as controls for displaying text, controls for displaying pictures, etc. The view system can be used to build applications. The display interface can be composed of one or more views. For example, a display interface including a text message notification icon can include a view for displaying text and a view for displaying pictures.

[0159] The phone manager is used to provide the communication functions of the electronic device 100. For example, the management of call states (including answering, hanging up, etc.).

[0160] The resource manager provides various resources for applications, such as localized strings, icons, pictures, layout files, video files, etc.

[0161] The notification manager enables applications to display notification information in the status bar. It can be used to convey informative messages, which can disappear automatically after a short stay without user interaction. For example, the notification manager is used to inform that the download is complete, message reminders, etc. The notification manager can also be a notification that appears in the system top status bar in the form of a chart or scroll bar text, such as notifications of background running applications, and can also be a notification that appears in the form of a dialogue window on the screen. For example, prompt text information in the status bar, emit a prompt sound, the electronic device vibrates, the indicator light flashes, etc.

[0162] The Android Runtime includes core libraries and a virtual machine. The Android runtime is responsible for the scheduling and management of the Android system.

[0163] The core libraries contain two parts: one part is the functional functions that need to be called by the Java language, and the other part is the core libraries of Android.

[0164] The application layer and the application framework layer run in the virtual machine. The virtual machine executes the Java files of the application layer and the application framework layer as binary files. The virtual machine is used to perform functions such as the management of object life cycles, stack management, thread management, security and exception management, and garbage collection.

[0165] The system libraries can include multiple functional modules. For example: surface manager, Media Libraries, 3D graphics processing libraries (such as: OpenGL ES), 2D graphics engines (such as: SGL), etc.

[0166] The surface manager is used to manage the display subsystem and provides the fusion of 2D and 3D layers for multiple applications.

[0167] The media library supports the playback and recording of multiple common audio and video formats, as well as static image files, etc. The media library can support multiple audio and video coding formats, such as: MPEG4, H.264, MP3, AAC, AMR, JPG, PNG, etc.

[0168] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, synthesis, and layer processing, etc.

[0169] The 2D graphics engine is the drawing engine for 2D drawing.

[0170] The HAL layer is the interface layer between the operating system kernel and the hardware circuit, which can abstract the hardware. The HAL layer includes an audio processing module. The audio processing module can be used to process the analog audio electrical signals obtained by the microphone according to the shooting perspective to generate the audio corresponding to different shooting perspectives and video frames. For example, for the wide-angle perspective, the audio processing module can include a timbre correction module, a stereo beamforming module, and a gain control module, etc. For the zoom perspective, the audio processing module can include a timbre correction module, a stereo / mono beamforming module, an ambient noise control module, and a gain control module, etc. For the front view perspective, the audio processing module can include a timbre correction module, a stereo / mono beam presentation module, a voice enhancement module, and a gain control module, etc.

[0171] The kernel layer is the layer between the hardware layer and the above software layers. The kernel layer at least includes a display driver, a camera driver, an audio driver, and a sensor driver. Among them, the hardware layer can include a camera, a display screen, a microphone, a processor, and a memory, etc.

[0172] In the embodiment of the present application, in the multi-channel video recording mode, the display screen in the hardware layer can display the shooting preview interface, the video recording preview interface, and the shooting interface during video recording. The camera in the hardware layer can be used to collect multi-channel video frames. The microphone in the hardware layer can be used to collect sound signals and generate analog audio electrical signals. The audio processing module in the HAL layer can be used to process the digital audio data converted from the analog audio electrical signals, so as to generate the audio corresponding to different shooting perspectives and video frames. During video playback, the display screen can display the video playback interface, and the speaker can play the audio corresponding to the shooting perspective and video frame that the user is concerned about, so as to improve the audio experience of the user's multi-channel video recording.

[0173] The following will take the electronic device as having Figure 1ATaking a mobile phone with the structure shown as an example, where the mobile phone has three or more microphones, the audio processing method provided by the embodiments of the present application will be described. Refer to Figure 3 , the method may include:

[0174] 301. After the mobile phone detects that the user opens the camera, it displays a shooting preview interface.

[0175] After the mobile phone detects that the user opens the camera (hereinafter also referred to as the camera application), it starts the camera application and displays a shooting preview interface. Among them, there are various operations for the user to open the camera. Exemplarily, after the mobile phone detects that the user clicks Figure 4 the operation of the camera icon 401 shown in (a) of Figure 4 it starts the camera application and displays the

[0176] shooting preview interface shown in (b) of Figure 4 .

[0177] 302. The mobile phone enters the multi-channel video recording mode.

[0178] After the mobile phone starts the camera application and displays the shooting preview interface, it can enter the multi-channel video recording mode, thereby displaying a video recording preview interface.

[0179] Among them, there are various ways for the mobile phone to enter the multi-channel video recording mode. For example, in some implementation manners, after the mobile phone starts the camera application, it defaults to enter a non-multi-channel video recording mode such as a photo-taking mode or a video recording mode (i.e., a single-channel video recording mode). After the mobile phone detects a preset operation 1 indicating to enter the multi-channel video recording mode by the user, it enters the multi-channel video recording mode. Exemplarily, this preset operation 1 may be the operation of the user clicking Figure 4 the control 402 shown in (b) of

[0180] to enter the multi-channel video recording mode. In some other implementation manners, in a non-multi-channel video recording mode such as a photo-taking mode or a video recording mode (i.e., a single-channel video recording mode), after the mobile phone detects that the user draws a preset trajectory 1 (such as an "M" trajectory) on the touch screen, it enters the multi-channel video recording mode.

[0181] In some other implementation manners, in the video recording mode (i.e., the single-channel video recording mode), the mobile phone can prompt the user on the video recording preview interface whether to enter the multi-channel video recording mode. The mobile phone can enter the multi-channel video recording mode according to the user's instruction.

[0182] In some other implementation manners, after the mobile phone starts the camera application, it defaults to enter the multi-channel video recording mode.

[0183] In some other embodiments, the mobile phone may directly execute step 302 without performing step 301, thereby directly entering the multi-channel video recording mode. For example, when the screen is on and the desktop is displayed, or when the screen is black, if the mobile phone detects that the user has drawn a preset trajectory (such as the "CM" trajectory) on the touch screen, the camera application is started and the multi-channel video recording mode is directly entered.

[0184] As described above, in Case 1, in the multi-channel video recording mode, the shooting angle corresponding to each video is fixed. First, an example of Case 1 will be described below.

[0185] 303. The mobile phone determines the target shooting mode and displays a video recording preview interface according to the target shooting mode.

[0186] There can be multiple shooting modes in the multi-channel video recording mode, and each shooting mode can correspond to a different combination of shooting angles. For example, examples of shooting modes in the multi-channel video recording mode can be seen in Table 1. The mobile phone can determine the target shooting mode in the multi-channel video recording mode, and thus display a video recording preview interface according to the target shooting mode.

[0187] In some implementation manners, the target shooting mode can be a preset shooting mode. Exemplarily, when the preset shooting mode is shooting mode 1 shown in Table 1, the video recording preview interface displayed by the mobile phone can be seen in Figure 4 in (d). As shown in Figure 4 in (d), shooting mode 1 includes line 1 on the left and line 2 on the right. Line 1 corresponds to the wide-angle view, and line 2 corresponds to the zoom view.

[0188] In some other implementation manners, the target shooting mode can be the shooting mode used last time recorded by the mobile phone, or the shooting mode pre-set by the user through the settings interface of the mobile phone before entering the multi-channel video recording mode.

[0189] In some other implementation manners, the target shooting mode can be the shooting mode indicated by the user after entering the multi-channel video recording mode. For example, after entering the multi-channel video recording mode, the mobile phone can prompt the user to select the target shooting mode. Exemplarily, after entering the multi-channel video recording mode, as seen in Figure 4 in (c), the mobile phone can display a prompt box 400 to prompt the user to select the target shooting mode. The prompt box 400 can include the following multiple shooting modes: wide-angle + zoom mode, wide-angle + front camera mode, zoom + front camera mode, and wide-angle + zoom + front camera mode. After the mobile phone detects the operation of the user selecting a certain shooting mode, it determines that this shooting mode is the target shooting mode. For example, after the mobile phone detects the operation of the user selecting the shooting mode of the wide-angle + zoom mode, it determines that the target shooting mode is shooting mode 1, and thus can display Figure 4 the video recording preview interface corresponding to shooting mode 1 shown in (d).

[0190] In some other implementations, after entering the multi-channel video recording mode, the target shooting mode defaults to a preset shooting mode, and then the mobile phone can switch the target shooting mode according to the user's instructions.

[0191] For example, the preset shooting mode is shooting mode 1. After entering the multi-channel video recording mode, the video preview interface displayed on the mobile phone includes the video picture in the wide-angle view of line 1 and the video picture in the zoom view of line 2. The mobile phone can display the above-mentioned prompt box 400 on the video preview interface for the user to select a new target shooting mode. Alternatively, the mobile phone can display the shooting mode options in the multi-channel video recording mode on the video preview interface for the user to select a new target shooting mode. Exemplarily, referring to Figure 5 (a) in, after the mobile phone detects the operation of the user clicking on the control 501, it determines that the target shooting mode is switched to shooting mode 2; then, the mobile phone can display the video preview interface corresponding to shooting mode 2.

[0192] Another example is that the preset shooting mode is shooting mode 1. After entering the multi-channel video recording mode, the mobile phone displays the video preview interface corresponding to shooting mode 1. The mobile phone can also display the shooting angle options in the multi-channel video recording mode. The mobile phone can determine the new target shooting mode according to the multiple shooting angles selected by the user. Among them, each shooting angle corresponds to a video picture. Exemplarily, referring to Figure 5 (b) in, after the mobile phone detects the operations of the user clicking on the control 502 and the control 503, it determines that the shooting angles include the wide-angle view and the front view, and thus determines that the new target shooting mode is shooting mode 2; then, the mobile phone can display the video preview interface corresponding to shooting mode 2.

[0193] Yet another example is that different cameras can correspond to different shooting angles. For example, the wide-angle camera can correspond to the wide-angle view, the telephoto camera can correspond to the zoom view, and the front camera can correspond to the front view. The preset shooting mode is shooting mode 1. After entering the multi-channel video recording mode, the mobile phone displays the video preview interface corresponding to shooting mode 1. The mobile phone can also display the identifiers of each camera. After the mobile phone detects the operation of the user selecting multiple cameras, it can determine the shooting angles respectively corresponding to each camera, and thus determine the new target shooting mode according to the shooting angles. Among them, each shooting angle corresponds to a video picture. Exemplarily, referring to Figure 5 (c) in, the mobile phone displays the identifier 504 of the wide-angle camera, the identifier 505 of the telephoto camera, and the identifier 506 of the front camera. After the mobile phone detects the user clicking on the identifier 504 of the wide-angle camera and the identifier 506 of the front camera, it determines that the target shooting mode is shooting mode 2 shown in Table 1; then, the mobile phone can display the video preview interface corresponding to shooting mode 2.

[0194] For another example, when the preset shooting mode is shooting mode 1, after entering the multi-channel video recording mode, the mobile phone displays a video preview interface corresponding to shooting mode 1. After the mobile phone detects that the user clicks the control 403 shown in (d) of Figure 4 , the shooting settings interface can be displayed. In response to the user's operation on the shooting settings interface, the mobile phone displays an interface for setting the target shooting mode as shown in (a), (b), or (c) of Figure 5 .

[0195] For still another example, when the preset shooting mode is shooting mode 1, after entering the multi-channel video recording mode, the mobile phone displays a video preview interface corresponding to shooting mode 1. The video preview interface may further include a control (such as a shooting mode control) for setting the target shooting mode. After the mobile phone detects that the user clicks this control, an interface for setting the target shooting mode as shown in (a), (b), or (c) of Figure 5 can be displayed.

[0196] In some embodiments, the mobile phone can also perform front / back switching according to the user's instruction and switch the target shooting mode. Exemplarily, referring to (a) of Figure 6 , the target shooting mode is shooting mode 1 in Table 1, and shooting mode 1 includes line 1 and line 2. Line 1 corresponds to a wide-angle view, and line 2 corresponds to a zoom view. After the mobile phone detects the user's operation of clicking the switching control 600, referring to (b) of Figure 6 , the target shooting mode is switched to shooting mode 2 in Table 1, and shooting mode 2 includes line 1 and line 2. Line 1 corresponds to a wide-angle view, and line 2 corresponds to a front view.

[0197] In addition, exemplarily, when the target shooting mode is shooting mode 4 in Table 1, the video preview interface can be referred to (c) of Figure 6 . The video preview interface includes line 1 - line 3, where line 1 corresponds to a wide-angle view, line 2 corresponds to a zoom view, and line 3 corresponds to a front view.

[0198] 304. After the mobile phone detects the user's operation indicating shooting, it collects multi-channel video images according to the target shooting mode and displays a shooting interface, which includes multi-channel video frames, and each video frame corresponds to a different shooting view.

[0199] For example, the user's operation indicating shooting can be the operation of the user clicking the shooting control 601 shown in (a) of Figure 6 , the user's voice instruction operation, or other operations, which are not limited in the embodiments of the present application. After the mobile phone detects the user's operation indicating shooting, it can display a corresponding shooting interface according to the line conditions and shooting views corresponding to the target shooting mode. The shooting interface includes multi-channel video frames, and each video frame corresponds to a different shooting view.

[0200] Among them, there can be multiple layout formats for multiple video images on the shooting interface, such as left / right splicing format, left / middle / right splicing format, up / down splicing format, up / middle / down splicing format, or picture-in-picture format, etc.

[0201] Exemplarily, taking the shooting mode 2 shown in Table 1 as the target shooting mode, the multiple video images on the shooting interface can be referred to Figure 7 in (a). Among them, Figure 7 the multiple video images shown in (a) in it adopt the left / right splicing format, the left image is the video image corresponding to the wide-angle view, and the right image is the video image corresponding to the front view.

[0202] Exemplarily again, taking the shooting mode 4 shown in Table 1 as the target shooting mode, the multiple video images on the shooting interface can be referred to Figure 7 in (b). Among them, Figure 7 the multiple video images shown in (b) in it adopt the picture-in-picture format, the small image in the lower left corner is the video image corresponding to the zoom view, the small image in the lower right corner is the video image corresponding to the front view, and the large image in the middle is the video image corresponding to the wide-angle view.

[0203] In some embodiments, during the multi-channel video recording process, the target shooting mode can also be switched. Exemplarily, after the mobile phone detects that the user clicks the front / back switching control 703 in (a) of Figure 7 , the line 2 can be switched from the front view to the zoom view, and the target shooting mode can be switched from shooting mode 2 to shooting mode 1.

[0204] It should be noted that in case 1, the zoom ratio corresponding to the video image in each line can be changed according to the user's instruction. And, the zoom ratio corresponding to the video image in each line can be changed within the zoom ratio range of the zoom view corresponding to that line. For example, the zoom ratio range of the wide-angle view is less than the above preset value K. The zoom ratio of the video image in the line corresponding to the wide-angle view can be changed within the range less than the above preset value K.

[0205] 305. The mobile phone records multiple video images according to the target shooting mode.

[0206] The mobile phone can record multiple video images according to the line situation and shooting view corresponding to the target shooting mode. This recording process can include that the mobile phone performs video encoding and other processing on the multiple video images collected, so as to generate a video file and save it.

[0207] In some embodiments, multiple video frames correspond to the same video file. In other embodiments, each video frame corresponds to a separate video file; in this way, when playing back the video subsequently, the mobile phone can play one of the video frames independently.

[0208] After the mobile phone detects an operation indicating shooting by the user, the method may further include step 306:

[0209] 306. The mobile phone collects a sound signal and records the audio corresponding to at least two video frames respectively according to the sound signal and the shooting angle.

[0210] Among them, the mobile phone recording the audio corresponding to at least two video frames respectively includes two solutions:

[0211] Solution 1: The mobile phone records the audio corresponding to each video frame among multiple video frames. That is to say, the mobile phone records N video frames and N audio tracks, where N is a positive integer greater than 1. That is to say, the mobile phone can record the audio corresponding to each video frame according to the shooting angle corresponding to each video frame among multiple video frames. That is, the mobile phone records the audio corresponding to the first video frame according to the shooting angle corresponding to the first video frame; the mobile phone records the audio corresponding to the second video frame according to the shooting angle corresponding to the second video frame.

[0212] Solution 2: The mobile phone records the audio corresponding to the video frames of some lines among multiple video frames. That is to say, the mobile phone records N video frames and M audio tracks, and N is greater than M, and both N and M are positive integers greater than 1. That is, in both Solution 1 and Solution 2, the mobile phone records multiple audio tracks according to the sound signal and the shooting angle.

[0213] The following first describes the situation described in Solution 1.

[0214] In Case 1, each video frame corresponds to a different shooting angle, and the mobile phone records the audio corresponding to each video frame respectively, that is, records the audio corresponding to the video frames under different shooting angles.

[0215] In Solution 1, referring to Figure 8 , this step 306 may include:

[0216] 801. The mobile phone obtains the audio data to be processed corresponding to each video frame in the target shooting mode.

[0217] In some embodiments, each microphone of the mobile phone can be turned on to collect a sound signal, and each microphone can convert the sound signal into an analog audio electrical signal. The audio module in the mobile phone can convert the analog audio electrical signal into initial audio data. The combination of the initial audio data corresponding to each microphone may include sound information in all directions.

[0218] In some technical solutions, since the combination of the initial audio data corresponding to each microphone can contain sound information in all directions, the mobile phone can obtain the audio data to be processed corresponding to different shooting perspectives based on the sound information in all directions. For example, the mobile phone weights the initial audio data corresponding to each microphone according to the preset weight strategy corresponding to different shooting perspectives, so as to obtain the audio data to be processed corresponding to different shooting perspectives respectively, that is, to obtain the audio data to be processed corresponding to the video pictures of different channels.

[0219] In other technical solutions, the audio data to be processed corresponding to each video picture in the target shooting mode is the audio data after the fusion of the initial audio data corresponding to each microphone, and the fused audio data contains sound information in all directions.

[0220] In other technical solutions, the audio data to be processed corresponding to each video picture in the target shooting mode is the audio data after the fusion of the initial audio data corresponding to each microphone. In the subsequent beamforming module or other audio processing modules, the mobile phone weights the audio data corresponding to each microphone according to the preset weight strategy corresponding to different shooting perspectives, so as to obtain the audio data corresponding to different shooting perspectives respectively.

[0221] In other embodiments, the mobile phone can use different microphone combinations to obtain the audio data to be processed corresponding to different shooting perspectives. For example, in Figure 1C the shown scenario, the audio data to be processed corresponding to the wide-angle perspective can be the audio data after weighting the initial audio data of microphone 1 and microphone 2. The audio data to be processed corresponding to the zoom perspective can be the audio data after weighting the initial audio data of microphone 1, microphone 2, and microphone 3. The audio data to be processed corresponding to the front perspective can be the audio data after weighting the initial audio data of microphone 1, microphone 2, and microphone 5.

[0222] In addition, the mobile phone can also generate different types of audio data to be processed, such as the audio data to be processed for the left channel, the audio data to be processed for the right channel, the audio data to be processed for the stereo, or the audio data to be processed for the mono, by weighting the initial audio data corresponding to different microphones. Different types of audio data to be processed can adopt different weight strategies.

[0223] 802. The mobile phone processes the audio data to be processed according to the shooting perspective corresponding to each video picture, so as to record the audio corresponding to each video picture respectively.

[0224] The audio processing module in the mobile phone can perform audio processing on the initial audio data according to the audio processing schemes corresponding to different shooting perspectives, so as to generate the audio corresponding to each shooting perspective in the target shooting mode, that is, generate the audio corresponding to each video frame in the target shooting mode.

[0225] The audio processing schemes corresponding to different shooting perspectives are described below respectively.

[0226] (1). Wide-angle perspective

[0227] In the wide-angle perspective, the user wants to record the video frame in a larger range. The mobile phone can record the video frame in a larger range. For example, it can record a panoramic video frame. Correspondingly, in the wide-angle perspective, the user also wants to obtain the sound in a larger range. Therefore, in the wide-angle perspective, the mobile phone can record the sound in a larger range. For example, it can record panoramic sound.

[0228] Exemplarily, the audio processing scheme corresponding to the wide-angle perspective can be referred to in Figure 9 in (a). As shown in (a) of Figure 9 , the audio data to be processed passes through a timbre correction module, a stereo sound beam forming module, a gain control module, etc.

[0229] Among them, the timbre correction module can perform timbre correction processing, which is used to correct the frequency response changes generated during the process of sound waves from the microphone hole to the analog-to-digital conversion, such as factors like the uneven frequency response of the microphone element, the resonance effect of the microphone pipeline, and the filter circuit.

[0230] The stereo sound beam forming module is used to form a stereo sound beam to retain the audio signal within the coverage of the stereo sound beam. Since the audio in the wide-angle perspective needs to pick up the sound in a larger range, and the mono sound beam usually cannot cover a large range, a stereo sound beam can be used to retain the audio signal in a larger coverage range. And, the angle between the left-channel sound beam and the right-channel sound beam in the stereo sound beam can be controlled within a large angle range, so that the synthesized stereo sound beam can cover a large range. The angle between the left-channel sound beam and the right-channel sound beam refers to the angle between the tangent line passing through the first tangent point and the tangent line passing through the second tangent point. Among them, the first tangent point is the tangent point of the left-channel sound beam and the circular boundary, and the second tangent point is the tangent point of the right-channel sound beam and the circular boundary. For example, the angle of the stereo sound beam can be controlled between 120 degrees and 180 degrees, and the stereo sound beam can cover the 360-degree panoramic direction. Exemplarily, the schematic diagram of the stereo sound beam can be referred to in Figure 9 in (b), where the solid line beam represents the beam of the left channel, the dashed line beam represents the beam of the right channel, and the 0-degree direction is aligned with the rear camera direction. The rear camera direction is perpendicular to the surface of the mobile phone at the back of the mobile phone, and the perpendicular point is the direction of the camera.

[0231] The gain control module can perform gain control processing to adjust the recording volume to an appropriate level so that the small-volume signal obtained by sound pickup can be clearly heard by the user, and the large-volume signal obtained by sound pickup will not be clipped and distorted. The mobile phone generates an audio file corresponding to the wide-angle view based on the audio data after audio processing, thereby recording the audio corresponding to the wide-angle view.

[0232] (2), Zoom view

[0233] In the zoom view, when the user wants to record the video picture within the zoom range, the mobile phone can record the video picture within the zoom range. Correspondingly, in the zoom view, the user also wants to obtain the sound within the zoom range. Therefore, in the zoom view, the mobile phone can record the sound within the zoom range and suppress the sound in other directions.

[0234] Exemplarily, the audio processing scheme corresponding to the zoom view can be referred to Figure 10 in (a). As Figure 10 shown in (a), the audio data to be processed passes through a timbre correction module, a stereo / mono beamforming module, an ambient noise control module, a gain control module, etc. Among them, the function of the timbre correction module is similar to that of this module in the wide-angle view. The stereo / mono beamforming module is used to generate a stereo beam or a mono beam to retain the audio signal within the coverage of the stereo beam or the mono beam. Since the zoom view needs to pick up the sound within the zoom range, and the zoom range is smaller than the sound pickup range in the wide-angle view, a stereo beam or a mono beam can be used in the zoom view. Compared with the wide-angle view, the stereo / mono beam module can narrow the mono beam or narrow the included angle of the stereo beam according to the zoom ratio. The beam gain of the stereo / mono beam within the zoom range is larger, which can retain the sound within the zoom range and suppress the sound outside the zoom range, so as to highlight the sound within the zoom range. Exemplarily, Figure 10(b) in it shows a schematic diagram of a three-dimensional sound beam when the zoom ratio is s, where s is greater than the above preset value K. The solid beam represents the beam of the left channel, and the dashed beam represents the beam of the right channel, with 0 degrees aligned with the rear shooting direction. The environmental noise suppression module can perform environmental noise suppression processing to suppress non-directional environmental noise, so as to improve the signal-to-noise ratio of the sound within the zoom range and make the sound within the zoom range more prominent. Moreover, the noise reduction intensity of the environmental noise suppression module also increases with the increase of the zoom ratio. In this way, the larger the zoom ratio, the weaker the environmental noise. The function of the gain control module is similar to that of this module under the above wide-angle view. The difference is that the volume gain of the sound within the zoom range increases with the increase of the zoom ratio. In this way, the larger the zoom ratio, the louder the sound within the zoom range, so that the volume of the sound in the distance can be increased and the auditory sense of the sound being brought closer can be enhanced. The mobile phone generates an audio file corresponding to the zoom view according to the audio data after audio processing, so as to record the audio corresponding to the zoom view.

[0235] (3), Front view

[0236] In the front view, when the user wants to record the front video picture, it is usually the picture of the user taking a selfie. Correspondingly, in the front view, the user usually wants to obtain the human voice during selfie. Therefore, in the front view, the mobile phone can record the human voice within the front range and suppress other sounds.

[0237] Exemplarily, the audio processing scheme corresponding to the front view can be referred to Figure 11 in (a). As Figure 11 shown in (a) in it, the audio data to be processed passes through a timbre correction module, a stereo / mono beamforming module, a human voice enhancement module, a gain control module, etc. respectively. Among them, the function of the timbre correction module is similar to that of this module under the wide-angle view. The stereo / mono beamforming module is used to generate a stereo beam or a mono beam to retain the audio signal within the coverage of the stereo beam or the mono beam. Since it is necessary to pick up the human voice within the front range in the front view, and the front range is smaller than the sound pickup range under the wide-angle view, a stereo beam or a mono beam can be used. Compared with the wide-angle view, the stereo / mono beam module can narrow the angle of the mono beam or narrow the angle of the stereo beam. The beam gain of the stereo / mono beam within the front range is relatively large, which can retain the sound within the front range and suppress the sound outside the front range, so as to highlight the sound within the front range. Exemplarily, Figure 11The schematic diagram of a three-dimensional sound beam corresponding to the front view is shown in (b) of . Among them, the solid beam represents the beam of the left channel, and the dashed beam represents the beam of the right channel. The 0-degree is aligned with the front camera direction, which is perpendicular to the surface of the mobile phone in front of the mobile phone, and the perpendicular point is the direction of the camera. The voice enhancement module is used to enhance the voice, for example, it can enhance the harmonic part of the voice and improve the clarity of the voice. The function of the gain control module is similar to that of this module under the above-mentioned wide-angle view. The mobile phone generates an audio file corresponding to the front view according to the audio data after audio processing, so as to record the audio corresponding to the front view.

[0238] For example, if the target shooting mode is shooting mode 1 shown in Table 1, the mobile phone can record the audio corresponding to the wide-angle view in line 1 and the audio corresponding to the zoom view in line 2.

[0239] In another example, if the target shooting mode is shooting mode 4 shown in Table 1, the mobile phone can record the audio corresponding to the wide-angle view in line 1, the audio corresponding to the zoom view in line 2, and the audio corresponding to the front view in line 3.

[0240] In some embodiments, when the mobile phone records the audio corresponding to multiple video frames, it can display a recording prompt message on each video frame to prompt the user that the audio corresponding to multiple video frames is being recorded. Exemplarily, the recording prompt message can be Figure 7 the microphone markers 701-702 shown in (a) of .

[0241] In case 2, during the multi-channel video recording process, the shooting angle corresponding to each video frame may change. Moreover, the mobile phone can determine the shooting angle corresponding to each video frame according to the size relationship between the zoom ratio corresponding to each video frame and the preset value and the front / back feature. When the multi-channel video frames include the first video frame and the second video frame, the mobile phone determines the shooting angle corresponding to the first video frame according to the size relationship between the zoom ratio corresponding to the first video frame and the preset value and the front / back feature; and determines the shooting angle corresponding to the second video frame according to the size relationship between the zoom ratio corresponding to the second video frame and the preset value and the front / back feature. Among them, the front / back feature corresponding to a certain video frame is used to indicate whether the video frame is a front video frame or a rear video frame.

[0242] The audio processing method in case 2 is similar to that in case 1. The following mainly explains the differences.

[0243] In Case 2, the target shooting mode in step 303 above can be the shooting modes shown in Table 2. The target shooting mode can be a preset shooting mode (such as shooting mode a), the shooting mode used last time, or the shooting mode indicated by the user, etc.

[0244] In Case 2, when the mobile phone records the audio corresponding to each video frame in step 306, for each video frame, since the shooting angle can be switched in real time, the mobile phone can adopt the audio processing solution corresponding to the current shooting angle for audio processing and record the audio matching the current shooting angle in real time. In this way, each video frame can correspond to the audio corresponding to different shooting angles.

[0245] For example, the target shooting mode is shooting mode a shown in Table 2, and the video frames corresponding to shooting mode a include Line 1 and Line 2. For Line 1, as the relationship between the zoom ratio corresponding to Line 1 and the preset value changes, the shooting angle can be switched in real time between the rear wide-angle view and the zoom view; for Line 2, the shooting angle is the front view.

[0246] In step 306, when the target shooting mode is shooting mode a shown in Table 2, for Line 1, when the shooting angle of the mobile phone is the wide-angle view, it records the audio according to the audio processing solution corresponding to the wide-angle view; when the shooting angle is the zoom view, it records the audio according to the audio processing solution corresponding to the wide-angle view. In this way, the audio corresponding to the video frame of Line 1 can include the audio corresponding to the wide-angle view and the audio corresponding to the zoom view. For example, for Line 1, if the shooting angle is first the wide-angle view, then switches to the zoom view, and then switches back to the wide-angle view, the audio corresponding to Line 1 includes a section of audio corresponding to the wide-angle view, a section of audio corresponding to the zoom view, and another section of audio corresponding to the wide-angle view. For Line 2, the mobile phone records the audio according to the audio processing solution corresponding to the front view, and the audio corresponding to Line 2 is the audio corresponding to the front view.

[0247] 307. After the mobile phone detects the operation indicating that the user stops shooting, it stops recording the video frames and audio and generates multiple video recordings.

[0248] Exemplarily, the operation indicating that the user stops shooting can be the operation of the user clicking on the control 700 shown in (a) of Figure 7 , the operation of the user's voice indicating to stop shooting, the gesture operation, or other operations, etc., which are not limited in the embodiments of the present application.

[0249] After the mobile phone detects the operation indicating that the user stops shooting, it generates multiple video recordings and returns to the video recording preview interface or the shooting preview interface. Among them, the multiple video recordings include multiple video frames and multiple audio. Exemplarily, for the thumbnail of the multiple video recordings generated by the mobile phone, reference can be made toFigure 12 the thumbnail 1201 shown in (a) of Figure 12 the thumbnail 1202 shown in (b) of

[0250] In some embodiments, for the multi-channel video recorded by using the method provided in the embodiments of the present application, the mobile phone may prompt the user that the video has multi-channel audio. Exemplarily, the multi-channel video thumbnail or the detailed information of the multi-channel video may include a prompt message for indicating multi-channel audio. For example, the prompt message may be Figure 12 the mark 1203 of multiple speakers shown in (b) of

[0251] In other embodiments, the mobile phone may save the video files corresponding to each channel of the video recording respectively. The video files corresponding to each channel of the video recording may also be displayed separately in the gallery. For example, referring to Figure 12 in (c), the file identified by the thumbnail 1204 in the gallery includes the files corresponding to the video frames of each channel of the video recording and the files corresponding to each channel of audio. Referring to Figure 12 in (c), the file identified by the thumbnail 1205 in the gallery includes the video frame and the file corresponding to the audio under the wide-angle view corresponding to line 1; the file identified by the thumbnail 1206 in the gallery includes the video frame and the file corresponding to the audio under the wide-angle view corresponding to line 2.

[0252] 308. After the mobile phone detects the operation of the user instructing to play the multi-channel video recording, it plays the video frames and audio of the multi-channel video recording.

[0253] Exemplarily, the operation of the user instructing to play the multi-channel video recording may be the operation of the user clicking on the thumbnail 1201 in the video preview interface shown in (a) of Figure 12 Again exemplarily, the operation of the user instructing to play the multi-channel video recording may be the operation of the user clicking on the thumbnail 1202 in the gallery shown in (b) of Figure 12 After the mobile phone detects the operation of the user instructing to play the multi-channel video recording, it plays the multi-channel video recording according to the multi-channel video frames and multi-channel audio recorded during the above multi-channel video recording process. That is, during video playback, the mobile phone plays the video frames and audio recorded during the above multi-channel video recording process.

[0254]

[0255] In some embodiments, during video playback, the mobile phone can play the recorded video images of each channel. That is, the mobile phone displays a video playback interface, and the video playback interface includes the video images of each channel. For audio playback, in some technical solutions, during video playback, the mobile phone can play the audio corresponding to the shooting perspective / route specified by the user. Or, during video playback, the mobile phone can default to playing the audio corresponding to the preset route / preset shooting perspective. For example, the preset route is the route on the left side of the video playback interface. For another example, the audio corresponding to the preset shooting perspective is the audio corresponding to the wide-angle perspective. During video playback, the mobile phone can play the audio corresponding to the wide-angle perspective. Subsequently, the mobile phone can also switch to play the audio corresponding to different routes / shooting perspectives according to the user's instructions.

[0256] In this way, the mobile phone can switch to play the audio corresponding to different routes / shooting perspectives according to the user's instructions. For example, when the user focuses on the video image in the wide-angle perspective, the mobile phone can play the audio corresponding to the wide-angle perspective according to the user's instructions, and the user can hear the panoramic sound. When the user focuses on the video image in the zoom perspective, the user focuses on the shooting object in the distance after zooming in, and the mobile phone can play the audio corresponding to the zoom perspective according to the user's instructions, and the user can mainly hear the sound after zooming in within the zoom range. When the user focuses on the video image in the front-facing perspective, the user focuses on the person in the front-facing direction, and the mobile phone can play the audio corresponding to the front-facing perspective according to the user's instructions, and the user can mainly hear the human voice in the front-facing direction.

[0257] For example, the target shooting mode is shooting mode 4 shown in Table 1, and the video playback interface can be referred to Figure 13 (a) in. The video playback interface includes the video image 1301 corresponding to the wide-angle perspective in route 1, the video image 1302 corresponding to the zoom perspective in route 2, and the video image 1303 corresponding to the front-facing perspective in route 3. The video playback interface also includes the audio playback control 1304 corresponding to the video image 1301, the audio playback control 1305 corresponding to the video image 1302, and the audio playback control 1306 corresponding to the video image 1303. Among them, the audio playback control corresponding to each video image can be used to control the playback / stop of the audio corresponding to that video image.

[0258] As Figure 13 (a) shown, during video playback, the mobile phone can default to playing the audio corresponding to the video image 1301 in the wide-angle perspective. The audio playback control 1304 is in the playing state, and the audio playback controls 1305 and 1306 are in the non-playing state. If the mobile phone detects that the user clicks Figure 13 (a) shown in the audio playback control 1305, then the mobile phone plays the audio 2 corresponding to the video image 1302 and the zoom perspective, as Figure 13As shown in (b) therein, the audio playback control 1305 changes to the playing state. In some implementations, such as Figure 13 As shown in (b) therein, the audio playback control 1304 automatically changes to the non-playing state, and the mobile phone plays audio 2 and stops playing audio 1. That is, the mobile phone plays only one audio stream at the same time to avoid audio mixing caused by playing multiple audio streams simultaneously. In other implementations, the audio playback control 1304 remains in the playing state, and the mobile phone plays audio 1 and audio 2 simultaneously.

[0259] It can be understood that when the target shooting mode is other shooting modes in Table 1, the video playback interface may include any two of the video frame 1301, the video frame 1302, or the video frame 1303; correspondingly, it may include any two of the audio playback control 1304, the audio playback control 1305, or the audio playback control 1306.

[0260] It should be noted that Figure 13 The audio playback control shown in (a)-(b) therein is in the form of a speaker. The audio playback control can also be in other forms and can also be located at other positions on the video playback interface, which is not limited in the embodiments of the present application. Exemplarily, referring to Figure 13 As shown in (c) therein, the audio playback control can also be the text information "Play Audio / Stop Playing Audio".

[0261] For another example, the mobile phone can play the audio corresponding to the specified shooting perspective / route according to the user's voice instruction.

[0262] In other technical solutions, during video playback, the mobile phone can preferentially play one audio stream corresponding to the shooting perspective / route with a higher priority. For example, the priority of the front view is higher than that of the zoom view, and the priority of the zoom view is higher than that of the wide-angle view. For another example, the route with a larger display area has a higher priority. Subsequently, the mobile phone can also switch to play the audio corresponding to different routes / shooting perspectives according to the user's instruction.

[0263] In other embodiments, during video playback, the mobile phone can play only one video frame, that is, only one video frame is displayed on the video playback interface, and the audio corresponding to this video frame is automatically played.

[0264] In some technical solutions, during video playback, the mobile phone defaults to displaying each video frame. After detecting the operation of the user instructing to display only a certain video frame, the mobile phone can magnify or full-screen display this video frame and automatically play the audio corresponding to this video frame.

[0265] Exemplarily, during video playback, referring to Figure 14 As shown in (a) therein, the mobile phone defaults to playing each video frame and the audio 1 corresponding to the video frame 1401. When the mobile phone detects that the user clicksFigure 14 After the operation of the full-screen display control 1400 shown in (a) in Figure 14 as shown in (b) in Figure 14 the mobile phone full-screen displays the video picture 1402 corresponding to the zoom view, stops displaying the video pictures corresponding to the wide-angle view and the front view, and automatically plays the audio 2 corresponding to the video picture 1402. Subsequently, after the mobile phone detects the operation of the user clicking on the control 1403, it can play the video picture 1401 and the audio 1 in full screen; after the mobile phone detects the operation of the user clicking on the control 1404, it can play the video picture corresponding to the front view and the audio 3 in full screen; after the mobile phone detects the operation of the user clicking on the control 1405, it can resume displaying the multi-channel video pictures as shown in

[0266] In the solution described in the above embodiment, in case 1 where the shooting angle of each video is fixed, the mobile phone can switch the video pictures and audio corresponding to different lines / shooting angles, so that the played audio is matched with the shooting angle and video picture that the user is concerned about in real time, thereby improving the user's audio experience.

[0267] In the solution described in the above embodiment, in case 2 where the shooting angle of each video is variable, when playing the audio corresponding to any line, the mobile phone can play the video picture and audio corresponding to the real-time changing shooting angle, so that the audio is matched with the shooting angle and video picture in real time, thereby improving the user's audio experience.

[0268] In addition, after the mobile phone detects that the user clicks on Figure 12 the thumbnail 1205 shown in (c) in Figure 12 it can play the single-channel video picture and audio corresponding to the wide-angle view. After the mobile phone detects that the user clicks on

[0269] the thumbnail 1206 shown in (c) in

[0270] The above is described by taking the example that each video picture corresponds to one audio in Solution 1. In Solution 2, the mobile phone can generate the audio corresponding to some lines according to the user's instructions without generating the audio corresponding to each video picture separately.

[0271] During video playback, the mobile phone can play the audio corresponding to a certain video frame according to the user's instructions. For example, when the mobile phone magnifies / fullscreens a certain video frame during video playback, the mobile phone can automatically play the audio corresponding to that video frame. If there is no corresponding audio for that video frame, the mobile phone can automatically play the audio corresponding to the preset shooting perspective (for example, the wide-angle perspective), or the mobile phone does not play audio.

[0272] In some other implementation manners, each shooting perspective corresponds to one audio, and each audio can correspond to one or more lines. For example, the target shooting mode is shooting mode 4 shown in Table 1, which includes 3 lines. The mobile phone only generates 2 audios. One audio corresponds to the front perspective, and the other audio corresponds to the wide-angle perspective and the zoom perspective. During video playback, the mobile phone can play the audio corresponding to a certain video frame according to the user's instructions.

[0273] In this way, for Solution 2, the mobile phone can switch the video frames and audios corresponding to different lines / shooting perspectives, so that the played audio is matched with the shooting perspective and video frame that the user is concerned about in real time, thereby improving the user's audio experience.

[0274] In some other embodiments, if the mobile phone includes a microphone at the top and two built-in microphones at the bottom, the mobile phone can also generate the audio corresponding to the wide-angle perspective. Exemplarily, in this scenario, the schematic diagram of the stereo sound beam can be seen in Figure 15 , and the included angle of the stereo sound beam is 180°. In this case, the audio corresponding to the zoom perspective and the front perspective can be picked up by an accessory microphone.

[0275] In some other embodiments, the mobile phone can use a directional microphone to collect the sound signals within the zoom range, thereby generating corresponding audio data to be processed according to the collected sound signals, and then processing the audio data to be processed and recording the audio corresponding to the zoom perspective. The directional microphone can be an internal component or an external accessory of the mobile phone. For example, the directional microphone is an external component and can point to the rear camera direction to cover the rear zoom range, thereby collecting the sound signals within the zoom range. In this case, the audio corresponding to the wide-angle perspective and the front perspective can be picked up by an internal microphone or an accessory microphone. For example, the audio corresponding to the wide-angle perspective can be picked up by an MS recording accessory, and the audio corresponding to the front perspective can be picked up by a microphone accessory with a cardioid directivity.

[0276] Among them, since the directional microphone has a cardioid or supercardioid directivity and can point to the direction included in the zoom range, the sound signals within the zoom range can be collected. In this scenario, see Figure 16 , and the audio processing scheme corresponding to the zoom perspective may not include a beamforming module.

[0277] In some other embodiments, the mobile phone can use a directional microphone to collect sound signals within the front range, so as to generate corresponding audio data to be processed based on the collected sound signals, and then process the audio data to be processed and record the audio corresponding to the front view. For example, the directional microphone is an external component and can point to the front camera direction to collect sound signals within the front range. Therefore, referring to Figure 17 , the audio processing solution corresponding to the front view may not include a beamforming module. In this case, the audio corresponding to the wide-angle view and the zoom view can be picked up by an internal microphone or an accessory microphone. For example, the audio corresponding to the wide-angle view can be picked up by an MS recording accessory, and the audio corresponding to the zoom view can be picked up by a directional microphone accessory.

[0278] Among them, the directional microphone can be a microphone on the earphone. For example, the user can place the earphone in front of the mobile phone and can also place the microphone towards the user. The earphone includes but is not limited to a wired earphone or a TWS earphone, etc. If the earphone provides stereo data, then referring to Figure 18 , the audio processing solution corresponding to the front view further includes a head related transfer function (HRTF) processing module for processing the stereo data to make the processed audio closer to the real listening experience of both ears at that time.

[0279] In some other embodiments, in the traditional video recording mode (i.e., the single-channel video recording mode), the shooting angle can also be switched. While the mobile phone is recording the video picture, it can also record the corresponding audio according to the real-time changing shooting angle. During video playback, the mobile phone can play the audio that is real-time matched with the shooting angle. In this way, the audio recorded by the mobile phone is real-time matched with the shooting angle that the user is concerned about and the recorded video picture, which can improve the user's audio experience. For example, the mobile phone initially records the video picture using the wide-angle view, and the mobile phone records the audio through the audio processing solution corresponding to the wide-angle view; then, during this video recording process, the mobile phone switches to using the zoom view to continue recording the video picture, and the mobile phone records the audio through the audio processing solution corresponding to the zoom view. During video playback, when the mobile phone plays the video picture in the wide-angle view, it plays the audio corresponding to the wide-angle view, and the user can hear the panoramic sound; when the mobile phone plays the video picture in the zoom view, it plays the audio corresponding to the zoom view, and the user can hear the sound of the distant object after zooming in within the zoom range.

[0280] For another example, the mobile phone initially records the video image with a wide-angle view, and the mobile phone records the audio through the audio processing solution corresponding to the wide-angle view. Then, during this video recording process, the mobile phone switches to using the front-facing view to continue recording the video image, and the mobile phone records the audio through the audio processing solution corresponding to the front-facing view. When playing back the video, when the mobile phone plays the video image in the wide-angle view, the audio corresponding to the wide-angle view is played, and the user can hear the panoramic sound; when playing the video image corresponding to the front-facing view, the audio corresponding to the front-facing view is played, and the user can hear the voices in the front-facing direction.

[0281] In the above embodiments, the audio in the multi-channel video recording is generated in real time during the video recording process. In some other embodiments, the mobile phone can collect the sound signal during the video recording process; after the video recording is completed, the audio corresponding to each video image is generated according to the collected sound signal and the shooting view used during the video recording process, so as to reduce the requirement for the processing ability of the mobile phone during the video recording process.

[0282] It can be understood that, in order to implement the above functions, the electronic device includes the corresponding hardware and / or software modules for executing each function. Combining the algorithm steps of each example described in the embodiments disclosed in this article, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving the hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application in combination with the embodiments, but such implementation should not be considered to exceed the scope of the present application.

[0283] This embodiment can divide the functional modules of the electronic device according to the above method examples. For example, each functional module can be divided corresponding to each function, or two or more functions can be integrated into one processing module. The above integrated module can be implemented in the form of hardware. It should be noted that the division of the modules in this embodiment is illustrative, only a logical function division, and there can be other division methods in actual implementation.

[0284] The embodiment of the present application also provides an electronic device, including one or more processors and one or more memories. The one or more memories are coupled to the one or more processors, and the one or more memories are used to store computer program code, and the computer program code includes computer instructions. When the one or more processors execute the computer instructions, the electronic device executes the above related method steps to implement the audio processing method in the above embodiments.

[0285] An embodiment of the present application further provides a computer-readable storage medium storing computer instructions, which, when running on an electronic device, cause the electronic device to execute the above-related method steps to implement the audio processing method in the above embodiment.

[0286] An embodiment of the present application further provides a computer program product, which, when running on a computer, causes the computer to execute the above-related steps to implement the audio processing method executed by the electronic device in the above embodiment.

[0287] In addition, an embodiment of the present application further provides a device, which may specifically be a chip, a component, a module, or a chip system. The device may include a processor and a memory connected to each other. The memory is used to store computer execution instructions. When the device runs, the processor may execute the computer execution instructions stored in the memory so that the chip executes the audio processing method executed by the electronic device in each of the above method embodiments.

[0288] Among them, the electronic device, the computer-readable storage medium, the computer program product, or the chip provided in this embodiment are all used to execute the corresponding method provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding method provided above, and will not be elaborated here.

[0289] Through the description of the above embodiments, those skilled in the art can understand that for the convenience and simplicity of description, only the above division of each functional module is used as an example. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.

[0290] In several embodiments provided by the present application, it should be understood that the disclosed device and method can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of the module or unit is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other may be through some interfaces. The indirect coupling or communication connection of the device or unit may be in an electrical, mechanical or other form.

[0291] The unit described as a separation component may or may not be physically separated. The component displayed as a unit may be a physical unit or multiple physical units, that is, it may be located in one place or distributed to multiple different places. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0292] In addition, each functional unit in various embodiments of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0293] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The software product is stored in a storage medium and includes several instructions for causing a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods described in various embodiments of the present application. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read only memory (ROM), random access memory (RAM), magnetic disks, or optical discs and other various media that can store program codes.

[0294] The above content is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in the present application, and all of them should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. An audio processing method, characterized in that, Including: After the electronic device detects that the user opens the camera, it displays a shooting preview interface; The electronic device enters a multi-channel video recording mode; After the electronic device detects the user's shooting operation, it displays a shooting interface, and the shooting interface includes multi-channel video images; The electronic device records multi-channel videos; The electronic device determines the shooting angle corresponding to each channel of video image according to the size relationship between the zoom ratio corresponding to each channel of video image and a preset value and the front / back camera feature; The electronic device records the audio corresponding to each channel of video image respectively according to the shooting angle corresponding to each channel of video image; After the electronic device detects the user's operation of stopping shooting, it generates a multi-channel video recording, and the multi-channel video recording also includes the audio corresponding to each channel of video image; After the electronic device detects the user's operation of playing the multi-channel video recording, it displays a video playback interface, and the video playback interface includes the multi-channel video images and the audio playback controls corresponding to each channel of video image respectively; When the electronic device detects the user's operation of clicking the audio playback control corresponding to the first channel of video image in the multi-channel video images, it plays the audio corresponding to the first channel of video image.

2. The method according to claim 1, characterized in that, One shooting angle corresponds to each channel of video image.

3. The method according to claim 1, characterized in that, The shooting angle corresponding to each channel of video image is variable.

4. The method according to any one of claims 1 - 3, characterized in that, The method further includes: If the electronic device plays other audio before playing the audio corresponding to the first channel of video image, the electronic device stops playing the other audio.

5. The method according to claim 4, characterized in that, Before the electronic device detects the user's operation of clicking the audio playback control corresponding to the first channel of video image in the multi-channel video images, the method further includes: The electronic device defaults to playing the audio corresponding to the second channel of video image in the multi-channel video images.

6. The method according to claim 5, characterized in that, The shooting angle corresponding to the second channel of video image is a preset shooting angle.

7. The method according to any one of claims 1 - 3, characterized in that, After the electronic device displays the video playback interface, the method further includes: After the electronic device detects the user's operation of playing the first channel of video image in the multi-channel video images, it displays the first channel of video image on the video playback interface and stops displaying the other channels of video images except the first channel of video image; The electronic device automatically plays the audio corresponding to the first channel of video image.

8. The method according to any one of claims 1 - 3, characterized in that, Before the electronic device displays the shooting interface, the method further includes: The electronic device determines a target shooting mode, and the target shooting mode is used to represent the number of lines of video images to be recorded; The electronic device displays the shooting interface, including: The electronic device displays the shooting interface according to the number of lines of video images to be recorded corresponding to the target shooting mode.

9. The method according to claim 8, characterized in that, The target shooting mode is further used to represent the corresponding relationship between each channel of video image and the shooting angle; the electronic device records the audio corresponding to each channel of video image respectively according to the shooting angle corresponding to each channel of video image in the multi-channel video images, including: The electronic device determines the shooting angle corresponding to each channel of video image according to the corresponding relationship between each channel of video image and the shooting angle corresponding to the target shooting mode; The electronic device records the audio corresponding to each video frame according to the shooting angle corresponding to each video frame.

10. The method according to claim 9, characterized in that, The target shooting mode includes: a wide-angle view and zoom view combination mode, a wide-angle view and front view combination mode, a zoom view and front view combination mode, or a wide-angle view, zoom view, and front view combination mode. The zoom ratio corresponding to the wide-angle view is less than or equal to a preset value, and the zoom ratio corresponding to the zoom view is greater than the preset value.

11. The method according to claim 10, characterized in that, The target shooting mode is a preset shooting mode.

12. The method according to claim 8, wherein, After the electronic device determines the target shooting mode, the method further includes: The electronic device detects an operation by the user to switch the target shooting mode. The electronic device switches the target shooting mode.

13. The method according to any one of claims 1 - 3, wherein, The shooting angle corresponding to the first video frame among the multiple video frames is the first shooting angle. When the electronic device records the audio corresponding to the first video frame according to the shooting angle corresponding to the first video frame, it includes: The electronic device obtains the audio data to be processed corresponding to the first video frame according to the collected sound signal. The electronic device performs timbre correction processing, forms a stereo sound beam, and performs gain control processing on the audio data to be processed according to the first shooting angle. The electronic device records the audio corresponding to the first video frame according to the audio data after the gain control processing.

14. The method according to any one of claims 1 - 3, wherein, The shooting angle corresponding to the first video frame among the multiple video frames is the second shooting angle. When the electronic device records the audio corresponding to the first video frame according to the shooting angle corresponding to the first video frame, it includes: The electronic device obtains the audio data to be processed corresponding to the first video frame according to the collected sound signal. The electronic device performs timbre correction processing, forms a stereo / mono sound beam, performs ambient noise control processing, and performs gain control processing on the audio data to be processed according to the second shooting angle. The electronic device records the audio corresponding to the first video frame according to the audio data after the gain control processing.

15. The method according to any one of claims 1 - 3, wherein, The shooting angle corresponding to the first video frame among the multiple video frames is the third shooting angle. When the electronic device records the audio corresponding to the first video frame according to the shooting angle corresponding to the first video frame, it includes: The electronic device obtains the audio data to be processed corresponding to the first video frame according to the collected sound signal. The electronic device performs timbre correction processing, forms a stereo / mono sound beam, performs human voice enhancement processing, and performs gain control processing on the audio data to be processed according to the third shooting angle. The electronic device records the audio corresponding to the first video frame according to the audio data after the gain control processing.

16. An electronic device, wherein, It includes: Multiple microphones for collecting sound signals; Multiple cameras for collecting video frames; A screen for displaying an interface; An audio playback component for playing audio; One or more processors; A memory; and one or more computer programs, wherein the one or more computer programs are stored in the memory, and the one or more computer programs include instructions that, when executed by the electronic device, cause the electronic device to perform the following steps: After detecting an operation by the user to turn on the camera, display a shooting preview interface; Enter the multi-channel video recording mode; After detecting a shooting operation by the user, display a shooting interface, the shooting interface including multi-channel video frames; Record multi-channel video frames; Determine the shooting angle corresponding to each channel of video frame according to the relationship between the zoom ratio corresponding to each channel of video frame and a preset value and the front / back camera feature; According to the shooting angle corresponding to each channel of video frame, record the audio corresponding to each channel of video frame respectively; after detecting an operation by the user to stop shooting, generate a multi-channel video recording, the multi-channel video recording further including the audio corresponding to each channel of video frame; After detecting an operation by the user to play the multi-channel video recording, display a video playback interface, the video playback interface including the multi-channel video frames and audio playback controls corresponding to each channel of video frame respectively; After detecting an operation by the user to click the audio playback control corresponding to the first channel of video frame in the multi-channel video frames, play the audio corresponding to the first channel of video frame.

17. The electronic device according to claim 16, wherein, The plurality of microphones include at least 3, and at least one microphone is provided on the top, bottom and back of the electronic device respectively.

18. The electronic device according to claim 16 or 17, characterized in that, The microphone is an internal component or an external accessory.

19. The electronic device according to claim 18, characterized in that, The microphone is an external accessory, and the microphone is a directional microphone.

20. The electronic device according to claim 16 or 17, characterized in that, When the instructions are executed by the electronic device, cause the electronic device to perform the audio processing method according to any one of claims 2-15.

21. A computer-readable storage medium, characterized in that, Includes computer instructions that, when the computer instructions run on a computer, cause the computer to perform the audio processing method according to any one of claims 1-15.

22. A computer program product, characterized in that, When the computer program product runs on a computer, cause the computer to perform the audio processing method according to any one of claims 1-15.

Citation Information

Patent Citations

  • Audio information processing method and device

    CN104699445A

  • Network broadcast method and device

    CN106954085A

  • Multi-channel video recording method and device

    CN110072070A