Video processing method and electronic device
Patent Information
- Application Number
- CN202210320689.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-12-27
- Filing Date
- 2022-03-29
- Publication Date
- 2026-09-15
- Estimated Expiration
- 2042-03-29
AI Technical Summary
[0002]电子设备在录像或者视频通话的场景中,经常会面临镜头转换的需求,需要对拍摄模式进行切换;例如,前置镜头与后置镜头的切换、多镜头录像或者单镜头录像的切换;目前,电子设备的镜头切换依赖于用户手动操作,因此需要拍摄者在拍摄过程中与电子设备之间的距离较近;若用户与电子设备之间的距离较远,则需要基于蓝牙技术实现电子设备的镜头切换;基于蓝牙技术实现电子设备的镜头切换时,需要通过控制设备对电子设备的镜头进行相应的操作,一方面操作较复杂;另一方面控制设备容易暴露在视频中,影响视频的美感,从而导致用户体验较差
[0053] In embodiments of this application, the electronic device can collect audio data in the shooting environment through at least two sound pickup devices (e.g., microphones); generate a switching command based on the audio data, and the electronic device automatically switches from the current first shooting mode to the second shooting mode based on the switching command, and displays the second image captured in the second shooting mode; without the user needing to switch the shooting mode of the electronic device, the electronic device can automatically switch shooting modes to complete video recording, thereby improving the user's shooting experience.
Smart Images

Figure CN116405774B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of video processing, and more specifically, to a video processing method and electronic device. Background Technology
[0002] In video recording or video call scenarios, electronic devices often need to switch camera angles, requiring switching shooting modes; for example, switching between front and rear cameras, or between multi-lens and single-lens recording. Currently, camera switching on electronic devices relies on manual operation by the user, thus requiring the photographer to be relatively close to the device during recording. If the distance between the user and the device is greater, Bluetooth technology is needed to enable camera switching. However, using Bluetooth technology for camera switching requires controlling the camera via a control device, which is complex and makes the control device easily visible in the video, affecting the video's aesthetics and resulting in a poor user experience.
[0003] Therefore, in video scenarios, how electronic devices can automatically switch cameras based on user needs has become an urgent problem to be solved. Summary of the Invention
[0004] This application provides a video processing method and an electronic device that can complete video recording without requiring the user to switch the shooting mode of the electronic device, thereby improving the user's shooting experience.
[0005] In a first aspect, a video processing method is provided, applied to an electronic device, the electronic device including at least two audio pickup devices, the video processing method comprising:
[0006] Run the camera application in the electronic device; Display a first image, which is an image captured by the electronic device when it is in a first shooting mode; Acquire audio data, wherein the audio data is data collected by the at least two microphones; A switching instruction is obtained based on the audio data; the switching instruction is used to instruct the electronic device to switch from the first shooting mode to the second shooting mode. Display a second image, which is an image captured by the electronic device when it is in the second shooting mode.
[0007] In embodiments of this application, the electronic device can collect audio data in the shooting environment through at least two sound pickup devices (e.g., microphones); generate a switching command based on the audio data, and the electronic device automatically switches from the current first shooting mode to the second shooting mode based on the switching command, and displays the second image captured in the second shooting mode; without the user needing to switch the shooting mode of the electronic device, the electronic device can automatically switch shooting modes to complete video recording, thereby improving the user's shooting experience.
[0008] It should be understood that, in the embodiments of this application, since the electronic device needs to determine the directionality of the audio data, the electronic device in the embodiments of this application includes at least two pickup devices, and there is no limitation on the specific number of pickup devices.
[0009] In one possible implementation, the first shooting mode can refer to either a single-camera mode or a multi-camera mode; wherein, the single-camera mode can include a front-facing single-camera mode or a rear-facing single-camera mode; the multi-camera mode can include a front / rear dual-camera mode, a rear / front dual-camera mode, a picture-in-picture front main camera mode, or a picture-in-picture rear main camera mode.
[0010] For example, in front-facing single-camera mode, video recording is performed using one of the front-facing cameras in the electronic device; in rear-facing single-camera mode, video recording is performed using one of the rear-facing cameras in the electronic device; in front and rear dual-camera mode, video recording is performed using one front-facing camera and one rear-facing camera; in picture-in-picture front-facing mode, video recording is performed using one front-facing camera and one rear-facing camera, with the image captured by the rear-facing camera placed within the image captured by the front-facing camera, and the image captured by the front-facing camera being the main image; in picture-in-picture rear-facing mode, video recording is performed using one front-facing camera and one rear-facing camera, with the image captured by the front-facing camera placed within the image captured by the rear-facing camera, and the image captured by the rear-facing camera being the main image.
[0011] Optionally, the multi-camera mode may also include a front dual-camera mode, a rear dual-camera mode, a front picture-in-picture mode, or a rear picture-in-picture mode.
[0012] It should be understood that the first shooting mode and the second shooting mode can refer to the same shooting mode or different shooting modes; if the switching command is the default current shooting mode, then the second shooting mode and the first shooting mode can be the same shooting mode; in other cases, the second shooting mode and the first shooting mode can be different shooting modes.
[0013] In conjunction with the first aspect, in some implementations of the first aspect, the electronic device includes a first camera and a second camera, the first camera and the second camera being located in different directions of the electronic device, and the step of obtaining the switching instruction based on audio data includes: Identify whether the audio data includes a target keyword, where the target keyword is the text information corresponding to the switching instruction; If the target keyword is identified in the audio data, the switching instruction is obtained based on the target keyword; If the target keyword is not identified in the audio data, the audio data is processed to obtain audio data in a first direction and / or audio data in a second direction. The first direction is used to represent a first preset angle range corresponding to the first camera, and the second direction is used to represent a second preset angle range corresponding to the second camera. Based on the audio data in the first direction and / or the audio data in the second direction, the switching instruction is obtained.
[0014] In the embodiments of this application, it can first be identified whether the audio data includes the target keyword; if the audio data includes the target keyword, the electronic device switches the shooting mode to the second shooting mode corresponding to the target keyword; if the audio data does not include the target keyword, the electronic device can obtain a switching instruction based on the audio data in the first direction and / or the audio data in the second direction; for example, if the user is in front of the electronic device, the image is generally captured by the front camera; if the user's audio information exists in the direction in front of the electronic device, it can be assumed that the user is in the direction in front of the electronic device, and the front camera can be turned on; if the user is behind the electronic device, the image is generally captured by the rear camera; if the user's audio information exists in the direction behind the electronic device, it can be assumed that the user is in the direction behind the electronic device, and the rear camera can be turned on.
[0015] In conjunction with the first aspect, in certain implementations of the first aspect, processing the audio data to obtain audio data in a first direction and / or audio data in a second direction includes: The audio data is processed based on a sound direction probability calculation algorithm to obtain audio data in the first direction and / or audio data in the second direction.
[0016] In the embodiments of this application, the probability of audio data in each direction can be calculated, thereby separating the audio data by direction to obtain audio data in a first direction and audio data in a second direction; a switching command can be obtained based on the audio data in the first direction and / or the audio data in the second direction; the electronic device can automatically switch the shooting mode based on the switching command.
[0017] In conjunction with the first aspect, in certain implementations of the first aspect, obtaining the switching instruction based on audio data from the first direction and / or audio data from the second direction includes: The switching instruction is obtained based on the energy of the first amplitude spectrum and / or the energy of the second amplitude spectrum, wherein the first amplitude spectrum is the amplitude spectrum of the audio data in the first direction, and the second amplitude spectrum is the amplitude spectrum of the audio data in the second direction.
[0018] It should be understood that in video recording scenarios, the direction with the greater energy of the audio data (e.g., the direction with the greater volume of the audio information) can generally be considered as the main shooting direction; the main shooting direction can be obtained based on the energy of the amplitude spectrum of the audio data in different directions; for example, if the energy of the amplitude spectrum of the audio data in the first direction is greater than the energy of the amplitude spectrum of the audio data in the second direction, then the first direction can be considered as the main shooting direction; at this time, the camera corresponding to the first direction in the electronic device can be turned on.
[0019] In conjunction with the first aspect, in some implementations of the first aspect, the switching instruction includes the current shooting mode, a first picture-in-picture mode, a second picture-in-picture mode, a first dual-view mode, a second dual-view mode, a single-camera mode of the first camera, or a single-camera mode of the second camera. The step of obtaining the switching instruction based on the energy of the first amplitude spectrum and / or the energy of the second amplitude spectrum includes: If the energy of the first amplitude spectrum and the energy of the second amplitude spectrum are both less than the first preset threshold, the switching instruction is to maintain the current shooting mode. If the energy of the first amplitude spectrum is greater than the second preset threshold, and the energy of the second amplitude spectrum is less than or equal to the second preset threshold, the switching instruction is to switch to the single-camera mode of the first camera; If the energy of the second amplitude spectrum is greater than the second preset threshold, and the energy of the first amplitude spectrum is less than or equal to the second preset threshold, the switching instruction is to switch to the single-camera mode of the second camera; If the energy of the first amplitude spectrum is greater than the second preset threshold, and the energy of the second amplitude spectrum is greater than or equal to the first preset threshold, the switching instruction is to switch to the first picture-in-picture mode; If the energy of the second amplitude spectrum is greater than the second preset threshold, and the energy of the first amplitude spectrum is greater than or equal to the first preset threshold, the switching instruction is to switch to the second picture-in-picture mode; If the energy of the first amplitude spectrum and the energy of the second amplitude spectrum are both greater than or equal to the second preset threshold, and the energy of the first amplitude spectrum is greater than the energy of the second amplitude spectrum, the switching instruction is to switch to the first dual-view mode; If the energy of the first amplitude spectrum and the energy of the second amplitude spectrum are both greater than or equal to the second preset threshold, and the energy of the second amplitude spectrum is greater than the energy of the first amplitude spectrum, the switching instruction is to switch to the second dual-view mode; Wherein, the second preset threshold is greater than the first preset threshold, the first picture-in-picture mode refers to the shooting mode in which the image captured by the first camera is the main picture, the second picture-in-picture mode refers to the shooting mode in which the image captured by the second camera is the main picture, the first dual-view mode refers to the shooting mode in which the image captured by the first camera is located on the top or left side of the display screen of the electronic device, and the second dual-view mode refers to the shooting mode in which the image captured by the second camera is located on the top or left side of the display screen of the electronic device.
[0020] In conjunction with the first aspect, in some implementations of the first aspect, the first amplitude spectrum is a first average amplitude spectrum obtained by averaging the amplitude spectra corresponding to each frequency point in the audio data in the first direction; and / or, The second amplitude spectrum is the second average amplitude spectrum obtained by averaging the amplitude spectra corresponding to each frequency point in the audio data in the second direction.
[0021] In the embodiments of this application, the amplitude spectrum obtained by averaging the amplitude spectra of different frequency points in the audio data in the first direction can be called the first average amplitude spectrum; the amplitude spectrum obtained by averaging the amplitude spectra of different frequency points in the audio data in the second direction can be called the second average amplitude spectrum; since the first average amplitude spectrum and / or the second average amplitude spectrum are amplitude spectra obtained by averaging the amplitude spectra of different frequency points, the accuracy of the information in the audio data in the first direction and / or the audio data in the first direction can be improved.
[0022] In conjunction with the first aspect, in some implementations of the first aspect, the first amplitude spectrum is an amplitude spectrum obtained by performing a first amplification process and / or a second amplification on the first average amplitude spectrum, and the first average amplitude spectrum is obtained by averaging the amplitude spectra corresponding to each frequency point in the audio data of the first direction.
[0023] In conjunction with the first aspect, in some implementations of the first aspect, the video processing method further includes: Speech detection is performed on the audio data in the first direction to obtain a first detection result; The direction of arrival (DOA) of the data collected by the at least two microphones is estimated to obtain the predicted angle information. If the first detection result indicates that the audio data in the first direction includes the user's audio information, the amplitude spectrum of the audio data in the first direction is subjected to the first amplification process; and / or If the predicted angle information includes angle information within the first preset angle range, the amplitude spectrum of the audio data in the first direction is subjected to the second amplification process.
[0024] In conjunction with the first aspect, in some implementations of the first aspect, the second amplitude spectrum is the amplitude spectrum obtained by performing a first amplification process and / or a second amplification on the second average amplitude spectrum, and the second average amplitude spectrum is obtained by averaging the amplitude spectra corresponding to each frequency point in the audio data of the second direction.
[0025] In conjunction with the first aspect, in some implementations of the first aspect, the video processing method further includes: Speech detection is performed on the audio data in the second direction to obtain a second detection result; The direction of arrival (DOA) of the data collected by the at least two microphones is estimated to obtain the predicted angle information. If the second detection result indicates that the audio data in the second direction includes the user's audio information, the amplitude spectrum of the audio data in the second direction is subjected to the first amplification process; and / or If the predicted angle information includes angle information within the second preset angle range, the amplitude spectrum of the audio data in the second direction is subjected to the second amplification process.
[0026] It should be understood that in a video recording scenario, the direction in which the user is located can usually be considered the main shooting direction; if the detection result indicates that the direction includes the user's audio information, then the user can be considered to be in that direction; at this time, the audio data in that direction can be subjected to a first amplification process, which can improve the accuracy of the acquired user audio information.
[0027] Direction of arrival (DOA) estimation refers to an algorithm that performs a spatial Fourier transform on the received signal, takes the square of the modulus to obtain the spatial spectrum, and estimates the signal's direction of arrival.
[0028] It should be understood that in video recording scenarios, the user's direction is usually considered the primary shooting direction. If the detection result indicates that the user's audio information is included in that direction, then the user can be considered to be in that direction. In this case, the audio data in that direction can be amplified in the first way, which can improve the accuracy of the acquired user audio information. When the predicted angle information includes a first preset angle range and / or a second preset angle range, it can be said that audio information exists in the first direction and / or the second direction of the electronic device. The accuracy of the first amplitude spectrum or the second amplitude spectrum can be improved through the second amplification process. With the improved accuracy of the amplitude spectrum and the user audio information, the switching command can be obtained accurately.
[0029] In conjunction with the first aspect, in some implementations of the first aspect, identifying whether the audio data includes target keywords includes: The audio data is separated using a blind signal separation algorithm to obtain N audio information pieces, which are audio information pieces from different users; Each of the N audio messages is identified to determine whether the target keyword is included in the N audio messages.
[0030] In the embodiments of this application, the audio data collected by at least two pickup devices can be separated and processed to obtain N audio information from different sources; whether the target keyword is included in each of the N audio information can be identified, thereby improving the accuracy of identifying the target keyword.
[0031] In conjunction with the first aspect, in some implementations of the first aspect, the first image is a preview image captured by the electronic device when it is recording with multiple cameras.
[0032] In conjunction with the first aspect, in some implementations of the first aspect, the first image is a video frame captured by the electronic device when it is recording with multiple cameras.
[0033] In conjunction with the first aspect, in some implementations of the first aspect, the audio data refers to the data collected by the sound pickup device in the shooting environment of the electronic device.
[0034] In a second aspect, an electronic device is provided, the electronic device including one or more processors, a memory, and at least two pickup devices; the memory is coupled to the one or more processors, the memory being used to store computer program code, the computer program code including computer instructions, and the one or more processors calling the computer instructions to cause the electronic device to execute: Run the camera application in the electronic device; Display a first image, which is an image captured by the electronic device when it is in a first shooting mode; Acquire audio data, wherein the audio data is data collected by the at least two microphones; A switching instruction is obtained based on the audio data; the switching instruction is used to instruct the electronic device to switch from the first shooting mode to the second shooting mode. Display a second image, which is an image captured by the electronic device when it is in the second shooting mode.
[0035] In conjunction with the second aspect, in some implementations of the second aspect, the electronic device includes a first camera and a second camera, the first camera and the second camera being located in different directions of the electronic device, and the one or more processors invoking computer instructions to cause the electronic device to execute: Identify whether the audio data includes a target keyword, where the target keyword is the text information corresponding to the switching instruction; If the target keyword is identified in the audio data, the switching instruction is obtained based on the target keyword; If the target keyword is not identified in the audio data, the audio data is processed to obtain audio data in a first direction and / or audio data in a second direction. The first direction is used to represent a first preset angle range corresponding to the first camera, and the second direction is used to represent a second preset angle range corresponding to the second camera. Based on the audio data in the first direction and / or the audio data in the second direction, the switching instruction is obtained.
[0036] In conjunction with the second aspect, in some implementations of the second aspect, the processing of the audio data to obtain audio data in the first direction and / or audio data in the second direction includes: The audio data is processed based on a sound direction probability calculation algorithm to obtain audio data in the first direction and / or audio data in the second direction.
[0037] In conjunction with the second aspect, in some implementations of the second aspect, the one or more processors invoke the computer instructions to cause the electronic device to execute: The switching instruction is obtained based on the energy of the first amplitude spectrum and / or the energy of the second amplitude spectrum, wherein the first amplitude spectrum is the amplitude spectrum of the audio data in the first direction, and the second amplitude spectrum is the amplitude spectrum of the audio data in the second direction.
[0038] In conjunction with the second aspect, in some implementations of the second aspect, the switching instruction includes the current shooting mode, a first picture-in-picture mode, a second picture-in-picture mode, a first dual-view mode, a second dual-view mode, a single-camera mode of the first camera, or a single-camera mode of the second camera, and the one or more processors invoke the computer instruction to cause the electronic device to execute: If the energy of the first amplitude spectrum and the energy of the second amplitude spectrum are both less than the first preset threshold, the switching instruction is to maintain the current shooting mode. If the energy of the first amplitude spectrum is greater than the second preset threshold, and the energy of the second amplitude spectrum is less than or equal to the second preset threshold, the switching instruction is to switch to the single-camera mode of the first camera; If the energy of the second amplitude spectrum is greater than the second preset threshold, and the energy of the first amplitude spectrum is less than or equal to the second preset threshold, the switching instruction is to switch to the single-camera mode of the second camera; If the energy of the first amplitude spectrum is greater than the second preset threshold, and the energy of the second amplitude spectrum is greater than or equal to the first preset threshold, the switching instruction is to switch to the first picture-in-picture mode; If the energy of the second amplitude spectrum is greater than the second preset threshold, and the energy of the first amplitude spectrum is greater than or equal to the first preset threshold, the switching instruction is to switch to the second picture-in-picture mode; If the energy of the first amplitude spectrum and the energy of the second amplitude spectrum are both greater than or equal to the second preset threshold, and the energy of the first amplitude spectrum is greater than the energy of the second amplitude spectrum, the switching instruction is to switch to the first dual-view mode; If the energy of the first amplitude spectrum and the energy of the second amplitude spectrum are both greater than or equal to the second preset threshold, and the energy of the second amplitude spectrum is greater than the energy of the first amplitude spectrum, the switching instruction is to switch to the second dual-view mode; Wherein, the second preset threshold is greater than the first preset threshold, the first picture-in-picture mode refers to the shooting mode in which the image captured by the first camera is the main picture, the second picture-in-picture mode refers to the shooting mode in which the image captured by the second camera is the main picture, the first dual-view mode refers to the shooting mode in which the image captured by the first camera is located on the top or left side of the display screen of the electronic device, and the second dual-view mode refers to the shooting mode in which the image captured by the second camera is located on the top or left side of the display screen of the electronic device.
[0039] In conjunction with the second aspect, in some implementations of the second aspect, the first amplitude spectrum is a first average amplitude spectrum obtained by averaging the amplitude spectra corresponding to each frequency point in the audio data in the first direction; and / or, The second amplitude spectrum is the second average amplitude spectrum obtained by averaging the amplitude spectra corresponding to each frequency point in the audio data in the second direction.
[0040] In conjunction with the second aspect, in some implementations of the second aspect, the first amplitude spectrum is an amplitude spectrum obtained by performing a first amplification process and / or a second amplification on the first average amplitude spectrum, and the first average amplitude spectrum is obtained by averaging the amplitude spectra corresponding to each frequency point in the audio data of the first direction.
[0041] In conjunction with the second aspect, in some implementations of the second aspect, the one or more processors invoke the computer instructions to cause the electronic device to execute: Speech detection is performed on the audio data in the first direction to obtain a first detection result; The direction of arrival (DOA) of the data collected by the at least two microphones is estimated to obtain the predicted angle information. If the first detection result indicates that the audio data in the first direction includes the user's audio information, the amplitude spectrum of the audio data in the first direction is subjected to the first amplification process; and / or If the predicted angle information includes angle information within the first preset angle range, the amplitude spectrum of the audio data in the first direction is subjected to the second amplification process.
[0042] In conjunction with the second aspect, in some implementations of the second aspect, the second amplitude spectrum is the amplitude spectrum obtained after performing a first amplification process and / or a second amplification on the second average amplitude spectrum, and the second average amplitude spectrum is obtained by averaging the amplitude spectra corresponding to each frequency point in the audio data of the second direction.
[0043] In conjunction with the second aspect, in some implementations of the second aspect, the one or more processors invoke the computer instructions to cause the electronic device to execute: Speech detection is performed on the audio data in the second direction to obtain a second detection result; The direction of arrival (DOA) of the data collected by the at least two microphones is estimated to obtain the predicted angle information. If the second detection result indicates that the audio data in the second direction includes the user's audio information, the amplitude spectrum of the audio data in the second direction is subjected to the first amplification process; and / or If the predicted angle information includes angle information within the second preset angle range, the amplitude spectrum of the audio data in the second direction is subjected to the second amplification process.
[0044] In conjunction with the second aspect, in some implementations of the second aspect, the one or more processors invoke the computer instructions to cause the electronic device to execute: The audio data is separated using a blind signal separation algorithm to obtain N audio information pieces, which are audio information pieces from different users; Each of the N audio messages is identified to determine whether the target keyword is included in the N audio messages.
[0045] In conjunction with the second aspect, in some implementations of the second aspect, the first image is a preview image captured by the electronic device when it is recording with multiple cameras.
[0046] In conjunction with the second aspect, in some implementations of the second aspect, the first image is a video frame captured by the electronic device when it is recording with multiple cameras.
[0047] In conjunction with the second aspect, in some implementations of the second aspect, the audio data refers to the data collected by the sound pickup device in the shooting environment of the electronic device.
[0048] Thirdly, an electronic device is provided, including a module / unit for performing the video processing method of the first aspect or any of the first aspects.
[0049] Fourthly, an electronic device is provided, the electronic device including one or more processors and a memory; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code including computer instructions, and the one or more processors call the computer instructions to cause the electronic device to perform the first aspect or any one of the methods in the first aspect.
[0050] Fifthly, a chip system is provided, the chip system being applied to an electronic device, the chip system including one or more processors, the processors being configured to invoke computer instructions to cause the electronic device to perform the first aspect or any of the methods in the first aspect.
[0051] In a sixth aspect, a computer-readable storage medium is provided, the computer-readable storage medium storing computer program code that, when executed by an electronic device, causes the electronic device to perform the first aspect or any one of the methods in the first aspect.
[0052] In a seventh aspect, a computer program product is provided, the computer program product comprising: computer program code, which, when executed by an electronic device, causes the electronic device to perform the method of the first aspect or any one of the first aspects.
[0053] In embodiments of this application, the electronic device can collect audio data in the shooting environment through at least two sound pickup devices (e.g., microphones); generate a switching command based on the audio data, and the electronic device automatically switches from the current first shooting mode to the second shooting mode based on the switching command, and displays the second image captured in the second shooting mode; without the user needing to switch the shooting mode of the electronic device, the electronic device can automatically switch shooting modes to complete video recording, thereby improving the user's shooting experience. Attached Figure Description
[0054] Figure 1 This is a schematic diagram of a hardware system for an electronic device applicable to this application; Figure 2 This is a schematic diagram of a software system applicable to an electronic device of this application; Figure 3 This is a schematic diagram illustrating an application scenario applicable to an embodiment of this application; Figure 4 This is a schematic diagram illustrating an application scenario applicable to an embodiment of this application; Figure 5 This is a schematic diagram illustrating an application scenario applicable to an embodiment of this application; Figure 6 This is a schematic diagram illustrating an application scenario applicable to an embodiment of this application; Figure 7 This is a schematic flowchart of a video processing method provided in an embodiment of this application; Figure 8 This is a schematic flowchart of a video processing method provided in an embodiment of this application; Figure 9 This is a schematic flowchart of a video processing method provided in an embodiment of this application; Figure 10 This is a schematic diagram of the target angle of an electronic device provided in an embodiment of this application; Figure 11 This is a schematic flowchart illustrating a method for identifying switching instructions provided in an embodiment of this application; Figure 12 This is a schematic diagram of a direction-of-arrival estimation provided in an embodiment of this application; Figure 13 This is a schematic diagram of a graphical user interface applicable to embodiments of this application; Figure 14 This is a schematic diagram of a graphical user interface applicable to embodiments of this application; Figure 15 This is a schematic diagram of a graphical user interface applicable to embodiments of this application; Figure 16 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application; Figure 17 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0055] In the embodiments of this application, the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this embodiment, unless otherwise stated, "multiple" means two or more.
[0056] To facilitate understanding of the embodiments of this application, the relevant concepts involved in the embodiments of this application will be briefly explained first.
[0057] 1. Fourier Transform The Fourier transform is a linear integral transform used to represent the transformation of a signal between the time domain (or, spatial domain) and the frequency domain.
[0058] 2. Fast Fourier Transform (FFT) FFT stands for Discrete Fourier Transform, a fast algorithm that can transform a signal from the time domain to the frequency domain.
[0059] 3. Blind signal separation (BSS) Blind signal separation refers to algorithms that recover individual source signals from acquired mixed signals (usually the outputs of multiple sensors).
[0060] 4. Beamforming Based on the frequency domain signal obtained by performing FFT transformation on the input signal acquired by a non-pickup device (e.g., a microphone) and the filter coefficients at different angles, beamforming results at different angles can be obtained.
[0061] For example, ;in, This indicates the beamforming results at different angles; Represents the filter coefficients at different angles; This represents the frequency domain signal obtained after performing an FFT transformation on the input signal acquired by the microphone; i represents the pickup signal of the i-th microphone; M represents the number of microphones.
[0062] 5. Voice activity detection (VAD) Speech activity detection is a technique used in speech processing to detect the presence of speech signals.
[0063] 6. Direction of arrival (DPA) estimation Direction of arrival (DOA) estimation is an algorithm that estimates the direction of arrival of a signal by performing a spatial Fourier transform on the received signal, taking the square of the modulus to obtain the spatial spectrum.
[0064] 7. Based on the time difference of arrival (TDOA) TDOA is used to represent the time difference between the arrival of a sound source at different microphones in an electronic device.
[0065] 8. Generalized cross correlation-phase transform (GCC-PHAT) GCC-PHAT is an algorithm for calculating the angle of arrival (AOA), such as... Figure 12 As shown.
[0066] 9. Estimating signal parameter via rotational invariance techniques (ESPRIT) ESRT refers to a rotation invariant technique algorithm, which is based on the principle of estimating signal parameters without writing code based on signal rotation.
[0067] 10. Positioning algorithm for controllable beamforming The principle of the localization algorithm of the controllable beamforming method is to filter and weight the signal received by the microphone to form a beam, and search for the sound source location according to a certain rule. When the microphone reaches its maximum output power, the searched sound source location is the true sound source location.
[0068] 11. Cepstral Algorithm The cepstral algorithm is a method used in signal processing and signal detection. The cepstral refers to the power spectrum of the logarithmic power spectrum of a signal. The principle of obtaining speech through the cepstral is as follows: since voiced signals are periodically excited, the voiced signals are periodic impulses on the cepstral, so the fundamental frequency can be obtained. Generally, the second impulse in the cepstral waveform (the first is the envelope information) is considered to be the fundamental frequency of the excitation source.
[0069] 12. Inverse Discrete Fourier Transform (IDFT) IDFT refers to the inverse transform, which is the reverse process of the Fourier transform.
[0070] 13. Complex angular central Gaussian mixture model (cACGMM) cACGMM is a Gaussian mixture model; a Gaussian mixture model refers to a model that uses a Gaussian probability density function (e.g., a normal distribution curve) to accurately quantify things, decomposing a thing into several models based on a Gaussian probability density function (e.g., a normal distribution curve).
[0071] 14. Amplitude Spectrum After transforming the signal to the frequency domain, the amplitude spectrum can be obtained by performing a modulus-taking operation on the signal.
[0072] 15. Multi-camera recording like Figure 4 As shown in (a), multi-lens video recording can refer to a camera mode in a camera application similar to video recording or shooting; multi-lens video recording can include various different shooting modes; for example, such as Figure 4As shown in (b), the shooting modes may include, but are not limited to: front / rear dual camera mode, rear / front dual camera mode, picture-in-picture 1 mode, picture-in-picture 2 mode, rear single camera mode, or front single camera mode, etc.
[0073] The video processing method and electronic device in the embodiments of this application will now be described with reference to the accompanying drawings.
[0074] Figure 1 A hardware system for an electronic device applicable to this application is shown.
[0075] Electronic device 100 can be a mobile phone, smart screen, tablet computer, wearable electronic device, in-vehicle electronic device, augmented reality (AR) device, virtual reality (VR) device, laptop computer, ultra-mobile personal computer (UMPC), netbook, personal digital assistant (PDA), projector, etc. This application embodiment does not impose any restrictions on the specific type of electronic device 100.
[0076] Electronic device 100 may include processor 110, external memory interface 120, internal memory 121, universal serial bus (USB) interface 130, charging management module 140, power management module 141, battery 142, antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, sensor module 180, button 190, motor 191, indicator 192, camera 193, display screen 194, and subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an accelerometer sensor 180E, a distance sensor 180F, a proximity sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.
[0077] For example, the audio module 170 is used to convert digital audio information into analog audio signal output, and can also be used to convert analog audio input into digital audio signal. The audio module 170 can also be used to encode and decode audio signals. In some embodiments, the audio module 170 or some functional modules of the audio module 170 may be located in the processor 110.
[0078] For example, in an embodiment of this application, the audio module 170 can send audio data collected by the microphone to the processor 110.
[0079] It should be noted that, Figure 1 The structure shown does not constitute a specific limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may include... Figure 1 The components shown may include more or fewer components, or the electronic device 100 may include... Figure 1 The components shown may be a combination of certain components, or the electronic device 100 may include... Figure 1 Sub-components of some of the components shown. Figure 1 The components shown can be implemented in hardware, software, or a combination of software and hardware.
[0080] Processor 110 may include one or more processing units. For example, processor 110 may include at least one of the following processing units: application processor (AP), modem processor, graphics processing unit (GPU), image signal processor (ISP), controller, video codec, digital signal processor (DSP), baseband processor, and neural network processing unit (NPU). These different processing units may be independent devices or integrated devices. The controller can generate operation control signals based on instruction opcodes and timing signals to control instruction fetching and execution.
[0081] The processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store instructions or data that the processor 110 has just used or that are used repeatedly. If the processor 110 needs to use the instruction or data again, it can retrieve it directly from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.
[0082] For example, the processor 110 can be used to execute the video processing method of the embodiments of this application; for example, running a camera application in an electronic device; displaying a first image, the first image being an image captured when the electronic device is in a first shooting mode; acquiring audio data, the audio data being data captured by at least two microphones; obtaining a switching instruction based on the audio data, the switching instruction being used to instruct the electronic device to switch from the first shooting mode to a second shooting mode; and displaying a second image, the second image being an image captured when the electronic device is in a second shooting mode.
[0083] Figure 1 The connection relationships between the modules shown are merely illustrative and do not constitute a limitation on the connection relationships between the modules of the electronic device 100. Optionally, the modules of the electronic device 100 may also adopt a combination of various connection methods described in the above embodiments.
[0084] The wireless communication function of electronic device 100 can be realized through devices such as antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modem processor and baseband processor.
[0085] Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in electronic device 100 can be used to cover one or more communication frequency bands. Different antennas can also be multiplexed to improve antenna utilization. For example, antenna 1 can be multiplexed as a diversity antenna for a wireless local area network. In some other embodiments, the antennas can be used in conjunction with tuning switches.
[0086] Electronic device 100 can implement display functions through a GPU, a display screen 194, and an application processor. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations and for graphics rendering. Processor 110 may include one or more GPUs, which execute program instructions to generate or modify display information.
[0087] Display screen 194 can be used to display images or videos.
[0088] Electronic device 100 can perform shooting functions through ISP, camera 193, video codec, GPU, display screen 194 and application processor.
[0089] The ISP (Image Signal Processor) is used to process data fed back from the camera 193. For example, when taking a picture, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, and the camera's photosensitive element transmits the electrical signal to the ISP for processing, transforming it into an image visible to the naked eye. The ISP can perform algorithmic optimization of image noise, brightness, and color. The ISP can also optimize parameters such as exposure and color temperature of the shooting scene. In some embodiments, the ISP can be set in the camera 193.
[0090] Camera 193 is used to capture still images or videos. An object is projected onto a photosensitive element by generating an optical image through the lens. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then passed to an ISP for conversion into a digital image signal. The ISP outputs the digital image signal to a DSP for processing. The DSP converts the digital image signal into image signals in standard red-green-blue (RGB), YUV, or other formats. In some embodiments, the electronic device 100 may include one or N cameras 193, where N is a positive integer greater than 1.
[0091] For example, in an embodiment of this application, the electronic device may include a plurality of cameras 193; the plurality of cameras may include a front-facing camera and a rear-facing camera.
[0092] Digital signal processors (DSPs) are used to process digital signals. Besides digital image signals, they can also process other digital signals. For example, when electronic device 100 selects a frequency, the DSP can perform Fourier transforms on the frequency energy.
[0093] Video codecs are used to compress or decompress digital video. Electronic device 100 may support one or more video codecs. Thus, electronic device 100 can play or record video in various encoding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, and MPEG4.
[0094] The gyroscope sensor 180B can be used to determine the motion attitude of the electronic device 100. In some embodiments, the gyroscope sensor 180B can determine the angular velocity of the electronic device 100 around three axes (i.e., the x-axis, y-axis, and z-axis). The gyroscope sensor 180B can be used for image stabilization. For example, when the shutter is pressed, the gyroscope sensor 180B detects the angle of the shake of the electronic device 100, calculates the distance that the lens module needs to compensate based on the angle, and allows the lens to counteract the shake of the electronic device 100 by moving in the opposite direction, thus achieving image stabilization. The gyroscope sensor 180B can also be used in scenarios such as navigation and motion-sensing games.
[0095] The accelerometer 180E can detect the magnitude of acceleration of the electronic device 100 in various directions (typically the x-axis, y-axis, and z-axis). When the electronic device 100 is stationary, it can detect the magnitude and direction of gravity. The accelerometer 180E can also be used to identify the attitude of the electronic device 100, serving as input parameters for applications such as screen orientation switching and pedometers.
[0096] The distance sensor 180F is used to measure distance. The electronic device 100 can measure distance using infrared or laser. In some embodiments, such as in a shooting scenario, the electronic device 100 can utilize the distance sensor 180F to measure distance for fast focusing.
[0097] The ambient light sensor 180L is used to sense the brightness of ambient light. The electronic device 100 can adaptively adjust the brightness of the display screen 194 based on the sensed ambient light brightness. The ambient light sensor 180L can also be used to automatically adjust the white balance when taking pictures. The ambient light sensor 180L can also work with the proximity sensor 180G to detect whether the electronic device 100 is in a pocket to prevent accidental touches.
[0098] The fingerprint sensor 180H is used to collect fingerprints. The electronic device 100 can use the collected fingerprint characteristics to perform functions such as unlocking, accessing application locks, taking photos, and answering calls.
[0099] Touch sensor 180K, also known as a touch device, can be disposed on display screen 194. The touch sensor 180K and display screen 194 together form a touchscreen, also known as a touch screen. Touch sensor 180K is used to detect touch operations applied to or near it. Touch sensor 180K can transmit the detected touch operation to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through display screen 194. In other embodiments, touch sensor 180K may also be disposed on the surface of electronic device 100, and in a different location from display screen 194.
[0100] The hardware system of electronic device 100 has been described in detail above. The software system of electronic device 100 will be introduced below.
[0101] Figure 2 This is a schematic diagram of the software system of the electronic device provided in the embodiments of this application.
[0102] like Figure 2 As shown, the system architecture may include an application layer 210, an application framework layer 220, a hardware abstraction layer 230, a driver layer 240, and a hardware layer 250.
[0103] Application layer 210 may include applications such as camera application, gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, video, and SMS; application layer 210 can be further divided into application interface and application logic; among them, the application interface of camera application may include single-view mode, dual-view mode, picture-in-picture mode, etc., corresponding to different video shooting modes.
[0104] The application framework layer 220 provides application programming interfaces (APIs) and programming frameworks for applications in the application layer; the application framework layer may include some predefined functions.
[0105] For example, the application framework layer 220 may include a camera access interface; the camera access interface may include camera management and camera devices. Specifically, camera management can be used to provide an access interface for managing cameras; camera devices can be used to provide an interface for accessing cameras.
[0106] The hardware abstraction layer 230 is used to abstract hardware. For example, the hardware abstraction layer can encompass the camera abstraction layer and other hardware device abstraction layers; the camera hardware abstraction layer can call camera algorithms.
[0107] For example, the hardware abstraction layer 230 includes a camera hardware abstraction layer and a camera algorithm; the camera algorithm may include software algorithms for video processing or image processing.
[0108] For example, the algorithm in a camera algorithm can refer to something that does not depend on specific hardware implementation; for example, code that can typically run on a CPU.
[0109] The driver layer 240 is used to provide drivers for different hardware devices. For example, the driver layer may include a camera driver.
[0110] Hardware layer 250 is located at the lowest level of the operating system; such as Figure 2 As shown, hardware layer 250 may include camera 1, camera 2, camera 3, etc. Camera 1, camera 2, and camera 3 may correspond to multiple cameras on an electronic device. For example, the video processing method and electronic device provided in the embodiments of this application can run in a hardware abstraction layer; or, can run in an application framework layer; or, can run in a digital signal processor.
[0111] Currently, switching shooting modes on electronic devices (e.g., cameras) relies on manual operation by the user, requiring the user to be relatively close to the electronic device during shooting. If the user is far from the electronic device, Bluetooth technology is needed to switch shooting modes. However, switching shooting modes on Bluetooth requires controlling the camera lens, which is complex and exposes the control device in the video, affecting its aesthetics and resulting in a poor user experience.
[0112] In view of this, embodiments of this application provide a video processing method. During the process of a user shooting video using an electronic device, the electronic device can obtain switching instructions based on audio data in the shooting environment. Based on the switching instructions, the electronic device can automatically switch the shooting mode of the electronic device; for example, it can automatically switch between different cameras in the electronic device; for example, the electronic device can automatically determine whether to switch cameras, or whether to start multi-camera recording, or whether to switch different shooting modes in multi-camera recording, etc., so that video recording can be completed without the user switching the shooting mode of the electronic device, achieving a one-shot recording experience.
[0113] It should be understood that "one-shot" means that after a user selects a certain shooting mode, the user does not need to perform any corresponding operation to switch shooting modes; the electronic device can automatically generate switching instructions based on the audio data collected in the shooting environment; and the electronic device automatically switches shooting modes based on the switching instructions.
[0114] The following is combined with Figures 3 to 15 The video processing method provided in the embodiments of this application will be described in detail.
[0115] For example, the video processing method in this application embodiment can be applied to the fields of video recording, video calling, or other image processing. In this application embodiment, audio data in the shooting environment is collected by at least two sound pickup devices (e.g., microphones) in the electronic device. A switching command is generated based on the audio data, and the electronic device automatically switches from a first shooting mode to a second shooting mode based on the switching command, displaying a second image captured in the second shooting mode. Without requiring the user to switch the shooting mode of the electronic device, the electronic device can automatically switch shooting modes to complete video recording, improving the user's shooting experience.
[0116] In one example, the video processing method in this application embodiment can be applied to the preview state of a recorded video.
[0117] like Figure 3As shown, the electronic device is in the preview state of multi-lens recording. The current shooting mode of the electronic device can be set to front / rear dual-view shooting mode by default. The foreground image can be as shown in image 251, and the background image can be as shown in image 252. The foreground image can refer to the image captured by the front camera of the electronic device, and the background image can refer to the image captured by the rear camera of the electronic device. The following is an example using image 251 as the foreground image and image 252 as the background image.
[0118] like Figure 4 As shown in (a), after the electronic device detects an operation on the control 260 for the multi-camera recording shooting mode, the electronic device can display various different shooting modes in the multi-camera recording; for example, the various different shooting modes may include, but are not limited to: front / rear dual-camera mode, rear / front dual-camera mode, picture-in-picture 1 mode (rear picture-in-picture mode), picture-in-picture 2 mode (front picture-in-picture mode), rear camera single-camera mode, or front camera single-camera mode, etc. Figure 4 As shown in (b) of this application, using the video processing method in this embodiment, when the electronic device is in the preview state of multi-camera recording, at least two sound pickup devices (e.g., microphones) in the electronic device collect audio data from the shooting environment; a switching command is generated based on the audio data, and the electronic device automatically switches from the first shooting mode to the second shooting mode based on the switching command, displaying the second image captured in the second shooting mode; for example, assuming that when the electronic device enters the preview state of multi-camera recording, the default shooting mode is as follows: Figure 3 The front / rear dual-camera mode shown is based on the switching instruction obtained from the audio data collected by at least two microphones in the electronic device, which switches the shooting mode of the electronic device to the rear picture-in-picture mode. Then, without user operation, the electronic device can automatically switch from the front / rear dual-camera mode to the rear picture-in-picture mode and display a second image, which is a preview image.
[0119] Among them, single-camera mode can include front single-camera mode, rear single-camera mode, etc.; multi-camera mode can include front / rear dual-camera mode, rear / front dual-camera mode, picture-in-picture 1 mode, picture-in-picture 2 mode, etc.
[0120] Optionally, the multi-camera mode may also include a front dual-camera mode or a rear dual-camera mode, etc.
[0121] It should be understood that in single-camera mode, one camera in the electronic device is used for video recording; in multi-camera mode, two or more cameras in the electronic device are used for video recording.
[0122] For example, in front-facing single-camera mode, one front-facing camera is used for video recording; in rear-facing single-camera mode, one rear-facing camera is used for video recording; in front-facing dual-camera mode, two front-facing cameras are used for video recording; in rear-facing dual-camera mode, two rear-facing cameras are used for video recording; in front-and-rear dual-camera mode, one front-facing camera and one rear-facing camera are used for video recording; in front-facing picture-in-picture mode, two front-facing cameras are used for video recording, and the image captured by one front-facing camera is placed within the image captured by the other front-facing camera; in rear-facing picture-in-picture mode, two rear-facing cameras are used for video recording, and the image captured by one rear-facing camera is placed within the image captured by the other rear-facing camera; in front-and-rear picture-in-picture mode, one front-facing camera and one rear-facing camera are used for video recording, and the image captured by either the front-facing or rear-facing camera is placed within the image captured by either the rear-facing or front-facing camera.
[0123] It should be understood that Figure 4 The image shows the shooting interface for different shooting modes of multi-camera video recording on an electronic device in portrait mode. Figure 5 The image shown is the shooting interface for different shooting modes of multi-camera video recording on an electronic device in landscape mode; among them, Figure 4 (a) and Figure 5 (a) in the text corresponds to, Figure 4 (b) and Figure 5 (b) corresponds to: The electronic device can determine whether to display in portrait or landscape mode based on the user's state when using the electronic device.
[0124] In one example, the video processing method in this application embodiment can be applied to the process of recording video.
[0125] like Figure 6 As shown, the electronic device is in multi-camera video recording mode. The current shooting mode of the electronic device can be set to front / rear dual-view shooting mode by default, such as... Figure 6 As shown in (a), when the electronic device detects an operation on the control 270 for the multi-camera recording shooting mode at the 5th second of video recording, the electronic device can display various different shooting modes in multi-camera recording, such as... Figure 6As shown in (b) of this application, when the electronic device is in a multi-camera recording state, at least two audio pickup devices (e.g., microphones) in the electronic device collect audio data in the shooting environment using the video processing method in this embodiment. A switching instruction is generated based on the audio data, and the electronic device automatically switches from the first shooting mode to the second shooting mode based on the switching instruction, displaying the second image captured in the second shooting mode. For example, assuming the electronic device is currently recording video, and the default front / rear dual-camera shooting mode is used when the electronic device starts recording video, the switching instruction obtained based on the audio data collected by at least two audio pickup devices in the electronic device is to switch the shooting mode of the electronic device to the rear picture-in-picture mode. Then, without user operation, the electronic device can automatically switch from the front / rear dual-camera mode to the rear picture-in-picture mode and display the second image, which is a video frame.
[0126] It should be understood that the above example of multi-camera recording can also be applied to video calls, video conferencing applications, long and short video applications, live video applications, online video courses, portrait intelligent camera movement applications, video recording by system camera function, video surveillance, or smart doorbell camera shooting scenarios.
[0127] In one example, the video processing method in this embodiment can also be applied to the recording state of an electronic device. For example, when the electronic device is in recording state, it can adopt the default rear single-camera shooting mode, and at least two audio pickup devices (e.g., microphones) in the electronic device collect audio data in the shooting environment. A switching command is generated based on the audio data, and the electronic device can automatically switch from the rear single-camera mode to the front single-camera mode based on the switching command; or, the electronic device can automatically switch from the single-camera mode to the multi-camera mode based on the switching command, and display the second image captured in the second shooting mode; the second image can be a preview image, or the second image can also be a video frame. Optionally, the video processing method in this embodiment can also be applied to the field of photography; for example, when the electronic device is in recording state, it can adopt the default rear single-camera shooting mode, and at least two audio pickup devices (e.g., microphones) in the electronic device collect audio data in the shooting environment. A switching command is generated based on the audio data, and the electronic device can automatically switch from the rear single-camera mode to the front single-camera mode based on the switching command, and display the second image captured in the second shooting mode; the second image can be a preview image, or the second image can also be a video frame.
[0128] It should be understood that the above are illustrative examples of application scenarios and do not limit the application scenarios of this application in any way.
[0129] Figure 7This is a schematic flowchart of a video processing method provided in an embodiment of this application. The video processing method 300 includes components that can be... Figure 1 The electronic device shown performs the video processing method, which includes steps S310 to S350. Steps S310 to S350 are described in detail below.
[0130] Step S310: Run the camera application on the electronic device.
[0131] For example, a user can instruct the electronic device to run the camera app by clicking the "Camera" app icon. Alternatively, when the electronic device is locked, the user can instruct the device to run the camera app by swiping right on the screen. Or, if the electronic device is locked and the lock screen includes a camera app icon, the user can instruct the device to run the camera app by clicking the icon. Alternatively, if the electronic device is running another app that has permission to access the camera app, the user can instruct the device to run the camera app by clicking the corresponding control. For example, if the electronic device is running an instant messaging application, the user can instruct the device to run the camera app by selecting the camera function control, and so on.
[0132] Step S320: Display the first image.
[0133] The first image is the image captured when the electronic device is in the first shooting mode.
[0134] For example, the first shooting mode can refer to either a single-camera mode or a multi-camera mode; wherein, the single-camera mode can include a front single-camera mode or a rear single-camera mode; the multi-camera mode can include a front / rear dual-camera mode, a rear / front dual-camera mode, a picture-in-picture front main camera mode, or a picture-in-picture rear main camera mode.
[0135] For example, in front-facing single-camera mode, video recording is performed using one of the front-facing cameras in the electronic device; in rear-facing single-camera mode, video recording is performed using one of the rear-facing cameras in the electronic device; in front and rear dual-camera mode, video recording is performed using one front-facing camera and one rear-facing camera; in picture-in-picture front-facing mode, video recording is performed using one front-facing camera and one rear-facing camera, with the image captured by the rear-facing camera placed within the image captured by the front-facing camera, and the image captured by the front-facing camera being the main image; in picture-in-picture rear-facing mode, video recording is performed using one front-facing camera and one rear-facing camera, with the image captured by the front-facing camera placed within the image captured by the rear-facing camera, and the image captured by the rear-facing camera being the main image.
[0136] Optionally, the multi-camera mode may also include a front dual-camera mode, a rear dual-camera mode, a front picture-in-picture mode, or a rear picture-in-picture mode.
[0137] Optionally, when the electronic device is in video preview mode, the first image is a preview image.
[0138] Optionally, when the electronic device is recording, the first image is a video frame.
[0139] Optionally, when the electronic device is in multi-camera recording preview mode, the first image is the preview image.
[0140] Optionally, when the electronic device is recording with multiple cameras, the first image is a video frame.
[0141] Step S330: Obtain audio data.
[0142] The audio data refers to data collected by at least two pickup devices in an electronic device; for example, data collected by at least two microphones.
[0143] It should be understood that, in the embodiments of this application, since the electronic device needs to determine the directionality of the audio data, the electronic device in the embodiments of this application includes at least two pickup devices, and there is no limitation on the specific number of pickup devices.
[0144] For example, as follows Figure 9 As shown, the electronic device includes three microphones.
[0145] For example, audio data can refer to data collected by a sound pickup device in the shooting environment where the electronic device is located.
[0146] Step S340: Obtain the switching instruction based on the audio data.
[0147] The switching command is used to instruct the electronic device to switch from the first shooting mode to the second shooting mode.
[0148] It should be understood that the first shooting mode and the second shooting mode can be the same shooting mode or different shooting modes; if the switching command is the default current shooting mode, the second shooting mode and the first shooting mode can be the same shooting mode, as shown by identifier 0 in Table 1; in other cases, the second shooting mode and the first shooting mode can be different shooting modes, as shown by identifiers 1 to 6 in Table 1.
[0149] Step 350: Display the second image.
[0150] The second image is the image captured when the electronic device is in the second shooting mode.
[0151] In embodiments of this application, the electronic device can collect audio data in the shooting environment through at least two sound pickup devices (e.g., microphones); generate a switching command based on the audio data, and the electronic device automatically switches from the current first shooting mode to the second shooting mode based on the switching command, and displays the second image captured in the second shooting mode; without the user needing to switch the shooting mode of the electronic device, the electronic device can automatically switch shooting modes to complete video recording, thereby improving the user's shooting experience.
[0152] For example, the electronic device may include a first camera (e.g., a front-facing camera) and a second camera (e.g., a rear-facing camera), the first camera and the second camera may be located in different directions of the electronic device; the aforementioned switching instruction based on audio data includes: Identify whether the audio data contains target keywords, where the target keywords are the text information corresponding to the switching command; If a target keyword is identified in the audio data, a switching instruction is obtained based on the target keyword. If no target keyword is identified in the audio data, the audio data is processed to obtain audio data in a first direction and / or audio data in a second direction. The first direction is used to represent a first preset angle range corresponding to the first camera, and the second direction is used to represent a second preset angle range corresponding to the second camera. Based on the audio data in the first direction and / or the audio data in the second direction, a switching instruction is obtained.
[0153] In the embodiments of this application, it can first be identified whether the audio data includes the target keyword; if the audio data includes the target keyword, the electronic device switches the shooting mode to the second shooting mode corresponding to the target keyword; if the audio data does not include the target keyword, the electronic device can obtain a switching instruction based on the audio data in the first direction and / or the audio data in the second direction; for example, if the user is in front of the electronic device, the image is generally captured by the front camera; if the user's audio information exists in the direction in front of the electronic device, it can be assumed that the user is in the direction in front of the electronic device, and the front camera can be turned on; if the user is behind the electronic device, the image is generally captured by the rear camera; if the user's audio information exists in the direction behind the electronic device, it can be assumed that the user is in the direction behind the electronic device, and the rear camera can be turned on.
[0154] The target keywords may include, but are not limited to: front camera, rear camera, front recording, rear recording, dual-view recording, picture-in-picture recording, etc.; the first direction may refer to the forward direction of the electronic device, and the first preset angle range may refer to -30 degrees to 30 degrees; the second direction may refer to the backward direction of the electronic device, and the second preset angle range may refer to 150 degrees to 210 degrees, such as... Figure 10 As shown.
[0155] For example, audio data can be processed based on a sound direction probability calculation algorithm to obtain audio data in a first direction (e.g., forward direction) and / or audio data in a second direction (e.g., backward direction). Specific details can be found later. Figure 9 Steps S507 to S510 and step S512 shown will not be described again here.
[0156] In the embodiments of this application, the probability of audio data in each direction can be calculated, thereby separating the audio data by direction to obtain audio data in a first direction and audio data in a second direction; a switching command can be obtained based on the audio data in the first direction and / or the audio data in the second direction; the electronic device can automatically switch the shooting mode based on the switching command.
[0157] For example, a switching instruction is obtained based on audio data from a first direction and / or audio data from a second direction, including: A switching instruction is obtained based on the energy of the first amplitude spectrum and / or the energy of the second amplitude spectrum, where the first amplitude spectrum is the amplitude spectrum of the audio data in the first direction and the second amplitude spectrum is the amplitude spectrum of the audio data in the second direction.
[0158] It should be understood that in video recording scenarios, the direction with the greater energy of the audio data (e.g., the direction with the greater volume of the audio information) can generally be considered as the main shooting direction; the main shooting direction can be obtained based on the energy of the amplitude spectrum of the audio data in different directions; for example, if the energy of the amplitude spectrum of the audio data in the first direction is greater than the energy of the amplitude spectrum of the audio data in the second direction, then the first direction can be considered as the main shooting direction; at this time, the camera corresponding to the first direction in the electronic device can be turned on.
[0159] For example, the switching instruction may include the current shooting mode, a first picture-in-picture mode, a second picture-in-picture mode, a first dual-view mode, a second dual-view mode, a single-camera mode of the first camera, or a single-camera mode of the second camera. Based on the energy of the first amplitude spectrum and / or the energy of the second amplitude spectrum, the switching instruction is obtained, including: If the energy of the first amplitude spectrum and the energy of the second amplitude spectrum are both less than the first preset threshold, the switching instruction is to maintain the current shooting mode. If the energy of the first amplitude spectrum is greater than the second preset threshold, and the energy of the second amplitude spectrum is less than or equal to the second preset threshold, the switching instruction is to switch to the single-camera mode of the first camera; If the energy of the second amplitude spectrum is greater than the second preset threshold, and the energy of the first amplitude spectrum is less than or equal to the second preset threshold, the switching instruction is to switch to the second camera single-lens mode; If the energy of the first amplitude spectrum is greater than the second preset threshold, and the energy of the second amplitude spectrum is greater than or equal to the first preset threshold, the switching instruction is to switch to the first picture-in-picture mode. If the energy of the second amplitude spectrum is greater than the second preset threshold, and the energy of the first amplitude spectrum is greater than or equal to the first preset threshold, the switching command is to switch to the second picture-in-picture mode. If the energy of the first amplitude spectrum and the energy of the second amplitude spectrum are both greater than or equal to the second preset threshold, and the energy of the first amplitude spectrum is greater than the energy of the second amplitude spectrum, the switching instruction is to switch to the first dual-view mode; If the energy of the first amplitude spectrum and the energy of the second amplitude spectrum are both greater than or equal to the second preset threshold, and the energy of the second amplitude spectrum is greater than the energy of the first amplitude spectrum, the switching command is to switch to the second dual-view mode. Among them, the second preset threshold is greater than the first preset threshold, the first picture-in-picture mode refers to the shooting mode in which the image captured by the first camera is the main picture, the second picture-in-picture mode refers to the shooting mode in which the image captured by the second camera is the main picture, the first dual-view mode refers to the shooting mode in which the image captured by the first camera is located on the top or left side of the display screen of the electronic device, and the second dual-view mode refers to the shooting mode in which the image captured by the second camera is located on the top or left side of the display screen of the electronic device.
[0160] Optionally, the specific implementation of the above process can be found in [reference needed]. Figure 9 The relevant description of step S515 shown.
[0161] For example, the first amplitude spectrum is the first average amplitude spectrum obtained by averaging the amplitude spectra corresponding to each frequency point in the audio data in the first direction; and / or, The second amplitude spectrum is the average amplitude spectrum obtained by averaging the amplitude spectra corresponding to each frequency point in the audio data of the second direction.
[0162] In the embodiments of this application, the amplitude spectrum obtained by averaging the amplitude spectra of different frequency points in the audio data in the first direction can be called the first average amplitude spectrum; the amplitude spectrum obtained by averaging the amplitude spectra of different frequency points in the audio data in the second direction can be called the second average amplitude spectrum; since the first average amplitude spectrum and / or the second average amplitude spectrum are amplitude spectra obtained by averaging the amplitude spectra of different frequency points, the accuracy of the information in the audio data in the first direction and / or the audio data in the first direction can be improved.
[0163] Optionally, the first amplitude spectrum is the amplitude spectrum obtained after performing a first amplification process and / or a second amplification on the first average amplitude spectrum, and the first average amplitude spectrum is obtained by averaging the amplitude spectra corresponding to each frequency point in the audio data in the first direction.
[0164] For example, the above video processing method further includes: Speech detection is performed on the audio data in the first direction to obtain the first detection result; The direction of arrival (DOA) is estimated from data collected by at least two microphones to obtain the predicted angle information. If the first detection result indicates that the audio data in the first direction includes the user's audio information, the amplitude spectrum of the audio data in the first direction is subjected to a first amplification process; and / or, if the predicted angle information includes angle information within a first preset angle range, the amplitude spectrum of the audio data in the first direction is subjected to a second amplification process.
[0165] Optionally, the second amplitude spectrum is the amplitude spectrum obtained after performing a first amplification process and / or a second amplification on the second average amplitude spectrum, and the second average amplitude spectrum is obtained by averaging the amplitude spectra corresponding to each frequency point in the audio data in the second direction.
[0166] For example, the above video processing method further includes: Speech detection is performed on the audio data from the second direction to obtain the second detection result; The direction of arrival (DOA) is estimated from data collected by at least two microphones to obtain the predicted angle information. If the second detection result indicates that the audio data in the second direction includes the user's audio information, the amplitude spectrum of the audio data in the second direction is subjected to a first amplification process; and / or, if the predicted angle information includes angle information within a second preset angle range, the amplitude spectrum of the audio data in the second direction is subjected to a second amplification process.
[0167] It should be understood that in a video recording scenario, the direction in which the user is located can usually be considered the main shooting direction; if the detection result indicates that the direction includes the user's audio information, then the user can be considered to be in that direction; at this time, the audio data in that direction can be subjected to a first amplification process, which can improve the accuracy of the acquired user audio information.
[0168] It should also be understood that in video recording scenarios, the user's direction can generally be considered the primary shooting direction. If the detection results indicate that the user's audio information is included in that direction, then the user can be considered to be in that direction. In this case, the audio data in that direction can be amplified in the first way, which can improve the accuracy of the acquired user audio information. When the predicted angle information includes a first preset angle range and / or a second preset angle range, it can be said that audio information exists in the first direction and / or the second direction of the electronic device. The accuracy of the first amplitude spectrum or the second amplitude spectrum can be improved through the second amplification process. With the improved accuracy of the amplitude spectrum and the user audio information, the switching command can be obtained accurately.
[0169] Optionally, the specific process of the above speech detection can be found in the following sections. Figure 9 The relevant descriptions of steps S511 or S513 in the text.
[0170] Optionally, the specific processes of the first amplification process and / or the second amplification process described above can be found in [reference needed]. Figure 9 The relevant description of step S515 in the process.
[0171] For example, direction-of-arrival estimation refers to an algorithm that performs a spatial Fourier transform on the received signal, then takes the square of the modulus to obtain the spatial spectrum, and estimates the direction of arrival of the signal.
[0172] Optionally, the specific process of direction-of-arrival estimation can be found in the following sections. Figure 8 The relevant description of step S407, or, Figure 9 Description of step S514.
[0173] For example, identifying whether the audio data includes target keywords includes: The audio data is separated based on the blind signal separation algorithm to obtain N audio information, which are audio information of different users; For each of the N audio messages, identify whether the target keyword is included.
[0174] In the embodiments of this application, the audio data collected by at least two pickup devices can be separated and processed to obtain N audio information from different sources; whether the target keyword is included in each of the N audio information can be identified, thereby improving the accuracy of identifying the target keyword.
[0175] For example, a blind signal separation algorithm is an algorithm that recovers individual source signals from an acquired mixed signal (typically the output of multiple sensors).
[0176] Optionally, the specific process of the blind signal separation algorithm can be found in the following sections. Figure 8 The relevant description of step S405 in the text, or Figure 9 The relevant description of step S504 in the text.
[0177] Optionally, the specific process for identifying target keywords in audio data can be found in the following sections. Figure 11 Related descriptions.
[0178] In embodiments of this application, the electronic device can collect audio data in the shooting environment through at least two sound pickup devices (e.g., microphones); generate a switching command based on the audio data, and the electronic device automatically switches from the current first shooting mode to the second shooting mode based on the switching command, and displays the second image captured in the second shooting mode; without the user needing to switch the shooting mode of the electronic device, the electronic device can automatically switch shooting modes to complete video recording, thereby improving the user's shooting experience.
[0179] Figure 8 This is a schematic flowchart of a video processing method provided in an embodiment of this application. The video processing method 400 includes components that can be... Figure 1 The electronic device shown performs the video processing method, which includes steps S401 to S410. Steps S401 to S410 are described in detail below.
[0180] Step S401: Obtain audio data collected by N pickup devices (e.g., microphones).
[0181] Step S402: Perform source separation processing on the audio data to obtain M audio information.
[0182] It should be understood that sound source separation can also be called audio source separation; for example, the N audio data collected can be subjected to Fourier transform, and then the frequency domain data of the N audio data can be added with hyperparameters and sent to the separator to separate the audio sources, so as to obtain M audio information.
[0183] Step S403: Determine whether each of the M audio messages contains a switching instruction (an example of a target keyword); if any one of the M audio messages contains a switching instruction, proceed to step S404; if none of the M audio messages contain a switching instruction, proceed to steps S405 to S410.
[0184] For example, the switching command may include, but is not limited to: switching to the front camera, switching to the rear camera, recording with the front camera, recording with the rear camera, dual-view recording, picture-in-picture recording, etc. Optionally, the method for recognizing the switching command can be subsequently... Figure 11 As shown.
[0185] Step S404: The electronic device executes the switching command.
[0186] It should be understood that an electronic device executing a switching command can mean that the electronic device can automatically switch cameras based on the switching command without the user having to manually switch camera applications.
[0187] Step S405: Perform directional separation processing on the audio data collected by N microphones to obtain forward audio information (an example of audio data in the first direction) and / or backward audio information (an example of audio data in the second direction).
[0188] In the embodiments of this application, if a switching command is detected in M audio information, the electronic device automatically executes the switching command; if no switching command is detected in M audio information, the electronic device can obtain forward audio information within the target angle in the forward direction of the electronic device and / or backward audio information within the target angle in the backward direction of the electronic device based on N audio data collected by the pickup device; based on the energy of the forward audio information and the energy of the backward audio information, the switching command can be obtained by analysis; thus enabling the electronic device to execute the corresponding switching command.
[0189] For example, such as Figure 10 As shown, the forward speech beam can refer to audio data in the forward direction of the electronic device; wherein, the target angle in the forward direction of the electronic device (an example of a first preset angle range) can be [-30, 30]; the backward speech beam can refer to audio data in the backward direction of the electronic device; wherein, the target angle in the backward direction of the electronic device (an example of a second preset angle range) can be [150, 210].
[0190] Optionally, the N audio data collected by the pickup device can be separated into forward audio data and / or backward audio data based on the sound direction probabilities in each direction of the electronic device; for example, specific implementation methods can be found in [reference needed]. Figure 6 Steps S507 to S511 are shown.
[0191] Step S406: Perform speech detection processing on the forward audio information and / or backward audio information to obtain the detection results.
[0192] In the embodiments of this application, speech detection processing is performed on the forward audio information and / or backward audio information to determine whether the forward audio information and / or backward audio information includes the user's audio information; if the forward audio information (or backward audio information) includes the user's audio information, the forward audio information (or backward audio information) can be amplified to ensure that the user's audio information can be accurately obtained.
[0193] For example, speech detection processing may include, but is not limited to, speech activity detection or other methods for detecting user audio information, and this application does not limit it in any way.
[0194] Step S407: Estimate the direction of arrival (DOA) of the audio data collected by N microphones to obtain the predicted angle information.
[0195] In the embodiments of this application, the N audio data collected by the pickup device can be divided into forward audio information and / or backward audio information through steps S405 and S406. Furthermore, by performing direction of arrival estimation on the N audio data collected by the pickup device, the angle information corresponding to the audio data can be obtained, thereby determining whether the audio data acquired by the pickup device is within the target angle range. For example, determining whether the audio data is within the target angle range in the forward direction of the electronic device, or within the target angle range in the backward direction.
[0196] Optionally, for details on how to estimate the direction of arrival (DOA) of the audio data collected by N microphones to obtain the predicted angle information, please refer to [link to relevant documentation]. Figure 9 The step S514 is shown.
[0197] Optionally, if the switching instruction is not included in each audio message, steps S405, S406, S408 to S410 can be executed.
[0198] Step S408: Amplify the amplitude spectrum of the forward audio information and / or the backward audio information.
[0199] For example, the amplitude spectrum of forward and / or backward audio information can be amplified based on the detection results of speech detection processing.
[0200] In the embodiments of this application, when the detection result of the speech detection corresponding to the forward audio information (or the backward audio information) includes the user's audio information, the amplitude spectrum of the forward audio information (or the backward audio information) can be amplified to improve the accuracy of the acquired user audio information.
[0201] For example, the amplitude spectrum of forward and / or backward audio information can be amplified based on the detection results and predicted angle information of speech detection processing.
[0202] In the embodiments of this application, when the predicted angle information includes the target angle of the electronic device in the forward or backward direction, the amplitude spectrum of the forward audio information (or, backward audio information) can be amplified to improve the accuracy of the amplitude spectrum. Furthermore, when the detection result of the speech detection corresponding to the forward audio information (or, backward audio information) includes the user's audio information, the amplitude spectrum of the forward audio information (or, backward audio information) can be amplified to improve the accuracy of the acquired user audio information. With the improved accuracy of both the amplitude spectrum and the user's audio information, the accuracy of the obtained switching command can be improved.
[0203] For example, the amplitude spectra of the forward audio information and the backward audio information are calculated separately; when the detection result of the speech detection processing indicates that the forward audio information includes the user's audio information, the amplitude spectrum of the forward audio information can be subjected to a first amplification process; or, when the speech activity detection result indicates that the backward audio information includes the user's audio information, the amplitude spectrum of the backward audio information can be subjected to a first amplification process; for example, the amplification factor of the first amplification process is α (1<α<2).
[0204] For example, when the predicted angle information obtained based on the direction of arrival estimation indicates that the N audio data collected by the pickup device includes the target angle in the forward direction, the amplitude spectrum of the forward audio information can be amplified in the second way; or, when the predicted angle information obtained based on the direction of arrival estimation indicates that the N audio data collected by the pickup device includes the target angle in the backward direction, the amplitude spectrum of the backward audio information can be amplified in the second way; for example, the amplification factor of the second amplification is β (1<β<2), and the amplitude spectrum of the amplified forward audio information and / or backward audio information is obtained.
[0205] Optionally, the specific implementation of the amplification process can be found in the following sections. Figure 9 Step S515 is shown.
[0206] Step S409: Obtain a switching command based on the energy of the amplitude spectrum of the amplified forward audio information and / or backward audio information.
[0207] In one example, if the energy of both the amplitude spectrum of the amplified forward audio information and the amplitude spectrum of the amplified backward audio information is less than a first preset threshold, it is considered that there is no audio data in either the forward or backward direction of the electronic device, and the electronic device will continue to record video using the default lens; for example, this switching instruction can correspond to the identifier 0.
[0208] In one example, if only one energy level in the amplitude spectrum of the amplified forward audio information or the amplitude spectrum of the amplified backward audio information is greater than the second preset threshold, the electronic device determines the direction corresponding to the amplitude spectrum with energy greater than the first preset threshold as the main sound source direction and switches the camera of the electronic device to that direction; for example, the switching instruction can be to switch to the rear camera, which can correspond to identifier 1; or, the switching instruction can be to switch to the front camera, which can correspond to identifier 2.
[0209] In one example, if only one energy in the amplitude spectrum of the amplified forward audio information or the amplitude spectrum of the amplified backward audio information is greater than or equal to a second preset threshold, and the other energy is greater than or equal to a first preset threshold, where the second preset threshold is greater than the first preset threshold, then the electronic device can determine that the direction corresponding to the amplitude spectrum with energy greater than the second preset threshold is the direction of the main sound source, and the direction corresponding to the amplitude spectrum with energy greater than the first preset threshold is the direction of the second sound source. At this time, the electronic device can start the picture-in-picture recording mode; the picture in the direction corresponding to the amplitude spectrum with energy greater than or equal to the second preset threshold is used as the main picture, and the picture in the direction corresponding to the amplitude spectrum with energy greater than or equal to the first preset threshold is used as the secondary picture.
[0210] For example, if the energy of the amplitude spectrum corresponding to the forward audio information is greater than or equal to the second preset threshold, and the energy of the amplitude spectrum corresponding to the backward audio information is greater than or equal to the first preset threshold, then the switching instruction of the electronic device can be the picture-in-picture front main picture, and this switching instruction can correspond to identifier 3.
[0211] For example, if the energy of the amplitude spectrum corresponding to the backward audio information is greater than or equal to the second preset threshold, and the energy of the amplitude spectrum corresponding to the forward audio information is greater than or equal to the first preset threshold, then the switching instruction of the electronic device can be the picture-in-picture rear main picture, and this switching instruction can correspond to identifier 4.
[0212] In one example, if the energy of the amplitude spectrum of the amplified forward audio information or the amplitude spectrum of the amplified backward audio information is greater than or equal to the second preset threshold, the electronic device can determine to start dual-view recording, that is, to start the front camera and the rear camera; optionally, the image captured by the camera corresponding to the direction with greater energy can be displayed on the top or left side of the display screen.
[0213] For example, if the energy of the amplitude spectrum corresponding to both the forward audio information and the backward audio information is greater than or equal to the second preset threshold, and the energy of the amplitude spectrum corresponding to the forward audio information is greater than the energy of the amplitude spectrum corresponding to the backward audio information, then the switching instruction of the electronic device can be dual-view recording, displaying the image captured by the front camera of the electronic device on the top or left side of the display screen. This switching instruction can correspond to identifier 5.
[0214] For example, if the energy of the amplitude spectrum corresponding to both the forward and backward audio information is greater than or equal to the second preset threshold, and the energy of the amplitude spectrum corresponding to the backward audio information is greater than the energy of the amplitude spectrum corresponding to the forward audio information, then the switching instruction of the electronic device can be dual-view recording, displaying the image captured by the rear camera of the electronic device on the top or left side of the display screen. This switching instruction can correspond to identifier 6.
[0215] Step S410: The electronic device executes the switching command.
[0216] For example, the electronic device can obtain a switching instruction based on the amplitude spectrum of the amplified forward audio information and / or the amplified backward audio information, and automatically execute the switching instruction; that is, the electronic device can automatically switch the camera of the electronic device based on the switching instruction without the user having to manually switch the camera application.
[0217] In the embodiments of this application, in the video shooting scenario, switching instructions can be obtained based on the audio data in the shooting environment, enabling the electronic device to automatically determine whether to switch lenses or start multi-lens recording, etc., so as to achieve a one-shot recording experience without the user having to operate manually, thereby improving the user experience.
[0218] Figure 9 This is a schematic flowchart of a video processing method provided in an embodiment of this application. The video processing method 500 includes components that can be... Figure 1 The electronic device shown performs the video processing method, which includes steps S501 to S515. Steps S501 to S515 are described in detail below.
[0219] It should be understood that Figure 8 The video processing method shown is illustrated by an example of an electronic device including three microphones. Since the electronic device needs to determine the directionality of the audio information, the electronic device in the embodiments of this application includes at least two microphones, and there is no limitation on the specific number of microphones.
[0220] Step S501: The pickup device 1 collects audio data.
[0221] Step S502: The pickup device 2 collects audio data.
[0222] Step S503: The pickup device 3 collects audio data.
[0223] For example, the pickup device 1, pickup device 2, or pickup device 3 may be located in different positions in the electronic device to collect audio information from different directions; for example, the pickup device 1, pickup device 2, or pickup device 3 may refer to a microphone.
[0224] Optionally, after the electronic device detects that the user has selected the recording mode and started recording video, it can activate the audio pickup device 1, audio pickup device 2, and audio pickup device 3 to start collecting audio data.
[0225] It should be understood that steps S501 to S503 can be performed simultaneously.
[0226] Step S504: Perform blind signal separation on the audio data collected by the pickup device to obtain M channels of audio information.
[0227] It should be understood that blind signal separation, also known as blind signal / source separation (BSS), refers to estimating the source signal based on the mixed signal without knowing the source signal or the signal mixing parameters. In the embodiments of this application, audio information from different sources, i.e., audio signals from different objects, can be obtained by performing blind signal separation on the acquired audio data.
[0228] For example, in the shooting environment of the electronic device, there are three users: user A, user B, and user C. By blind signal separation, the audio information of user A, user B, and user C in the audio data can be obtained.
[0229] Step S505: Determine whether the M audio channels contain a switching instruction; if the M audio channels contain a switching instruction, proceed to step S506; if the M audio channels do not contain a switching instruction, proceed to steps S507 to S515.
[0230] For example, M audio information can be obtained through step S504; by identifying the switching instruction of each audio signal in the M audio information, it is determined whether each audio signal in the M audio information includes a switching instruction; wherein, the switching instruction may include, but is not limited to: switching to the front camera, switching to the rear camera, front camera recording, rear camera recording, dual-view recording, picture-in-picture recording, etc.
[0231] Optionally, Figure 11 This is a schematic flowchart illustrating a method for identifying switching instructions provided in an embodiment of this application. The identification method 600 includes steps S601 to S606, which are described in detail below.
[0232] Step S601: Obtain the M audio signals after separation and processing.
[0233] Optionally, step S601 can also involve acquiring audio data collected by the pickup device, such as... Figure 8 The step S401 shown.
[0234] Step S602: Perform noise reduction processing on each of the M audio information items.
[0235] For example, the noise reduction process can employ any noise reduction algorithm; for instance, the noise reduction algorithm may include spectral subtraction or Wiener filtering. The principle of spectral subtraction is to subtract the spectrum of the noisy signal from the spectrum of the noisy signal to obtain the spectrum of the clean signal. The principle of Wiener filtering is to approximate the original signal by passing the noisy signal through a linear filter and to find the linear filter parameters that minimize the mean square error.
[0236] Step S603: Input the M audio information after noise reduction into the acoustic model, wherein the acoustic model is a pre-trained deep neural network.
[0237] Step S604: For each of the M audio messages, output a confidence score. The confidence score is used to represent the confidence level that a certain switching instruction is included in an audio message.
[0238] Step S605: Compare the confidence level with a preset threshold; if the confidence level is greater than the preset threshold, proceed to step S606.
[0239] Step S606: Obtain the switching instruction.
[0240] It should be understood that steps S601 to S606 above are illustrative examples; other identification methods can also be used to identify whether the audio information includes a switching instruction, and this application does not impose any limitations on this.
[0241] Step S506: Execute the switching command.
[0242] For example, the electronic device automatically executes the switching instruction based on the switching instruction identified in step S505.
[0243] It should be understood that the electronic device automatically executing the switching command can mean that the electronic device can automatically switch the camera based on the switching command without the user having to manually switch the camera application.
[0244] It should be understood that steps S507 to S509 are used to output the directionality of M audio information, that is, to determine the forward audio signal and the backward audio signal among the M audio information; wherein, the forward audio signal may refer to the audio signal within the preset angle range of the front camera of the electronic device; the backward audio signal may refer to the audio signal within the preset angle range of the rear camera of the electronic device.
[0245] Step S507: Calculate the sound direction probability for the M audio information items if the switching instruction is not included.
[0246] For example, based on cACGMM and audio data collected by three pickup devices, the probability value of the frequency point of the current input audio data existing in each direction can be calculated.
[0247] It should be understood that cACGMM is a Gaussian mixture model; a Gaussian mixture model refers to the precise quantification of things using Gaussian probability density functions (e.g., normal distribution curves), decomposing a thing into several models based on Gaussian probability density functions (e.g., normal distribution curves).
[0248] For example, the probability values of frequency points in audio data in each direction satisfy the following constraints: ; in, represents the probability value in the k direction; t represents the speech frame (e.g., a frame of audio data); and f represents the frequency point (e.g., the frequency angle of a frame of audio data).
[0249] It should be understood that in the embodiments of this application, frequency point may refer to time frequency point; time frequency point may include time information, frequency range information and energy information corresponding to audio data.
[0250] For example, in the embodiments of this application, K can be 36; since the circumference of an electronic device is 360 degrees, K of 36 can be set to a direction every 10 degrees.
[0251] It should be understood that the above constraints can be interpreted as the sum of probabilities of a certain frequency point in all directions being 1.
[0252] Step S508: Spatial clustering.
[0253] It should be understood that, in the embodiments of this application, spatial clustering can be used to determine the probability value of audio data within the field of view of the camera of an electronic device.
[0254] For example, the area directly in front of the screen of an electronic device is typically at a 0-degree angle. To ensure that audio data from the camera of the electronic device is not lost, such as... Figure 10 As shown, the target angle in the forward direction can be set to [-30, 30]; the target angle in the backward direction of the electronic device can be set to [150, 210]; the corresponding angle direction indices are k1~k2, and the spatial clustering probability is: ; in, This indicates the probability that a frequency point in the audio data falls within the target angle. This represents the probability value of the frequency point of the audio data in the k direction.
[0255] Step S509: Gain calculation.
[0256] For example, ; in, This indicates the frequency gain of the audio data; Indicates the first probability threshold; This represents the second probability threshold; This represents the frequency gain of audio data at non-target angles.
[0257] It should be understood that when the probability of an audio data frequency point being within the target angle is greater than a first probability threshold, it can be said that the frequency point is within the target angle range; when the probability of an audio data frequency point being within the target angle is less than or equal to a second probability threshold, it can be said that the frequency point is within the non-target angle range; for example, the first probability threshold can be 0.8; the frequency gain of the audio data in the non-target angle can be a pre-configured parameter; for example, 0.2; the second probability threshold can be 0.1.
[0258] It should also be understood that the gain calculation of the audio data described above can achieve smoothing of the audio data, thereby enhancing the frequency points of the audio data within the target angle range and weakening the frequency points of the audio data outside the target angle range.
[0259] Step S510: Based on the frequency gain and Fourier transform processing of the audio data, the backward audio data can be obtained.
[0260] For example, such as Figure 10 As shown, backward audio data can refer to audio data in the backward direction of an electronic device; wherein, the target angle in the backward direction of the electronic device can be [150, 210].
[0261] For example, ;in, It can represent backward audio data; Indicates the frequency gain of the backward audio data; This represents the Fourier transform of the backward audio data.
[0262] Step S511: Perform speech activity detection on the backward audio data.
[0263] For example, semantic detection of backward audio data can be performed using a cepstral algorithm to obtain speech activity detection results; if the fundamental frequency is detected, it is determined that the backward speech beam includes the user's audio information; if the fundamental frequency is not detected, it is determined that the backward speech beam does not include the user's speech information.
[0264] It should be noted that rearward audio data refers to audio data collected by the electronic device within the angular range of the rearward direction; the rearward audio data may include audio information in the shooting environment (e.g., vehicle horn sounds) or user voice information; voice detection is performed on the rearward audio data to determine whether the rearward audio data includes user voice information; when the rearward audio data includes user voice information, the rearward audio data can be amplified during the subsequent step S515, thereby improving the accuracy of acquiring user voice information.
[0265] It should be understood that cepstral algorithms are methods used in signal processing and signal detection; the cepstral spectrum refers to the power spectrum of the logarithmic power spectrum of a signal. The principle of obtaining speech through the cepstral spectrum is as follows: since voiced signals are periodically excited, they exhibit periodic impulses on the cepstral spectrum, thus allowing the determination of the fundamental frequency. Generally, the second impulse in the cepstral waveform (the first being envelope information) is considered the fundamental frequency of the excitation source. The fundamental frequency is one of the characteristics of speech; the presence of a fundamental frequency indicates the presence of speech in the current audio data.
[0266] Step S512: Based on the frequency gain and energy of the audio data, the forward audio data can be obtained.
[0267] For example, such as Figure 10 As shown, forward audio data can refer to audio data in the forward direction of the electronic device; where the target angle in the forward direction of the electronic device can be [-30, 30].
[0268] For example, ;in, It can represent a forward speech beam; This indicates the frequency gain of the forward audio data; This represents the Fourier transform of the forward audio data.
[0269] Step S513: Perform speech activity detection on the forward audio data.
[0270] For example, semantic detection of forward audio data can be performed using a cepstral algorithm to obtain speech activity detection results; if the fundamental frequency is detected, it is determined that the forward speech beam includes the user's audio information; if the fundamental frequency is not detected, it is determined that the forward speech beam does not include the user's speech information.
[0271] It should be noted that forward audio data refers to audio data collected by the electronic device within the angular range of the forward direction; forward audio data may include audio information in the shooting environment (e.g., vehicle horn sounds) or user voice information; voice detection is performed on the forward audio data to determine whether the forward audio data includes user voice information; when the forward audio data includes user voice information, the forward audio data can be amplified during the subsequent step S515, thereby improving the accuracy of acquiring user voice information.
[0272] Step S514: Estimate the direction of arrival (DOA) of the audio data collected by the pickup device.
[0273] It should be understood that, in the embodiments of this application, the angle information corresponding to the audio data can be obtained by performing direction of arrival estimation on the audio data collected by the pickup device, thereby determining whether the audio data acquired by the pickup device is within the target angle range; for example, determining whether the audio data is within the target angle range in the forward direction of the electronic device, or in the target angle range in the backward direction.
[0274] For example, a high-resolution spectral estimation localization algorithm (e.g., estimating signal parameter via rotational invariance techniques (ESPRIT)), a controlled beamforming localization algorithm, or a localization algorithm based on time difference of arrival (TDOA) can be used to estimate the direction of arrival of the audio data collected by the pickup device.
[0275] ESRT refers to a rotation-invariant technique algorithm, whose principle is mainly based on the rotation invariance of signals to estimate signal parameters. The localization algorithm of the controllable beamforming method works by filtering and weighting the signals received by the microphone to form a beam, and then searching for the sound source location according to a certain rule. When the microphone reaches its maximum output power, the searched sound source location is the true sound source location. TDOA is used to represent the time difference between the arrival of a sound source at different microphones in the electronic device.
[0276] In one example, the localization algorithm for TDOA may include the GCC-PHAT algorithm; taking the GCC-PHAT algorithm as an example, the method of direction-of-arrival estimation based on audio data is explained; such as Figure 12 As shown, the audio pickup device 1 and the audio pickup device 2 collect audio data; the distance between the audio pickup device 1 and the audio pickup device 2 is d, and the angle information between the audio data and the electronic device can be obtained according to the GCC-PHAT algorithm.
[0277] For example, Figure 12 The angle θ shown can be obtained based on the following formula: ; Where IDFT represents the inverse operation of the discrete inverse Fourier transform; This represents the frequency domain information obtained after performing a Fourier transform on the audio data collected by the pickup device 1. This represents the frequency domain information obtained after Fourier transforming the audio data collected by the pickup device 2; arg represents the variable (i.e., the English abbreviation for argument); arg max represents the value of the variable that makes the following formula reach its maximum value.
[0278] Step S515: Based on the angle information obtained from the direction of arrival estimation and the speech activity detection results, data analysis can be performed on the forward speech beam and the backward speech beam to obtain the switching command.
[0279] For example, the average amplitude spectra of the forward speech beam and the backward speech beam are calculated respectively; when the speech activity detection result indicates that the forward speech beam includes the user's audio information, the average amplitude spectra of the forward speech beam can be subjected to a first amplification process respectively; or when the speech activity detection result indicates that the backward speech beam includes the user's audio information, the average amplitude spectra of the backward speech beam can be subjected to a first amplification process respectively; for example, the amplification factor of the first amplification process is α (1<α<2).
[0280] It should be understood that the amplitude spectrum obtained by averaging the amplitude spectra of different frequency points in the forward speech beam can be called the average amplitude spectrum of the forward beam; the amplitude spectrum obtained by averaging the amplitude spectra of different frequency points in the backward speech beam can be called the average amplitude spectrum of the backward beam; data analysis based on the average amplitude spectrum of the forward speech beam and / or the average amplitude spectrum of the backward speech beam can improve the accuracy of information in the forward speech beam and / or the backward speech beam.
[0281] Furthermore, when the forward target angle range of the forward speech beam is determined based on the angle information obtained from the direction of arrival estimation, a second amplification process can be performed on the average amplitude spectrum of the forward speech beam; or, when the backward target angle range of the backward speech beam is determined based on the angle information obtained from the direction of arrival estimation, a second amplification process can be performed on the average amplitude spectrum of the backward speech beam; for example, the amplification factor of the second amplification process is β (1<β<2), and the amplitude spectra of the forward speech beam and the backward speech beam after amplification are obtained.
[0282] It should be understood that, in the embodiments of this application, the amplification of the forward or backward speech beam is to adjust the accuracy of the amplitude spectrum; furthermore, when the speech beam (e.g., the forward and / or backward speech beam) includes the user's audio information, amplifying the amplitude spectrum of the speech beam can improve the accuracy of the acquired user audio information; with the improved accuracy of the amplitude spectrum and the user audio information, the switching instructions in the speech beam can be accurately obtained.
[0283] For example, the amplitude spectrum corresponding to a frequency point in audio data can be calculated using the following formula: ; Where Mag(i) represents the amplitude spectrum corresponding to the i-th frequency point; i represents the i-th frequency point; and K represents the frequency range. ~ This indicates the frequency range required for averaging; it should be understood that it is not necessary to average all frequencies, but rather to obtain the average value of a subset of frequencies.
[0284] For example, when the speech activity detection result indicates that the forward speech beam includes the user's audio information, and the forward speech beam is within the forward target angle range, the average amplitude spectrum of the amplified forward speech beam is: ; in, This represents the average amplitude spectrum of the forward speech beam after amplification. α represents the average amplitude spectrum of the original forward speech beam; α represents the preset first amplification factor; β represents the preset second amplification factor.
[0285] It should be understood that the amplitude spectrum obtained by averaging the amplitude spectra of different frequency points in the forward speech beam can be called the average amplitude spectrum of the forward beam.
[0286] For example, when the speech activity detection result indicates that the backward speech beam includes the user's audio information, and the backward speech beam is within the backward target angle range, the average amplitude spectrum of the amplified backward speech beam is: ; in, This represents the average amplitude spectrum of the amplified backward speech beam. α represents the average amplitude spectrum of the original backward speech beam; α represents the preset first amplification factor; β represents the preset second amplification factor.
[0287] It should be understood that the amplitude spectrum obtained by averaging the amplitude spectra of different frequency points in the backward speech beam can be called the average amplitude spectrum of the backward beam.
[0288] In one example, if and If the energy is less than the first preset threshold, it is considered that there is no audio data in the forward and backward directions of the electronic device, and the electronic device will continue to record video with the default lens; for example, as shown in Table 1, this switching instruction can correspond to the identifier 0.
[0289] In one example, if or If only one energy level is greater than the second preset threshold, the electronic device determines the direction corresponding to the amplitude spectrum of the energy level greater than the first preset threshold as the main sound source direction and switches the camera of the electronic device to that direction. For example, as shown in Table 1, the switching instruction can be to switch to the rear camera, which can correspond to identifier 1; or, the switching instruction can be to switch to the front camera, which can correspond to identifier 2.
[0290] In one example, if or If only one energy level is greater than or equal to a second preset threshold, and the other energy level is greater than or equal to a first preset threshold, and the second preset threshold is greater than the first preset threshold, then the electronic device can determine that the direction corresponding to the amplitude spectrum of the energy level greater than the second preset threshold is the direction of the main sound source, and the direction corresponding to the amplitude spectrum of the energy level greater than the first preset threshold is the direction of the second sound source. At this time, the electronic device can start the picture-in-picture recording mode; the picture in the direction corresponding to the amplitude spectrum of the energy level greater than or equal to the second preset threshold is used as the main picture, and the picture in the direction corresponding to the amplitude spectrum of the energy level greater than or equal to the first preset threshold is used as the secondary picture.
[0291] For example, if The energy is greater than or equal to the second preset threshold, and If the energy is greater than or equal to the first preset threshold, the switching instruction of the electronic device can be the picture-in-picture front main picture, and this switching instruction can be identified by identifier 3.
[0292] For example, if the energy of the amplitude spectrum corresponding to the backward speech beam is greater than or equal to the second preset threshold, and the energy of the amplitude spectrum corresponding to the forward speech beam is greater than or equal to the first preset threshold, then the switching instruction of the electronic device can be the picture-in-picture rear main picture; for example, as shown in Table 1, this switching instruction can correspond to identifier 4.
[0293] In one example, if or If the energy values are all greater than or equal to a second preset threshold, the electronic device can determine to activate dual-view recording, i.e., activate both the front and rear cameras. Optionally, the image captured by the camera corresponding to the direction with higher energy can be displayed on the top or left side of the screen.
[0294] For example, if and The energy is greater than or equal to the second preset threshold, and Energy greater than If the energy is sufficient, the switching command of the electronic device can be dual-view recording, displaying the image captured by the front camera of the electronic device on the top or left side of the screen; for example, as shown in Table 1, this switching command can correspond to label 5.
[0295] For example, if and The energy is greater than or equal to the second preset threshold, and Energy greater than If the energy is sufficient, the switching command of the electronic device can be dual-view recording, displaying the image captured by the rear camera of the electronic device on the top or left side of the screen; for example, as shown in Table 1, this switching command can correspond to label 6.
[0296] Table 1
[0297] It should be understood that Table 1 is an example of the labels corresponding to the recording scenarios, and this application does not impose any limitations on them; the electronic device can automatically switch between different cameras in different recording scenarios.
[0298] For example, the electronic device can obtain a switching instruction based on the amplitude spectrum of the amplified forward audio information and / or the amplified backward audio information, and automatically execute the switching instruction; that is, the electronic device can automatically switch the camera of the electronic device based on the switching instruction without the user having to manually switch the camera application.
[0299] In the embodiments of this application, in the video shooting scenario, switching instructions can be obtained based on the audio data in the shooting environment, enabling the electronic device to automatically determine whether to switch lenses or start multi-lens recording, etc., so as to achieve a one-shot recording experience without the user having to operate manually, thereby improving the user experience.
[0300] Figure 13 This illustrates a graphical user interface (GUI) for an electronic device.
[0301] like Figure 13 As shown in (a), the preview interface for multi-camera recording may include a control 600 for indicating settings; upon detecting a user's click on the control 600, a settings interface is displayed in response to the user's action, such as... Figure 13As shown in (b) of the diagram; the settings interface includes a voice-activated photo-taking control 610, which detects when the user activates voice-activated photo-taking; the voice-activated photo-taking also includes an automatic shooting mode switching control 620, which detects when the user clicks the automatic shooting mode switching control 620, allowing the electronic device to enable the camera application's automatic shooting mode switching; that is, the video processing method provided in this application embodiment can be executed, and in the video shooting scenario, a switching instruction can be obtained based on the audio data in the shooting environment, enabling the electronic device to automatically determine whether to switch shooting modes; video recording is completed without the user needing to switch the electronic device's shooting mode, improving the user's shooting experience.
[0302] In one example, such as Figure 14 As shown, the preview interface for multi-camera recording may include a control 630 that indicates the automatic switching of shooting modes. After the user clicks the control 630, the electronic device can enable the automatic switching of shooting modes in the camera application; that is, the video processing method provided in this application embodiment can be executed. In the video shooting scenario, a switching instruction can be obtained based on the audio data in the shooting environment, so that the electronic device can automatically determine whether to switch shooting modes; the video recording is completed without the user having to switch the shooting modes of the electronic device, thereby improving the user's shooting experience.
[0303] Figure 15 This illustrates a graphical user interface (GUI) for an electronic device.
[0304] Figure 15 The GUI shown in (a) is the desktop 640 of the electronic device; when the electronic device detects that the user clicks the settings icon 650 on the desktop 640, it can display as follows: Figure 15 Another GUI is shown in (b) above; Figure 15 The GUI shown in (b) can be the settings display interface, which may include options such as wireless network, Bluetooth, or camera; clicking the camera option will take you to the camera's settings interface, which displays as follows: Figure 15 The camera settings interface shown in (c) can include a control 660 for automatically switching shooting modes. After the user clicks the control 660, the electronic device can enable the camera application to automatically switch shooting modes. That is, the video processing method provided in this application embodiment can be executed. In the video shooting scenario, a switching instruction can be obtained based on the audio data in the shooting environment, so that the electronic device can automatically determine whether to switch shooting modes. The video recording is completed without the user having to switch the shooting mode of the electronic device, thus improving the user's shooting experience.
[0305] It should be understood that the above examples are provided to help those skilled in the art understand the embodiments of this application, and are not intended to limit the embodiments of this application to the specific values or scenarios illustrated. Those skilled in the art can obviously make various equivalent modifications or changes based on the above examples, and such modifications or changes also fall within the scope of the embodiments of this application.
[0306] The above text combined Figures 1 to 15 The video processing method provided in the embodiments of this application is described in detail below; the following will be combined with Figure 16 and Figure 17 The apparatus embodiments of this application are described in detail below. It should be understood that the apparatus in the embodiments of this application can perform the various methods described in the foregoing embodiments of this application, that is, the specific working processes of the various products described below can be referred to the corresponding processes in the foregoing method embodiments.
[0307] Figure 16 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device 700 includes a processing module 710 and a display module 720; the electronic device 700 may also include at least two sound pickup devices; for example, at least two microphones.
[0308] The processing module 710 is used to launch the camera application in the electronic device; the display module 720 is used to display a first image, which is an image captured by the electronic device when it is in a first shooting mode; the processing module 710 is also used to acquire audio data, which is data collected by the at least two microphones; based on the audio data, a switching instruction is obtained, which is used to instruct the electronic device to switch from the first shooting mode to a second shooting mode; the display module 720 is also used to display a second image, which is an image captured by the electronic device when it is in the second shooting mode.
[0309] Optionally, as an embodiment, the electronic device includes a first camera and a second camera, the first camera and the second camera being located in different directions of the electronic device, and the processing module 710 is specifically used for: Identify whether the audio data includes a target keyword, where the target keyword is the text information corresponding to the switching instruction; If the target keyword is identified in the audio data, the switching instruction is obtained based on the target keyword; If the target keyword is not identified in the audio data, the audio data is processed to obtain audio data in a first direction and / or audio data in a second direction. The first direction is used to represent a first preset angle range corresponding to the first camera, and the second direction is used to represent a second preset angle range corresponding to the second camera. Based on the audio data in the first direction and / or the audio data in the second direction, the switching instruction is obtained.
[0310] Optionally, as an embodiment, the processing module 710 is specifically used for: The audio data is processed based on a sound direction probability calculation algorithm to obtain audio data in the first direction and / or audio data in the second direction.
[0311] Optionally, as an embodiment, the processing module 710 is specifically used for: The switching instruction is obtained based on the energy of the first amplitude spectrum and / or the energy of the second amplitude spectrum, wherein the first amplitude spectrum is the amplitude spectrum of the audio data in the first direction, and the second amplitude spectrum is the amplitude spectrum of the audio data in the second direction.
[0312] Optionally, as an embodiment, the switching instruction includes the current shooting mode, a first picture-in-picture mode, a second picture-in-picture mode, a first dual-view mode, a second dual-view mode, a single-camera mode of the first camera, or a single-camera mode of the second camera, and the processing module 710 is specifically used for: If the energy of the first amplitude spectrum and the energy of the second amplitude spectrum are both less than the first preset threshold, the switching instruction is to maintain the current shooting mode. If the energy of the first amplitude spectrum is greater than the second preset threshold, and the energy of the second amplitude spectrum is less than or equal to the second preset threshold, the switching instruction is to switch to the single-camera mode of the first camera; If the energy of the second amplitude spectrum is greater than the second preset threshold, and the energy of the first amplitude spectrum is less than or equal to the second preset threshold, the switching instruction is to switch to the single-camera mode of the second camera; If the energy of the first amplitude spectrum is greater than the second preset threshold, and the energy of the second amplitude spectrum is greater than or equal to the first preset threshold, the switching instruction is to switch to the first picture-in-picture mode; If the energy of the second amplitude spectrum is greater than the second preset threshold, and the energy of the first amplitude spectrum is greater than or equal to the first preset threshold, the switching instruction is to switch to the second picture-in-picture mode; If the energy of the first amplitude spectrum and the energy of the second amplitude spectrum are both greater than or equal to the second preset threshold, and the energy of the first amplitude spectrum is greater than the energy of the second amplitude spectrum, the switching instruction is to switch to the first dual-view mode; If the energy of the first amplitude spectrum and the energy of the second amplitude spectrum are both greater than or equal to the second preset threshold, and the energy of the second amplitude spectrum is greater than the energy of the first amplitude spectrum, the switching instruction is to switch to the second dual-view mode; Wherein, the second preset threshold is greater than the first preset threshold, the first picture-in-picture mode refers to the shooting mode in which the image captured by the first camera is the main picture, the second picture-in-picture mode refers to the shooting mode in which the image captured by the second camera is the main picture, the first dual-view mode refers to the shooting mode in which the image captured by the first camera is located on the top or left side of the display screen of the electronic device, and the second dual-view mode refers to the shooting mode in which the image captured by the second camera is located on the top or left side of the display screen of the electronic device.
[0313] Optionally, as an embodiment, the first amplitude spectrum is a first average amplitude spectrum obtained by averaging the amplitude spectra corresponding to each frequency point in the audio data in the first direction; and / or, The second amplitude spectrum is the second average amplitude spectrum obtained by averaging the amplitude spectra corresponding to each frequency point in the audio data in the second direction.
[0314] Optionally, as an embodiment, the first amplitude spectrum is an amplitude spectrum obtained by performing a first amplification process and / or a second amplification on the first average amplitude spectrum, and the first average amplitude spectrum is obtained by averaging the amplitude spectra corresponding to each frequency point in the audio data in the first direction.
[0315] Optionally, as an embodiment, the processing module 710 is specifically used for: Speech detection is performed on the audio data in the first direction to obtain a first detection result; The direction of arrival (DOA) of the data collected by the at least two microphones is estimated to obtain the predicted angle information. If the first detection result indicates that the audio data in the first direction includes the user's audio information, the amplitude spectrum of the audio data in the first direction is subjected to the first amplification process; and / or If the predicted angle information includes angle information within the first preset angle range, the amplitude spectrum of the audio data in the first direction is subjected to the second amplification process.
[0316] Optionally, as an embodiment, the second amplitude spectrum is the amplitude spectrum obtained after performing a first amplification process and / or a second amplification on the second average amplitude spectrum, and the second average amplitude spectrum is obtained by averaging the amplitude spectra corresponding to each frequency point in the audio data of the second direction.
[0317] Optionally, as an embodiment, the processing module 710 is specifically used for: Speech detection is performed on the audio data in the second direction to obtain a second detection result; The direction of arrival (DOA) of the data collected by the at least two microphones is estimated to obtain the predicted angle information. If the second detection result indicates that the audio data in the second direction includes the user's audio information, the amplitude spectrum of the audio data in the second direction is subjected to the first amplification process; and / or If the predicted angle information includes angle information within the second preset angle range, the amplitude spectrum of the audio data in the second direction is subjected to the second amplification process.
[0318] Optionally, as an embodiment, the processing module 710 is specifically used for: The audio data is separated using a blind signal separation algorithm to obtain N audio information pieces, which are audio information pieces from different users; Each of the N audio messages is identified to determine whether the target keyword is included in the N audio messages.
[0319] Optionally, as an embodiment, the first image is a preview image captured when the electronic device is in multi-camera recording mode.
[0320] Optionally, as an embodiment, the first image is a video frame captured when the electronic device is recording multiple cameras.
[0321] Alternatively, as an embodiment, the audio data refers to the data collected by the sound pickup device in the shooting environment where the electronic device is located.
[0322] It should be noted that the aforementioned electronic device 700 is embodied in the form of functional modules. The term "module" here can be implemented in software and / or hardware, without specific limitations.
[0323] For example, a "module" can be a software program, a hardware circuit, or a combination of both that implements the above functions. The hardware circuit may include an application-specific integrated circuit (ASIC), electronic circuitry, a processor (e.g., a shared processor, a proprietary processor, or a group processor) and memory for executing one or more software or firmware programs, integrated logic circuitry, and / or other suitable components that support the described functions.
[0324] Therefore, the units of the various examples described in the embodiments of this application can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0325] Figure 17 A schematic diagram of the structure of an electronic device provided in this application is shown. Figure 17 The dashed lines in the diagram indicate that the unit or module is optional; the electronic device 800 can be used to implement the methods described in the above method embodiments.
[0326] The electronic device 800 includes one or more processors 801, which can support the video processing method implemented in the method embodiments of the electronic device 800. The processor 801 can be a general-purpose processor or a special-purpose processor. For example, the processor 801 can be a central processing unit (CPU), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, such as discrete gates, transistor logic devices, or discrete hardware components.
[0327] The processor 801 can be used to control the electronic device 800, execute software programs, and process data from the software programs. The electronic device 800 may also include a communication unit 805 for inputting (receiving) and outputting (transmitting) signals.
[0328] For example, electronic device 800 may be a chip, communication unit 805 may be the input and / or output circuit of the chip, or communication unit 805 may be the communication interface of the chip, and the chip may be a component of terminal device or other electronic device.
[0329] For example, electronic device 800 can be a terminal device, communication unit 805 can be the transceiver of the terminal device, or communication unit 805 can be the transceiver circuit of the terminal device.
[0330] The electronic device 800 may include one or more memories 802, which store a program 804. The program 804 can be executed by the processor 801 to generate instructions 803, causing the processor 801 to execute the video processing method described in the above method embodiments according to the instructions 803.
[0331] Optionally, the memory 802 may also store data.
[0332] Optionally, the processor 801 can also read data stored in the memory 802, which may be stored at the same memory address as the program 804, or the data may be stored at a different memory address than the program 804.
[0333] The processor 801 and memory 802 can be configured separately or integrated together, for example, integrated on the system-on-chip (SOC) of the terminal device.
[0334] For example, the memory 802 can be used to store the relevant program 804 of the video processing method provided in the embodiments of this application, and the processor 801 can be used to call the relevant program 804 of the video processing method stored in the memory 802 when executing the video processing method, and execute the video processing method of the embodiments of this application; for example, starting the camera application in the electronic device; displaying a first image, the first image being an image captured when the electronic device is in a first shooting mode; acquiring audio data, the audio data being data captured by at least two microphones in the electronic device; obtaining a switching instruction based on the audio data, the switching instruction being used to instruct the electronic device to switch from the first shooting mode to a second shooting mode; and displaying a second image, the second image being an image captured when the electronic device is in a second shooting mode.
[0335] This application also provides a computer program product that, when executed by processor 801, implements the video processing method of any method embodiment in this application.
[0336] The computer program product can be stored in memory 802, for example, program 804. Program 804 is finally converted into an executable object file that can be executed by processor 801 after processing such as preprocessing, compilation, assembly and linking.
[0337] This application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a computer, implements the video processing method described in any of the method embodiments of this application. The computer program may be a high-level language program or an executable object program.
[0338] The computer-readable storage medium is, for example, memory 802. Memory 802 can be volatile memory or non-volatile memory, or memory 802 can include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (SLDRAM), and direct rambus RAM (DR RAM).
[0339] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0340] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0341] In the embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the embodiments of the electronic devices described above are merely illustrative. For instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection through some interfaces, devices, or units, and may be electrical, mechanical, or other forms.
[0342] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0343] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0344] It should be understood that in the various embodiments of this application, the sequence number of each process does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0345] Furthermore, the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.
[0346] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0347] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of protection of the claims. In conclusion, the above description is merely a preferred embodiment of the technical solution of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A video processing method, characterized in that, Applied to an electronic device, the electronic device including at least two microphones, a first camera and a second camera, the video processing method includes: Run the camera application in the electronic device; Display a first image, which is an image captured by the electronic device when it is in a first shooting mode; Acquire audio data, wherein the audio data is data collected by the at least two microphones; A switching instruction is obtained based on the audio data. The switching instruction is used to instruct the electronic device to switch from the first shooting mode to the second shooting mode. The first shooting mode and the second shooting mode are different shooting modes in multi-camera recording. The switching instruction includes the current shooting mode, the first picture-in-picture mode, the second picture-in-picture mode, the first dual-view mode, the second dual-view mode, the single-camera mode of the first camera, or the single-camera mode of the second camera. Display a second image, which is an image captured by the electronic device when it is in the second shooting mode; The first camera and the second camera are located in different directions of the electronic device, and the step of obtaining the switching instruction based on the audio data includes: Identify whether the audio data includes a target keyword, where the target keyword is the text information corresponding to the switching instruction; If the target keyword is not identified in the audio data, the audio data is processed to obtain audio data in a first direction and / or audio data in a second direction. The first direction is used to represent a first preset angle range corresponding to the first camera, and the second direction is used to represent a second preset angle range corresponding to the second camera. If the energy of the first amplitude spectrum and the energy of the second amplitude spectrum are both less than the first preset threshold, the switching instruction is to maintain the current shooting mode; wherein, the first amplitude spectrum is the amplitude spectrum of the audio data in the first direction, and the second amplitude spectrum is the amplitude spectrum of the audio data in the second direction; If the energy of the first amplitude spectrum is greater than the second preset threshold, and the energy of the second amplitude spectrum is greater than or equal to the first preset threshold, the switching instruction is to switch to the first picture-in-picture mode; If the energy of the second amplitude spectrum is greater than the second preset threshold, and the energy of the first amplitude spectrum is greater than or equal to the first preset threshold, the switching instruction is to switch to the second picture-in-picture mode; If the energy of the first amplitude spectrum and the energy of the second amplitude spectrum are both greater than or equal to the second preset threshold, and the energy of the first amplitude spectrum is greater than the energy of the second amplitude spectrum, the switching instruction is to switch to the first dual-view mode; If the energy of the first amplitude spectrum and the energy of the second amplitude spectrum are both greater than or equal to the second preset threshold, and the energy of the second amplitude spectrum is greater than the energy of the first amplitude spectrum, the switching instruction is to switch to the second dual-view mode.
2. The video processing method as described in claim 1, characterized in that, The method further includes: If the target keyword is identified in the audio data, the switching instruction is obtained based on the target keyword.
3. The video processing method as described in claim 2, characterized in that, The step of processing the audio data to obtain audio data in a first direction and / or audio data in a second direction includes: The audio data is processed based on a sound direction probability calculation algorithm to obtain audio data in the first direction and / or audio data in the second direction.
4. The video processing method according to any one of claims 1 to 3, characterized in that, The method further includes: If the energy of the first amplitude spectrum is greater than the second preset threshold, and the energy of the second amplitude spectrum is less than or equal to the second preset threshold, the switching instruction is to switch to the single-camera mode of the first camera; If the energy of the second amplitude spectrum is greater than the second preset threshold, and the energy of the first amplitude spectrum is less than or equal to the second preset threshold, the switching instruction is to switch to the single-camera mode of the second camera; Wherein, the second preset threshold is greater than the first preset threshold, the first picture-in-picture mode refers to the shooting mode in which the image captured by the first camera is the main picture, the second picture-in-picture mode refers to the shooting mode in which the image captured by the second camera is the main picture, the first dual-view mode refers to the shooting mode in which the image captured by the first camera is located on the top or left side of the display screen of the electronic device, and the second dual-view mode refers to the shooting mode in which the image captured by the second camera is located on the top or left side of the display screen of the electronic device.
5. The video processing method according to any one of claims 1 to 3, characterized in that, The first amplitude spectrum is the first average amplitude spectrum obtained by averaging the amplitude spectra corresponding to each frequency point in the audio data in the first direction; and / or, The second amplitude spectrum is the second average amplitude spectrum obtained by averaging the amplitude spectra corresponding to each frequency point in the audio data in the second direction.
6. The video processing method according to any one of claims 1 to 3, characterized in that, The first amplitude spectrum is obtained by amplifying the first average amplitude spectrum, and the first average amplitude spectrum is obtained by averaging the amplitude spectra corresponding to each frequency point in the audio data in the first direction.
7. The video processing method as described in claim 6, characterized in that, The magnification process includes a first magnification process and / or a second magnification process; the video processing method further includes: Speech detection is performed on the audio data in the first direction to obtain a first detection result; The direction of arrival (DOA) of the data collected by the at least two microphones is estimated to obtain the predicted angle information. If the first detection result indicates that the audio data in the first direction includes the user's audio information, the amplitude spectrum of the audio data in the first direction is subjected to the first amplification process; and / or If the predicted angle information includes angle information within the first preset angle range, the amplitude spectrum of the audio data in the first direction is subjected to the second amplification process.
8. The video processing method according to any one of claims 1 to 3, characterized in that, The second amplitude spectrum is the amplitude spectrum obtained by amplifying the second average amplitude spectrum. The second average amplitude spectrum is obtained by averaging the amplitude spectra corresponding to each frequency point in the audio data of the second direction.
9. The video processing method as described in claim 8, characterized in that, The magnification process includes a first magnification process and / or a second magnification process; the video processing method further includes: Speech detection is performed on the audio data in the second direction to obtain a second detection result; The direction of arrival (DOA) of the data collected by the at least two microphones is estimated to obtain the predicted angle information. If the second detection result indicates that the audio data in the second direction includes the user's audio information, the amplitude spectrum of the audio data in the second direction is subjected to the first amplification process; and / or If the predicted angle information includes angle information within the second preset angle range, the amplitude spectrum of the audio data in the second direction is subjected to the second amplification process.
10. The video processing method according to any one of claims 1 to 3, characterized in that, The step of identifying whether the audio data includes the target keyword includes: The audio data is separated using a blind signal separation algorithm to obtain N audio information pieces, which are audio information pieces from different users; Each of the N audio messages is identified to determine whether the target keyword is included in the N audio messages.
11. The video processing method according to any one of claims 1 to 3, characterized in that, The first image is a preview image captured when the electronic device is in multi-camera recording mode.
12. The video processing method according to any one of claims 1 to 3, characterized in that, The first image is a video frame captured when the electronic device is recording multiple cameras.
13. The video processing method according to any one of claims 1 to 3, characterized in that, The audio data refers to the data collected by the sound pickup device in the shooting environment of the electronic device.
14. An electronic device, characterized in that, include: One or more processors and memory; The memory is coupled to the one or more processors, the memory being used to store computer program code, the computer program code including computer instructions, the one or more processors invoking the computer instructions to cause the electronic device to perform the video processing method as described in any one of claims 1 to 13.
15. A chip system, characterized in that, The chip system is applied to an electronic device, the chip system including one or more processors, the processors being configured to invoke computer instructions to cause the electronic device to perform the video processing method as described in any one of claims 1 to 13.
16. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, causes the processor to perform the video processing method according to any one of claims 1 to 13.
Citation Information
Patent Citations
Shooting mode switching method and device, storage medium and program product
CN113422903A
Automatic video stream selection
US20110164105A1