Video processing method and electronic equipment
By implementing simultaneous zooming of audio and images in electronic devices and using filter coefficients and zoom ratios to process audio, the problem of insufficient audio processing during image zooming is solved, thereby improving the viewing and listening experience of the video.
Patent Information
- Application Number
- CN202110927102.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-08-12
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2041-08-12
AI Technical Summary
During video recording, existing electronic devices are unable to process audio accordingly when the image is zoomed, resulting in the video's visual and auditory experience not meeting user needs.
By implementing a method of simultaneous zooming of audio and image in an electronic device, the audio is processed using filter coefficients and zoom ratios, so that the target sound is enhanced and the non-target sound is suppressed, achieving an audio zoom effect.
When the image is zoomed, the sound of the subject displayed in the picture is enhanced and the sound not displayed in the picture is suppressed, thereby improving the video quality.
Smart Images

Figure CN115942108B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of terminals and communication technologies, and in particular to a video processing method and electronic equipment. Background Art
[0002] With the development of electronic devices, more and more electronic devices have the function of image zoom when recording videos. In the process of electronic devices recording videos, the image zoom involved refers to the change in the size of the subject in the captured image by changing the zoom ratio. Even if the position of the subject relative to the electronic device does not change, if the zoom ratio is increased, the subject will appear larger when the electronic device displays the subject in the video, giving the user the impression that the subject is relatively closer; if the zoom ratio is decreased, the subject will appear smaller when the electronic device displays the subject in the video, giving the user the impression that the subject is relatively farther away. In this way, the subject that really needs to be displayed can be highlighted in the recorded video, making the recorded video more in line with the needs in terms of visual perception.
[0003] However, some electronic devices cannot process the audio accordingly when zooming in on a video. Thus, the video recorded by the electronic device can be made more suitable for viewing and hearing.
[0004] Therefore, how electronic devices perform audio zoom is the key to improving video quality and is the direction of research. Summary of the Invention
[0005] The present application provides a video processing method and an electronic device, which enable the audio and image to be zoomed simultaneously in the video recorded by the electronic device.
[0006] In a first aspect, the present application provides a video processing method, applied to an electronic device, the method comprising:
[0007] The electronic device starts a camera; displays a preview interface, the preview interface includes a first control; detects a first operation on the first control; starts shooting in response to the first operation; displays a shooting interface, the shooting interface includes a second control, the second control is used to adjust the zoom ratio; at a first moment, the zoom ratio is a first zoom ratio, and a first shot image is displayed, the first shot image includes a first target object and a second target object; detects a second operation on the second control; in response to the second operation, the zoom ratio is adjusted to a second zoom ratio, the second zoom ratio is greater than the first zoom ratio; at a second moment, displays a second shot image, the second shot image includes the first target object, and the second shot image does not include the second target object; at the second moment, At the second moment, the microphone collects the first audio, the first audio includes a first sound and a second sound, the first sound corresponds to the first target object, and the second sound corresponds to the second target object; a third operation on the third control is detected; in response to the third operation, shooting is stopped and the first video is saved, wherein the first video includes the first captured image, and the first video includes the second captured image and second audio at the second moment, the second audio is obtained by processing the first audio according to the second zoom ratio, the second audio includes a third sound and a fourth sound, the third sound corresponds to the first target object, the fourth sound corresponds to the second target object, the third sound is enhanced relative to the first sound, and the fourth sound is suppressed relative to the second sound.
[0008] In the above embodiment, during video recording, the electronic device can simultaneously zoom the image and audio. When the zoom ratio increases, the subject who can still be displayed on the screen becomes larger, and the subject's voice is enhanced, making the subject's voice louder. For subjects who are not displayed on the screen, the subject's voice is suppressed, making the subject's voice soft or inaudible.
[0009] In combination with the first aspect, in one embodiment, the method further includes: the electronic device processes the first audio according to the second zoom ratio to obtain a first output audio, wherein the first sound in the first output audio remains unchanged and the second sound is suppressed; according to the second zoom ratio, the first output audio is enhanced to obtain a second output audio; in the second output audio, both the first sound and the second sound are enhanced; according to the second zoom ratio, in combination with the first audio, the second sound in the second output audio is suppressed to obtain a second audio.
[0010] In the above embodiment, the electronic device determines the degree to enhance the target sound and the degree to suppress the non-target sound in the first audio according to the second zoom ratio, thereby achieving audio zoom.
[0011] In combination with the first aspect, in one embodiment, the first output audio includes one audio channel; according to the second zoom ratio, the first audio is processed to obtain the first output audio, specifically including: the electronic device obtains a first filter coefficient corresponding to the first direction, a second filter coefficient corresponding to the second direction, and a third filter coefficient corresponding to the third direction; the first direction is any direction within the range of 10° clockwise in front of the electronic device to 70° clockwise in front of the electronic device; the second direction is any direction within the range of 10° counterclockwise in front of the electronic device to 10° clockwise in front of the electronic device; the third direction is any direction within the range of 10° counterclockwise in front of the electronic device to 70° counterclockwise in front of the electronic device; the first filter coefficient is combined with the first audio to obtain a first beam corresponding to the first direction; the second filter coefficient is combined with the first audio to obtain a second beam corresponding to the second direction; the third filter coefficient is combined with the first audio to obtain a third beam corresponding to the third direction; the electronic device obtains the first output audio using the first beam, the second beam, and the third beam according to the second zoom ratio, wherein the first sound in the first output audio remains unchanged and the second sound is suppressed.
[0012] In the above embodiment, the electronic device processes the first audio through the filter coefficient, so that after the first audio is processed by the filter coefficient, a first beam, a second beam and a third beam are obtained, and at the same time, the first beam, the second beam and the third beam are fused using the second zoom ratio to obtain the first output audio, so that the target sound remains unchanged and the non-target sound is suppressed.
[0013] In combination with the first aspect, in one embodiment, the second output audio includes one audio channel; according to the second zoom ratio, the first output audio is enhanced to obtain the second output audio, specifically including: according to the second zoom ratio, determining the adjustment parameter corresponding to the second zoom ratio, and the adjustment parameter is used to enhance the audio; converting the first output audio from the frequency domain to the time domain to obtain the first output audio in the time domain; using the adjustment parameter to enhance the first output audio in the time domain to obtain the second output audio in the time domain; converting the second output audio in the time domain to the frequency domain as the second output audio.
[0014] In the above embodiment, the target sound in the first output audio may be changed from unchanged to enhanced, so that the target sound in the second output audio becomes louder as the zoom factor changes.
[0015] In combination with the first aspect, in one embodiment, the second audio includes one audio channel; according to the second zoom ratio, combined with the first audio, the second sound in the second output audio is suppressed to obtain the second audio, specifically including: using the first audio to perform Zelinsky filtering to obtain a first target gain; the first target gain is used to filter out the sound corresponding to the high-frequency frequency point in the second sound in the second output audio; using the first audio, based on the coherent diffusion power ratio algorithm, obtain the second target gain; the second target gain is used to filter out the sound corresponding to the low-frequency frequency point in the second sound in the second output audio; according to the frequency of the second output audio, a part of the first target gain and the second target gain are combined to obtain a third target gain; the second sound in the second output audio is suppressed using the third target gain to obtain the second audio.
[0016] In the above embodiment, both non-target sounds and target sounds in the second output audio are enhanced. The second output audio can be filtered to suppress the non-target sounds. The Zelinsky filter is effective for filtering high-frequency sounds, while the Coherent Spread Power Ratio algorithm is effective for filtering low-frequency sounds. Combining these two algorithms to filter the second output audio can improve the filtering effect.
[0017] In combination with the first aspect, in one embodiment, the second audio includes one audio channel; according to the second zoom ratio, in combination with the first audio, the second sound in the second output audio is suppressed to obtain the second audio, specifically including: using the first audio to perform Zelinsky filtering to obtain a first target gain; the first target gain is used to filter out the sound corresponding to the high-frequency frequency points in the second sound in the second output audio; the electronic device uses the audio to obtain a second target gain based on the coherent diffusion power ratio algorithm; the second target gain is used to filter out the sound corresponding to the low-frequency frequency points in the second sound in the second output audio; the electronic device suppresses the sound corresponding to the high-frequency frequency points in the second audio in the second audio by using the first target gain, and suppresses the sound corresponding to the low-frequency frequency points in the second audio in the second audio by using the second target gain to obtain the second audio.
[0018] In the above embodiment, both non-target sounds and target sounds in the second output audio are enhanced. The second output audio can be filtered to suppress the non-target sounds. The Zelinsky filter is effective for filtering high-frequency sounds, while the Coherent Spread Power Ratio algorithm is effective for filtering low-frequency sounds. Combining these two algorithms to filter the second output audio can improve the filtering effect.
[0019] In combination with the first aspect, in one embodiment, the first output audio includes left-channel audio and right-channel audio;
[0020] According to the second zoom ratio, the first audio is processed to obtain a first output audio, specifically including: obtaining a first filter coefficient corresponding to a first direction, a second filter coefficient corresponding to a second direction, and a third filter coefficient corresponding to a third direction; the first direction is any direction within a range of 10° clockwise in front of the electronic device to 70° clockwise in front of the electronic device; the second direction is any direction within a range of 10° counterclockwise in front of the electronic device to 10° clockwise in front of the electronic device; the third direction is any direction within a range of 10° counterclockwise in front of the electronic device to 70° counterclockwise in front of the electronic device. Any direction within a range of 70° clockwise; using the first filter coefficient in combination with the first audio to obtain a first beam corresponding to the first direction; using the second filter coefficient in combination with the first audio to obtain a second beam corresponding to the second direction; using the third filter coefficient in combination with the first audio to obtain a third beam corresponding to the third direction; according to the second zoom ratio, using the first beam and the second beam to obtain the left channel audio in the first output audio; using the second beam and the third beam to obtain the right channel audio in the first output audio, the first sound in the left channel audio and the right channel audio remains unchanged, and the second sound is suppressed.
[0021] In the above embodiment, the electronic device generates left-channel audio and right-channel audio using the first audio, suppresses non-target sounds in the left-channel audio and the right-channel audio signals, and leaves the target sound unchanged. In this way, the electronic device can achieve a stereo effect when playing the second audio.
[0022] In combination with the first aspect, in one embodiment, the second output audio includes left-channel audio and right-channel audio; according to the second zoom ratio, the first output audio is enhanced to obtain the second output audio, specifically including: according to the second zoom ratio, determining the adjustment parameters corresponding to the second zoom ratio, and the adjustment parameters are used to enhance the audio; converting the left-channel audio and the right-channel audio in the first output audio from the frequency domain to the time domain respectively, to obtain the left-channel audio and the right-channel audio in the first output audio in the time domain; using the adjustment parameters and enhancing the left-channel audio and the right-channel audio in the first output audio in the time domain respectively, to obtain the left-channel audio and the right-channel audio in the second output audio in the time domain; converting the left-channel audio and the right-channel audio in the second output audio in the time domain to the frequency domain as the second output audio.
[0023] In the above embodiment, the electronic device processes the left channel audio and the right channel audio respectively, so that the target sound in the left channel audio signal and the right channel audio signal is enhanced.
[0024] In combination with the first aspect, in one embodiment, the second audio includes left-channel audio and right-channel audio; according to the second zoom ratio, combined with the first audio, the second sound in the second output audio is suppressed to obtain the second audio, specifically including: using the first audio to perform Zelinsky filtering to obtain a first target gain; the first target gain is used to filter out the sound corresponding to the high-frequency frequency point in the second sound in the second output audio; using the first audio, based on the coherent diffusion power ratio algorithm, obtain the second target gain; the second target gain is used to filter out the sound corresponding to the low-frequency frequency point in the second sound in the second output audio; according to the frequency of the second output audio, a part of the first target gain and the second target gain are combined to obtain a third target gain; using the third target gain to suppress the second sounds in the left-channel audio and the right-channel audio in the second output audio, respectively, to obtain the left-channel audio and the right-channel audio in the second audio.
[0025] In the above embodiment, both non-target sounds and target sounds in the left and right channel audio signals are enhanced. The left and right channel audio signals can be filtered separately to suppress the non-target sounds. The Zelinsky filter is effective for filtering high-frequency sounds, while the Coherent Spread Power Ratio algorithm is effective for filtering low-frequency sounds. Combining these two algorithms to filter the left and right channel audio signals can improve the filtering effect.
[0026] In combination with the first aspect, in one embodiment, the second audio includes left channel audio and right channel audio; according to the second zoom ratio, in combination with the audio, the second sound in the second output audio is suppressed to obtain the second audio, specifically including: using the first audio to perform Zelinsky filtering to obtain a first target gain; the first target gain is used to filter out the sound corresponding to the high-frequency frequency points in the second sound in the second output audio; using the first audio, based on the coherent diffusion power ratio algorithm, obtain the second target gain; the second target gain is used to filter out the sound corresponding to the low-frequency frequency points in the second sound in the second output audio; according to the first target gain, the sound corresponding to the high-frequency frequency points in the left channel audio and the right channel audio of the second audio is suppressed, and the sound corresponding to the low-frequency frequency points in the left channel audio and the right channel audio of the second audio is suppressed using the second target gain to obtain the left channel audio and the right channel audio of the second audio.
[0027] In the above embodiment, both non-target sounds and target sounds in the left and right channel audio signals are enhanced. The left and right channel audio signals can be filtered separately to suppress the non-target sounds. The Zelinsky filter is effective for filtering high-frequency sounds, while the Coherent Spread Power Ratio algorithm is effective for filtering low-frequency sounds. Combining these two algorithms to filter the left and right channel audio signals can improve the filtering effect.
[0028] In combination with the first aspect, in one embodiment, the first filter, the second filter coefficients and the third filter are pre-set in the electronic device; among the first filter coefficients, the coefficient corresponding to the sound signal in the first direction is 1, indicating that the sound signal in the first direction is not suppressed; the closer the sound signal is to the first direction, the closer the corresponding coefficient is to 1, and the degree of suppression increases successively; among the second filter coefficients, the coefficient corresponding to the sound signal in the second direction is 1, indicating that the sound signal in the second direction is not suppressed; the closer the sound signal is to the second direction, the closer the corresponding coefficient is to 1, and the degree of suppression increases successively; among the third filter coefficients, the coefficient corresponding to the sound signal in the third direction is 1, indicating that the sound signal in the third direction is not suppressed; the closer the sound signal is to the third direction, the closer the corresponding coefficient is to 1, and the degree of suppression increases successively.
[0029] In combination with the first aspect, in one embodiment, the adjustment parameter corresponding to the second zoom ratio is pre-set in the electronic device; the value of the adjustment parameter is directly related to the zoom ratio, and any second zoom ratio uniquely corresponds to one adjustment parameter.
[0030] In a second aspect, the present application provides an electronic device, which includes: one or more processors and a memory; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code includes computer instructions, and the one or more processors call the computer instructions to enable the electronic device to: start a camera; display a preview interface, the preview interface includes a first control; detect a first operation on the first control; start shooting in response to the first operation; display a shooting interface, the shooting interface includes a second control, the second control is used to adjust the zoom ratio; at a first moment, the zoom ratio is a first zoom ratio, and a first shot image is displayed, the first shot image includes a first target object and a second target object; detect a second operation on the second control; in response to the second operation, the zoom ratio is adjusted to a second zoom ratio, the second zoom ratio is greater than the first zoom ratio; at a second moment At the second moment, a second captured image is displayed, the second captured image includes the first target object, and the second captured image does not include the second target object; at the second moment, the microphone collects a first audio, the first audio includes a first sound and a second sound, the first sound corresponds to the first target object, and the second sound corresponds to the second target object; a third operation on a third control is detected; in response to the third operation, shooting is stopped and a first video is saved, wherein the first video includes the first captured image, and the first video includes the second captured image and the second audio at the second moment, the second audio is obtained by processing the first audio according to the second zoom ratio, the second audio includes a third sound and a fourth sound, the third sound corresponds to the first target object, the fourth sound corresponds to the second target object, the third sound is enhanced relative to the first sound, and the fourth sound is suppressed relative to the second sound.
[0031] In the above embodiment, during video recording, the electronic device can simultaneously zoom the image and audio. When the zoom ratio increases, the subject who can still be displayed on the screen becomes larger, and the subject's voice is enhanced, making the subject's voice louder. For subjects who are not displayed on the screen, the subject's voice is suppressed, making the subject's voice soft or inaudible.
[0032] In combination with the second aspect, in one embodiment, the one or more processors are also used to call the computer instructions to enable the electronic device to perform: processing the first audio according to the second zoom ratio to obtain a first output audio, wherein the first sound in the first output audio remains unchanged and the second sound is suppressed; enhancing the first output audio according to the second zoom ratio to obtain a second output audio; both the first sound and the second sound in the second output audio are enhanced; suppressing the second sound in the second output audio according to the second zoom ratio in combination with the first audio to obtain a second audio.
[0033] In the above embodiment, the electronic device determines the degree to enhance the target sound and the degree to suppress the non-target sound in the first audio according to the second magnification, thereby achieving audio zoom.
[0034] In combination with the second aspect, in one embodiment, the first output audio includes one audio channel; the one or more processors are specifically used to call the computer instructions to enable the electronic device to execute: obtaining a first filter coefficient corresponding to a first direction, a second filter coefficient corresponding to a second direction, and a third filter coefficient corresponding to a third direction; the first direction is any direction within the range of 10° clockwise in front of the electronic device to 70° clockwise in front of the electronic device; the second direction is any direction within the range of 10° counterclockwise in front of the electronic device to 10° clockwise in front of the electronic device; the third direction is any direction within the range of 10° counterclockwise in front of the electronic device to 70° counterclockwise in front of the electronic device; using the first filter coefficient in combination with the first audio to obtain a first beam corresponding to the first direction; using the second filter coefficient in combination with the first audio to obtain a second beam corresponding to the second direction; using the third filter coefficient in combination with the first audio to obtain a third beam corresponding to the third direction; the electronic device obtains the first output audio according to the second zoom ratio using the first beam, the second beam, and the third beam, wherein the first sound in the first output audio remains unchanged and the second sound is suppressed.
[0035] In the above embodiment, the electronic device processes the first audio through the filter coefficient, so that after the first audio is processed by the filter coefficient, the first beam, the second beam and the third beam are obtained, and the first beam, the second beam and the third beam are fused using the second zoom ratio to obtain the first output audio, so that the target sound remains unchanged and the non-target sound is suppressed.
[0036] In combination with the second aspect, in one embodiment, the second output audio includes one audio channel; the one or more processors are specifically used to call the computer instructions to enable the electronic device to execute: determining the adjustment parameters corresponding to the second zoom ratio according to the second zoom ratio, and the adjustment parameters are used to enhance the audio; converting the first output audio from the frequency domain to the time domain to obtain the first output audio in the time domain; using the adjustment parameters to enhance the first output audio in the time domain to obtain the second output audio in the time domain; converting the second output audio in the time domain to the frequency domain as the second output audio.
[0037] In the above embodiment, the target sound in the first output audio may be changed from unchanged to enhanced, so that the target sound in the second output audio becomes louder as the zoom factor changes.
[0038] In combination with the second aspect, in one embodiment, the second audio includes one audio channel; the one or more processors are specifically used to call the computer instructions to enable the electronic device to execute: using the first audio to perform Zelinsky filtering to obtain a first target gain; the first target gain is used to filter out the sound corresponding to the high-frequency frequency point in the second sound in the second output audio; using the first audio, based on the coherent diffusion power ratio algorithm, to obtain a second target gain; the second target gain is used to filter out the sound corresponding to the low-frequency frequency point in the second sound in the second output audio; according to the frequency of the second output audio, a part of the first target gain and the second target gain are combined to obtain a third target gain; and the third target gain is used to suppress the second sound in the second output audio to obtain the second audio.
[0039] In the above embodiment, both non-target sounds and target sounds in the second output audio are enhanced. The second output audio can be filtered to suppress the non-target sounds. The Zelinsky filter is effective for filtering high-frequency sounds, while the Coherent Spread Power Ratio algorithm is effective for filtering low-frequency sounds. Combining these two algorithms to filter the second output audio can improve the filtering effect.
[0040] In combination with the second aspect, in one embodiment, the second audio includes one audio channel; the one or more processors are specifically used to call the computer instructions to enable the electronic device to execute: using the first audio to perform Zelinsky filtering to obtain a first target gain; the first target gain is used to filter out the sound corresponding to the high-frequency frequency points in the second sound in the second output audio; the electronic device uses the audio to obtain a second target gain based on a coherent diffusion power ratio algorithm; the second target gain is used to filter out the sound corresponding to the low-frequency frequency points in the second sound in the second output audio; the electronic device suppresses the sound corresponding to the high-frequency frequency points in the second audio in the second audio according to the first target gain, and suppresses the sound corresponding to the low-frequency frequency points in the second audio in the second audio according to the second target gain, to obtain the second audio.
[0041] In the above embodiment, both non-target sounds and target sounds in the second output audio are enhanced. The second output audio can be filtered to suppress the non-target sounds. The Zelinsky filter is effective for filtering high-frequency sounds, while the Coherent Spread Power Ratio algorithm is effective for filtering low-frequency sounds. Combining these two algorithms to filter the second output audio can improve the filtering effect.
[0042] In combination with the second aspect, in one embodiment, the first output audio includes left channel audio and right channel audio; the one or more processors are specifically used to call the computer instruction to enable the electronic device to execute: obtain a first filter coefficient corresponding to a first direction, a second filter coefficient corresponding to a second direction, and a third filter coefficient corresponding to a third direction; the first direction is any direction within a range of 10° clockwise in front of the electronic device to 70° clockwise in front of the electronic device; the second direction is any direction within a range of 10° counterclockwise in front of the electronic device to 10° clockwise in front of the electronic device; the third direction is the direction in front of the electronic device. Any direction within the range of 10° counterclockwise to 70° counterclockwise in front of the electronic device; using the first filter coefficient in combination with the first audio to obtain a first beam corresponding to the first direction; using the second filter coefficient in combination with the first audio to obtain a second beam corresponding to the second direction; using the third filter coefficient in combination with the first audio to obtain a third beam corresponding to the third direction; according to the second zoom ratio, using the first beam and the second beam to obtain the left channel audio in the first output audio; using the second beam and the third beam to obtain the right channel audio in the first output audio, the first sound in the left channel audio and the right channel audio remains unchanged, and the second sound is suppressed.
[0043] In the above embodiment, the electronic device generates left-channel audio and right-channel audio using the first audio, suppresses non-target sounds in the left-channel audio and the right-channel audio signals, and leaves the target sound unchanged. In this way, the electronic device can achieve a stereo effect when playing the second audio.
[0044] In combination with the second aspect, in one embodiment, the second output audio includes left channel audio and right channel audio; the one or more processors are specifically used to call the computer instructions to enable the electronic device to execute: according to the second zoom ratio, determine the adjustment parameters corresponding to the second zoom ratio, and the adjustment parameters are used to enhance the audio; convert the left channel audio and the right channel audio in the first output audio from the frequency domain to the time domain, respectively, to obtain the left channel audio and the right channel audio in the first output audio in the time domain; use the adjustment parameters to enhance the left channel audio and the right channel audio in the first output audio in the time domain, respectively, to obtain the left channel audio and the right channel audio in the second output audio in the time domain; convert the left channel audio and the right channel audio in the second output audio in the time domain to the frequency domain as the second output audio.
[0045] In the above embodiment, the electronic device processes the left channel audio and the right channel audio respectively, so that the target sound in the left channel audio signal and the right channel audio signal is enhanced.
[0046] In combination with the second aspect, in one embodiment, the second audio includes left channel audio and right channel audio; the one or more processors are specifically used to call the computer instructions to enable the electronic device to execute: using the first audio to perform Zelinsky filtering to obtain a first target gain; the first target gain is used to filter out the sound corresponding to the high-frequency frequency point in the second sound in the second output audio; using the first audio, based on the coherent diffusion power ratio algorithm, obtain the second target gain; the second target gain is used to filter out the sound corresponding to the low-frequency frequency point in the second sound in the second output audio; according to the frequency of the second output audio, a part of the first target gain and the second target gain are combined to obtain a third target gain; using the third target gain to suppress the second sound in the left channel audio and the right channel audio in the second output audio, respectively, to obtain the left channel audio and the right channel audio in the second audio.
[0047] In the above embodiment, both non-target sounds and target sounds in the left and right channel audio signals are enhanced. The left and right channel audio signals can be filtered separately to suppress the non-target sounds. The Zelinsky filter is effective for filtering high-frequency sounds, while the Coherent Spread Power Ratio algorithm is effective for filtering low-frequency sounds. Combining these two algorithms to filter the left and right channel audio signals can improve the filtering effect.
[0048] In combination with the second aspect, in one embodiment, the second audio includes left channel audio and right channel audio; the one or more processors are specifically used to call the computer instructions to enable the electronic device to execute: using the first audio to perform Zelinsky filtering to obtain a first target gain; the first target gain is used to filter out the sound corresponding to the high-frequency frequency points in the second sound in the second output audio; using the first audio, based on the coherent diffusion power ratio algorithm, obtain the second target gain; the second target gain is used to filter out the sound corresponding to the low-frequency frequency points in the second sound in the second output audio; according to the first target gain, the sound corresponding to the high-frequency frequency points in the left channel audio and the right channel audio of the second audio is suppressed, and the sound corresponding to the low-frequency frequency points in the left channel audio and the right channel audio of the second audio is suppressed using the second target gain to obtain the left channel audio and the right channel audio of the second audio.
[0049] In the above embodiment, both non-target sounds and target sounds in the left and right channel audio signals are enhanced. The left and right channel audio signals can be filtered separately to suppress the non-target sounds. The Zelinsky filter is effective for filtering high-frequency sounds, while the Coherent Spread Power Ratio algorithm is effective for filtering low-frequency sounds. Combining these two algorithms to filter the left and right channel audio signals can improve the filtering effect.
[0050] In a third aspect, the present application provides an electronic device comprising: one or more processors and a memory; the memory is coupled to the one or more processors, the memory being used to store computer program code, the computer program code comprising computer instructions, the one or more processors calling the computer instructions to enable the electronic device to execute the method described in the first aspect or any one of the embodiments of the first aspect.
[0051] In the above embodiment, during video recording, the electronic device can simultaneously zoom the image and audio. When the zoom ratio increases, the subject who can still be displayed on the screen becomes larger, and the subject's voice is enhanced, making the subject's voice louder. For subjects who are not displayed on the screen, the subject's voice is suppressed, making the subject's voice soft or inaudible.
[0052] In a fourth aspect, an embodiment of the present application provides a chip system, which is applied to an electronic device. The chip system includes one or more processors, which are used to call computer instructions to enable the electronic device to execute the method described in the first aspect or any one of the embodiments of the first aspect.
[0053] In the above embodiment, during video recording, the electronic device can simultaneously zoom the image and audio. When the zoom ratio increases, the subject who can still be displayed on the screen becomes larger, and the subject's voice is enhanced, making the subject's voice louder. For subjects who are not displayed on the screen, the subject's voice is suppressed, making the subject's voice soft or inaudible.
[0054] In a fifth aspect, an embodiment of the present application provides a computer program product comprising instructions, which, when executed on an electronic device, enables the electronic device to execute the method described in the first aspect or any one of the embodiments of the first aspect.
[0055] In the above embodiment, during video recording, the electronic device can simultaneously zoom the image and audio. When the zoom ratio increases, the subject who can still be displayed on the screen becomes larger, and the subject's voice is enhanced, making the subject's voice louder. For subjects who are not displayed on the screen, the subject's voice is suppressed, making the subject's voice soft or inaudible.
[0056] In a sixth aspect, an embodiment of the present application provides a computer-readable storage medium comprising instructions, which, when executed on an electronic device, enables the electronic device to execute the method described in the first aspect or any one of the embodiments of the first aspect.
[0057] In the above embodiment, during video recording, the electronic device can simultaneously zoom the image and audio. As the zoom factor increases, subjects who are still visible in the image become larger, and their voices are amplified, making them louder. Subjects who are not visible in the image are suppressed, making their voices soft or inaudible. The key to quality lies in the direction of research. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] Figure 1a-Figure 1e A set of exemplary user interfaces for zooming an image when an electronic device is recording a video is shown;
[0059] Figure 2-Figure 4 A set of exemplary user interfaces for video recorded by an electronic device with image zoom but no audio zoom;
[0060] Figure 5a-5d A set of exemplary user interfaces for an electronic device when previewing a video recorded by the electronic device;
[0061] Figure 6a-6f A set of exemplary user interfaces for post-processing recorded videos by an electronic device provided in an embodiment of the present application;
[0062] Figure 7 A schematic diagram of a video obtained by real-time processing of a current frame image and a current frame input audio signal set provided in an embodiment of the present application;
[0063] Figure 8 An exemplary flow chart of an electronic device processing a current frame input audio signal in real time provided by an embodiment of the present application;
[0064] Figures 9-11 A schematic diagram of beamforming technology is shown;
[0065] Figure 12 An exemplary flow chart for generating a left channel audio signal or a right channel audio signal for an electronic device;
[0066] Figure 13 An example diagram of the first direction, the second direction, and the third direction provided in an embodiment of the present application;
[0067] Figure 14 An exemplary flow chart of the electronic device provided in an embodiment of the present application generating a first filter corresponding to the first direction;
[0068] Figure 15 An exemplary flow chart of adjusting the amplitude of a first output audio signal for an electronic device;
[0069] Figure 16A schematic flow chart of an electronic device suppressing a non-target sound signal in a second output audio signal;
[0070] Figure 17 A schematic diagram of real-time processing of a current frame image and a current frame input audio signal set and then playing them;
[0071] Figure 18 An exemplary flow chart of post-processing a frame of input audio signal by an electronic device provided in an embodiment of the present application;
[0072] Figure 19 It is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0073] The terms used in the following examples of the present application are for the purpose of describing specific embodiments only and are not intended to limit the present application. As used in the specification and appended claims of the present application, the singular expressions "a," "an," "said," "above," "the," and "this" are intended to include plural expressions as well, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in the present application refers to and encompasses any or all possible combinations of one or more of the listed items.
[0074] In the following, the terms "first" and "second" are used for descriptive purposes only and should not be understood to imply or suggest relative importance or implicitly indicate the number of the technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of the embodiments of this application, unless otherwise specified, "plurality" means two or more.
[0075] The term "user interface (UI)" in the following embodiments of this application refers to a medium interface for interaction and information exchange between an application or operating system and a user, which realizes the conversion between the internal form of information and the form acceptable to the user. The user interface is a source code written in a specific computer language such as Java and extensible markup language (XML). The interface source code is parsed and rendered on an electronic device and finally presented as content that the user can recognize. The commonly used form of user interface is graphical user interface (GUI), which refers to a user interface related to computer operations that is displayed in a graphical manner. It can be a visual interface element such as text, icons, buttons, menus, tabs, text boxes, dialog boxes, status bars, navigation bars, widgets, etc. displayed on the display screen of an electronic device.
[0076] To facilitate understanding, the following first introduces the relevant terms and concepts involved in the embodiments of this application.
[0077] (1) Focal length and field of view
[0078] In the embodiment of the present application, the focal length refers to the focal length used by the electronic device when recording a video or taking an image.
[0079] The field of view (FOV) is the angle between the two edges of the electronic device's lens, with the lens at its apex, and the maximum range through which a subject can pass. The angle determines the electronic device's field of view: subjects within the FOV are included in the image, while subjects outside the FOV are not.
[0080] Specifically, when an electronic device is recording a video or capturing an image, for the same subject whose relative position to the electronic device does not change, if the focal length is different, the electronic device can acquire different images. For example, in one case, the larger the focal length used by the electronic device, the smaller the field of view of the electronic device. At this time, the larger the subject appears in the image acquired by the electronic device. Since the display screen of the electronic device is limited, sometimes only a portion of the subject can be displayed. In another case, the smaller the focal length used by the electronic device, the larger the field of view of the electronic device. The smaller the subject appears in the image acquired by the electronic device. Generally, the larger the field of view of the electronic device, the more other subjects will be displayed in the acquired image.
[0081] In some embodiments, when recording a video or capturing an image, the electronic device may change the focal length according to user settings.
[0082] In other embodiments, the electronic device may change the focal length according to some preset rules when recording a video or capturing an image. For example, when recording an interesting video, the electronic device may change the focal length according to preset rules.
[0083] The change in focal length includes increasing the focal length and decreasing the focal length. In some embodiments, the electronic device can change the focal length by adjusting the zoom ratio. The user can select the zoom ratio through a zoom ratio control in the user interface or by inputting a gesture command in the user interface to select the zoom ratio.
[0084] Among them, the zoom ratio control can be Figure 1b, the zoom ratio control 111 shown in the user interface 11 is shown. Please refer to the following description of the zoom ratio control 111. By adjusting the zoom ratio, the user continuously enlarges the object being photographed on the preview interface. The user can select the zoom ratio using the zoom ratio button on the camera device; or by inputting gesture commands through the camera device's display screen to select the zoom ratio. Generally, zoom photography includes two methods: optical zoom photography and digital zoom photography. Both methods can change the size of the object in the preview image displayed by the electronic device.
[0085] (2) Image zoom and audio zoom
[0086] Image zoom refers to the change of focal length of an electronic device during the process of shooting an image. The electronic device can change the focal length and complete the image zoom by adjusting the zoom ratio. For example, when a user shoots a distant object through an electronic device, the object will inevitably appear smaller in the preview image displayed. Without changing its own position, the user can choose to increase the zoom ratio to make the object displayed on the electronic device interface larger, thereby achieving image zoom. In the embodiment of the present application, it can also be understood that by adjusting the zoom ratio, the object displayed on the electronic device interface is enlarged or reduced; it can be applied in the process of recording a video, and can also be applied in the process of playing a video.
[0087] Audio zoom can be analogous to image zoom. As the zoom ratio increases, the subject displayed on the electronic device in the video becomes larger, giving the user the impression that the subject is closer. Consequently, the subject's sound correspondingly becomes louder. As the zoom ratio decreases, the subject displayed on the electronic device in the video becomes smaller, giving the user the impression that the subject is farther away. Consequently, the subject's sound can also be correspondingly quieter. If both the image and the corresponding audio can be zoomed, the effect of simultaneous audio and image zooming can be achieved, enhancing the user's sensory experience and increasing interest.
[0088] Figure 1a-Figure 1e A set of exemplary user interfaces for performing image zooming when an electronic device is recording a video is shown.
[0089] Figure 1a-Figure 1e The electronic device shown has three microphones. Figure 1a For example, the microphones of an electronic device may include a first microphone, a second microphone, and a third microphone. When the electronic device records a video, the microphones capture multiple frames of audio signals from the environment and process them to generate an audio stream. Simultaneously, the camera captures multiple frames of images and processes them to generate an image stream. The audio stream is then mixed with the image stream to produce the recorded video.
[0090] It should be understood that when recording a video, the electronic device may use N microphones, where N is a positive integer greater than or equal to 2, and is not limited to the first microphone, second microphone, and third microphone mentioned above.
[0091] Figure 1a-Figure 1d In the video, the subjects may include subject 101 (dog), subject 102 (man near the dog in the picture), and subject 103 (child in the picture), etc. At this time, the electronic device uses the rear camera to record the video.
[0092] like Figure 1a As shown, the user interface 10 of the electronic device is a preview interface when recording a video. The user interface 11 may include a recording control 101, which can be used to receive an instruction to record a video. In response to a user's operation on the first control 101 (such as a click operation), the electronic device can start recording a video and display the video. Figure 1b The user interface shown.
[0093] like Figure 1b As shown, the user interface 11 of the electronic device can be a user interface used when recording a video. In this case, the electronic device captures the image corresponding to the first second of the video and the audio signal corresponding to the first second. The user interface may include a zoom control 111, a zoom increase control 112, and a zoom decrease control 113. Zoom control 111 is used to receive instructions for changing the zoom factor and to inform the user of the current zoom factor of the electronic device. For example, 1.0 indicates a zoom factor of 1x, and 5.0 indicates a zoom factor of 5x, where a zoom factor of 5x is greater than a zoom factor of 1x. Zoom increase control 112 is used to receive instructions for increasing the zoom factor. Zoom decrease control 113 is used to receive instructions for decreasing the zoom factor. As can be seen from zoom control 11, during the video recording process, the image corresponding to the first second was captured at a zoom factor of 1x, and the image includes subjects 101, 102, and 103. In response to the user sliding up the zoom magnification control 111, the electronic device can change the zoom magnification when recording the video. The electronic device can display Figure 1c The user interface shown.
[0094] like Figure 1c As shown, the user interface 12 is a user interface for the electronic device when recording a video. At this time, the electronic device obtains the image corresponding to the 2nd second in the video and the audio signal corresponding to the 2nd second. At this time, the positions of all the subjects relative to the electronic device have not changed. However, since the zoom ratio is increased from 1x to 5x, the field of view of the electronic device becomes smaller. Compared with Figure 1bIn the image captured by the electronic device at a 1x zoom ratio, it can be seen that the subject 101 is no longer displayed in the image displayed in the user interface 12, and the other subjects displayed are enlarged, for example, the subjects 102 and 103 are enlarged. At this time, in response to the user sliding the zoom ratio control 111 upward, the electronic device can change the zoom ratio when recording the video. The electronic device can display Figure 1d The user interface shown.
[0095] like Figure 1d As shown, the user interface 13 is a user interface for an electronic device when recording a video. The user interface 13 may include a stop recording control 131, and the stop recording control 131 can be used to receive an instruction to stop recording the video. At this time, the electronic device obtains the image corresponding to the 3rd second in the video and the audio signal corresponding to the 3rd second. At this time, the positions of all the subjects relative to the electronic device have not changed. However, since the zoom ratio is increased from 1x zoom ratio to 10x zoom ratio, the field of view of the electronic device becomes smaller. Compared to Figure 1c In the image captured by the electronic device at a 5x zoom ratio, it can be seen that the image displayed in the user interface 13 no longer displays the subjects 101 and 103, and the other displayed subjects are enlarged, for example, the subject 102 is enlarged. In response to the user stopping the operation on the recording control 131 (for example, a click operation), the electronic device can display the following Figure 1e The user interface shown.
[0096] like Figure 1e The user interface 14 shown in FIG is a user interface after the electronic device completes video recording. The electronic device can save the recorded video.
[0097] It should be understood that the above Figure 1a-Figure 1e The user interfaces shown here illustrate a set of exemplary user interfaces for an electronic device that changes the field of view due to a change in zoom ratio during video recording, thereby changing the captured image. These user interfaces should not limit the embodiments of the present application. The electronic device may also change the zoom ratio in other ways. The embodiments of the present application are not limited thereto.
[0098] In the embodiments of the present application, when an electronic device generates a video, it can zoom in or out according to the change in zoom ratio, which is called image zoom. It can also process audio according to the change in zoom ratio, which is called audio zoom. The following term (2) will introduce image zoom and audio zoom in detail.
[0099] (3) Inhibition and enhancement
[0100] In the embodiment of the present application, suppression refers to reducing the energy of the audio signal so that the audio signal sounds smaller or even inaudible. Suppression of the audio signal can be achieved by reducing the amplitude of the audio signal.
[0101] Enhancement refers to increasing the energy of an audio signal so that the audio signal sounds louder. Enhancement of the audio signal can be achieved by increasing the amplitude of the audio signal.
[0102] The amplitude is used to indicate the voltage corresponding to the audio signal; it can also indicate the energy of the audio signal; or the decibel level.
[0103] (4) Beamforming and gain coefficient
[0104] In an embodiment of the present application, beamforming can be used to describe the correspondence between the audio collected by the microphone of an electronic device and the audio when it is transmitted to the speaker for playback. The correspondence is a set of gain coefficients, which are used to indicate the degree of suppression of the audio signals in various directions collected by the microphone. Suppression refers to reducing the energy of the audio signal so that the audio signal sounds smaller or even inaudible. The degree of suppression is used to describe the degree to which the audio signal is reduced. The greater the degree of suppression, the more the energy of the audio signal is reduced. For example, a gain coefficient of 0.0 means that the audio signal is completely removed, and a gain coefficient of 1.0 means no suppression. The closer it is to 0.0, the greater the degree of suppression, and the closer it is to 1.0, the smaller the degree of suppression.
[0105] In one solution, when generating a video, an electronic device can zoom the image based on the zoom factor, but it doesn't zoom the audio based on the zoom factor. As a result, the video recorded by the electronic device may show the subject's image moving closer or farther away, while the subject's voice remains unchanged. Furthermore, in addition to the subject's voice, the recorded video may also contain sounds from other objects not shown in the image.
[0106] In this solution, the process of recording video by electronic equipment refers to the aforementioned Figure 1a-Figure 1e If the electronic device does not perform audio zoom during the video recording process, the process of playing the video by the electronic device can refer to the following Figure 2-Figure 4 Description.
[0107] Figure 2-Figure 4 A set of exemplary user interfaces for video recorded for an electronic device with image zoom but no audio zoom.
[0108] Figure 2 (b)- Figure 4In (b), in icon 301, the sound signal is a solid line, which indicates that in the embodiment of the present application, the sound of the person being photographed does not belong to the suppressed object, but belongs to the target object, and the sound of the person being photographed 102 can be heard when the video is played. In icon 302, the person being photographed 101 is drawn with a dotted line, indicating that the person being photographed 101 will not appear in the video screen, and is a non-target object being photographed. In icon 303, the person being photographed 102 is drawn with a solid line, indicating that the person being photographed will appear in the video screen, and is a target object being photographed. Optionally, in the embodiment of the present application, the sound of the non-target object belongs to the suppressed object. It can be understood that the target object is an object that appears in the video screen, and its sound does not need to be suppressed; correspondingly, the non-target object is an object that does not appear in the video screen, and its sound needs to be suppressed.
[0109] It should be understood that Figure 2 (b)- Figure 4 In (b), icons with similar shapes have the same meaning and will not be explained one by one. For example, when the subject is drawn with a dotted line, it means that the subject will not appear in the video and is a non-target object of the video. Figure 4 In (b), the photographed persons 101 and 103 are both drawn with dotted lines, indicating that the photographed persons 101 and 103 will not appear in the video and are non-target objects of the shooting.
[0110] like Figure 2 As shown in (a) of FIG, the user interface 20 is a user interface when the electronic device plays a video. At this time, the electronic device plays Figure 1b The image and audio corresponding to the first second of the recording are shown in FIG. The zoom magnification of the electronic device is 1. In the user interface 20, it can be seen that the currently played image includes the subject 101, the subject 102, and the subject 103.
[0111] like Figure 2 As shown in (b), it is the beamforming diagram corresponding to the audio when the electronic device plays the audio corresponding to the first second. Beamforming can be used to describe the correspondence between the audio collected by the microphone of the electronic device and the audio when it is transmitted to the speaker for playback. This correspondence is a set of gain coefficients, which are used to represent the degree of suppression of the audio signals in various directions collected by the microphone. Suppression refers to reducing the energy of the audio signal so that the audio signal sounds smaller or even inaudible. The degree of suppression is used to describe the degree of reduction of the audio signal. The greater the degree of suppression, the more the energy of the audio signal is reduced. For example, Figure 2In (b), a gain factor of 0.0 completely removes the audio signal, and a gain factor of 1.0 does not suppress it. The closer to 0.0, the greater the suppression, and the closer to 1.0, the less suppression.
[0112] For a detailed description of the gain coefficient, please refer to the relevant description of the gain coefficient in step S302 below, which will not be repeated here.
[0113] The electronic device can suppress the audio signal collected by the microphone according to the gain coefficient, and then transmit it to the speaker for playback. Figure 2 It can be seen from the beamforming diagram shown in (b) that the gain coefficient of the electronic device for the audio in each direction collected by the microphone is 1. This means that the electronic device does not suppress the collected audio signal. For example, the gain coefficients corresponding to the directions of the photographed person 101, the photographed person 102, and the photographed person 103 are all 1 (or close to 1), then the audio signal collected by the electronic device includes the voices of the photographed person 101, the photographed person 102, and the photographed person 103, and will not be suppressed. That is, during playback, the voices of the photographed person 101, the photographed person 102, and the photographed person 103 in the audio are not suppressed, and the user can hear the voices of the photographed person 101, the photographed person 102, and the photographed person 103.
[0114] After the electronic device finishes playing the video corresponding to the first second, it can play the video corresponding to the second second.
[0115] like Figure 3 As shown in (a) of FIG, the user interface 30 is a user interface when the electronic device plays the video corresponding to the second second. At this time, the electronic device plays Figure 1c At this time, the zoom ratio of the electronic device increases from 1x to 5x, and it can be seen in the user interface 30 that the currently played image includes the photographed person 102 and the photographed person 103, but no longer includes the photographed person 101.
[0116] like Figure 3(b) in the figure shows the beamforming diagram corresponding to the audio when the electronic device plays the audio corresponding to the second second. The gain coefficients corresponding to the directions of the subjects 101, 102, and 103 are all 1 (or close to 1), so the audio signal collected by the electronic device includes the voices of the subjects 101, 102, and 103, and they are not suppressed. That is, during playback, the voices of the subjects 101, 102, and 103 in the audio are not suppressed, and the user can hear the voices of the subjects 101, 102, and 103. That is, the corresponding audio still includes the voice of the subject 101, but the played image no longer includes the subject 101. The images of the subjects 102 and 103 become larger, but the voices of the subjects 102 and 103 in the audio do not become larger.
[0117] After the electronic device finishes playing the video corresponding to the 2nd second, it can play the video corresponding to the 3rd second.
[0118] Reference to the aforementioned Figure 2 as well as Figure 3 Description, combined with Figure 4 (a) and Figure 4 In (b), the user interface 40 is a user interface when the electronic device plays the video corresponding to the 3rd second. Figure 1d The image and audio corresponding to the 3rd second of the recording are shown in FIG40 . At this time, the zoom ratio of the electronic device has increased from 1x to 10x. In the user interface 40 , it can be seen that the currently played image includes subject 102, but no longer includes subjects 101 and 103. However, the corresponding audio still includes the voices of subjects 101 and 103, but the played image no longer includes subjects 101 and 103. Moreover, the image of subject 102 has become larger, but the voice of subject 102 has not become larger in the audio.
[0119] When playing back a video recorded using this solution, there will be a mismatch between the user's visual and auditory perception. That is, the user's hearing of the subject's voice will not change in volume as the image size increases or decreases. Furthermore, the user can also hear the sounds of other objects not shown in the image.
[0120] When implementing the video processing method of the present application, the electronic device can perform image zoom according to the change in zoom ratio when generating a video, and can also perform audio zoom according to the change in zoom ratio. The electronic device performing audio zoom on the audio includes: increasing the zoom ratio, decreasing the field of view angle, suppressing the sound of objects outside the shooting range, and enhancing the sound of the subject within the shooting range. Decreasing the zoom ratio, increasing the field of view angle, suppressing the sound of objects outside the shooting range, and reducing the sound of the subject within the shooting range.
[0121] Enhancement refers to increasing the energy of an audio signal so that it sounds louder. Suppression refers to reducing the energy of an audio signal so that it sounds quieter or even inaudible. This energy change can be achieved by adjusting the amplitude of the audio signal. The enhancement and suppression of the audio signal can be found in the description of step S105 below and will not be repeated here.
[0122] In this way, the size of the image of the subject and the sound of the subject can be changed in the video recorded by the electronic device. Furthermore, in the recorded video, except for the sound of the subject, the sounds of other objects not shown in the image are suppressed, so that the user can hear the sounds of other objects not shown in the image.
[0123] In this application, the process of recording video by electronic equipment refers to the aforementioned Figure 1a-Figure 1e The electronic device performs audio zoom during the video recording process, achieving the effect of simultaneous zooming of the image and audio.
[0124] The following introduces three usage scenarios involved in the embodiments of the present application. In the embodiment of the present application, in the process of the electronic device generating a video, the image can be enlarged or reduced according to the change of the zoom ratio (for the sake of convenience, image zoom is used for explanation in this application, and image zoom means the enlargement or reduction of the image presented on the mobile phone interface), and the audio can also be processed according to the change of the zoom ratio (for the sake of convenience, it will be referred to as audio zoom later). In this way, when the zoom ratio increases, the field of view angle becomes smaller, and the electronic device can suppress the sound of objects outside the field of view angle and enhance the sound of the subject within the field of view angle. When the zoom ratio decreases, the field of view angle increases, and the electronic device can suppress the sound of objects outside the field of view angle and weaken the sound of the subject within the field of view angle. The process of audio zoom by the electronic device can occur in different scenarios, and three of the scenarios are described in detail below.
[0125] Scenario 1: When recording a video, an electronic device can use the video processing method described in this application to perform real-time image zoom on each captured image frame according to the change in zoom ratio, and simultaneously perform real-time audio zoom on each captured audio frame according to the change in zoom ratio. Finally, an image stream is generated based on multiple image frames, and an audio stream is generated based on multiple audio frames. The image stream and audio stream are then mixed to obtain a recorded video. The recorded video is then played.
[0126] The exemplary user interface involved in recording the video in scene 1 can refer to the aforementioned Figure 1a-Figure 1e The process of playing this video can be referred to the following description. Figures 9-11 The description is not repeated here.
[0127] The methods involved in video processing in scenario 1 can be referred to as follows Figure 8 The description of steps S101 to S107 is not repeated here.
[0128] Scenario 2: While recording video, the electronic device can capture audio through a microphone while displaying the image. This audio can then be processed and played through connected headphones. This means that each captured image frame is zoomed in real time based on the zoom factor, and each captured audio frame is zoomed in real time based on the zoom factor. Each generated image and audio frame is played.
[0129] The exemplary user interface involved in scenario 2 can refer to the following Figure 5a-5d Description.
[0130] Figure 5a-5d In the figure, the subjects may include subject 101 (dog), subject 102 (the man closest to the dog in the figure), and subject 103 (the child in the figure). At this time, the electronic device uses the rear camera to record the video. It is assumed that from the time the electronic device starts recording the video to the time it ends recording the video, the positions of all the subjects relative to the electronic device do not change, and the volume of the sounds made by all the subjects does not change. In order to avoid the audio played during the preview of the video being collected by the electronic device and affecting the audio that needs to be collected later, the electronic device can play the audio through the connected headphones.
[0131] In some embodiments, the electronic device may not play audio through headphones, but may directly play audio through the local speakers, and then use acoustic echo cancellation (AEC) to eliminate the audio played by the speakers of the electronic device.
[0132] like Figure 5aAs shown, the user interface 80 is a user interface for an electronic device to preview a video. The user interface 80 may include a recording control 801, a zoom ratio control 802, a zoom ratio increase control 803, and a zoom ratio decrease control 804. The zoom ratio of the electronic device is 1x zoom ratio. The electronic device can capture images, process them, and then display them using a display screen. At the same time, the electronic device can process the captured audio and play it through a pair of headphones 805. For the user, the voices of the subject 101, the subject 102, and the subject 103 can be heard. In response to the user sliding the zoom ratio control 802 upward, the electronic device can change the zoom ratio when recording the video. The electronic device can display Figure 5b The user interface shown.
[0133] like Figure 5b As shown, the user interface 81 is another user interface for the electronic device to preview the video. The zoom ratio of the electronic device is changed from 1x zoom ratio to 5x zoom ratio. The electronic device can capture images, process them, and then display them using the display screen. At the same time, the electronic device can process the captured audio and play it through the earphones 805. For the user, the user can see the subject 102 and the subject 103 become larger from the image displayed in the user interface 81, but the subject 101 cannot be seen. At the same time, the user can hear the voices of the subject 102 and the subject 103 and the voices become louder, but the voice of the subject 101 cannot be heard. In response to the user sliding the zoom ratio control 802 upward, the electronic device can change the zoom ratio when recording the video. The electronic device can display Figure 5c The user interface shown.
[0134] like Figure 5c As shown, the user interface 82 is another user interface for the electronic device to preview the video. The zoom ratio of the electronic device changes from 1x zoom ratio to 10x zoom ratio. The electronic device can capture images, process them, and then display them using the display screen. At the same time, the electronic device can process the captured audio and play it through the earphones 805. For the user, the user can see that the subject 102 becomes larger in the image displayed in the user interface 82, but the subject 101 and the subject 102 cannot be seen. At the same time, the user can hear the voice of the subject 102 and the voice becomes louder, but the voices of the subject 101 and the subject 103 cannot be heard. In response to the user's operation on the recording control 801 (such as a click operation), the electronic device can start recording the video. Display Figure 5d The user interface shown.
[0135] Figure 5d A user interface 83 for recording a video of an electronic device. At this time, the zoom magnification of the electronic device is 10 times the zoom magnification.
[0136] In this way, the electronic device can debug the optimal zoom ratio when recording the video by previewing the video it records.
[0137] In some embodiments, in addition to playing each generated frame of image and audio in real time, the electronic device can also save each frame of image and audio, and finally generate an image stream based on multiple frames of image and an audio stream based on multiple frames of audio, and then mix the image stream and audio stream to obtain a video.
[0138] The method involved in video processing in scenario 2 can be referred to the following description of steps S101 to S107, which will not be repeated here.
[0139] Scenario 3: The electronic device can perform audio zoom processing on the audio stream in a recorded video. The electronic device can save the zoom ratio corresponding to each frame of audio when recording the video. Then, the electronic device can obtain any frame of audio in the audio stream and the zoom ratio of the audio object in the frame, and then perform audio zoom processing on the audio frame according to the zoom ratio. After performing audio zoom processing on each frame of audio in the audio stream, the electronic device re-encodes it to obtain a new audio stream.
[0140] Assume that the electronic device does not perform audio zoom on the video when recording it, and only saves the zoom ratio of the electronic device when capturing each frame of audio.
[0141] The process of recording video by electronic devices can refer to the above Figure 1a-Figure 1e Description.
[0142] The exemplary user interface involved in scenario 3 can refer to the following Figure 6a-6f Description.
[0143] Figure 6a As shown, the user interface 90 is a user interface when the electronic device has finished recording a video. The user interface 90 includes an echo control 901. The echo control 901 can be used to display the most recently recorded video or captured image of the electronic device. In response to the user's operation (such as a click operation) on the echo control 901, the electronic device can display the following information: Figure 6b The user interface shown.
[0144] like Figure 6b As shown, the user interface 91 can be a user interface for the electronic device to set the video. The user interface 91 includes a more control 911, and the more control 911 can be used to display more setting items for the video. In response to the user's operation on the more control 911 (such as a click operation), the electronic device can display the following Figure 6c The user interface shown.
[0145] like Figure 6c As shown, the user interface 92 may display setting items for the video. This includes a zoom mode setting item 921. The zoom mode setting item may be used to receive an instruction to perform audio zoom on the video. In response to a user operation (e.g., a click operation) on the zoom mode setting item 921, the electronic device may perform audio zoom on the audio in the video, displaying the video as shown in FIG. Figure 6d The user interface shown.
[0146] like Figure 6d As shown, the user interface 93 is an exemplary user interface in the process of the electronic device performing audio zoom on the audio in the video. The user interface 93 may display a prompt box 931, and the prompt box 931 may display a prompt text: "Performing audio zoom on the file 'Video 1', please wait". The prompt text can be used to remind the user that the current electronic device is performing audio zoom on the video. After the electronic device completes the audio zoom on the video, it may display the following Figure 6e The user interface shown.
[0147] like Figure 6e As shown, the user interface 94 is a user interface after the electronic device completes audio zoom. The user interface 94 may include a prompt box 941, and the prompt box 941 may include prompt text: "The audio zoom of the file 'Video 1' has been completed, do you want to replace the original file?" In response to the user's operation on the control 941 (such as a click operation), the electronic device can store the video after audio zoom to replace the original video without audio zoom. The electronic device can display Figure 6f The user interface shown.
[0148] like Figure 6f As shown, user interface 95 is a user interface after the video after audio zoom is replaced with the video without audio zoom. User interface 95 may include prompt text 951: "Replacement successful". Prompt text 951 is used to prompt the user that the electronic device has successfully replaced the video without audio zoom with the video after audio zoom. User interface 95 may also include a play control 952. In response to the user's operation on the play control 952 (e.g., a click operation), the electronic device can play the video after audio zoom.
[0149] The method involved in the video processing in scenario 3 can be referred to the following description of steps S601 to S609, which will not be repeated here.
[0150] The video processing method involved in this application is applicable to an electronic device with N microphones, where N is an integer greater than or equal to 2. Taking an electronic device with three microphones as an example, the processes of the video processing method involved in the above three scenarios are described in detail.
[0151] Scenario 1: A set of exemplary user interfaces when using the audio processing method of the present application in Scenario 1 can refer to the aforementioned Figure 1a-Figure 1e Description of user interface 10-user interface 14 in . For the real-time video processing method involved in scenario 1, from the start of video recording, the electronic device can process the collected current frame image in real time, and at the same time perform audio zoom on the collected current frame input audio signal set in real time according to the change of zoom magnification. Assume that from the start of video recording to the completion of video recording, there are a total of N frames of image and N frames of input audio signal set. Then the electronic device can generate an image stream based on the N frames of image and an audio stream based on the N frames of audio, and then mix the image stream and the audio stream to obtain a recorded video.
[0152] The current frame input audio signal set includes multiple frames of input audio signals. Any frame input audio signal refers to an audio signal collected by any microphone of the electronic device.
[0153] Figure 7 A schematic diagram showing a video obtained by real-time processing of a current frame image and a current frame input audio signal set in scenario 1.
[0154] Stop the following until the electronic device finishes video recording Figure 7 Involved processes. During the process of generating the image stream and audio stream, the electronic device processes the captured current frame image in the order of acquisition and stores it in the image stream cache. Simultaneously, the captured current frame input audio signal set is processed in the order of acquisition and stores it in the audio stream cache. The image stream in the image stream cache is then encoded and processed to generate the image stream. The image stream in the image stream cache is then encoded and processed to generate the image stream.
[0155] Among them, for the current frame image, the electronic device can use the processing technology in the existing technology to process it, and this application will not go into details.
[0156] For the current frame input audio signal set, the electronic device can process it using the audio zoom method involved in this application. The process will be described in the following Figure 8 Steps S101 to S107 are described in detail and will not be repeated here.
[0157] Specifically, the electronic device first begins recording a video, capturing the first image frame and the first input audio signal set. The first image frame is then processed and cached in area 1 of the image stream cache. Simultaneously, the first input audio signal set is processed and cached in area 1 of the audio stream cache. During playback, the electronic device can simultaneously play the processed first image frame and the processed first input audio signal set.
[0158] Then, after the electronic device has captured the first frame of image and the first frame of input audio signal set, it can continue to capture the second frame of image and the second frame of input audio signal set while processing them, and the processing process is similar to that of the first frame of image and the first frame of input audio signal set. The electronic device can cache the processed second frame of image in area 2 of the image stream cache and cache the processed second frame of input audio signal set in area 2 of the audio stream cache. During playback, the electronic device can also play the processed second frame of input audio signal set while playing the processed second frame of image.
[0159] Similarly, after the electronic device has collected the N-1th frame image and the N-1th frame input audio signal set, it can continue to collect the Nth frame image and the Nth frame input audio signal set during the processing process, and the processing process is similar to the first frame image and the first frame input audio signal set. The electronic device can cache the processed Nth frame image in area N of the image stream cache and cache the processed Nth frame input audio signal set in area N of the audio stream cache. During playback, the electronic device can also play the processed Nth frame input audio signal set while playing the processed Nth frame image.
[0160] In some embodiments, the time it takes for an electronic device to play one frame of image is 30ms, and the time it takes to play one frame of audio is 10ms. Figure 7 In the embodiment, when the electronic device plays a frame of image, a frame of input audio signal set includes 3 frames of audio.
[0161] The following describes in detail the process of processing any frame input audio signal set involved in the above scenario 1.
[0162] Taking an electronic device with three microphones as an example, Figure 7In the video processing process involved, any set of input audio signals collected by the electronic device includes three frames of input audio signals. The electronic device can convert the three frames of input audio signals into the frequency domain to obtain three frames of audio signals. The electronic device can then, based on the zoom factor corresponding to the set of input audio signals, maintain the target sound signal in the three frames of audio signals and suppress the non-target sound signals to generate a first output audio signal.
[0163] In some embodiments, if the zoom ratio increases, the first output audio signal is enhanced; if the zoom ratio decreases, the first output audio signal is suppressed; the enhanced or suppressed first output audio signal is output to the audio stream buffer, and the processing of the frame input audio signal set can be completed.
[0164] In other embodiments, the electronic device uses the enhanced or suppressed first output audio signal as the second output audio signal, suppresses the non-target sound signal in the second output audio signal, generates a third output audio signal, and then caches the third output audio signal in the audio stream cache area, thereby completing the processing of the frame input audio signal set.
[0165] The above process can refer to the following Figure 8 Detailed description of steps S101 to S107:
[0166] S101. The electronic device collects a first input audio signal, a second input audio signal, and a third input audio signal;
[0167] An exemplary user interface for an electronic device to collect a first input audio signal, a second input audio signal, and a third input audio signal can refer to the above Figure 1b-1d The user interface shown.
[0168] The first input audio signal, the second input audio signal and the third input audio signal collected by the electronic device are the above Figure 7 Any frame input audio signal set involved in .
[0169] The first input audio signal is a current frame audio signal converted from a sound signal collected by a first microphone of the electronic device during a first time period. The second input audio signal is a current frame audio signal converted from a sound signal collected by a second microphone of the electronic device during the first time period. The third input audio signal is a current frame audio signal converted from a sound signal collected by a third microphone of the electronic device during the first time period.
[0170] Take the electronic device collecting a first input audio signal as an example.
[0171] Specifically, during the first time period, the first microphone of the electronic device can collect a sound signal and then convert the sound signal into an analog electrical signal. The electronic device then samples the analog electrical signal and converts it into an audio signal in the time domain. The audio signal in the time domain is a digital audio signal, which is W sampling points of the analog electrical signal. The electronic device can use an array to represent the first input audio signal, where any element in the array is used to represent a sampling point, and any element includes two values, one of which represents time and the other represents the amplitude of the audio signal corresponding to the time, and the amplitude is used to represent the voltage corresponding to the audio signal.
[0172] It can be understood that the process of the electronic device collecting the second input audio signal and the third input audio signal can refer to the description of the first input audio signal, and will not be repeated here.
[0173] S102. The electronic device converts the first input audio signal, the second input audio signal, and the third audio signal into the frequency domain to obtain a first audio signal, a second audio signal, and a third audio signal;
[0174] The first input audio signal, the second input audio signal, and the third audio signal involved in step S101 are audio signals in the time domain. To facilitate processing, the electronic device may convert the first input audio signal into the frequency domain to obtain a first audio signal, convert the second input audio signal into the frequency domain to obtain a second audio signal, and convert the third input audio signal into the frequency domain to obtain a third audio signal.
[0175] Take the example of an electronic device converting a first input audio signal into a first audio signal in the frequency domain.
[0176] Specifically, the electronic device may divide the first input audio signal in the time domain into the frequency domain by using Fourier transform (FT), such as discrete Fourier transform (DFT).
[0177] In some embodiments, the electronic device may divide the first input audio signal into first audio signals corresponding to N frequency points using a 2N-point DFT. N is an integer power of 2, and the value of N is determined by the computing power of the electronic device. The greater the processing speed of the electronic device, the larger the value of N can be.
[0178] The following description is based on an example in which an electronic device divides a first input audio signal into first audio signals corresponding to 1024 frequency points using a 2048-point DFT. The electronic device can then represent the first audio signal using an array comprising 1024 elements. Each element represents a frequency point and includes two values: one representing the frequency (Hz) of the audio signal corresponding to the frequency point, and the other representing the amplitude of the audio signal corresponding to the frequency point. The amplitude is expressed in decibels (dB), indicating the decibel level of the audio signal corresponding to the time.
[0179] It should be understood that, in addition to an array, the electronic device may also express the first audio signal in other ways, such as a matrix, etc., and this embodiment of the present application does not limit this.
[0180] It can be understood that the electronic device converts the second input audio signal into the frequency domain to obtain the second audio signal, and converts the third input audio signal into the frequency domain to obtain the third audio signal. The process is the same as the above-mentioned method of converting the first input audio signal into the frequency domain to obtain the first audio signal, and will not be repeated here.
[0181] S103. The electronic device obtains a first zoom ratio;
[0182] An exemplary user interface for an electronic device to collect a first input audio signal, a second input audio signal, and a third input audio signal can refer to the above Figure 1b-1d The user interface shown.
[0183] The first zoom ratio refers to the zoom ratio used by the electronic device when capturing the current frame image. When the electronic device starts recording a video, the default zoom ratio used is 1x. The electronic device can change the zoom ratio used when capturing the current frame image according to the user's settings. One way to change the zoom ratio can refer to the above-mentioned Figure 1a-Figure 1e For the description of the zoom ratio, please refer to the description of term (1) above.
[0184] It should be understood that there is no particular order in which step S102 and step S103 are executed. The electronic device may execute step S102 first and then step S103, or may execute step S103 first and then step S102. They may also be executed simultaneously, which is not limited in the present embodiment.
[0185] S104. The electronic device generates a first output audio signal based on the first zoom ratio using the first audio signal, the second audio signal, and the third audio signal, wherein the target sound signal is retained and non-target sound signals are suppressed in the first output audio signal;
[0186] The target sound signal refers to an audio signal corresponding to a sound emitted by a subject within the field of view of the electronic device at the first zoom ratio. The non-target sound signal refers to an audio signal corresponding to a sound emitted by a subject outside the field of view.
[0187] In step S104, the electronic device may filter and synthesize the first audio signal, the second audio signal, and the third audio signal according to the first zoom ratio to generate a first output audio signal. The first output audio signal may include multiple audio signals, for example, two audio signals, namely a left channel audio signal and a right channel audio signal. The first output audio signal may also include one audio signal. The specific number of audio signals included in the first output audio signal may be determined by the number of speakers of the electronic device and does not limit the embodiments of the present application.
[0188] The purpose of filtering is to suppress non-target sound signals in the first audio signal, the second audio signal, and the third audio signal while keeping the target sound signal unchanged.
[0189] Optionally, the process may employ beamforming technology.
[0190] Specifically, to provide a stereo effect in the generated first audio signal, the electronic device may filter the first audio signal, the second audio signal, and the third audio signal in different directions using filter coefficients corresponding to different directions to obtain audio signals corresponding to the directions. The electronic device may then synthesize the audio signals corresponding to the directions to generate the first audio signal.
[0191] In some embodiments, the first audio signal includes a single-channel audio signal, ie, a mono audio signal.
[0192] In some other embodiments, the first audio signal includes two audio signals, namely a left-channel audio signal and a right-channel audio signal. Exemplarily, the electronic device includes two speakers.
[0193] Figures 9-11 A schematic diagram showing beamforming technology.
[0194] In order to compare the above Figure 2-Figure 4 The difference between the solution involved in the embodiment of the present application is as follows: Figure 2-Figure 4 The shooting scene in is taken as an example to illustrate the beamforming diagrams involved in the electronic device at a zoom ratio of 1x, 5x, and 10x.
[0195] Figure 9 (b)- Figure 11In (b), in icon 601, the sound signal is a solid line, which indicates that in the embodiment of the present application, the sound of the subject is not a suppressed object, but a target object, and the sound of the subject 103 can be heard when the video is played. In 602, a cross is drawn on the sound signal, which indicates that in the embodiment of the present application, the sound of the subject is a suppressed object, a non-target object, and the sound of the subject 101 cannot be heard when the video is played. In icon 603, the subject 101 is drawn with a dotted line, indicating that the subject 101 will not appear in the video screen and is a non-target object of the shooting. In icon 604, the subject 103 is drawn with a solid line, indicating that the subject 103 will appear in the video screen and is a target object of the shooting. Optionally, in the embodiment of the present application, the sound of the non-target object is a suppressed object. It can be understood that the target object is an object that appears in the video screen, and its sound does not need to be suppressed; correspondingly, the non-target object is an object that does not appear in the video screen, and its sound needs to be suppressed.
[0196] The above Figure 9 (b)- Figure 11 The description of each icon in (b) is in Figure 9 (c)- Figure 11 The same applies to (c).
[0197] It should be understood that Figure 9 (b)- Figure 11 In (b), icons with similar shapes have the same meaning and will not be explained one by one. For example, when the subject is drawn with a solid line, it means that the subject can appear in the video and is a non-target object of the video. Figure 9 In (b), the person being photographed 101 and the person being photographed 102 are both drawn with solid lines, which means that the person being photographed 101 and the person being photographed 102 can both appear in the video frame and are the target objects of the shooting.
[0198] like Figure 9 As shown in (a) of FIG, the user interface 50 is a user interface when the electronic device plays a video. At this time, the electronic device plays Figure 1b The image and audio corresponding to the first second of the recording are shown in FIG. 1. At this time, the zoom ratio of the electronic device is 1. As can be seen in the user interface 50, the currently played image includes the photographed person 101, the photographed person 102, and the photographed person 103.
[0199] In the video generated by the electronic device, if the audio corresponding to the first second is mono audio, you can Figure 9 The mono beamforming diagram shown in (b) generates the mono audio.
[0200] In the video generated by the electronic device, when the audio corresponding to the first second includes a left channel audio signal and a right channel audio signal, you can use Figure 9 The beamforming diagram of the left channel shown in (c) generates the left channel audio signal, using Figure 9 The beamforming pattern of the right channel shown in (c) generates the right channel audio signal.
[0201] Optionally, the beamforming diagram of the left channel and the beamforming diagram of the right channel are symmetrical.
[0202] In some embodiments, the electronic device may utilize Figure 9 The beamforming diagram of the left channel and the beamforming diagram of the right channel shown in (c) are fused to obtain Figure 9 The mono beamforming diagram shown in (b) is shown in FIG.
[0203] like Figure 9 As shown in (b) in the figure, it is a mono beamforming diagram when the zoom ratio is 1x. The symmetry line of the beamforming diagram is in the 0° direction. The electronic device can use the mono beamforming diagram to generate mono audio. It can be seen from the beamforming diagram that the gain coefficients corresponding to the directions of the photographed persons 101, 102 and 103 are all 1 (or close to 1). Then, the audio signal collected by the electronic device includes the voices of the photographed persons 101, 102 and 103, and will not be suppressed (it can be understood that since the gain coefficient is close to 1, the suppression effect is weak and it can be considered that no suppression is performed. The same applies to other embodiments). That is, during playback, the voices of the photographed persons 101, 102 and 103 in the mono audio are not suppressed, and the user can hear the voices of the photographed persons 101, 102 and 103.
[0204] like Figure 9 As shown in (c) of FIG. 1 , the beamforming pattern for the left channel and the beamforming pattern for the right channel at a zoom ratio of 1x are shown. The symmetry line of the beamforming pattern for the left channel is in the 45° direction, and the electronic device can use the beamforming pattern for the left channel to generate the left channel audio signal. The symmetry line of the beamforming pattern for the right channel is in the 315° direction, and the electronic device can use the beamforming pattern for the right channel to generate the right channel audio signal.
[0205] It is understandable that the angles of 45° and 315° in this embodiment are merely examples and can be adjusted to other angles as needed, and this application does not limit this. Figure 10 and Figure 11 The same applies to the angles in the embodiments.
[0206] At this stage, whether the electronic device suppresses the sounds of subjects 101, 102, and 103 should be the effect presented after the left channel audio signal and the right channel audio signal are output. For example, for subject 101, the gain coefficient in the beamforming diagram of the left channel and the gain coefficient in the beamforming diagram of the right channel are both 1 (or close to 1), so the electronic device does not suppress the sound of subject 101. For subject 102, the gain coefficient in the beamforming diagram of the left channel and the gain coefficient in the beamforming diagram of the right channel are both 1 (or close to 1), so the electronic device does not suppress the sound of subject 102. For subject 103, the gain coefficient in the beamforming diagram of the left channel and the gain coefficient in the beamforming diagram of the right channel are both 1 (or close to 1), so the electronic device does not suppress the sound of subject 103. In this way, when the electronic device plays the image and audio corresponding to the first second, the photographed persons 102, 103 and 101 can be seen, and the voices of the photographed persons 102, 103 and 101 can be heard.
[0207] like Figure 10 As shown in (a) of FIG, the user interface 60 is a user interface when the electronic device plays the video corresponding to the second second. At this time, the electronic device plays Figure 1c At this time, the zoom ratio of the electronic device has increased from 1x to 5x. In the user interface 60, it can be seen that the currently played image includes the subject 102 and the subject 103, but no longer includes the subject 101.
[0208] In the video generated by the electronic device, if the audio corresponding to the second second is mono audio, you can Figure 10 The mono beamforming diagram shown in (b) generates the mono audio.
[0209] In the video generated by the electronic device, when the audio corresponding to the second second includes a left channel audio signal and a right channel audio signal, you can use Figure 10 The beamforming diagram of the left channel shown in (c) generates the left channel audio signal, using Figure 10 The beamforming pattern of the right channel shown in (c) generates the right channel audio signal.
[0210] Optionally, the beamforming diagram of the left channel and the beamforming diagram of the right channel are symmetrical.
[0211] In some embodiments, the electronic device may utilize Figure 10The beamforming diagram of the left channel and the beamforming diagram of the right channel shown in (c) are fused to obtain Figure 10 The mono beamforming diagram shown in (b) is shown in FIG.
[0212] like Figure 10 As shown in (b) in the figure, it is a mono beamforming diagram when the zoom ratio is 5 times. The symmetry line of the beamforming diagram is in the 0° direction. The electronic device can use the mono beamforming diagram to generate mono audio. It can be seen from the beamforming diagram that the corresponding gain coefficients in the directions of the subjects 102 and 103 are both 1 (or close to 1), so the electronic device will not suppress the sounds of the subjects 102 and 103. However, the corresponding gain coefficients in the direction of the subject 101 are both 0 (or close to 0), so the electronic device can suppress the sound of the subject 101. The audio signal collected by the electronic device includes the sounds of the subjects 101, 102 and 103, but the sound of the subject 101 is suppressed in the played audio. From the auditory point of view, the sound of the subject 101 cannot be heard or the sound of the subject 101 sounds smaller.
[0213] like Figure 10 As shown in (c) of FIG. 1 , the beamforming pattern for the left channel and the beamforming pattern for the right channel at a zoom ratio of 5x are shown. The symmetry line of the beamforming pattern for the left channel is in the 45° direction, and the electronic device can use the beamforming pattern for the left channel to generate the left channel audio signal. The symmetry line of the beamforming pattern for the right channel is in the 315° direction, and the electronic device can use the beamforming pattern for the right channel to generate the right channel audio signal.
[0214] At this time, whether the electronic device suppresses the sounds of subjects 101, 102, and 103 should be the effect presented after the left channel audio signal and the right channel audio signal are output. For example, for subject 101, the gain coefficient in the beamforming diagram of the left channel and the gain coefficient in the beamforming diagram of the right channel are both 0 (or close to 0), so the electronic device can suppress the sound of subject 101. For subject 102, the gain coefficient in the beamforming diagram of the left channel and the gain coefficient in the beamforming diagram of the right channel are both 1 (or close to 1), so the electronic device will not suppress the sound of subject 102. For subject 103, although the gain coefficient in the beamforming diagram of the left channel is 1 (or close to 1) and the gain coefficient in the beamforming diagram of the right channel is both 1 (or close to 1), the electronic device will not suppress the sound of subject 103. In this way, when the electronic device plays the image and audio corresponding to the second second, the subjects 102 and 103 can be seen but the subject 101 cannot be seen, and the voices of the subjects 102 and 103 can be heard, but the voice of the subject 101 cannot be heard or sounds very small.
[0215] like Figure 11 As shown in (a) of FIG, the user interface 70 is a user interface when the electronic device plays the video corresponding to the 3rd second. At this time, the electronic device plays Figure 1d The image and audio corresponding to the 3rd second of the recording are displayed. At this time, the zoom ratio of the electronic device is increased from 5x to 10x. In the user interface 70, it can be seen that the currently played image includes the subject 102, but no longer includes the subjects 101 and 103.
[0216] In the video generated by the electronic device, if the audio corresponding to the 3rd second is mono audio, you can Figure 11 The mono beamforming diagram shown in (b) generates the mono audio.
[0217] In the video generated by the electronic device, when the audio corresponding to the 3rd second includes the left channel audio signal and the right channel audio signal, you can use Figure 11 The beamforming diagram of the left channel shown in (c) generates the left channel audio signal, using Figure 11 The beamforming pattern of the right channel shown in (c) generates the right channel audio signal.
[0218] Optionally, the beamforming diagram of the left channel and the beamforming diagram of the right channel are symmetrical.
[0219] In some embodiments, the electronic device may utilize Figure 11The beamforming diagram of the left channel and the beamforming diagram of the right channel shown in (c) are fused to obtain Figure 11 The mono beamforming diagram shown in (b) is shown in FIG.
[0220] like Figure 11 As shown in (b), it is a mono beamforming diagram when the zoom ratio is 10 times. The symmetry line of the beamforming diagram is in the 0° direction. The electronic device can use the mono beamforming diagram to generate mono audio. It can be seen from the beamforming diagram that the corresponding gain coefficients in the directions of the subjects 101 and 103 are both 0 (or close to 0), and the electronic device will suppress the sounds of the subjects 101 and 103. However, the corresponding gain coefficients in the direction of the subject 102 are both 1 (or close to 1), and the electronic device will not suppress the sound of the subject 102. The audio signal collected by the electronic device includes the sounds of the subjects 101, 102 and 103, but in the played audio, the sounds of the subjects 101 and 103 are suppressed. From the auditory point of view, the sound of the first subject is inaudible or sounds smaller.
[0221] like Figure 11 As shown in (c) of FIG. 1 , the beamforming pattern for the left channel and the beamforming pattern for the right channel at a 10x zoom ratio are shown. The symmetry line of the beamforming pattern for the left channel is in the 10° direction, and the electronic device can use the beamforming pattern for the left channel to generate the left channel audio signal. The symmetry line of the beamforming pattern for the right channel is in the 350° direction, and the electronic device can use the beamforming pattern for the right channel to generate the right channel audio signal.
[0222] At this time, whether the electronic device suppresses the sounds of subjects 101, 102, and 103 should be the effect presented after the left channel audio signal and the right channel audio signal are output. For example, for subject 101, the gain coefficient in the beamforming diagram of the left channel and the gain coefficient in the beamforming diagram of the right channel are both 0 (or close to 0), so the electronic device can suppress the sound of subject 101. For subject 102, the gain coefficient in the beamforming diagram of the left channel and the gain coefficient in the beamforming diagram of the right channel are both 1 (or close to 1), so the electronic device will not suppress the sound of subject 102. For subject 103, although the gain coefficient in the beamforming diagram of the left channel and the gain coefficient in the beamforming diagram of the right channel are 0 (or close to 0), the electronic device will suppress the sound of subject 103. In this way, when the electronic device plays the image and audio corresponding to the 3rd second, the subject 102 can be seen but the subjects 101 and 103 cannot be seen, and the subject 102 can be heard but the voice of the subject 101 and the voice of the subject 103 cannot be heard or the voice of the subject 101 and the voice of the subject 103 sound smaller.
[0223] The first output audio signal includes two audio signals: a left-channel audio signal and a right-channel audio signal. The electronic device generates the left-channel audio signal and the right-channel audio signal using the first audio signal, the second audio signal, and the third audio signal according to the first zoom factor, as described below in steps S201 to S203.
[0224] Figure 12 An exemplary flow chart for generating a left channel audio signal or a right channel audio signal for an electronic device.
[0225] S201. The electronic device obtains a first filter coefficient corresponding to a first direction, a second filter coefficient corresponding to a second direction, and a third filter coefficient corresponding to a third direction;
[0226] For detailed descriptions of the first direction, the second direction, and the third direction, please refer to the following:
[0227] The direction facing the rear camera is regarded as the front of the electronic device. The first direction is any direction within the range of 10° clockwise from the front of the electronic device to 70° clockwise from the front of the electronic device, for example, 45° clockwise from the front of the electronic device. The second direction can be any direction within the range of 10° counterclockwise from the front of the electronic device to 10° counterclockwise from the front of the electronic device, for example, the front of the electronic device. The third direction is any direction within the range of 10° counterclockwise from the front of the electronic device to 70° counterclockwise from the front of the electronic device, for example, 45° counterclockwise from the front of the electronic device.
[0228] The electronic device can adjust the angles of the first direction, the second direction, and the third direction as needed.
[0229] For example, at 1x zoom ratio, it can be adjusted to the above Figure 9 In the beamforming diagram (c), the second direction is directly in front of the electronic device, i.e., the 0° direction in the diagram. The first direction is 45° clockwise from the positive direction, i.e., the 45° direction in the beamforming diagram for the left channel in the diagram. The first direction is 45° counterclockwise from the positive direction, i.e., the 315° direction in the beamforming diagram for the right channel in the diagram.
[0230] Understandably, Figure 9 The angles of 45° and 315° in (c) are only examples and can be adjusted to other angles as needed, and this application does not limit this. Figure 10 and Figure 11 The same applies to the angles in the embodiments.
[0231] For a description of the above angles, see Figure 13 , Figure 13 This is an example diagram of the first direction, the second direction, and the third direction.
[0232] Combine Figure 13 A front view of the electronic device shown in (a) and Figure 13 The top view of the electronic device shown in (b) can be seen. Figure 13 As shown in (c), the direction in front of the electronic device ranges from 90° clockwise to 270°. The direction directly in front of the electronic device is 0°, 45° clockwise from the direction directly in front of the electronic device is 45° in the figure, and 45° counterclockwise from the direction directly in front of the electronic device is 315° in the figure. Therefore, we can define the first direction as 45°, the second direction as 0°, and the third direction as 315°.
[0233] For example, as mentioned above Figure 9 As shown in (c), the first direction may be 45°, the second direction may be 0°, and the third direction may be 315°. Figure 10As shown in (c), the first direction may be 30°, the second direction may be 0°, and the third direction may be 330°. Figure 11 As shown in (c), the first direction can be 10°, the second direction can be 0°, and the third direction can be 350°.
[0234] It is understandable that the angles mentioned above are only examples and can be adjusted to other angles as needed, and this application does not limit this.
[0235] To achieve stereo sound for the left and right channel audio signals, the retained and suppressed audio signals in the left and right channel audio signals must be different. That is, in the left channel audio signal, the audio signal collected from the left side of the front of the electronic device is retained, while the audio signal collected from the right side is suppressed. In the right channel audio signal, the audio signal collected from the right side of the front of the electronic device is retained, while the audio signal collected from the left side is suppressed.
[0236] In the embodiment of the present application, deviating to the left means deviating to the first direction, directly in front means deviating to the second direction, and deviating to the right means deviating to the third direction.
[0237] The first direction is to the left of the front of the electronic device, the third direction is to the right of the front of the electronic device, and the second direction is relatively in front of the electronic device.
[0238] The first filter coefficient corresponding to the first direction and the second filter coefficient corresponding to the second direction are used to generate a left-channel audio signal. This can retain audio signals collected from the left side of the front of the electronic device and suppress audio signals collected from the right side. The first filter coefficient corresponding to the second direction and the third filter coefficient corresponding to the third direction are used to generate a right-channel audio signal. This can retain audio signals collected from the right side of the front of the electronic device and suppress audio signals collected from the left side.
[0239] The first filter coefficient corresponding to the first direction, the second filter coefficient corresponding to the second direction, and the third filter coefficient corresponding to the third direction are pre-configured in the electronic device before the electronic device leaves the factory.
[0240] Taking the electronic device generating the first filter corresponding to the first direction as an example, this is described in detail. The process can refer to the following Figure 14 Description of steps S301-S303 in .
[0241] The first filter is described in detail as follows:
[0242] The first filter coefficients corresponding to the first direction include the first filter coefficients corresponding to the first microphone in the first direction, the first filter coefficients corresponding to the second microphone in the first direction, and the first filter coefficients corresponding to the third microphone in the first direction. The first filter coefficients corresponding to the first microphone in the first direction can be used to retain the audio signal collected in the first audio signal from the left direction relative to the front of the electronic device, and suppress the audio signal collected from the front and right directions. The first filter coefficients corresponding to the second microphone in the first direction can be used to retain the audio signal collected in the second audio signal from the left direction relative to the front of the electronic device, and suppress the audio signal collected from the front and right directions. The third filter coefficients corresponding to the first microphone in the third direction can be used to retain the audio signal collected in the third audio signal from the left direction relative to the front of the electronic device, and suppress the audio signal collected from the front and right directions. For details of this process, please refer to the description of step S202 below.
[0243] If the first audio signal includes N frequency points, the first filter coefficient corresponding to the first microphone in the first direction should also have N elements (coefficients), where the jth element represents the degree of suppression of the jth frequency point among the N frequency points corresponding to the first audio signal.
[0244] Specifically, when the jth element is equal to or close to 1, the electronic device does not suppress the audio signal corresponding to the jth frequency point (when it is close to 1, the degree of suppression is very low and almost no suppression is performed, which is considered to be retained), that is, it is retained, and the direction of the audio signal corresponding to the jth frequency point is considered to be left-leaning. In other cases, the audio signal corresponding to the jth frequency point is suppressed. For example, when the jth element is equal to or close to 0, the greater the degree of suppression of the audio signal corresponding to the jth frequency point by the electronic device, that is, suppression, the more right-leaning the direction of the audio signal corresponding to the jth frequency point is considered to be.
[0245] Figure 14 An exemplary flowchart of generating a first filter corresponding to the first direction for an electronic device.
[0246] In some embodiments, the specific process of the electronic device generating the first filter coefficient corresponding to the first direction can refer to the following description of steps S301 to S303:
[0247] S301. The electronic device obtains a first test audio signal, a second test audio signal, and a third test audio signal at different distances in multiple directions.
[0248] Direction refers to the horizontal angle between the sound source and the electronic device, and distance refers to the Euclidean distance between the sound source and the electronic device. The sound source is single.
[0249] The purpose of acquiring audio signals at different distances in multiple directions is to ensure that the generated first filter coefficients are universal. That is, when the electronic device is manufactured and recording video, if the direction of the first, second, and third input audio signals is the same as or similar to one of the multiple directions, then the first filter coefficients are applicable to the first, second, and third input audio signals.
[0250] In some embodiments, the multiple directions may include 36 directions, and around the electronic device, each direction is every 10 degrees. The multiple distances may include 3 distances of 1m, 2m, and 3m.
[0251] The first test audio signal is a collection of input audio signals at different distances collected by a first microphone of the electronic device in multiple directions.
[0252] The second test audio signal is a collection of input audio signals at different distances collected by a second microphone of the electronic device in multiple directions.
[0253] The third test audio signal is a collection of input audio signals at different distances collected by the third microphone of the electronic device in multiple directions.
[0254] S302. The electronic device obtains a first target beam corresponding to a first direction.
[0255] The first target beam is used by the electronic device to generate a first filter coefficient corresponding to a first direction, which describes the degree of filtering of the electronic device in multiple directions. The first target beam is an expected beam or a beam that is expected to be formed and can be set.
[0256] In some embodiments, when the number of directions is 36, the first target beam has 36 gain coefficients. The i-th gain coefficient represents the degree of suppression in the i-th direction, and each direction corresponds to a gain coefficient. The corresponding gain coefficient in the first direction is 1. Then, for each direction that differs by 10° from the first direction, the gain coefficient is reduced by 1 / 36. Thus, the elements corresponding to directions closer to the first direction are closer to 1, and the elements corresponding to directions further from the first direction are closer to 0.
[0257] S303. The electronic device generates corresponding first filter coefficients in a first direction using the first test audio, the second test audio, the third test audio, and the first target beam through a device-dependent transfer function.
[0258] The electronic device generates the first filter coefficient corresponding to the first direction using the following formula (1):
[0259]
[0260] In formula (1), w1(ω) is the first filter coefficient, which includes 3 elements, where the i-th element can be expressed as w 1i (ω), w 1i (ω) is the first filter coefficient corresponding to the i-th microphone in the first direction, H1(ω) represents the first test audio signal, H2(ω) represents the second test audio signal, and H3(ω) represents the third test audio signal. G(H1(ω), H2(ω), H3(ω)) represents the processing of the first test audio signal, the second test audio signal, and the third test audio signal through a device-dependent transfer function, which can be used to describe the correlation between the first test audio signal, the second test audio signal, and the third test audio signal. H1 represents the first target beam, w1 represents the filter coefficient that can be obtained in the first direction, and argmin represents w1 obtained using the least squares frequency-invariant fixed beamforming method as the first filter coefficient corresponding to the first direction.
[0261] The second filter coefficients corresponding to the second direction include the second filter coefficients corresponding to the first microphone in the second direction, the second filter coefficients corresponding to the second microphone in the second direction, and the second filter coefficients corresponding to the third microphone in the second direction. The second filter coefficients corresponding to the first microphone in the second direction can be used to retain the audio signal collected relative to the front of the electronic device in the first audio signal, and suppress the audio signal collected from the left and right directions. The second filter coefficients corresponding to the second microphone in the second direction can be used to retain the audio signal collected relative to the front of the electronic device in the second audio signal, and suppress the audio signal collected from the left and right directions. The third filter coefficients corresponding to the first microphone in the third direction can be used to retain the audio signal collected relative to the front of the electronic device in the third audio signal, and suppress the audio signal collected from the left and right directions.
[0262] For a detailed description of the second filter, reference may be made to the detailed description of the first filter, which will not be repeated here.
[0263] The electronic device generates the second filter coefficient corresponding to the second direction using the following formula (2):
[0264]
[0265] The description of formula (2) can refer to the description of steps S401 and S402 above. The difference is that w2(ω) is the second filter coefficient, which includes 3 elements, where the i-th element can be expressed as w 2i (ω), w2i (ω) is the second filter coefficient corresponding to the i-th microphone in the second direction, H2 represents the second target beam corresponding to the second direction, w2 represents the filter coefficient that can be obtained in the second direction, and argmin represents w2 obtained by the least squares frequency-invariant fixed beamforming method as the second filter coefficient corresponding to the second direction.
[0266] The second target beam is used by the electronic device to generate a second filter corresponding to a second direction, which describes the filtering degree of the electronic device in multiple directions.
[0267] In some embodiments, when the number of directions is 36, the second target beam has 36 gain coefficients. The i-th gain coefficient represents the degree of filtering in the i-th direction, with each direction corresponding to a gain coefficient. The corresponding gain coefficient in the second direction is 1. Then, for each direction that differs by 10° from the second direction, the gain coefficient is reduced by 1 / 36. Thus, the elements corresponding to directions closer to the second direction are closer to 1, and the elements corresponding to directions further from the second direction are closer to 0.
[0268] The third filter coefficient corresponding to the third direction includes the third filter coefficient corresponding to the third direction of the first microphone, the third filter coefficient corresponding to the third direction of the second microphone, and the third filter coefficient corresponding to the third direction of the third microphone. The third filter coefficient corresponding to the third direction of the first microphone can be used to retain the audio signal collected from the right direction relative to the front of the electronic device in the first audio signal, and suppress the audio signal collected from the front and left directions. The third filter coefficient corresponding to the third direction of the second microphone can be used to retain the audio signal collected from the right direction relative to the front of the electronic device in the second audio signal, and suppress the audio signal collected from the front and left directions. The third filter coefficient corresponding to the third direction of the first microphone can be used to retain the audio signal collected from the right direction relative to the front of the electronic device in the third audio signal, and suppress the audio signal collected from the front and left directions.
[0269] For a detailed description of the third filter, reference may be made to the detailed description of the first filter, which will not be repeated here.
[0270] The electronic device generates a third filter coefficient corresponding to the third direction using the following formula (3):
[0271]
[0272] The description of formula (3) can refer to the description of steps S401 and S402 above. The difference is that w3(ω) is the third filter coefficient, which includes 3 elements, where the i-th element can be expressed as w 3i (ω), w 3i (ω) is the third filter coefficient corresponding to the i-th microphone in the third direction, H3 represents the third target beam corresponding to the third direction, w3 represents the filter coefficient that can be obtained in the third direction, and argmin represents w3 obtained by the least squares frequency-invariant fixed beamforming method as the third filter coefficient corresponding to the third direction.
[0273] The third target beam is used by the electronic device to generate a third filter corresponding to a third direction, which describes the degree of filtering of the electronic device in multiple directions.
[0274] In some embodiments, when the number of directions is 36, the third target beam has 36 gain coefficients. The i-th gain coefficient represents the degree of filtering in the i-th direction, with each direction corresponding to a gain coefficient. The gain coefficient corresponding to the third direction is 1. Then, for each direction that differs by 10° from the third direction, the gain coefficient is reduced by 1 / 36. Thus, the elements corresponding to directions closer to the third direction are closer to 1, and the elements corresponding to directions further from the third direction are closer to 0.
[0275] S202. Generate a first beam corresponding to a first direction, a second beam corresponding to a second direction, and a third beam corresponding to a third direction by combining the first audio signal, the second audio signal, and the third audio signal using the first filter coefficient, the second filter coefficient, and the third filter coefficient, respectively;
[0276] The first beam corresponding to the first direction is the audio signal generated by the electronic device synthesizing the first, second, and third audio signals. During the synthesis process, the electronic device may retain the audio signal collected from the left direction relative to the front of the electronic device, and suppress the audio signals collected from the front and right directions.
[0277] The second beam corresponding to the second direction is the audio signal generated by the electronic device synthesizing the first, second, and third audio signals. During the synthesis process, the electronic device may retain the audio signal collected directly in front of the electronic device, and suppress the audio signal collected to the left and right of the electronic device.
[0278] The third beam corresponding to the third direction is the audio signal generated by the electronic device synthesizing the first, second, and third audio signals. During the synthesis process, the electronic device may retain the audio signal collected from the first, second, and third audio signals slightly to the right of the front of the electronic device, and suppress the audio signals collected from the front and slightly to the left of the electronic device.
[0279] The electronic device uses the first filter coefficient, combines the first input audio signal, the second input audio signal, and the third audio input signal to generate a first beam corresponding to the first direction. The formula involved is as follows:
[0280]
[0281] y1 represents a first beam corresponding to a first direction, and includes N elements. Each element represents a frequency point. The number of frequency points corresponding to the first beam is the same as the number of frequency points corresponding to the first audio signal, the second audio signal, and the third audio signal.
[0282] Where w 1i (ω) is the first filter coefficient corresponding to the i-th microphone in the first direction, w 1i The jth element in (ω) represents the degree of suppression of the audio signal corresponding to the jth frequency point in the audio signal. i (ω) is the audio signal corresponding to the i-th microphone, x i The jth element in (ω) represents the complex domain of the jth frequency point, which represents the amplitude and phase information of the sound signal corresponding to the frequency point.
[0283] For example, the jth element in the first filter coefficient corresponding to the i-th microphone in the first direction is recorded as a ji , the jth element in the audio signal corresponding to the i-th microphone is recorded as b ij The above formula can be expressed as the following formula (5):
[0284]
[0285] Then the first beam corresponding to the first direction can be specifically expressed as the following formula (6):
[0286]
[0287] Formula (6) above shows the process by which the electronic device synthesizes the first audio signal, the second audio signal, and the third audio signal. Furthermore, during the synthesis process, the audio signal collected from the left direction relative to the front of the electronic device is retained, while the audio signals collected from the front and the right direction are suppressed.
[0288] Specifically, according to the above formulas (4)-(6), the electronic device cross-multiplies the complex domain corresponding to the N frequency points in the audio signals (including the first audio signal, the second audio signal, and the third audio signal) collected by the M microphones with the N elements in the first filter coefficients corresponding to the M microphones in the first direction, obtains M cross-multiplication results, and then adds the M cross-multiplication results point-to-point to finally obtain the first beam. The first beam includes N new frequency points. The cross-multiplication process is to retain the audio signal collected in the left direction in front of the electronic device and to suppress the audio signal collected in the front and right directions. The point-to-point addition process is the synthesis process.
[0289] It should be understood that when the jth element in the first filter coefficients corresponding to the M microphones in the first direction is equal to or close to 1, the electronic device does not suppress the audio signal corresponding to the frequency point multiplied by the jth element, that is, retains the audio signal, and the direction of the audio signal corresponding to the jth frequency point is considered to be close to the first direction. In other cases, the audio signal corresponding to the frequency point multiplied by the jth element is suppressed. For example, when the jth element is equal to or close to 0, the greater the degree to which the electronic device suppresses the audio signal corresponding to the jth frequency point, the more the direction of the audio signal corresponding to the jth frequency point is considered to be away from the first direction.
[0290] The electronic device uses the second filter coefficient to combine the first input audio signal, the second input audio signal, and the third audio input signal to generate a second beam corresponding to the second direction. The formulas involved are as follows: Formula (7) to Formula (9). The description of Formula (7) to Formula (9) can refer to the description of Formula (4) to Formula (6) above:
[0291]
[0292] y2 represents a second beam corresponding to the second direction, which includes N elements. Each element is used to represent a frequency point. The number of frequency points corresponding to the second beam is the same as the number of frequency points corresponding to the first audio signal, the second audio signal, and the third audio signal.
[0293] Where w 2i (ω) is the second filter coefficient corresponding to the i-th microphone in the second direction, w 2i The j-th element in (ω) represents the degree of suppression of the audio signal corresponding to the j-th frequency point in the audio signal.
[0294] For example, the jth element in the second filter coefficient corresponding to the i-th microphone in the second direction is recorded as c ji , the jth element in the audio signal corresponding to the i-th microphone is recorded as bij The above formula can be expressed as the following formula (8):
[0295]
[0296] Then the second beam corresponding to the second direction can be specifically expressed as the following formula (9):
[0297]
[0298] Formula (9) above shows the process by which the electronic device synthesizes the first audio signal, the second audio signal, and the third audio signal. Furthermore, during the synthesis process, the audio signal collected from the front of the electronic device is retained, while the audio signals collected from the left and right directions are suppressed. A detailed description of this process can be found in the description of formulas (4)-(6) above and will not be repeated here.
[0299] The electronic device uses the third filter coefficient, combines the first input audio signal, the third input audio signal, and the third audio input signal, and generates a third beam corresponding to the third direction. Formulas involved are as follows: (10) to (12). The description of formulas (10) to (12) can refer to the description of formulas (4) to (6) above:
[0300]
[0301] y3 represents a third beam corresponding to the third direction, which includes N elements. Each element is used to represent a frequency point. The number of frequency points corresponding to the third beam is the same as the number of frequency points corresponding to the first audio signal, the second audio signal, and the third audio signal.
[0302] Where w 3i (ω) is the third filter coefficient corresponding to the i-th microphone in the third direction, w 3i The jth element in (ω) represents the degree of suppression of the audio signal corresponding to the jth frequency point in the audio signal.
[0303] For example, the jth element in the third filter coefficient corresponding to the i-th microphone in the third direction is recorded as d ji , the jth element in the audio signal corresponding to the i-th microphone is recorded as b ij The above formula can be expressed as the following formula (11):
[0304]
[0305] Then the third beam corresponding to the third direction can be specifically expressed as the following formula (12):
[0306]
[0307] Formula (12) above shows the process by which the electronic device synthesizes the first audio signal, the second audio signal, and the third audio signal. Furthermore, during the synthesis process, the audio signal collected from the first audio signal, the second audio signal, and the third audio signal, which are collected from the right direction relative to the front of the electronic device, are retained, while the audio signals collected from the front direction and the left direction are suppressed. A detailed description of this process can be found in the description of formulas (4)-(6) above and will not be repeated here.
[0308] S203. According to the first zoom ratio, using the first beam and the second beam, generate a left channel audio signal and using the second beam and the third beam, generate a right channel audio signal
[0309] The left-channel audio signal is retained for audio signals collected from the front and slightly to the left of the front of the electronic device, while audio signals collected from the right are suppressed. This is an audio signal in the frequency domain, including N frequency points (the same number of frequency points as the frequency points corresponding to the first audio signal, the second audio signal, and the third audio signal).
[0310] The right-channel audio signal is retained for audio signals collected directly in front of the electronic device and slightly to the right of the front, while audio signals collected slightly to the left of the electronic device are suppressed. This is an audio signal in the frequency domain, including N frequency points (the same number of frequency points as the frequency points corresponding to the first, second, and third audio signals).
[0311] Specifically, the electronic device may fuse the first beam and the second beam into a left channel audio signal according to the fusion coefficient. The formula involved in this process may refer to the following formula (13):
[0312] y ; =αy1+(1-α)y2 formula (13)
[0313] In the above formula (13), y l represents the left channel audio signal, y1 represents the first beam, y2 represents the second beam, and α is the fusion coefficient, the specific value of which is pre-set in the electronic device. The value of the fusion coefficient is directly related to the first zoom factor. Any first zoom factor uniquely corresponds to a fusion coefficient. The electronic device can determine the fusion coefficient based on the first zoom factor. The value of the fusion coefficient is [0, 1]. The larger the zoom factor, the smaller the fusion coefficient. For example, when the zoom factor is 1, the fusion coefficient can be 1. When the zoom factor is the maximum zoom factor, the fusion coefficient can be 0.
[0314] The electronic device fuses the first beam and the second beam into a left channel audio signal according to the fusion coefficient. The formula involved in this process can refer to the following formula (14):
[0315] y r =αy3+(1-α)y2 Formula (14)
[0316] In the above formula (14), y r represents the right channel audio signal. For the specific description of formula (14), please refer to the above description of formula (13), which will not be repeated here.
[0317] Among them, the fusion coefficient is used to determine whether the left channel audio signal is more to the left or more to the front and the right channel audio signal is more to the right or more to the front. The value of the fusion coefficient is directly related to the zoom ratio. The principle is that when the zoom ratio is larger, the field of view is smaller (the degree to which the left side of the field of view is less to the left relative to the front, and the degree to which the right side of the field of view is less to the right relative to the front), then the left channel audio signal and the right channel audio signal should be more concentrated in the front, that is, the second direction. Combining formula (13) and formula (14), it can be seen that α should be smaller at this time. It should be understood that the smaller the zoom ratio, the larger the field of view, and the larger α. In this way, the left channel audio signal can retain more audio signals to the left of the front (that is, the first direction), and the right channel audio signal can retain more audio signals to the right of the front (that is, the second direction).
[0318] It is understandable that different zoom ratios correspond to different fusion coefficients α, which results in different beamforming diagrams when fusion occurs. Figure 9 (b) shows the mono beamforming diagram at 1x zoom factor, Figure 10 (b) shows the mono beamforming diagram at a 5x zoom factor, and Figure 11 The shape of the mono beamforming pattern at 10x zoom ratio shown in (b) is different. Figure 9 (c) shows the beamforming diagram of the left channel and the beamforming diagram of the right channel at a 1x zoom factor, Figure 10 (c) shows the beamforming diagram of the left channel and the beamforming diagram of the right channel at a 5x zoom factor, and Figure 11 The beamforming diagram of the left channel and the beamforming diagram of the right channel at a 10x zoom factor shown in (c) in FIG. 1 are different in shape.
[0319] The above step S203 is optional.
[0320] In some other embodiments, the first output audio signal includes only one audio signal.
[0321] Then the electronic device generates the first output audio signal using the following formula:
[0322]
[0323] The parameters in formula (15) can refer to the description of formula (13) and formula (14) above, and will not be repeated here. Formula (15) indicates that the audio signal collected within the field of view of the electronic device is retained, and the audio signals collected in other directions are suppressed.
[0324] S105. The electronic device enhances or suppresses the first output audio signal according to the first zoom ratio to obtain a second output audio signal, wherein both the target sound signal and the non-target sound signal in the second output audio signal are enhanced or suppressed compared to the first output audio signal;
[0325] Enhancing the first output audio signal refers to adjusting the amplitude of the first output audio signal to increase it, so that the decibel level of the first output audio signal can be increased. Suppressing the first output audio signal refers to adjusting the amplitude of the first output audio signal to decrease it, so that the decibel level of the first output audio signal can be decreased.
[0326] When the first output audio signal includes a left-channel audio signal and a right-channel audio signal, the second output audio signal also includes a left-channel audio signal and a right-channel audio signal. The enhancement or suppression of the first output audio signal in step S105 is to enhance or suppress the left-channel audio signal and the right-channel audio signal in the first output audio signal, respectively. The process of enhancing or suppressing the left-channel audio signal and the right-channel audio signal in the first output audio signal can be referred to the description of steps S401 to S404 below.
[0327] If the first output audio signal includes only one audio signal, the second output audio signal also includes only one audio signal. Therefore, the enhancement or suppression of the first output audio signal in step S105 is equivalent to the enhancement or suppression of the audio signal. This process can also be referred to the description of steps S401 to S404 below.
[0328] When the first zoom ratio increases, the closer the subject is during imaging, the louder the sound should be, and the electronic device can enhance the first output audio signal. When the first zoom ratio decreases, the farther the subject is during imaging, the quieter the sound should be, and the electronic device can suppress the first output audio signal.
[0329] Figure 15An exemplary flow chart of adjusting the amplitude of a first output audio signal for an electronic device:
[0330] The process of the electronic device adjusting the amplitude of the first output audio signal according to the first zoom ratio may refer to the following steps S401 to S404.
[0331] S401. The electronic device determines an adjustment parameter corresponding to the first zoom ratio according to the first zoom ratio, the adjustment parameter being used to adjust the amplitude of the audio signal, the adjustment including one of enhancement or suppression;
[0332] The specific value of the adjustment parameter is pre-set in the electronic device. The value of the adjustment parameter is directly related to the zoom factor. Each first zoom factor uniquely corresponds to an adjustment parameter. The electronic device can then determine the adjustment parameter corresponding to the first zoom factor based on the first zoom factor. The adjustment parameter is a numerical value whose unit is the same as the amplitude, namely dB.
[0333] The adjustment parameter is used to adjust the amplitude of the audio signal, and the adjustment includes one of increasing or suppressing.
[0334] Specifically, when the first zoom ratio is greater than 1x, the larger the first zoom ratio, the larger the adjustment parameter and the positive number, in which case the electronic device uses the adjustment coefficient to enhance the first audio signal, and the larger the adjustment parameter, the greater the degree of enhancement. When the first zoom ratio is less than 1x, the smaller the first zoom ratio, the smaller the adjustment parameter and the negative number, in which case the electronic device uses the adjustment coefficient to suppress the first audio signal, and the smaller the adjustment parameter, the greater the degree of suppression.
[0335] S402. Convert the first output audio signal from the frequency domain to the time domain to obtain a first output audio signal in the time domain;
[0336] The electronic device may adjust the amplitude of the first audio signal in the time domain, and may use an inverse Fourier transform (IFT) to convert the first output audio signal from the frequency domain to the time domain to obtain the first output audio signal in the time domain.
[0337] The first output audio signal in the time domain is a digital audio signal, which may be W sampling points of an analog electrical signal. In an electronic device, the first input audio signal may be represented by an array, where each element in the array represents a sampling point. Each element includes two values: one representing a time and the other representing the amplitude of the audio signal corresponding to that time. The amplitude is expressed in decibels (dB), indicating the decibel level of the audio signal corresponding to that time.
[0338] S403. Using the adjustment parameter to adjust the amplitude of the first output audio signal in the time domain to obtain a second output audio signal in the time domain;
[0339] In some embodiments, the electronic device may adjust the amplitude of the first output audio signal in the time domain by using methods such as automatic gain control (AGC) or dynamic range compression (DRC).
[0340] Taking the electronic device using a dynamic range control algorithm to adjust the amplitude of the first output audio signal in the time domain as an example, a detailed description is given:
[0341] In one possible case, when the first zoom factor is greater than 1, the electronic device may increase the amplitudes of all sampling points in the first output audio signal in the time domain. The electronic device uses the adjustment parameter to increase the amplitude of the first output audio signal in the time domain according to the following formula:
[0342] A ′ i =A i +|D|i∈(1,M) Formula (16)
[0343] When the first zoom ratio is less than 1, the electronic device uses the adjustment parameter to adjust the amplitude of the first output audio signal in the time domain using the adjustment parameter according to the following formula:
[0344] A i i =A i -|D|i∈(1,M) Formula (17)
[0345] Formula (16)-Formula (17), A i Represents the amplitude of the i-th sampling point, A i i represents the adjusted amplitude. D is the adjustment parameter. M is the total number of sampling points.
[0346] S404: Convert the second output audio signal in the time domain into the frequency domain as the second output audio signal.
[0347] The process of step S404 is similar to the process of the above-mentioned step S102. Please refer to the above-mentioned description of step S102 and will not be repeated here.
[0348] The second audio signal is an audio signal in the frequency domain.
[0349] When the second audio signal includes a left-channel audio signal and a right-channel audio signal, the left-channel audio signal and the right-channel audio signal correspond to N frequency points respectively, and the number of corresponding frequency points is the same as that of the first audio signal, the second audio signal, and the third audio signal.
[0350] When the second audio signal includes only one audio signal, the second output audio signal corresponds to N frequency points, and the number of corresponding frequency points is the same as that of the first audio signal, the second audio signal, and the third audio signal.
[0351] S106. The electronic device suppresses the non-target sound signal in the second output audio signal again according to the first zoom ratio, in combination with the first audio signal, the second audio signal, and the third audio signal, to generate a third output audio signal.
[0352] If the second output audio signal includes a left-channel audio signal and a right-channel audio signal, the third output audio signal also includes a left-channel audio signal and a right-channel audio signal. The suppression of non-target sound signals in the second output audio signal in step S106 is equivalent to suppressing non-target sound signals in the left-channel audio signal and the right-channel audio signal in the second output audio signal. The process of suppressing non-target sound signals in the left-channel audio signal and the right-channel audio signal in the second output audio signal can be referred to the description of steps S501 to S504 below.
[0353] If the second output audio signal includes only one audio signal, the third output audio signal also includes only one audio signal. The process of suppressing non-target sound signals in the second output audio signal in step S106 can also be referred to the following description of steps S501 to S504.
[0354] When the first zoom ratio is greater than the default zoom ratio, in step S105, the electronic device may enhance the first output audio signal to obtain a second output audio signal. During this process, the non-target sound signal in the first output audio signal will also be enhanced, resulting in the second output audio signal including the non-target sound signal. In order to prevent the audio signal played by the electronic device from being affected by the non-target sound signal, the electronic device may further suppress the non-target sound signal.
[0355] In some embodiments, the electronic device may filter the first output audio signal using a filtering method in the prior art to further suppress non-target sound signals. Common filtering methods include spectral subtraction and Wiener filtering.
[0356] In other embodiments, because a single filtering algorithm cannot fully filter both high-frequency and low-frequency audio signals in an audio signal, the electronic device may use a first filtering algorithm that performs better at high frequencies to calculate a first target gain, and then use this first target gain to filter the high-frequency audio signals in the second output audio signal. Simultaneously, a second filtering algorithm that performs better at low frequencies may be used to calculate a second target gain, and then use this second target gain to filter the low-frequency audio signals in the second output audio signal.
[0357] The first filtering algorithm may include a Zelinski filter algorithm, and the second filtering algorithm may include a coherent-to-diffuse power ratio (CDR) algorithm.
[0358] Figure 16 The present invention is a schematic flowchart of an electronic device suppressing a non-target sound signal in a second output audio signal.
[0359] S501. Performing Zelinsky filtering on the first audio signal, the second audio signal, and the third audio signal to obtain a first target gain;
[0360] The first target gain is calculated by the Zelinsky filtering algorithm. For a specific description, please refer to the following formulas (18) and (19), which are used to filter out the non-target sound signals corresponding to the high-frequency frequency points in the second output audio signal.
[0361] The first target gain is N target gain values (N is the same as the number of frequency points in the second output audio signal). If the value of any gain is not greater than the first threshold, it is equal to the first threshold, where the first threshold is pre-set in the electronic device. The first threshold can be obtained by the difference (Δ) between the degree of suppression of the non-target sound signal in step S104 and the degree of enhancement of the non-target sound in step S105, and the corresponding relationship between the first threshold and the difference is Δ=20logx, where x is the first threshold. A target gain value is used to filter a frequency point in the second output audio signal. When the target gain value is greater than the first threshold, it will be close to 1 or equal to 1, indicating that the sound signal corresponding to the frequency point is retained. When the target gain value is equal to the first threshold, it means that the sound signal corresponding to the frequency point is suppressed but not removed.
[0362] Among them, the first threshold is pre-set in the electronic device, and any first zoom ratio uniquely corresponds to a first threshold, which can be obtained by the difference (Δ) between the degree of suppression of the non-target sound signal in step S104 and the degree of enhancement of the non-target sound in step S105. The corresponding relationship between the difference is Δ=20logx, where x is the first threshold. The difference (Δ) is determined as follows: the non-target sound signal can be suppressed in step S104, and the degree of suppression is directly related to the first zoom ratio. Any first zoom ratio uniquely corresponds to a degree of suppression, which is a suppression parameter, a numerical value, and its unit is the same as the unit of the amplitude, which is dB. The degree of enhancement of the non-target sound signal in S105 is the adjustment parameter mentioned above. For example, if the degree of suppression is m (dB) and the degree of enhancement is n (dB), then Δ=|mn|.
[0363] Specifically, the electronic device can first calculate a first gain using the first audio signal, the second audio signal, and the third audio signal. The first gain is N gain values (N is the same as the number of frequency points in the second output audio signal). If the value of any gain is not equal to 1 or close to 1, it is equal to 0 or close to 0. When the gain value is equal to 1 or close to 1, it indicates that the sound signal corresponding to the frequency point is retained. When the target gain value is equal to 0 or close to 0, it indicates that the sound signal corresponding to the frequency point is suppressed and removed. The process of the electronic device generating the first gain can refer to the following description of formula (18).
[0364] Then, the electronic device adjusts the N gains in the first gain to obtain a first target gain. This process can be referred to the following description of formula (19).
[0365] The electronic device generates the first target gain using the following formulas (18) and (19):
[0366]
[0367] In formula (18), gain(ω) represents the first gain, h1 represents the first threshold, R is the operator for taking the real part, N represents the number of microphones, i represents the i-th microphone, j represents the j-th microphone, and {i, j}∈N represents that two microphones among the N microphones perform calculations on each other. ij (ω) represents the cross power spectrum between the audio signal in the frequency domain corresponding to the i-th microphone (for example, the first audio signal corresponding to the first microphone) and the audio signal in the frequency domain corresponding to the j-th microphone at the frequency point, φ ii (ω) represents the autopower spectrum of the first audio signal at this frequency point, φ jj (ω) represents the autopower spectrum of the second output audio signal at this frequency point.
[0368] In formula (19), gain1(ω) represents the first target gain, gain i (ω) represents the i-th gain value in the first gain. The electronic device can use formula (19) to adjust the N gain values in the first gain so that the gain values greater than the first threshold remain unchanged and the gain values less than the first threshold are adjusted to the first gain.
[0369] S502. The electronic device uses the first audio signal, the second audio signal, and the third audio signal to obtain a second target gain based on a coherent spread power ratio algorithm;
[0370] The second target gain is calculated by the coherent diffusion power ratio algorithm. For a specific description, please refer to the following formulas (20) and (21), which are used to filter out the non-target sound signals corresponding to the low-frequency points in the second output audio signal.
[0371] The second target gain is N target gain values (N is the same as the number of frequency points in the second output audio signal). If the value of any gain is greater than the first threshold, it is equal to the first threshold. For a detailed description of the first threshold, please refer to the description of the related content above and will not be repeated here.
[0372] Specifically, the electronic device may first calculate a second gain using the first audio signal, the second audio signal, and the third audio signal. The second gain comprises N gain values (N being the same as the number of frequency points in the second output audio signal). If the value of any gain is not equal to 1 or close to 1, it is equal to 0 or close to 0. For a detailed description, please refer to the description of the related content above and will not be repeated here.
[0373] Then, the electronic device adjusts the N gains in the second gain to obtain a second target gain. This process can be referred to the following description of formula (21).
[0374] The electronic device generates the second target gain using the following formulas (20) and (21):
[0375]
[0376] In formula (20), gain′(ω) represents the second gain, R is the operator for taking the real part, The parameters in this expression can refer to the above description of formula (18). Wherein, π represents pi, f represents the frequency corresponding to the frequency point, d represents the distance between the i-th microphone and the j-th microphone, and c represents the speed of sound propagation.
[0377] The relevant description in formula (21) can refer to the description of formula (19) above, gain2(ω) represents the second target gain, gain i ′(ω) represents the i-th gain value in the second gain. The electronic device can use formula (21) to adjust the N gain values in the second gain so that the gain values greater than the first threshold remain unchanged and the gain values less than the first threshold are adjusted to the first gain.
[0378] S503. The electronic device combines the first target gain and the second target gain according to the frequency of the second output audio signal to generate a third target gain;
[0379] Since the first target gain is used to filter out non-target sound signals corresponding to high-frequency points in the second output audio signal, the second target gain is used to filter out non-target sound signals corresponding to low-frequency points in the second output audio signal.
[0380] The electronic device may determine that a frequency greater than a (khz) is a high frequency, and a frequency less than a (khz) is a low frequency. The value of a may be 0.5-1.5, for example, 1.
[0381] The electronic device determines that the first K frequency points of the N frequency points corresponding to the second output audio signal are low frequencies and the last NK frequency points are high frequencies. The electronic device can obtain the first K gain values of the first target gain and the last NK gain values of the second target gain. The K gain values are placed before the NK gain values to obtain N gain values as the third target gain.
[0382] S504. The electronic device suppresses non-target sound signals in the second output audio signal using the third target gain to generate a third output audio signal.
[0383] The electronic device uses the N gain values corresponding to the third target gain to adjust the N frequency points corresponding to the second output audio signal, respectively, to suppress non-target sound sources in the second output audio signal. Specifically, the electronic device multiplies the i-th gain value in the third target gain by the i-th frequency point in the second output audio signal. When the i-th gain value is greater than a first threshold, the sound signal corresponding to the frequency point is retained. When the target gain value is equal to the first threshold, the sound signal corresponding to the frequency point is suppressed.
[0384] The electronic device uses the N frequency points corresponding to the second output audio signal after adjusting the N frequency points as the third output audio signal.
[0385] In some embodiments, the above steps S503 to S504 are optional. The electronic device may not perform steps S503 to S504. Perform the following operations to suppress non-target sound signals in the second output audio signal: the electronic device determines that the first K frequency points of the N frequency points corresponding to the second output audio signal are low frequencies, and then the electronic device can obtain the first K gain values in the first target gain, and the electronic device uses the K gain values to adjust the K frequency points. At the same time, the electronic device determines that the last NK frequency points corresponding to the second output audio signal are high frequencies, and then the electronic device can obtain the last NK gain values in the second target gain, and the electronic device uses the NK gain values to adjust the NK frequency points. In this way, the N frequency points corresponding to the second output audio signal after the adjustment are used as the third output audio signal. Among them, the process of adjusting the frequency points can refer to the aforementioned description of step S504, which will not be repeated here.
[0386] S107. The electronic device stores the third output audio signal.
[0387] In some embodiments, the electronic device can cache the third output audio signal in a cache area. The cache area can be the aforementioned Figure 7 In the audio stream cache involved.
[0388] Scenario 2: A set of exemplary user interfaces when using the audio processing method of the present application in Scenario 2 can refer to the aforementioned Figure 5a-5d Description of user interface 80-user interface 83 in
[15] . Regarding the real-time video processing method involved in scenario 2, the electronic device can process the current frame of the captured image in real time from the start of video recording, and simultaneously perform audio zoom on the current frame of the captured input audio signal set according to the change in zoom factor. Each processed frame of the image and input audio signal set is played.
[0389] It should be understood that because audio processing takes a certain amount of time, when the electronic device plays the processed current frame, the audio played may not be the processed set of input audio signals for the current frame, but may be the processed set of input audio signals for the previous N frames. Where N is a positive integer greater than or equal to 1, and the specific value of N may be determined by factors such as the processing speed of the electronic device. However, for the user, the difference between these N frames of audio is not noticeable.
[0390] like Figure 17 A schematic diagram showing a scenario 2 in which a current frame image and a current frame input audio signal set are processed in real time and then played.
[0391] The process of the electronic device processing the collected current frame image and the current frame input audio signal set can refer to the aforementioned Figure 7 The description in , will not be repeated here.
[0392] like Figure 17 As shown, when the electronic device finishes processing the first frame of image, it can play the processed first frame of image for preview. At this point, the electronic device has not yet finished processing the first frame of input audio signal set. Then, the electronic device finishes processing the second frame of image signal and plays the processed second frame of image. Simultaneously, the electronic device finishes processing the first frame of input audio signal set and plays the processed first frame of input audio signal set for preview. Then, when the electronic device plays the processed current frame of image, it plays the previously processed input audio signal set. For example, the electronic device finishes processing the third frame of image signal and plays the processed third frame of image. Simultaneously, the electronic device finishes processing the second frame of input audio signal set and plays the processed second frame of input audio signal set for preview. Similarly, the electronic device finishes processing the Nth frame of image signal and plays the processed Nth frame of image. Simultaneously, the electronic device finishes processing the N-1th frame of input audio signal set and plays the processed N-1th frame of input audio signal set for preview.
[0393] In some embodiments, the time it takes for an electronic device to play one frame of image is 30ms, and the time it takes to play one frame of audio is 10ms. Figure 17 In the embodiment, when the electronic device plays a frame of image, a frame of input audio signal set includes 3 frames of audio.
[0394] The process of processing any frame input audio signal set involved in the above scenario 2 by the electronic device is similar to the real-time video processing process involved in scenario 1. Please refer to the above description of steps S101 to S107, which will not be repeated here.
[0395] Scenario 3: A set of exemplary user interfaces when using the audio processing method of the present application in Scenario 3 can refer to the aforementioned Figure 6a-6f The electronic device can use the audio processing method involved in this application to perform post-processing on the audio signal.
[0396] When any microphone of an electronic device collects the current frame of input audio signal, it can save the zoom factor used to collect the input audio signal of that frame. One frame of input audio signal corresponds to one zoom factor. If the electronic device collects N frames of input audio signal, it can obtain N zoom factors. At the same time, the electronic device can save the N frames of input audio signal collected by any microphone separately to obtain input audio streams. If the electronic device has M microphones, it can obtain M input audio streams.
[0397] The electronic device may obtain the M input audio streams, and starting from the first frame of input audio signal in the M input audio streams, sequentially obtain N frames of input audio signals in the M input audio streams. For example, the electronic device may first obtain the M first frame of input audio signals, then obtain the M second frame of input audio signals, and so on. For each of the M i-th frames of input audio signals in the M input audio streams, the electronic device may perform audio zoom using the method of steps S101 to S107 described above.
[0398] Figure 18 This is an exemplary flow chart of an electronic device performing post-processing on an i-th frame of audio signal using the audio processing method involved in this application.
[0399] For this process, reference may be made to the following description of steps S601 to S609 .
[0400] S601. The electronic device obtains a first input audio stream, a second input audio stream, and a third input audio stream, wherein any one of the input audio streams includes multiple frames of input audio signals;
[0401] An exemplary user interface for an electronic device to obtain a first input audio stream, a second input audio stream, and a third input audio stream may be as follows: Figure 6c The user interface 92 is shown.
[0402] The first input audio stream refers to a set of N frames of input audio signals collected by the first microphone of the electronic device.
[0403] The second input audio stream refers to a set of N frames of input audio signals collected by the second microphone of the electronic device.
[0404] The third input audio stream refers to a set of N frames of input audio signals collected by the second microphone of the electronic device.
[0405] S602. The electronic device obtains zoom information, where the zoom information includes a plurality of first zoom ratios;
[0406] The zoom information may include N first zoom magnifications, wherein the i-th first zoom magnification corresponds to the i-th frame of input audio signal collected by the microphone of the electronic device.
[0407] S603. The electronic device determines a first input audio signal from the first audio stream, determines a second input audio signal from the second audio stream, and determines a third input audio signal from the third audio stream;
[0408] The first input audio signal is the input audio signal frame with the earliest acquisition time among all input audio signals in the first audio stream that are not currently subjected to audio zoom.
[0409] The second input audio signal is the input audio signal frame with the earliest acquisition time among all input audio signals in the second audio stream that are not currently undergoing audio zoom.
[0410] The third input audio signal is the input audio signal frame with the earliest acquisition time among all input audio signals in the third audio stream that are not currently undergoing audio zoom.
[0411] S604. The electronic device obtains a first zoom ratio corresponding to when collecting the first input audio signal, the second input audio signal, and the third input audio signal from the zoom information;
[0412] The first zoom magnification is the first zoom magnification among the zoom magnifications in the zoom information that have not been currently acquired by the electronic device.
[0413] S605. The electronic device converts the first input audio signal, the second input audio signal, and the third audio signal into the frequency domain to obtain a first audio signal, a second audio signal, and a third audio signal;
[0414] This process is the same as the description of the aforementioned step S102. Please refer to the above description of step S102 and will not be repeated here.
[0415] S606. The electronic device generates a first output audio signal based on the first zoom ratio using the first audio signal, the second audio signal, and the third audio signal, wherein the target sound signal in the first output audio signal remains unchanged and the non-target sound signal is suppressed;
[0416] This process is the same as the description of the aforementioned step S104. Please refer to the above description of step S104 and will not be repeated here.
[0417] S607. The electronic device enhances or suppresses the first output audio signal according to the first zoom ratio to obtain a second output audio signal, wherein both the target sound signal and the non-target sound signal in the second output audio signal are enhanced or suppressed compared to the first output audio signal;
[0418] This process is the same as the description of the aforementioned step S105. Please refer to the above description of step S105 and will not be repeated here.
[0419] S608. The electronic device suppresses non-target sound signals in the second output audio signal according to the first zoom ratio, in combination with the first audio signal, the second audio signal, and the third audio signal, to generate a third output audio signal.
[0420] This process is the same as the description of the aforementioned step S106. Please refer to the above description of step S106 and will not be repeated here.
[0421] S609. The electronic device saves the third output audio signal.
[0422] The user interface involved in step S609 can be as follows Figure 6e The user interface 94 is shown.
[0423] This process is the same as the description of the aforementioned step S107. Please refer to the above description of step S107 and will not be repeated here.
[0424] In the embodiment of the present application, any audio signal may be referred to as audio or sound.
[0425] The audio signal collected by the electronic device may be referred to as a first audio signal, and the first audio signal may include a first input audio signal, a second input audio signal, a third input audio signal and other audio signals collected by the electronic device.
[0426] The third output audio signal may also be referred to as a second audio signal.
[0427] The zoom ratio control may be referred to as a second control.
[0428] The first zoom magnification may be referred to as a second zoom magnification.
[0429] The video processing method in the embodiment of the present application can be used to process the collected input audio signal in real time when the electronic device is recording a video, for example, the processes involved in scene 1 and scene 2. It can also be used for post-processing of the audio stream, such as scene 3. Of course, it is not limited to the aforementioned scenes 1-3. Implementing the video processing method involved in the present application in these scenarios can increase the zoom ratio of the video played by the electronic device, enhance the sound of the subject within the field of view, and suppress the sound of objects outside the field of view. When the zoom ratio decreases and the field of view increases, the sound of the subject within the field of view is weakened, and the sound of objects outside the field of view is suppressed.
[0430] The following describes an exemplary electronic device provided by an embodiment of the present application.
[0431] Figure 19 It is a structural diagram of an electronic device provided in an embodiment of the present application.
[0432] The following embodiments are described in detail using an electronic device as an example. It should be understood that the electronic device may have more or fewer components than shown in the figures, may combine two or more components, or may have different component configurations. The various components shown in the figures may be implemented in hardware, including one or more signal processing and / or application-specific integrated circuits, software, or a combination of hardware and software.
[0433] The electronic device may include: a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, an earphone interface 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, an air pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.
[0434] It is understood that the structures illustrated in the embodiments of the present invention do not constitute specific limitations on the electronic device. In other embodiments of the present application, the electronic device may include more or fewer components than shown, or may combine or separate certain components, or arrange the components differently. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0435] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU). The different processing units may be independent devices or integrated into one or more processors.
[0436] The controller can be the nerve center and command center of the electronic device. The controller can generate operation control signals based on instruction operation codes and timing signals to complete the control of instruction fetching and execution.
[0437] Processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in processor 110 is a cache memory. This memory can store instructions or data that have just been used or are being recycled by processor 110. If processor 110 needs to use the same instruction or data again, it can directly access the memory. This avoids duplicate accesses, reduces processor 110 latency, and thus improves system efficiency.
[0438] In some embodiments, the processor 110 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface.
[0439] It is understood that the interface connection relationship between the modules illustrated in the embodiments of the present invention is only a schematic illustration and does not constitute a structural limitation of the electronic device. In other embodiments of the present application, the electronic device may also adopt different interface connection methods in the above embodiments, or a combination of multiple interface connection methods.
[0440] The electronic device implements display functionality through a GPU, display screen 194, and an application processor. A GPU is a microprocessor for image processing that connects display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. Processor 110 may include one or more GPUs that execute program instructions to generate or modify display information.
[0441] The display screen 194 is used to display images, videos, etc.
[0442] The electronic device can realize the shooting function through the ISP, camera 193, video codec, GPU, display 194 and application processor.
[0443] The ISP processes data fed back by camera 193. For example, when taking a photo, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, which is then passed to the ISP for processing and converted into a visible image. The ISP can also perform algorithmic optimization on image noise, brightness, and skin tone. It can also optimize parameters such as exposure and color temperature of the captured scene. In some embodiments, the ISP can be located within camera 193.
[0444] The camera 193 is used to capture still images or videos. The object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, and then passes the electrical signal to the ISP for conversion into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in a standard RGB, YUV or other format. In some embodiments, the electronic device may include 1 or N cameras 193, where N is a positive integer greater than 1.
[0445] Digital signal processors (DSPs) are used to process digital signals. Besides digital image signals, they can also process other digital signals. For example, when an electronic device selects a frequency, the DSP performs a Fourier transform on the frequency energy.
[0446] Video codecs are used to compress or decompress digital video. Electronic devices may support one or more video codecs. This allows them to play or record videos in a variety of encoding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, and MPEG4.
[0447] The NPU is a neural network (NN) computing processor. Drawing on the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, it rapidly processes input information and can continuously self-learn. The NPU enables intelligent cognitive applications in electronic devices, such as image recognition, face recognition, speech recognition, and text comprehension.
[0448] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device. The external memory card communicates with the processor 110 via the external memory interface 120 to implement data storage functions. For example, files such as music and videos can be stored on the external memory card.
[0449] The internal memory 121 can be used to store computer executable program codes, which include instructions. The processor 110 executes various functional applications and data processing of the electronic device by running the instructions stored in the internal memory 121. The internal memory 121 can include a program storage area and a data storage area.
[0450] The electronic device can implement audio functions such as music playback and recording through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the headphone jack 170D, and the application processor.
[0451] The audio module 170 is used to convert digital audio information into analog audio signal output, and is also used to convert analog audio input into digital audio signals. The audio module 170 can also be used to encode and decode audio signals. In some embodiments, the audio module 170 can be provided in the processor 110, or some functional modules of the audio module 170 can be provided in the processor 110.
[0452] The speaker 170A, also called a "speaker," is used to convert audio electrical signals into sound signals. The electronic device can listen to music or make hands-free calls through the speaker 170A.
[0453] The receiver 170B, also called a "handset", is used to convert audio electrical signals into sound signals. When the electronic device receives a call or voice message, the voice can be heard by placing the receiver 170B close to the human ear.
[0454] Microphone 170C, also known as "microphone" or "microphone", is used to convert sound signals into electrical signals. When making a call or sending a voice message, the user can speak by putting their mouth close to the microphone 170C to input the sound signal into the microphone 170C. The electronic device can be provided with at least one microphone 170C. In other embodiments, the electronic device can be provided with two microphones 170C, which can not only collect sound signals but also realize noise reduction function. In other embodiments, the electronic device can also be provided with three, four or more microphones 170C to realize sound signal collection, noise reduction, and identification of sound sources, and realize directional recording function, etc.
[0455] The headphone jack 170D is used to connect a wired headphone.
[0456] The gyro sensor 180B can be used to determine the motion posture of the electronic device. In some embodiments, the gyro sensor 180B can be used to determine the angular velocity of the electronic device around three axes (i.e., the x, y, and z axes). The gyro sensor 180B can also be used for anti-shake photography.
[0457] The touch sensor 180K is also called a “touch panel.” The touch sensor 180K can be disposed on the display screen 194 . The touch sensor 180K and the display screen 194 form a touch screen, also called a “touch screen.”
[0458] In the embodiment of the present application, the processor 110 can call the computer instructions stored in the internal memory 121 to enable the electronic device to execute the audio processing method in the embodiment of the present application.
[0459] In an embodiment of the present application, relevant instructions related to the video processing method involved in the embodiment of the application can be stored in the internal memory 121 of the electronic device or in a storage device external to the storage interface 120, so that the electronic device executes the video processing method involved in the embodiment of the present application.
[0460] The following describes the workflow of the electronic device by combining steps S101 to S107 and the hardware structure of the electronic device.
[0461] 1. The electronic device collects a first input audio signal, a second input audio signal, and a third input audio signal;
[0462] In some embodiments, when the touch sensor 180K of the electronic device receives a touch operation (triggered by a user touching a capture control), a corresponding hardware interrupt is sent to the kernel layer. The kernel layer processes the touch operation into a raw input event (including touch coordinates, a timestamp of the touch operation, and other information). The raw input event is stored in the kernel layer. The application framework layer obtains the raw input event from the kernel layer and identifies the control corresponding to the input event.
[0463] For example, the above touch operation is a touch single-click operation, and the control corresponding to the single-click operation is a shooting control in the camera application. The camera application calls the interface of the application framework layer to start the camera application, and then starts the microphone driver by calling the kernel layer, and collects the first input audio signal through the first microphone, the second input audio signal through the second microphone, and the third input audio signal through the third microphone.
[0464] Specifically, the microphone 170C (first microphone, second microphone, and third microphone) of the electronic device can convert the collected sound signal into an analog electrical signal. The electrical signal is then converted into an audio signal in the time domain. The audio signal in the time domain is a digital audio signal, which is stored in the form of 0 and 1. The processor of the electronic device can process the audio signal in the time domain. The audio signal refers to the first input audio signal, the second input audio signal, and also refers to the third input audio signal.
[0465] The electronic device may store the first input audio signal, the second input audio signal, and the third input audio signal in the internal memory 121 or in a storage device externally connected to the storage interface 120 .
[0466] 2. The electronic device converts the first input audio signal, the second input audio signal, and the third audio signal into the frequency domain to obtain a first audio signal, a second audio signal, and a third audio signal;
[0467] The digital signal processor of the electronic device obtains the first input audio signal, the second input audio signal, and the third audio signal from the internal memory 121 or the storage device connected to the storage interface 120, and converts the signals from the time domain to the frequency domain using a DFT to obtain the first audio signal, the second audio signal, and the third audio signal.
[0468] The electronic device may store the first audio signal, the second audio signal, and the third audio signal in the internal memory 121 or in a storage device externally connected to the storage interface 120 .
[0469] 3. The electronic device obtains a first zoom ratio;
[0470] In some embodiments, when the touch sensor 180K of the electronic device receives a touch operation (triggered by a user touching a zoom control), a corresponding hardware interrupt is sent to the kernel layer. The kernel layer processes the touch operation into a raw input event (including touch coordinates, a timestamp of the touch operation, and other information). The raw input event is stored in the kernel layer. The application framework layer obtains the raw input event from the kernel layer and identifies the control corresponding to the input event.
[0471] For example, if the touch operation is a sliding operation and the control corresponding to the sliding operation is a zoom magnification control in a camera application, the camera application calls an interface of the application framework layer to obtain a parameter corresponding to the zoom magnification control, namely, a first zoom magnification.
[0472] 4. The electronic device generates a first output audio signal using the first audio signal, the second audio signal, and the third audio signal according to the first zoom ratio;
[0473] The electronic device can obtain the first audio signal, the second audio signal, and the third audio signal stored in the memory 121 or in a storage device externally connected to the storage interface 120 through the processor 110. The processor 110 of the electronic device calls relevant computer instructions to generate a first output audio signal based on the first audio signal, the second audio signal, and the third audio signal.
[0474] 5. The electronic device enhances or suppresses the first output audio signal according to the first zoom ratio to obtain a second output audio signal;
[0475] The processor 110 of the electronic device calls relevant computer instructions and enhances or suppresses the first output audio signal according to the first zoom ratio to obtain a second output audio signal.
[0476] 6. The electronic device suppresses non-target sound signals in the second output audio signal to generate a third output audio signal;
[0477] The processor 110 of the electronic device calls relevant computer instructions to suppress the non-target sound signal in the second output audio signal to generate a third output audio signal.
[0478] 7. The electronic device stores the third output audio signal
[0479] The electronic device buffers the fourth output audio signal in a buffer area.
[0480] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.
[0481] As used in the above embodiments, the term “when…” may be interpreted to mean “if…” or “after…” or “in response to determining…” or “in response to detecting…”, depending on the context. Similarly, the phrases “upon determining…” or “if (stated condition or event) is detected” may be interpreted to mean “if determining…” or “in response to determining…” or “upon detecting (stated condition or event)” or “in response to detecting (stated condition or event)”, depending on the context.
[0482] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrations. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state hard disk).
[0483] Those skilled in the art will appreciate that all or part of the process steps in the above-described method embodiments can be implemented by a computer program instructing the relevant hardware. The program can be stored in a computer-readable storage medium, and when executed, the program can include the process steps in the above-described method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM or random access memory (RAM), magnetic disks, or optical disks.
Claims
1. A video processing method, characterized in that: Applied to electronic equipment, the method includes: The electronic device activates a camera; Display a preview interface, wherein the preview interface includes a first control; detecting a first operation on the first control; In response to the first operation, starting shooting; Displaying a shooting interface, wherein the shooting interface includes a second control, and the second control is used to adjust the zoom ratio; At a first moment, the zoom magnification is a first zoom magnification, and a first captured image is displayed, where the first captured image includes a first target object and a second target object; detecting a second operation on the second control; In response to the second operation, the zoom ratio is adjusted to a second zoom ratio, the second zoom ratio being greater than the first zoom ratio; At a second moment, displaying a second captured image, where the second captured image includes the first target object and the second captured image does not include the second target object; At the second moment, the microphone collects first audio, where the first audio includes a first sound and a second sound, where the first sound corresponds to the first target object and the second sound corresponds to the second target object; detecting a third operation on a third control; In response to the third operation, the shooting is stopped and the first video is saved, wherein, The first video includes the first captured image, and the first video includes a second captured image and a second audio at a second moment, the second audio is obtained by processing the first audio according to the second zoom ratio, and the second audio includes a third sound and a fourth sound, the third sound corresponds to the first target object, the fourth sound corresponds to the second target object, the third sound is enhanced relative to the first sound, and the fourth sound is suppressed relative to the second sound; the enhancement is achieved by adjusting the amplitude of the first output audio, and the first output audio is obtained by first suppressing the second sound in the first audio; after the enhancement, the second sound is suppressed again; the first suppression is achieved by using the second variable The zoom ratio is achieved by fusing beams corresponding to three directions, and the beams corresponding to the three directions are determined by the first audio, and the first audio includes audio collected by at least two microphones, and the beam in one direction includes the audio in the one direction collected by the at least two microphones; the second zoom ratio determines the fusion coefficient when the beams in the three directions are fused to obtain the first output audio, and the fusion coefficient is used to indicate the degree to which the second sound is suppressed for the first time; the re-suppression is achieved by a first filtering algorithm and a second filtering algorithm, and the first filtering algorithm is used to suppress the sound corresponding to the high-frequency points in the second sound, and the second filtering algorithm is used to suppress the sound corresponding to the low-frequency points in the second sound.
2. The method according to claim 1, characterized in that The method further comprises: enhancing the first output audio according to the second zoom ratio to obtain a second output audio; wherein both the first sound and the second sound are enhanced in the second output audio; and wherein the first sound remains unchanged and the second sound is suppressed in the first output audio; According to the second zoom ratio and in combination with the first audio, the second sound in the second output audio is suppressed to obtain the second audio.
3. The method according to claim 2, characterized in that The first output audio includes one audio channel; and the method further includes: The electronic device obtains a first filter coefficient corresponding to a first direction, a second filter coefficient corresponding to a second direction, and a third filter coefficient corresponding to a third direction; the first direction is any direction within a range of 10° clockwise in front of the electronic device to 70° clockwise in front of the electronic device; the second direction is any direction within a range of 10° counterclockwise in front of the electronic device to 10° clockwise in front of the electronic device; and the third direction is any direction within a range of 10° counterclockwise in front of the electronic device to 70° counterclockwise in front of the electronic device; Using the first filter coefficient in combination with the first audio, a first beam corresponding to a first direction is obtained; using the second filter coefficient in combination with the first audio, a second beam corresponding to a second direction is obtained; and using the third filter coefficient in combination with the first audio, a third beam corresponding to a third direction is obtained; The electronic device determines the fusion coefficient according to the second zoom ratio, and uses the fusion coefficient to fuse the first beam, the second beam and the third beam to obtain the first output audio, in which the first sound remains unchanged and the second sound is suppressed for the first time.
4. The method according to claim 2 or 3, characterized in that The second output audio includes one audio channel; Enhancing the first output audio according to the second zoom ratio to obtain a second output audio specifically includes: determining, according to the second zoom ratio, an adjustment parameter corresponding to the second zoom ratio, wherein the adjustment parameter is used to enhance the amplitude of the audio; Converting the first output audio from the frequency domain to the time domain to obtain a first output audio in the time domain; Using the adjustment parameter to enhance the first output audio in the time domain to obtain a second output audio in the time domain; The second output audio in the time domain is converted into the frequency domain as the second output audio.
5. The method according to claim 2 or 3, characterized in that The second audio includes one audio channel; Suppressing a second sound in the second output audio according to the second zoom ratio and in combination with the first audio to obtain the second audio specifically includes: Performing a Zelinsky filter on the first audio to obtain a first target gain; the first target gain is used to filter out sounds corresponding to high-frequency points in the second sound in the second output audio; Using the first audio, based on a coherent diffusion power ratio algorithm, a second target gain is obtained; the second target gain is used to filter out the sound corresponding to the low frequency point in the second sound in the second output audio; combining a portion of the first target gain and a portion of the second target gain according to the frequency of the second output audio to obtain a third target gain; The second sound in the second output audio is suppressed again using the third target gain to obtain the second audio.
6. The method according to claim 2 or 3, characterized in that The second audio includes one audio channel; Suppressing a second sound in the second output audio according to the second zoom ratio and in combination with the first audio to obtain the second audio specifically includes: Performing a Zelinsky filter on the first audio to obtain a first target gain; the first target gain is used to filter out sounds corresponding to high-frequency points in the second sound in the second output audio; Using the first audio, based on a coherent diffusion power ratio algorithm, a second target gain is obtained; the second target gain is used to filter out the sound corresponding to the low frequency point in the second sound in the second output audio; The second audio is obtained by suppressing the sound corresponding to the high-frequency point in the second audio by using the first target gain and suppressing the sound corresponding to the low-frequency point in the second audio by using the second target gain.
7. The method according to claim 2, characterized in that The first output audio includes left-channel audio and right-channel audio; and the method further includes: Obtaining a first filter coefficient corresponding to a first direction, a second filter coefficient corresponding to a second direction, and a third filter coefficient corresponding to a third direction; the first direction is any direction within a range of 10° clockwise in front of the electronic device to 70° clockwise in front of the electronic device; the second direction is any direction within a range of 10° counterclockwise in front of the electronic device to 10° clockwise in front of the electronic device; and the third direction is any direction within a range of 10° counterclockwise in front of the electronic device to 70° counterclockwise in front of the electronic device; Using the first filter coefficient in combination with the first audio, a first beam corresponding to a first direction is obtained; using the second filter coefficient in combination with the first audio, a second beam corresponding to a second direction is obtained; and using the third filter coefficient in combination with the first audio, a third beam corresponding to a third direction is obtained; According to the second zoom ratio, the first beam and the second beam are used to obtain the left channel audio in the first output audio; the second beam and the third beam are used to obtain the right channel audio in the first output audio, the first sound in the left channel audio and the right channel audio remains unchanged, and the second sound is suppressed.
8. The method according to claim 7, characterized in that The second output audio includes a left channel audio and a right channel audio; Enhancing the first output audio according to the second zoom ratio to obtain a second output audio specifically includes: determining, according to the second zoom ratio, an adjustment parameter corresponding to the second zoom ratio, wherein the adjustment parameter is used to enhance the audio; Convert the left channel audio and the right channel audio in the first output audio from the frequency domain to the time domain to obtain the left channel audio and the right channel audio in the first output audio in the time domain; Using the adjustment parameters, respectively enhancing the left channel audio and the right channel audio in the first output audio in the time domain to obtain the left channel audio and the right channel audio in the second output audio in the time domain; The left channel audio and the right channel audio in the second output audio in the time domain are converted into the frequency domain as the second output audio.
9. The method according to claim 7 or 8, characterized in that The second audio includes a left channel audio and a right channel audio; Suppressing a second sound in the second output audio according to the second zoom ratio and in combination with the first audio to obtain the second audio specifically includes: Performing a Zelinsky filter on the first audio to obtain a first target gain; the first target gain is used to filter out sounds corresponding to high-frequency points in the second sound in the second output audio; Using the first audio, based on a coherent diffusion power ratio algorithm, a second target gain is obtained; the second target gain is used to filter out the sound corresponding to the low frequency point in the second sound in the second output audio; combining a portion of the first target gain and a portion of the second target gain according to the frequency of the second output audio to obtain a third target gain; The third target gain is used to suppress the second sound in the left channel audio and the right channel audio in the second output audio, respectively, to obtain the left channel audio and the right channel audio in the second audio.
10. The method according to claim 7 or 8, characterized in that The second audio includes a left channel audio and a right channel audio; Suppressing a second sound in the second output audio according to the second zoom ratio and in combination with the first audio to obtain the second audio specifically includes: Performing a Zelinsky filter on the first audio to obtain a first target gain; the first target gain is used to filter out sounds corresponding to high-frequency points in the second sound in the second output audio; Using the first audio, based on a coherent diffusion power ratio algorithm, a second target gain is obtained; the second target gain is used to filter out the sound corresponding to the low frequency point in the second sound in the second output audio; According to the first target gain, the sound corresponding to the high-frequency points in the second sound in the left channel audio and the right channel audio in the second audio is suppressed, and the sound corresponding to the low-frequency points in the second sound in the left channel audio and the right channel audio in the second audio is suppressed using the second target gain to obtain the left channel audio and the right channel audio in the second audio.
11. The method according to any one of claims 3, 7 or 8, characterized in that The first filter coefficient, the second filter coefficient, and the third filter coefficient are pre-set in the electronic device; among the first filter coefficients, the coefficient corresponding to the sound signal in the first direction is 1, indicating that the sound signal in the first direction is not suppressed; the closer the sound signal is to the first direction, the closer the corresponding coefficient is to 1, and the degree of suppression increases successively; among the second filter coefficients, the coefficient corresponding to the sound signal in the second direction is 1, indicating that the sound signal in the second direction is not suppressed; the closer the sound signal is to the second direction, the closer the corresponding coefficient is to 1, and the degree of suppression increases successively; among the third filter coefficients, the coefficient corresponding to the sound signal in the third direction is 1, indicating that the sound signal in the third direction is not suppressed; The closer the sound signal is to the third direction, the closer the corresponding coefficient is to 1, and the degree of suppression increases accordingly.
12. The method according to claim 8, characterized in that The adjustment parameter corresponding to the second zoom ratio is pre-set in the electronic device; the value of the adjustment parameter is directly related to the zoom ratio, and any second zoom ratio uniquely corresponds to one adjustment parameter.
13. An electronic device, characterized in that: The electronic device includes one or more processors and one or more memories; wherein the one or more memories are coupled to the one or more processors, and the one or more memories are used to store computer program code, and the computer program code includes computer instructions, and when the one or more processors execute the computer instructions, the electronic device executes the method according to any one of claims 1 to 12.
14. A chip system, applied to an electronic device, comprising one or more processors, wherein the processors are configured to call computer instructions so that the electronic device executes the method according to any one of claims 1 to 12.
15. A computer program product comprising instructions, characterized in that When the computer program product is run on an electronic device, the electronic device is enabled to perform the method according to any one of claims 1 to 12.
16. A computer-readable storage medium comprising instructions, characterized in that: When the instructions are executed on an electronic device, the electronic device is caused to execute the method according to any one of claims 1 to 12.
Citation Information
Patent Citations
Information processing method and device and electronic device
CN106157986A