Video processing method and electronic device
By automatically extracting visual semantic features from videos and separating the voice of the target person through electronic devices, the problem of obtaining voiceprints when users shoot videos is solved, improving the user experience and reducing the computational load.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HUAWEI TECH CO LTD
- Filing Date
- 2021-07-29
- Publication Date
- 2026-04-24
AI Technical Summary
In existing technologies, users need to manually obtain the voiceprint of the target person to separate the voice when shooting videos, resulting in a poor user experience in scenarios such as interviews, casual shooting, and street photography. In addition, professional video processing software has a high learning curve.
By separating a person's voice from a video based on visual semantic features, electronic devices can automatically identify and separate the voiceprint of the target person without prior acquisition of the voiceprint. They can directly extract visual semantic features from the video and generate the corresponding voice.
It enables the separation and processing of the target person's voice in a video without the need for specialized software, improving the user experience, having a wide range of applications, and reducing the computational load.
Smart Images

Figure CN115691538B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of electronic technology, and in particular to video processing methods and electronic devices. Background Technology
[0002] With the development of electronic technology and the increasing demand for entertainment, more and more users are using the shooting functions of electronic devices to record and share their lives. In daily life, when most users shoot videos, the shooting device inevitably captures sounds from other people, animals, and objects in the environment. Users need to perform complex processing on the captured video files to suppress noise. However, professional video processing software has a high learning curve, requiring users to learn additional skills, resulting in a poor user experience.
[0003] To enhance the user experience and allow users to easily process their own videos and reduce noise, a voiceprint-based sound separation method includes: pre-acquiring the voice of the person being filmed, and obtaining the voiceprint of the target person corresponding to that voice; after obtaining the voiceprint of the target person, processing the audio signal in the video based on the voiceprint to separate the voice of the target person in the video; and performing a series of processing steps after obtaining the voice of the target person in the video, such as enhancement processing and noise reduction processing.
[0004] However, it is clear that voiceprint-based voice separation methods require pre-acquiring the target person's voice to determine their voiceprint. But in many scenarios, such as interviews, casual snapshots, and street photography, the target person is often a stranger, and users do not record a separate segment of the target person's voice before or after filming for later video processing to obtain their voiceprint. Summary of the Invention
[0005] This application provides a video processing method and an electronic device. The video processing method provided by this application can separate a person's voice from the audio in a video based on the person's visual semantic features, thereby enabling the user to hear the person's voice more clearly.
[0006] In a first aspect, this application provides a video processing method, the method comprising: an electronic device acquiring a first video, wherein at least a portion of the image frames displayed in the first video include a first object, the first object satisfying the first condition, and the first video including first audio; the electronic device generating a second audio based on the first video, wherein the second audio corresponds to the first object; the electronic device displaying a first interface, the first interface including a first control for processing the second audio; and, in response to an operation performed on the first control, the electronic device playing a third audio, the third audio being audio information processed from the second audio.
[0007] In the above embodiments, the electronic device can separate the sound of any object in the video and, in response to the user's operation, play the sound processed based on the sound of that object, eliminating the need for the user to use professional video processing software to process the video, thus greatly improving the user experience. Furthermore, since the electronic device has already separated the object's sound, it can perform various processing operations, such as voice changing, enhancement, and reduction.
[0008] In conjunction with some embodiments of the first aspect, in some embodiments, the method further includes: the first interface further includes a second control, the second control being used to play only the second audio; in response to an operation performed on the second control, the electronic device plays the second audio.
[0009] In the above embodiments, since the electronic device has separated the sound of the first object, the electronic device can also play the picture of the first video and the sound of the first object during the first video.
[0010] In conjunction with some embodiments of the first aspect, in some embodiments, the method further includes: at least another portion of image frames of the first video further includes a second object, the second object satisfying the first condition, and the first interface further includes a third control; the electronic device generates a fourth audio based on the first audio, wherein the fourth audio corresponds to the second object; in response to an operation performed on the third control, the electronic device plays a fifth audio, the fifth audio including audio information processed from the fourth audio.
[0011] In the above embodiments, a portion of the image frames corresponding to the first object may partially overlap, fully overlap, or not overlap at all with another portion of the image frames corresponding to the second image; furthermore, the sound of the first object may fully overlap, partially overlap, or not overlap at all with the sound of the second object; in all three cases, the electronic device can separate the sound of the second object, and then play different sounds based on the user's interaction.
[0012] In conjunction with some embodiments of the first aspect, in some embodiments, the first condition includes: the pitch angle of the object's face is within a preset pitch angle range and / or the rotation of the object's face is within a preset rotation angle range and / or the tilt angle of the object's face is within a preset tilt angle range.
[0013] In the above embodiments, by determining whether the pitch angle, rotation angle, and tilt angle of the object's face are within a preset range, the number of subsequent image frames processed is reduced, thereby reducing the computational load on the electronic device.
[0014] In conjunction with some embodiments of the first aspect, in some embodiments, the electronic device generates a second audio based on the first video, specifically including: the electronic device determining a corresponding first segmented audio based on the at least a portion of the image frames; the electronic device determining visual semantic features of the first object based on the at least a portion of the image frames, the visual semantic features being voice-related and sound-related facial morphological features; the electronic device determining the sound of the first object in the first segmented audio based on the visual semantic features of the first object and the first segmented audio; the electronic device determining the voiceprint of the first object based on the sound of the first object in the first segmented audio; and the electronic device generating the second audio based on the voiceprint of the first object and the first audio.
[0015] In the above embodiments, the electronic device can first determine the visual semantic features of the first object, then determine part of the sound of the first object, and determine the voiceprint of the first object based on the part of the sound, and then determine all the sound of the first object in the first video based on the voiceprint of the first object, and finally complete the separation of the sound of the first object.
[0016] In conjunction with some embodiments of the first aspect, in some embodiments, the electronic device generates a second audio based on the first video, specifically including: the electronic device determining that all image frames of the first video include a first object, the first object satisfying a first condition; the electronic device determining visual semantic features of the first object based on all image frames of the first video, the visual semantic features being features of a person's facial shape related to speech or sound; and the electronic device generating the second audio based on the visual semantic features of the first object and the first audio.
[0017] In the above embodiments, when all image frames in the video include the first object, all sounds of the first object in the video can be determined by determining the visual semantic features of the first object. Furthermore, compared to methods that separate the sound of the first object using voiceprints, the visual semantic features contain more information, allowing for the separation of the sound of the first object with higher quality.
[0018] In conjunction with some embodiments of the first aspect, in some embodiments, the electronic device determines the voiceprint of the first object based on the sound of the first object in the first segmented audio, specifically including: the electronic device filtering out sound segments from the sound of the first object in the first segmented audio where the signal-to-noise ratio is greater than a signal-to-noise ratio threshold and the duration is greater than a duration threshold; the electronic device determining the voiceprint of the first object based on the sound segments.
[0019] In the above embodiments, when determining the voiceprint of the first object from a portion of the first object's sound, the electronic device can filter the portion of the sound, select the sound of the first object with a high signal-to-noise ratio and a long duration, and thus obtain the voiceprint of the first object with higher confidence.
[0020] In conjunction with some embodiments of the first aspect, in some embodiments, the first video is stored locally on the electronic device; or, the first video is a video received during a video call.
[0021] Secondly, this application provides a video processing method, the method comprising: the electronic device acquiring a first video, the first video including a first audio; the electronic device determining a first portion of image frames based on the first video, the first portion of image frames including a first object, the first object in the first portion of image frames satisfying a first condition; the electronic device determining a first segmented audio corresponding to the first portion of image frames; the electronic device determining visual semantic features of the first object based on the first portion of image frames, the visual semantic features being voice-related and sound-related facial morphological features; the electronic device determining the sound of the first object in the first segmented audio based on the visual semantic features of the first object and the first segmented audio; the electronic device determining the voiceprint of the first object based on the sound of the first object in the first segmented audio; and the electronic device determining a second audio corresponding to the first object based on the voiceprint of the first object and the first audio.
[0022] In the above embodiments, the electronic device first determines a portion of image frames corresponding to the first object, and then determines the visual semantic features of the first object from the portion of image frames, thereby determining a portion of the sound of the first object, then determining the voiceprint of the first object, and finally determining the sound of the first object. Obviously, the electronic device does not need to obtain the voiceprint of the first object in advance, and it does not require that the first object exists in all image frames, and the first object satisfies the first condition. It has a wide range of applications and strong robustness.
[0023] In conjunction with some embodiments of the second aspect, in some embodiments, the method further includes: the electronic device determining a second portion of image frames based on the first video, the second portion of image frames including a second object, the second object in the second portion of image frames satisfying the first condition; the electronic device determining a second segmented audio corresponding to the second portion of image frames; the electronic device determining visual semantic features of the second object based on the second portion of image frames, the visual semantic features being voice-related and sound-related facial morphological features; the electronic device determining the sound of the first object in the second segmented audio based on the visual semantic features of the first object and the first segmented audio; the electronic device determining the voiceprint of the second object based on the sound of the second object in the second segmented audio; and the electronic device determining a third audio based on the voiceprint of the second object and the first audio, the third audio corresponding to the second object.
[0024] In the above embodiments, when the first video includes multiple objects, the sounds of different objects can be separated. The image frames corresponding to the first object and the image frames corresponding to the second object can completely overlap, partially overlap, or not overlap at all.
[0025] In conjunction with some embodiments of the second aspect, in some embodiments, the first condition includes: the pitch angle of the object's face is within a preset pitch angle range and / or the rotation of the object's face is within a preset rotation angle range and / or the tilt angle of the object's face is within a preset tilt angle range.
[0026] In the above embodiments, by determining whether the pitch angle, rotation angle, and tilt angle of the object's face are within a preset range, the number of subsequent image frames processed is reduced, thereby reducing the computational load on the electronic device.
[0027] In conjunction with some embodiments of the second aspect, in some embodiments, the electronic device determines the voiceprint of the first object based on the sound of the first object in the first segmented audio, specifically including: the electronic device filtering out sound segments from the sound of the first object in the first segmented audio where the signal-to-noise ratio is greater than a signal-to-noise ratio threshold and the duration is greater than a duration threshold; the electronic device determining the voiceprint of the first object based on the sound segments.
[0028] In the above embodiments, when determining the voiceprint of the first object from a portion of the first object's sound, the electronic device can filter the portion of the sound, select the sound of the first object with a high signal-to-noise ratio and a long duration, and thus obtain the voiceprint of the first object with higher confidence.
[0029] In conjunction with some embodiments of the second aspect, in some embodiments, the first video is stored locally on the electronic device; or, the first video is a video received during a video call.
[0030] Thirdly, this application provides a video processing method, comprising: the electronic device acquiring a first video, the first video including a first audio; the electronic device determining all image frames based on the first video, the all image frames including a first object, the first object in the all image frames satisfying a first condition; the electronic device determining visual semantic features of the first object based on the all image frames, the visual semantic features being features of facial morphology related to speech or sound; and the electronic device determining a second audio based on the visual semantic features of the first object and the first audio, the second audio corresponding to the first object.
[0031] In the above embodiments, when the electronic device determines that all image frames in the video include the first object, and the first object satisfies the first condition, it can extract the visual semantic features of the object from all image frames of the video, and thus determine the sound of the object. Clearly, the electronic device does not need to obtain the object's voiceprint beforehand, making it widely applicable.
[0032] In conjunction with some embodiments of the third aspect, in some embodiments, the first condition includes: the pitch angle of the object's face is within a preset pitch angle range and / or the rotation of the object's face is within a preset rotation angle range and / or the tilt angle of the object's face is within a preset tilt angle range.
[0033] In the above embodiments, by determining whether the pitch angle, rotation angle, and tilt angle of the object's face are within a preset range, the number of subsequent image frames processed is reduced, thereby reducing the computational load on the electronic device.
[0034] Fourthly, this application provides a video processing method, the method comprising: an electronic device acquiring a first video, wherein at least a portion of the image frames of the first video include a first object, the first object in the at least a portion of the image frames satisfies the first condition, and the sound of the first video is a first audio, the first audio consisting of the sound of the first object and other sounds; the electronic device displaying a first interface displaying images of the first video, the first interface including one or more controls, the one or more controls including a first control for instructing the electronic device to play a second audio; in response to an operation performed on the first control, the electronic device generating and playing the second audio based on the first video; the second audio being the sound of the first object; or, the second audio differing from the first audio in that the sound of the first object is different; or, the second audio and the first audio being identical only in that the sound of the first object is the same.
[0035] In the above embodiments, in response to the user's operation, the sound played when the first video is played is different from the sound of the first video. The difference is that the sound played is the sound of any object in the first video, or only the sound of the first object is changed or not changed.
[0036] In conjunction with some embodiments of the fourth aspect, in some embodiments, before the electronic device generates and plays the second audio based on the first video in response to an operation performed on the first control, the method further includes: the electronic device determining a third audio; the electronic device generating the second audio based on the third audio; and the second audio corresponding to the first object.
[0037] In the above embodiments, the electronic device can separate the sound of any object in the video and, in response to the user's operation, play the sound processed based on the sound of that object, eliminating the need for the user to use professional video processing software to process the video, thus greatly improving the user experience. Furthermore, since the electronic device has already separated the object's sound, it can perform various processing operations, such as voice changing, enhancement, and reduction.
[0038] Fifthly, this application provides an electronic device comprising: one or more processors and a memory; the memory is coupled to the one or more processors, the memory being used to store computer program code, the computer program code including computer instructions, and the one or more processors calling the computer instructions to cause the electronic device to execute:
[0039] The system acquires a first video, wherein at least a portion of the image frames displayed in the first video include a first object, the first object satisfies a first condition, and the first video includes a first audio; generates a second audio based on the first video, wherein the second audio corresponds to the first object; displays a first interface, the first interface including a first control for processing the second audio; and plays a third audio in response to an operation performed on the first control, the third audio being the processed audio information of the second audio.
[0040] In conjunction with some embodiments of the fifth aspect, in some embodiments, the one or more processors are further configured to invoke the computer instructions to cause the electronic device to perform: the first interface further includes a second control for playing only the second audio; and, in response to an operation performed on the second control, playing the second audio.
[0041] In conjunction with some embodiments of the fifth aspect, in some embodiments, the one or more processors are further configured to invoke the computer instructions to cause the electronic device to perform: the first video further includes a second object in at least another portion of the image frames, the second object satisfying the first condition, the first interface further including a third control; generate a fourth audio based on the first audio, wherein the fourth audio corresponds to the second object; and play a fifth audio in response to an operation performed on the third control, the fifth audio including audio information processed from the fourth audio.
[0042] In conjunction with some embodiments of the fifth aspect, in some embodiments, the first condition includes: the pitch angle of the object's face is within a preset pitch angle range and / or the rotation of the object's face is within a preset rotation angle range and / or the tilt angle of the object's face is within a preset tilt angle range.
[0043] In conjunction with some embodiments of the fifth aspect, in some embodiments, the one or more processors are specifically configured to invoke the computer instructions to cause the electronic device to perform: determining a corresponding first captured audio based on the at least a portion of the image frames; determining visual semantic features of the first object based on the at least a portion of the image frames, the visual semantic features being voice-related or sound-related facial morphological features; determining the voice of the first object in the first captured audio based on the visual semantic features of the first object and the first captured audio; determining the voiceprint of the first object based on the voiceprint of the first object in the first captured audio; and generating the second audio based on the voiceprint of the first object and the first audio.
[0044] In conjunction with some embodiments of the fifth aspect, in some embodiments, the one or more processors specifically invoke the computer instructions to cause the electronic device to perform: determining that all image frames of the first video include a first object, the first object satisfying a first condition; determining visual semantic features of the first object based on all image frames of the first video, the visual semantic features being features of a voice-related, sound-related facial morphology; and generating the second audio based on the visual semantic features of the first object and the first audio.
[0045] In conjunction with some embodiments of the fifth aspect, in some embodiments, the one or more processors are specifically configured to invoke the computer instructions to cause the electronic device to perform: filtering from the sound of the first object in the first captured audio a sound segment with a signal-to-noise ratio greater than a signal-to-noise ratio threshold and a duration greater than a duration threshold; and determining the voiceprint of the first object based on the sound segment.
[0046] In conjunction with some embodiments of the fifth aspect, in some embodiments, the first video is stored locally on the electronic device; or, the first video is a video received during a video call.
[0047] Sixthly, this application provides an electronic device comprising: one or more processors and a memory; the memory being coupled to the one or more processors, the memory storing computer program code including computer instructions, the one or more processors calling the computer instructions to cause the electronic device to execute:
[0048] The process involves: acquiring a first video, which includes a first audio; determining a first portion of image frames based on the first video, wherein the first portion of image frames includes a first object, and the first object in the first portion of image frames satisfies a first condition; determining a first segmented audio corresponding to the first portion of image frames; determining visual semantic features of the first object based on the first portion of image frames, wherein the visual semantic features are facial morphological features related to speech or sound; determining the sound of the first object in the first segmented audio based on the visual semantic features of the first object and the first segmented audio; determining the voiceprint of the first object based on the sound of the first object in the first segmented audio; and determining a second audio corresponding to the first object based on the voiceprint of the first object and the first audio.
[0049] In conjunction with some embodiments of the sixth aspect, in some embodiments, the one or more processors are further configured to invoke the computer instructions to cause the electronic device to perform: determining a second portion of an image frame based on the first video, the second portion of the image frame including a second object, the second object in the second portion of the image frame satisfying the first condition; determining a second segmented audio corresponding to the second portion of the image frame; determining visual semantic features of the second object based on the second portion of the image frame, the visual semantic features being voice-related and sound-related facial morphological features; determining the voice of the first object in the second segmented audio based on the visual semantic features of the first object and the first segmented audio; determining the voiceprint of the second object based on the voiceprint of the second object in the second segmented audio; and determining a third audio based on the voiceprint of the second object and the first audio, the third audio corresponding to the second object.
[0050] In conjunction with some embodiments of the sixth aspect, in some embodiments, the first condition includes: the pitch angle of the object's face is within a preset pitch angle range and / or the rotation of the object's face is within a preset rotation angle range and / or the tilt angle of the object's face is within a preset tilt angle range.
[0051] In conjunction with some embodiments of the sixth aspect, in some embodiments, the one or more processors are specifically configured to invoke the computer instructions to cause the electronic device to perform: filtering from the sound of the first object in the first captured audio a sound segment with a signal-to-noise ratio greater than a signal-to-noise ratio threshold and a duration greater than a duration threshold; and determining the voiceprint of the first object based on the sound segment.
[0052] In conjunction with some embodiments of the sixth aspect, the first video is stored locally on the electronic device; or, the first video is a video received during a video call.
[0053] In a seventh aspect, this application provides an electronic device comprising: one or more processors and a memory; the memory being coupled to the one or more processors, the memory storing computer program code including computer instructions, the one or more processors calling the computer instructions to cause the electronic device to execute:
[0054] A first video is acquired, the first video including a first audio; all image frames are determined based on the first video, the all image frames including a first object, the first object in the all image frames satisfying a first condition; visual semantic features of the first object are determined based on the all image frames, the visual semantic features being features of facial morphology related to speech and sound; a second audio is determined based on the visual semantic features of the first object and the first audio, the second audio corresponding to the first object.
[0055] In conjunction with some embodiments of the seventh aspect, in some embodiments, the first condition includes: the pitch angle of the object's face is within a preset pitch angle range and / or the rotation of the object's face is within a preset rotation angle range and / or the tilt angle of the object's face is within a preset tilt angle range.
[0056] Eighthly, this application provides an electronic device comprising: one or more processors and a memory; the memory being coupled to the one or more processors, the memory storing computer program code including computer instructions, the one or more processors calling the computer instructions to cause the electronic device to execute:
[0057] The system acquires a first video, wherein at least a portion of the image frames of the first video include a first object, the first object in the at least a portion of the image frames satisfies a first condition, and the sound of the first video is a first audio, which consists of the sound of the first object and other sounds; displays a first interface that displays images of the first video, the first interface including one or more controls, the one or more controls including a first control for instructing the electronic device to play a second audio; in response to an operation performed on the first control, generates and plays the second audio based on the first video; the second audio is the sound of the first object; or, the second audio differs from the first audio in that the sound of the first object is different; or, the second audio and the first audio are identical only in that the sound of the first object is the same.
[0058] In conjunction with some embodiments of the eighth aspect, in some embodiments, the one or more processors are further configured to invoke the computer instructions to cause the electronic device to perform: determining a third audio; generating a second audio based on the third audio; the second audio corresponding to the first object.
[0059] Ninthly, this application provides a chip system applied to an electronic device, the chip system including one or more processors, the processors being configured to invoke computer instructions to cause the electronic device to perform the methods described in the first aspect, the second aspect, the third aspect, the fourth aspect, and any possible implementation of the first aspect, the second aspect, the third aspect, and the fourth aspect.
[0060] In a tenth aspect, embodiments of this application provide a computer program product containing instructions that, when the computer program product is run on an electronic device, cause the electronic device to perform the method described in the first aspect, the second aspect, the third aspect, the fourth aspect, and any possible implementation of the first aspect, the second aspect, the third aspect, and the fourth aspect.
[0061] Eleventhly, embodiments of this application provide a computer-readable storage medium including instructions that, when executed on an electronic device, enable the electronic device to perform the method described in the first aspect, second aspect, third aspect, fourth aspect, and any possible implementation of the first aspect, second aspect, third aspect, and fourth aspect.
[0062] It is understood that the electronic devices provided in the fifth, sixth, seventh, and eighth aspects, the chip system provided in the ninth aspect, the computer program product provided in the tenth aspect, and the computer-readable storage medium provided in the eleventh aspect are all used to execute the methods provided in the embodiments of this application. Therefore, the beneficial effects they can achieve can be referred to the beneficial effects in the corresponding methods, and will not be repeated here. Attached Figure Description
[0063] Figure 1A This is an exemplary schematic diagram of the interface of an electronic device for acquiring voiceprints as described in this application;
[0064] Figure 1B This is an exemplary schematic diagram illustrating the principle of voiceprint acquisition in the electronic device involved in this application;
[0065] Figure 2A , Figure 2B This is an exemplary schematic diagram of the face detection involved in this application;
[0066] Figure 3A , Figure 3B An exemplary schematic diagram illustrating the acquisition of visual semantic features provided in this application embodiment;
[0067] Figure 4 An exemplary schematic diagram illustrating the operation of the decoupled network provided in this application embodiment;
[0068] Figure 5A An exemplary schematic diagram of the audio-video fusion module architecture provided in the embodiments of this application;
[0069] Figure 5B Another exemplary schematic diagram of the audio-video fusion module architecture provided in the embodiments of this application;
[0070] Figure 6A , Figure 6B This is an exemplary schematic diagram of a voiceprint-based sound separation method involved in this application;
[0071] Figure 6C This is an exemplary schematic diagram of another voiceprint-based sound separation method involved in this application;
[0072] Figure 6D This is an exemplary schematic diagram of the text-based sound separation method involved in this application;
[0073] Figure 7 An exemplary schematic diagram of the video processing method architecture provided in the embodiments of this application;
[0074] Figure 8 A schematic diagram illustrating an example of the method flow of the video processing method provided in this application embodiment;
[0075] Figure 9A , Figure 9B An exemplary schematic diagram of a target image frame for determining a target person provided in an embodiment of this application;
[0076] Figure 9C An exemplary schematic diagram of target image frame filtering provided in the embodiments of this application;
[0077] Figure 10 An exemplary schematic diagram of the data flow of the video processing method provided in the embodiments of this application;
[0078] Figure 11A An exemplary schematic diagram of the data flow for extracting the voice of a single person using the video processing method provided in this application embodiment;
[0079] Figure 11B Another exemplary schematic diagram illustrating the data flow for extracting the voice of a single person using the video processing method provided in this application embodiment;
[0080] Figure 12 An exemplary schematic diagram of the electronic device interface during shooting provided in the embodiments of this application;
[0081] Figure 13A , Figure 13B Another exemplary schematic diagram of an electronic device interface provided in the embodiments of this application;
[0082] Figure 13CAnother exemplary schematic diagram of an electronic device interface provided in the embodiments of this application;
[0083] Figure 13D , Figure 13E An exemplary schematic diagram illustrating how an electronic device alters a person's voice, as provided in this application embodiment;
[0084] Figure 14A , Figure 14B An exemplary schematic diagram of the electronic device interface when playing video provided in the embodiments of this application;
[0085] Figure 14C Another exemplary schematic diagram of the video playback interface provided in the embodiments of this application;
[0086] Figure 15 An exemplary schematic diagram of the electronic device interface during a video call provided in this application embodiment;
[0087] Figure 16A , Figure 16B An exemplary schematic diagram of an audio focusing scenario for an in-vehicle voice assistant provided in an embodiment of this application;
[0088] Figure 17 An exemplary schematic diagram of the hardware structure of the electronic device 100 provided in this application embodiment;
[0089] Figure 18 An exemplary schematic diagram of the software structure of the electronic device 100 provided in the embodiments of this application;
[0090] Figure 19 Another exemplary schematic diagram of the software structure of the electronic device 100 provided in the embodiments of this application. Detailed Implementation
[0091] The terminology used in the following embodiments of this application is for the purpose of describing particular embodiments only and is not intended to be limiting of this application. As used in the specification and appended claims of this application, the singular expressions “a,” “an,” “the,” “the,” “the,” and “this” are intended to include the plural expressions as well, unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in this application refers to and includes any or all possible combinations of one or more of the listed items.
[0092] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as implying or suggesting relative importance or implicitly indicating the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature, and in the description of the embodiments of this application, unless otherwise stated, "multiple" means two or more.
[0093] First, for ease of understanding, the relevant terms and concepts involved in the embodiments of this application will be introduced below. The terminology used in the embodiments of this invention is only used to explain specific embodiments of the invention and is not intended to limit the invention.
[0094] (1) Voiceprint
[0095] Voiceprint is a behavioral characteristic of a person, and it can be used to distinguish the voices of different people.
[0096] Different people have innate differences in their vocal organs, and different people also have different acquired vocal habits, resulting in different voiceprints. That is, voiceprints can be used as a personal identifier to distinguish different people.
[0097] Voiceprints include one or more acoustic features, such as linear prediction cepstral coefficients (LPCC) and melfrequency cepstral coefficients (MFCC), etc., which are not limited here.
[0098] Considering that a person's voice also includes text information, in order to reduce the computational complexity of the device, electronic devices often require the person to speak a specific phrase when acquiring a person's voiceprint, thereby reducing the influence of text information on determining the person's voiceprint. The following example... Figure 1A , Figure 1B The following example illustrates the process of obtaining a person's voiceprint.
[0099] Figure 1A This is an exemplary schematic diagram of the interface of an electronic device for acquiring voiceprints as described in this application.
[0100] like Figure 1A As shown, the electronic device can display the text that the character needs to say on the main interface, and call audio input devices such as microphones to obtain the character's voice.
[0101] Figure 1B This is an exemplary schematic diagram illustrating the principle of voiceprint acquisition in the electronic device involved in this application.
[0102] like Figure 1BAs shown, after acquiring the voice of a person, A / D sampling is performed to obtain a digital signal. The sampling frequency is related to the frequency of the audio signal. After acquiring the digital signal, preprocessing is performed. Preprocessing generally includes pre-emphasis and frame-by-frame windowing, etc.
[0103] Pre-emphasis primarily enhances the high-frequency components of the audio signal to reduce transmission loss of high-frequency signals in space. Considering that audio signals are non-stationary but possess short-term stationary characteristics, the audio signal can be processed by framing, for example, with a frame duration of 10ms to 30ms. Each frame of the audio signal is then windowed, such as with a Hamming window or a Chebyshev window, to reduce frequency domain leakage.
[0104] After preprocessing, feature extraction is performed on the preprocessed signal. The set of extracted features is equivalent to the voiceprint of the target person.
[0105] It is worth noting that after obtaining the voiceprint of the target person, it is possible to filter out other information in the audio signal besides the target person's voice, or to determine whether the target person's voice is in the audio signal, or to extract the target person's voice from a segment of audio signal.
[0106] (2) Face detection
[0107] Face detection refers to the process of using a specific strategy to process digital signals from a given image frame to determine whether it contains a face. If a face is detected, its position, size, pose, and other features within the image frame can be further determined. Image frames do not contain audio information, but the correspondence between image frames and audio information can be established based on temporal relationships.
[0108] Face images contain a wealth of feature information, such as facial contour features, structural features (symmetry, shadows, etc.), and motion information in consecutive image frames. Based on these features, various methods can be used to achieve face detection, such as methods based on prior knowledge, template matching, and feature-based methods, which are not limited here.
[0109] The following is based on Figure 2A , Figure 2B The following example illustrates the process of face detection.
[0110] Figure 2A , Figure 2B This is an exemplary schematic diagram of the face detection involved in this application.
[0111] like Figure 2AAs shown, after obtaining the original video data, face detection can be performed on each frame of the original video.
[0112] like Figure 2B As shown, face detection is performed on the image frame to obtain the face region. Figure 2B In the middle, it is represented by a rectangular frame.
[0113] After face detection, a complete image / region containing facial information can be obtained.
[0114] After obtaining the complete image (e.g.) Figure 2B The area selected by the rectangle in the image / region can be further extracted and analyzed to extract and analyze the features of the faces contained in the image / region.
[0115] Understandably, face detection uses general features to distinguish faces from other image content, thereby extracting the image corresponding to the face from the image frame. These general features can be calculated manually or determined through methods such as artificial intelligence and deep learning.
[0116] (3) Visual semantic features
[0117] Visual semantic features, provided in the embodiments of this application, are used to link the facial information of a person in an image frame with the person's voice, thereby helping to extract the person's voice from the audio information of the video.
[0118] In this embodiment of the application, visual semantic features are specific facial features related to a person's voice and sound.
[0119] Specifically, visual semantic features include static, isolated morphological features of the mouth, cheeks, and other parts of the face when a person pronounces words (keywords, consonants, voiced sounds, etc.); and dynamic, continuous morphological features of specific mouth, cheeks, and other parts of the face when a person speaks one or more words or sentences.
[0120] In a voice recognition system, the smallest distinguishable unit can be called a phoneme. Similarly, in a visual recognition system, the lowest distinguishable unit can be called a visem.
[0121] In an image frame, facial features such as the lips, chin, and nose are visible, but vocal cords and tongue are often invisible. Therefore, a view position can be associated with multiple phonemes. The concept of a view position can be found in the relevant description in the MPEG4 standard, and will not be elaborated upon here.
[0122] It is worth noting that, considering the semantic and auditory continuity when a person speaks, the visual speech unit (VSU) can also be used as a unit in the visual recognition system.
[0123] It is worth noting that regardless of whether the person is speaking Chinese, English, or other languages, visual semantic features exist in the image corresponding to their face.
[0124] It is understandable that, considering that visual semantic features are related to the content of a person's voice, the voices of different people can be independently separated from the mixed voices of multiple people by using the visual semantic features of the person.
[0125] The following is based on Figure 3A , Figure 3B The content shown illustrates, by way of example, the process of obtaining visual semantic features provided in the embodiments of this application.
[0126] Figure 3A , Figure 3B This is an exemplary schematic diagram illustrating the acquisition of visual semantic features provided in an embodiment of this application.
[0127] like Figure 3A As shown, face detection can be performed before obtaining the visual semantic features of a certain image frame in the video.
[0128] like Figure 3B As shown, after acquiring the face in the image frame, visual semantic features can be extracted from the content and information contained in the face. In the process of visual semantic feature extraction, several relatively obvious visual features of the person can be separated and analyzed first, such as the lips, nose, and cheeks.
[0129] Taking lips as an example, different lip images can be obtained in different image frames, and then visual semantic features can be extracted.
[0130] It's worth noting that there are many methods for extracting features such as the nose, lips, and cheeks from the face, and no specific method will be used here. Taking the mouth as an example, a region within a fixed height of the face can be designated as the lip region, such as 1 / 4 to 2 / 5 of the height; alternatively, considering that the color of the lips differs from that of the cheeks, resulting in a color segmentation, edge-based methods can be used to obtain the lip region; or, considering that the shape of the lips has certain commonalities, region-based methods can be used to obtain the lip region, such as the active contour model (ACM), active shape model (ASM), and active appearance model (AAM).
[0131] Understandably, performing face detection before or during the extraction of visual semantic features, using the detected face regions as the extracted visual semantic features, reduces the amount of data for these features and thus significantly reduces computational load. Furthermore, after extracting features such as lips, nose, and cheeks from the face region, or after extracting corresponding images of lips, nose, and cheeks, visual semantic features can be extracted from these features / images.
[0132] The method for extracting visual semantic features can be found in the textual description of the decoupling network in the terminology explanation below (4), which will not be repeated here.
[0133] (4) Decoupling Network
[0134] The decoupling network is a functional module provided in the embodiments of this application for extracting visual semantic features.
[0135] The input to a decoupling network can be one or more, continuous or discontinuous image frames, or it can be a face region in that image frame, or the features corresponding to that face region, such as lips or nose. The output of the decoupling network is the visual semantic features of the person in the image information of the image frame.
[0136] The decoupling network is a pre-built functional module. The following section first introduces the construction / training process of the decoupling network by example, and then introduces the working process of the decoupling network by example.
[0137] Figure 4 This is an exemplary schematic diagram illustrating the construction / training process of the decoupled network provided in the embodiments of this application.
[0138] like Figure 4 As shown, when the input to the decoupling network can be one or more, continuous or discontinuous image frames and the corresponding audio information, the training process of the decoupling network includes:
[0139] First, multiple consecutive image frames and their corresponding audio information are input into the decoupling network. The correspondence between image frames and audio information means that the image frames and audio information are synchronized in time, i.e., audio-visual synchronization.
[0140] Secondly, the image frames are converted into visual representations; the audio information is converted into audio representations.
[0141] Among them, the representation refers to data that can be processed by electronic devices or data that is easier to process; the visual representation refers to data containing image information that can be processed by electronic devices or data that is easier to process; and the audio representation refers to data containing audio information that can be processed by electronic devices or data that is easier to process.
[0142] Understandably, converting video data into corresponding visual and audio representations reduces the computational load of subsequent processing.
[0143] It is worth noting that there are various methods to convert video data into corresponding visual and audio representations, which are not limited here. For example, the visual encoder E provided in the embodiments of this application... v Audio encoder E a The corresponding visual and audio representations can be obtained, or the visual and audio representations of the video can be obtained through various methods such as Principal Component Analysis (PCA).
[0144] Among them, visual encoder E v and audio encoder E a It can minimize the distance between the visual and audio representations of a pronunciation, a word, or a sentence in high-dimensional space. That is, the visual encoder E... v and audio encoder E a It can match audio information of a pronunciation, a word, or a sentence with image information.
[0145] Furthermore, using minimizing the distance between visual and audio representations as the optimization criterion, the parameters of the functional modules used to convert image frames into visual representations are continuously adjusted, or the parameters of both the functional modules used to convert image frames into visual representations and the functional modules used to convert audio information into audio representations are simultaneously adjusted.
[0146] There are many convex and non-convex optimization methods that can minimize the distance between visual and audio representations, and there are many loss functions to choose from (there are many ways to calculate the distance between visual and audio representations, such as L2 norm, L∞ norm, etc.), which are not limited here.
[0147] For example, the distance between visual representation and audio representation can be directly chosen as the loss function; or the output of the classifier on visual representation and audio representation can be used as the loss function (i.e., the classifier cannot effectively distinguish between visual representation and audio representation); or the output of the binary classifier on visual representation and audio representation can be used as the loss function (i.e., the binary classifier cannot effectively distinguish between visual representation and audio representation), etc.
[0148] Understandably, by minimizing the distance between the visual representation and the audio representation, and thus making full use of the changes in facial features when a person speaks different pronunciations, words, and sentences, a high-confidence visual representation that can be used for target person voice separation in subsequent processing steps is obtained.
[0149] Finally, the visual representation can be directly used as the visual semantic feature provided in the embodiments of this application; or the visual representation can be further trained so that the visual representation does not have identity information, that is, it cannot distinguish or identify people in the image frame through the visual representation, thereby obtaining the visual semantic feature.
[0150] It is worth noting that when visual representation is directly used as the visual semantic feature provided in the embodiments of this application, different decoupling networks can be used to extract the visual semantic features of the corresponding person for different people; when visual representation does not have identity information, the same decoupling network can be used to extract the visual semantic features of the corresponding person for different people.
[0151] After introducing the construction / training process of the decoupled network, the following is an example of how the decoupled network extracts visual semantic features.
[0152] First, after acquiring the video, the decoupling network segments the video to independently acquire one or more image frames and audio information.
[0153] Optionally, in some embodiments of this application, the image frame can also be processed, for example, cropped into a face detection box or into one or more images with facial features, thereby reducing the computational load.
[0154] It is worth noting that after the decoupling network is built / trained, when the electronic device uses the decoupling network, the input of the decoupling network may not include audio information.
[0155] Secondly, the initial input of the decoupled network is transformed into the same format as when the decoupled network was built / trained. The decoupled network then processes this input and transforms it into visual semantic feature output.
[0156] (5) Audio and video fusion module
[0157] The audio-video fusion module is a functional module provided in this application embodiment for separating the voice of a target person. Alternatively, the audio-video fusion module is a functional module provided in this application embodiment for separating the audio representation of the voice of a target person.
[0158] The audio-video fusion module provided in this application takes as input the audio data of the video and the visual semantic features of the video image data, performs joint processing on the audio data and visual semantic features, and outputs the voices of different people.
[0159] Optionally, in some embodiments of this application, the audio-video fusion module provided in this application can also obtain the voiceprints of different individuals by jointly processing the mixed acoustic features and visual semantic features obtained from the audio data of the video and the visual semantic features of the image data of the video. The mixed voiceprint is a mixed acoustic feature from multiple users.
[0160] Since the audio data contains the voice of the target person, the voices of other people besides the target person, and background noise, the acoustic features extracted from this audio data are mixed acoustic features.
[0161] It's worth noting that electronic devices cannot directly process audio data based on mixed voiceprints to extract the voice of a single target person from multiple people. While it's possible to extract the combined voices of multiple people using only mixed voiceprints, it's impossible to separate the voice of a single person.
[0162] However, the audio-visual fusion module provided in this application embodiment can extract and separate the voices of people in the video separately and independently by combining hybrid voiceprints and visual semantic features.
[0163] In conjunction with the above text Figure 3B The content described herein is an example of how an audio-visual fusion module separates a person's voice.
[0164] For example, in Figure 3B In the scene shown, the person on the right says "I'm sorry," and their voice overlaps with the person on the left saying "You're late."
[0165] By decoupling the network, the visual semantic features of the person on the right saying "I'm sorry" can be obtained. These visual semantic features can correspond to multiple words, including "I'm sorry," and the facial features corresponding to their pronunciations. Based on the visual semantic features, the audio signal can be processed to obtain the sound of the person on the right saying "I'm sorry."
[0166] It is worth noting that the audio and video fusion module provided in this application embodiment can not only separate Chinese, but also separate the voice of the target person in multiple languages such as English.
[0167] It's worth noting that the signal-to-noise ratio (SNR) of the "sorry" voice extracted by the audio-video fusion module may be higher or lower than that extracted from the voiceprint. Compared to simply using voiceprints to extract the target person's voice, visual semantic features provide more information. Therefore, the confidence level of visual semantic features directly affects the SNR of the extracted target person's voice.
[0168] There are many ways to build an audio-video fusion module, and no specific method is specified here.
[0169] For example, the following combination Figure 5A , Figure 5B The content shown exemplifies a method for constructing an audio-video fusion module using a Temporal Convolutional Network (TCN).
[0170] Figure 5A This is an exemplary schematic diagram of the audio and video fusion module architecture provided in the embodiments of this application.
[0171] like Figure 5A As shown, the audio-video fusion module provided in this embodiment may include: a temporal convolutional network 1, a temporal convolutional network 2, a temporal convolutional network 3, a regular convolutional unit, a modality fusion unit, an activation unit, an activation convolutional unit, and an output unit. The input to the audio-video fusion module architecture is visual semantic features and the corresponding audio information, and the output is the voice of a person, where the person is the one from whom the visual semantic features have been extracted. The correspondence between visual semantic features and audio information can be determined based on temporal relationships (in the case of audio-visual synchronization, all visual semantic features extracted from the image at any given time correspond to the audio information at that time).
[0172] The system comprises the following components: Temporal Convolutional Network 1 (TCN1) is used to acquire deep visual semantic features, which are deeper visual semantic features that describe trends over time, taking into account the temporal dependence of visual semantic features; Regularization Convolutional Unit (RCN2) is used to perform regularization and one-dimensional convolution processing on audio information; Temporal Convolutional Network 2 (TCN2) is used to acquire deep audio features, which are deeper deep audio features that describe trends over time, taking into account the temporal dependence of audio; Modality Fusion Unit (MFU) is used to connect deep visual semantic features and deep audio features and perform dimensionality transformation through a linear layer to obtain fused audiovisual features; Temporal Convolutional Network 3 (TCN3) and Activation-Convolutional Unit (ACU) are used to determine the mask value of a person's voice based on the fused audiovisual features; Activation Unit (ACU) is used to introduce non-linear effects to map the mask value; and Output Unit (ACU) is used to determine the person's voice based on the mixed acoustic features and the mask value of the person's voice.
[0173] Regularization is used to prevent the audio-video fusion module from overfitting, thereby improving the generalization ability of the audio-video fusion module; the masking value is the threshold for different sounds to overlap in the acoustic field, and determining the masking value helps to restore the masked voice of the person.
[0174] Figure 5B This is another exemplary schematic diagram of the audio and video fusion module architecture provided in the embodiments of this application.
[0175] like Figure 5B As shown, optionally, in some embodiments of this application, with Figure 5A In contrast, the audio-visual fusion module architecture takes visual semantic features and audio representations of the corresponding audio information as input, and outputs the voiceprint / audio representation of a person, where the person is the one whose visual semantic features have been extracted.
[0176] The audio representation of the audio information can be obtained through an audio encoder.
[0177] Understandably, the audio-video fusion module can separate the voice of the person corresponding to the semantic features of the video by jointly processing the semantic features of the video and the features that can represent audio information, such as audio representation and the audio information itself.
[0178] Secondly, the following combination Figure 6A , Figure 6B The content shown exemplarily describes a voiceprint-based sound separation method related to this application, combined with... Figure 6C This application provides an exemplary description of another voiceprint-based sound separation method and its combination with... Figure 6D The content shown in this application describes a text-based sound separation method and a combination thereof. Figure 7 The content shown exemplarily illustrates the video processing method provided in the embodiments of this application.
[0179] Figure 6A , Figure 6B This is an exemplary schematic diagram of a voiceprint-based sound separation method involved in this application.
[0180] like Figure 6A As shown, when at least two people are speaking, the sound waves of the different people are superimposed in space and then captured by the audio input device. The time-domain waveform of person 1's voice is shown below. Figure 6A As shown, the content of Person 1's speech is "You're late"; the time-domain waveform of Person 2's voice is as follows. Figure 6A As shown, character 2 says "I'm sorry".
[0181] like Figure 6B As shown, if the electronic device stores the voiceprint of Person 1, Person 1's voice can be separated from the mixed voices of multiple persons obtained from the audio input device.
[0182] There are many ways to separate a person's voice from a mixed voice of multiple people using an electronic device based on the voiceprint of a specific person, and no specific method is limited here. For example, template matching methods, such as dynamic time warping (DTW) and vector quantization (VQ) techniques, can be used; or probabilistic statistical models, such as hidden markov models (HMM) and Gaussian mixture models (GMM), can be used.
[0183] but, Figure 6A , Figure 6B The voiceprint-based voice separation method described in the text requires obtaining the target person's voiceprint beforehand in order to separate the target person's voice from a mixture of multiple people's voices. In other words, before using an electronic device to record video, the user needs to separately record the voice of the person being filmed to obtain their voiceprint. Clearly, obtaining the target person's voiceprint before filming increases the user's workload. Secondly, in some scenarios, electronic devices cannot obtain the target person's voiceprint in advance, such as in real-time interviews, vlog recordings, and post-production video processing.
[0184] Furthermore, since voiceprints are a form of personal privacy, storing the voiceprints of individuals other than the owner of the electronic device locally is actually insecure, just as most people don't want their fingerprints or payment passwords stored on other people's electronic devices. Moreover, storing the voiceprints of other individuals locally requires an additional set of privacy protection rules, increasing the workload for developers and the complexity of electronic devices accessing other individuals' voiceprints.
[0185] Figure 6C This is an exemplary schematic diagram of another voiceprint-based sound separation method involved in this application.
[0186] like Figure 6C As shown, in the video, there is a period of time when only one person speaks. This means that by capturing the video corresponding to this period of time and extracting the voiceprint of the person from the audio stream corresponding to the video, and then performing sound separation based on the voiceprint, the voiceprint of the person can be obtained.
[0187] For example, from second 0 to second 15, the audio stream only includes the voice of person 1 and background noise, and the image stream only shows person 1 speaking; from second 15 to second 30, the audio stream only includes the voice of person 2 and background noise, and the image stream only shows person 2 speaking.
[0188] Then, voiceprint extraction can be performed on the audio stream from second 0 to second 15 to obtain the voiceprint of person 1, and then the voice of person 1 can be extracted from the audio stream after second 30 based on the voiceprint of person 1; similarly, voiceprint extraction can be performed on the audio stream from second 15 to second 30 to obtain the voiceprint of person 2, and then the voice of person 2 can be extracted from the audio stream after second 30 based on the voiceprint of person 2.
[0189] However, it is clear that relying on data from an image stream to determine whether only one person is speaking is highly unreliable and only applicable in a few scenarios. For example, if only one person is speaking in the video image, but other people are speaking in areas not covered by the video image, the voiceprint of that person cannot be correctly extracted.
[0190] Figure 6D This is an exemplary schematic diagram of the text-based sound separation method involved in this application.
[0191] like Figure 6D As shown, when the voices of multiple people do not completely overlap in the time domain, electronic devices can identify the audio stream data of the video, obtain the text information of the audio stream, and directly separate the voices of different people based on the text information.
[0192] For example, Person 1's voice content is "that's good advice," and Person 2's voice content is "what time can we check in." Although there is some overlap in the temporal domain between Person 1's and Person 2's voice content, each word in Person 1's and Person 2's voice content is not completely overlapping. Therefore, after the electronic device acquires the mixed voice of Person 1 and Person 2, it can perform text information recognition to obtain the content of the mixed voice as "what that's good advice we check in." After obtaining the content of the mixed voice, the text information is recombined and split based on semantics to obtain text information 1 "that's good advice" and text information 2 "what time can we check in." Based on the obtained text information, the voices corresponding to "that's good advice" and "what time can we check in" can be separated from the mixed voice.
[0193] It is worth noting that, Figure 6DThe text-based sound separation method shown can also separate sounds from other languages, such as Chinese and Italian. However, it is clear that this method requires prior knowledge of the number of people in the mixed sound in order to re-splitting and recombining the text information; secondly, the text information may be split and recombined in multiple ways, and cannot be uniquely mapped to the correct text information, often requiring manual intervention.
[0194] Figure 7 This is an exemplary schematic diagram of the video processing method architecture provided in the embodiments of this application.
[0195] like Figure 7 As shown, the video processing method provided in this application embodiment does not require prior acquisition of the voiceprint of the person being filmed. Instead, it extracts visual semantic features from real-time / non-real-time video and uses these visual semantic features to non-explicitly acquire the voiceprint of the person being filmed, thereby extracting the voice of the person being filmed. This allows for better noise reduction and voice enhancement of the audio information in the video.
[0196] The video processing method provided in this application embodiment may include multiple functional modules, such as a decoupling network, an audio-video fusion module, and a speech separation module. Taking the determination of a single target person's voice by an electronic device as an example, the functions and collaboration between different functional modules are exemplarily described as follows:
[0197] (1) Electronic devices acquire real-time video stream data or non-real-time video files.
[0198] For the sake of clarity in the following discussion, Figure 7 In the corresponding content, the real-time captured video stream data and the non-real-time video file are referred to as the raw video.
[0199] (2) Synchronize the audio and video streams of the original video to determine the correspondence between the image frames and audio information of the original video. The audio-video synchronization can be determined by an electronic device based on parameters such as the frame rate and time of the image frames and parameters such as the sampling rate and time of the audio information, or it can be accomplished by the decoupled network provided in the embodiments of this application.
[0200] (3) The electronic device performs face detection on all image frames in the video stream to identify image frames containing complete faces. After identifying image frames containing complete faces, further filtering can be performed based on facial features to select image frames from which visual semantic features can be extracted by the decoupling network. For example, filtering can be performed based on the pitch angle and / or tilt angle and / or rotation angle of the face, where the pitch angle, tilt angle, and rotation angle can be determined based on facial features. Whether further filtering is needed, and the filtering conditions, can be related to the data in the training set of the decoupling network, or to the internal implementation of the decoupling network.
[0201] For the sake of clarity in the following discussion, Figure 7 In the corresponding content, the complete face is the face from which visual semantic features can be extracted.
[0202] (4) The electronic device determines the image frame corresponding to the face of the target person. The target person is either selected by the user or determined by the electronic device according to preset rules. When an image frame includes the face of the target person, the face of the target person is said to correspond to that image frame.
[0203] For example, in the following text Figure 12 , Figure 13C , Figure 14A , Figure 14B , Figure 14C , Figure 15 As shown in the corresponding text description, when an electronic device displays video image frames on the screen, after the user selects a person in the video image frame through various interactive methods such as clicking and long-pressing, the electronic device determines whether the person's face is complete, that is, whether the electronic device can extract the person's visual semantic features; if it can, the electronic device can respond to the user's interaction and determine that the person is the target person; if it cannot, the electronic device can not respond to the user's interaction, and the person is not the target person.
[0204] For example, in the following text Figure 12 , Figure 16A , Figure 16B As shown. In Figure 12 As shown, the electronic device can select a clear person in the image frame acquired after the camera's autofocus as the target person, or the electronic device can select a person appearing in the center of the shooting frame as the target person, or the electronic device can select a person occupying a large proportion of the image frame as the target person. Figure 16A , Figure 16B In this system, electronic devices can identify either the driver or the person in the passenger seat as the target. Specifically, the electronic devices can identify the target based on information such as the steering wheel and the person's position within the image frame.
[0205] (5) Electronic devices extract the visual semantic features of the target person from the image frames corresponding to the target person's face through a decoupled network.
[0206] (6) The electronic device separates the audio representation / sound / voiceprint of the target person from the audio information corresponding to the image frame corresponding to the face of the target person through the audio-video fusion module and the visual semantic features of the target person.
[0207] The output format of the audio-video fusion module can be related to its internal implementation or to the visual semantic features of the target person. Understandably, since audio representation / sound / voiceprint are different ways of describing the target person's voice, the three formats can be converted to each other. For more specific details, please refer to the textual description of the audio-video fusion module in (5) of the terminology explanation above; it will not be repeated here.
[0208] (7) After the electronic device determines the voiceprint of the target person, the electronic device extracts the voice of the target person from the audio stream of the original video through the target person's audio representation / sound / voiceprint and speech separation module.
[0209] If the audio-visual fusion module outputs the audio representation / voice of the target person, it can be converted into the voiceprint of the target person.
[0210] The voice separation module determines the voice of the target person from the audio stream of the original video. This can be done by determining the voice of the target person from the audio stream of the original video based on voiceprint; or by splicing the voice of the target person 1 determined from the audio information corresponding to other image frames besides the image frame corresponding to the target person's face, based on voiceprint, with the voice of the target person 2 determined in the audio-video fusion module, thereby determining the complete voice of the target person.
[0211] In some embodiments of this application, if all image frames in the video stream of the original video contain the face of the target person, the electronic device can obtain the voice of the target person directly through the audio-video fusion module without determining the voice of the target person through the voice separation module. That is, the voiceprint of the target person can be determined without knowing it.
[0212] It's obvious, in comparison Figure 6A , Figure 6B The present application relates to a voiceprint-based voice separation method. Firstly, the video processing method provided in this application does not require pre-saving the voiceprint of the person being filmed locally; secondly, the voiceprint extracted by the video processing method provided in this application comes from the filmed video, and this voiceprint feature is compared to… Figure 6A , Figure 6B The pre-stored voiceprints can better characterize and distinguish the features of the voices of the people being filmed in that shooting context, thereby improving parameters such as the signal-to-noise ratio and fidelity of the separated voices.
[0213] Obviously, in comparison Figure 6CThe other voiceprint-based voice separation method involved in this application, as shown in the embodiment of this application, firstly, the video processing method provided in this application extracts visual semantic features from the image stream data of the video. These visual semantic features not only represent whether the person being filmed is speaking, but also represent the content of the person being filmed speaking to a certain extent, thereby improving the accuracy of voice separation. Secondly, the video processing method provided in this application does not require that there must be a period of time in the video during which a person speaks alone.
[0214] Obviously, in comparison Figure 6D The text-based audio separation method shown in this application embodiment does not require parsing or processing of text information.
[0215] Next, the method flow for implementing the video processing method provided in the embodiments of this application will be described below by way of example.
[0216] Figure 8 This is a schematic diagram illustrating an example of the method flow of the video processing method provided in an embodiment of this application.
[0217] like Figure 8 As shown, the video processing method provided in this application includes:
[0218] S801: Get real-time / non-real-time video.
[0219] Specifically, when an electronic device is recording video, it can acquire video stream data through video input devices such as cameras / webcams and audio input devices such as microphones, and use this video stream data as the object of subsequent processing steps. Alternatively, pre-recorded videos, downloaded videos, and shared videos obtained through near-field communication services can also be used as the object of subsequent processing steps.
[0220] Specifically, in this embodiment of the application, real-time video also includes video data transmitted over the network by the electronic device on the other end in a video chat scenario.
[0221] For ease of description and to distinguish it from related terms in subsequent steps, the real-time / non-real-time video acquired electronically in step S801 will be referred to as the raw video. It is worth noting that the raw video can also be processed through editing, rendering, and other operations.
[0222] S802: Perform face detection on the image frames of the video and obtain the image frame containing the face as the target image frame.
[0223] Specifically, face detection is performed on the image frames of the original video, and the image frames in which faces are detected are selected as target image frames.
[0224] After initially determining the target image frames based on the results of face detection, the target image frames corresponding to a specific target person can be further filtered out.
[0225] Electronic devices can divide target image frames into multiple groups based on whether they include the face of a specific target person. Image frames containing a face are then grouped into multiple groups, with each group corresponding to a single target person. Once the electronic device identifies the target person, it can directly determine the group of target image frames corresponding to that person.
[0226] Alternatively, the electronic device can first identify a target person, and then determine the target image frame of that person from the image frames that detect faces. The method for the electronic device to identify the target person can be found in [reference needed]. Figure 7 The text description in (4) is not repeated here.
[0227] Figure 9A , Figure 9B This is an exemplary schematic diagram of a target image frame for determining a target person, provided in an embodiment of this application.
[0228] like Figure 9A As shown in this embodiment, image frames containing detected faces and identical faces can be used as target image frames. Specifically, after acquiring the target image frames and the faces within them, the faces are distinguished based on their features, and different faces are selected to divide the target image frames into multiple groups. Each image frame in each group of target image frames includes the face corresponding to that group, and the faces corresponding to different groups of target image frames are different.
[0229] For example, if the original video is 30 seconds long and has a frame rate of 60 FPS, then the original video consists of 1800 image frames.
[0230] First, face detection was performed on 1800 image frames, and 1500 of them were found to contain faces. Each image frame contained at least one face. For example, the 1500 image frames included 3548 face images.
[0231] Secondly, the facial images were clustered into five groups. The first group corresponds to person 1, the second group to person 2, the third group to person 3, the fourth group to person 4, and the fifth group to person 5. The first group contains 354 facial images, the second group contains 787, the third group contains 659, the fourth group contains 162, and the fifth group contains 1373.
[0232] Finally, based on the clustering results of the face images, five groups of target image frames were determined. That is, each group of target image frames corresponds to a group of face images.
[0233] For example, if the original video is 30 seconds long and has a frame rate of 60 FPS, then the original video consists of 1800 image frames.
[0234] First, face detection can be performed on 1800 image frames in chronological order. When the fifth image frame is detected, the first face is identified as face 1. Based on face 1, all image frames after the fifth image frame can be traversed to determine the image frames containing face 1 as the first group of image frames.
[0235] Secondly, similarly, face detection continues from the sixth image frame. Face 2 is detected in the eighth image frame. Based on face 2, all image frames after the eighth image frame can be traversed to determine the image frames including face 2 as the second group of image frames.
[0236] Finally, the previous steps were repeated until face detection was performed on all 1800 image frames, and five groups of target image frames were identified.
[0237] Understandably, by further dividing the target image frame into multiple groups, subsequent steps can be performed independently, thereby obtaining multiple sets of voices corresponding to different faces. Furthermore, determining multiple groups of target image frames reduces the number of image frames required for subsequent processing, thus reducing the computational load on subsequent steps.
[0238] like Figure 9B As shown in the embodiments of this application, the electronic device can first determine the target person, and then determine the target image frame of the target person from the image frames of the original video stream based on the target person's face.
[0239] For example, if the original video is 30 seconds long and has a frame rate of 60 FPS, then the original video consists of 1800 image frames.
[0240] Electronic devices can identify a target person from 1800 image frames or the first 200 image frames based on preset rules. For example, the electronic device may select the face in the center of an image frame as the target person's face, or identify the person holding a steering wheel as the target person. As another example, in video editing and playback scenarios, when the user displays frame 254 or frames 244 to 264 of the original video on the electronic device, the electronic device responds to the user's interaction and identifies the person specified by the user as the target person.
[0241] Since the electronic device can detect faces in 1,800 image frames of the original video, it can know the position of the person's face in the image frame. Therefore, it can know which person in the image frame the user specified by clicking, long-pressing, etc., and thus determine the face of the target person specified by the user.
[0242] When identifying a target person, the electronic device iterates through 1800 image frames based on the target person's face, filtering out image frames containing the target person's face. The filtered result is 787 image frames. These 787 image frames are the target image frames corresponding to the target person. Optionally, in some embodiments of this application, after performing face detection on the video image frames, further filtering can be performed to obtain image frames more suitable for determining visual semantic features, which helps reduce the computational burden of subsequent steps. The step of determining whether an image frame is suitable for extracting visual semantic features and the filtering to determine the image frames corresponding to the target person can be independent; that is, the order can be variable, or they can be performed simultaneously.
[0243] For example, facial features can be extracted from image frames containing faces, and image frames with facial features such as lips and nose that are not incomplete or missing can be selected as target image frames; or, for example, pose recognition can be performed on image frames containing faces, and image frames with the pitch angle and / or tilt angle and / or rotation angle of the face within a preset range can be selected as target image frames.
[0244] The following is based on Figure 9C The content shown exemplifies the process of selecting target image frames that are more suitable for extracting visual semantic features.
[0245] Figure 9C This is an exemplary schematic diagram of target image frame filtering provided in an embodiment of this application.
[0246] like Figure 9C As shown, the electronic device first performs face detection on the image information in the image frames of the video, and performs a first screening to obtain image frames containing faces. Secondly, the electronic device obtains the pitch angle and / or tilt angle and / or rotation angle of the face, and further filters to obtain target image frames more suitable for extracting video semantic features in step S803.
[0247] Understandably, when the electronic device executes step S802, it filters the image frames to identify those containing faces, and can further filter these face-containing image frames. Filtering image frames effectively reduces the number of image frames that need to be processed in subsequent steps, thus reducing the computational load on those steps.
[0248] S803: Extract visual semantic features of the target person from the target image frame.
[0249] Specifically, after the electronic device executes step S802, it has determined that there is one or more sets of target image frames. The electronic device can then independently extract the visual semantic features of the one or more sets of target image frames to obtain the visual semantic features corresponding to different people or different groups of faces.
[0250] More specifically, for any set of target image frames, consecutive image frames can be selected, meaning temporally discontinuous (isolated) image frames are not selected as the objects for extracting visual semantic features. For example, multiple consecutive target image frames with a duration greater than 1.5 seconds can be selected as the objects for extracting visual semantic features. The duration is related to the video's image sampling rate.
[0251] The one or more, continuous or discontinuous image frames may contain different subjects.
[0252] The terms for concepts such as visual semantic features and extracting visual semantic features can be found in the textual descriptions of (3) visual semantic features and (4) decoupling network in the terminology explanation, and will not be repeated here.
[0253] It is worth noting that the electronic device can repeat step S803 multiple times to extract the visual semantic features of different people.
[0254] S804: Extract the voice of the target person from the audio corresponding to the target image frame based on visual semantic features, and determine the voiceprint of the target person.
[0255] Specifically, after determining the visual semantic features of the target image frame in step S803, the electronic device extracts the voice of the target person from the audio information corresponding to the target image frame based on the visual semantic features. The target person is the person corresponding to the target image frame, and this correspondence can be determined according to step S802.
[0256] The audio corresponding to the target image frame can also be called the extracted audio.
[0257] Alternatively, electronic devices can extract the audio representation of a target person from mixed acoustic features based on visual semantic features. An audio encoder can then decode this audio representation of the target person's voice to obtain the target person's voice.
[0258] After obtaining the target person's voice, the person's voiceprint can be determined based on the target person's voice.
[0259] Among them, the audio and video fusion module provided in the embodiments of this application can be used to extract the voice of a single person from the target image frame. For details, please refer to the text description in the terminology explanation (5) audio and video fusion module, which will not be repeated here.
[0260] The correspondence between target image frames and audio information can be determined by parameters such as image frame rate and audio sampling rate, which are not limited here. When the electronic device achieves audio-visual synchronization, the correspondence between image frames and audio information can be directly determined based on time information. For example, if the original video is 30 seconds long and has a frame rate of 60 FPS, then the original video includes 1800 image frames. If the audio information sampling rate of the original video is 8000 Hz, then the sampled digital audio information consists of 64000 points. The first image frame can correspond to the 1st to the 1067th point of the digital audio information. If the audio and video of the original video are not synchronized, audio-visual synchronization of the original video can be performed in step S801. Optionally, in some embodiments of this application, when the target image frames selected in step S802 can form multiple discontinuous video segments, in step S804, after extracting the voice of a person in each video segment, the electronic device can acquire multiple segments of the person's voice. Electronic devices can select high-quality human voices based on the results of voice activity detection (VAD), thereby extracting more accurate voiceprints. In each video segment, the voice of the person being filmed is that of the same individual.
[0261] It's worth noting that electronic devices can select one or more (less than M) segments of a person's voice from multiple segments (e.g., M segments) using other parameters such as signal-to-noise ratio, sound duration, or other digital processing methods, and extract the person's voiceprint based on that segment. For example, a signal-to-noise ratio greater than SNR can be selected. threshold And the duration is greater than T threshold The voice clips of the characters are selected as high-quality voice clips, and the voiceprints of these high-quality voice clips are extracted. Among these, SNR... threshold To preset the signal-to-noise ratio threshold, T threshold This is the duration threshold.
[0262] S805: Extracts the voice of the target person from the original video based on the target person's voiceprint.
[0263] Specifically, after obtaining the voiceprints of different individuals, the voices of different individuals in the original video can be extracted to achieve functions such as noise reduction, voice changing, and voice enhancement / reduction. Among these, voice changing can alter the pitch, loudness, timbre, and other characteristics of a person's voice. Electronic devices can achieve voice changing by altering the time domain, frequency domain, and time-frequency domain characteristics of the sound.
[0264] Similarly, by repeating steps S803 and S804, voiceprints of multiple different people can be obtained. After obtaining the voiceprints of different people, the voices of different people in the original video can be extracted.
[0265] Optionally, in some embodiments of this application, the voice of a person can be extracted from the audio stream corresponding to other image frames besides the target image frame based on the voiceprint, and then spliced with the voice of the person extracted in step 804 to obtain the complete voice of the person.
[0266] The following is combined with Figure 10 The content shown is an exemplary introduction. Figure 8 The video processing method provided in the embodiments of this application is shown.
[0267] Figure 10 This is an exemplary schematic diagram of the data flow of the video processing method provided in the embodiments of this application.
[0268] like Figure 10 As shown, corresponding to step S801, after obtaining the original video, the extraction of target image frames can begin. The audio information of the original video includes the voices of N people, where N is a positive integer.
[0269] Corresponding to step S802, the electronic device can determine target image frames including the face of person 1, such as image frames from 1s to 5s and image frames from 13s to 16s. The electronic device can also determine target image frames including the face of person 2, such as image frames from 3s to 4s and image frames from 20s to 45s.
[0270] Corresponding to step S802, when the electronic device extracts target image frames, some people may not be the subjects being photographed, or some people may not have corresponding image frames with extractable visual semantic features. For example, person N may be a person outside the video image. Since the video image frames do not include this person, person N's face is not detected during face detection. Or, for example, the person may be facing away from the camera of the electronic device or have an obstruction on their face, so person N's face is not detected during face detection. Or, for example, due to limitations such as shooting angle and lighting, although person N's face is detected during face detection, the rotation angle and pitch angle of person N's face do not meet the conditions, resulting in the inability to extract features from person N's face.
[0271] Corresponding to steps S803 and S804, the electronic device can extract the visual semantic features of different people, thereby obtaining the voices of different people, and further obtaining the voiceprints of different people. For example, the electronic device extracts the visual semantic features of person 1 from the target image frame of person 1's face, then extracts the voice of person 1 in the corresponding audio of the target image frame, and then obtains the voiceprint of person 1.
[0272] The following is combined with Figure 11A , Figure 11BThe content shown exemplarily illustrates the process of extracting the voice of a single person using the video processing method provided in the embodiments of this application.
[0273] Figure 11A This is an exemplary schematic diagram of the data flow for extracting the voice of a single person using the video processing method provided in this application embodiment.
[0274] like Figure 11A As shown, corresponding to steps S801 and S802, after acquiring the original video, the electronic device performs audio and image segmentation on the video to obtain an audio stream and an image stream, respectively. Face detection is then performed on the image frames in the video stream to obtain a set of target image frames. Each of the target image frames contains the face of the same person.
[0275] Corresponding to step S803, after acquiring the target image frame of the target person, the electronic device performs visual semantic feature extraction and audio segmentation corresponding to the target image frame. After extracting the visual semantic features from the target image frame, it combines the segmented audio corresponding to the target image frame to perform sound separation, thereby obtaining the voice of the specified person.
[0276] Corresponding to step S804, after obtaining the voice of the person, the voiceprint of the specified person can be determined from the voice, and the audio stream of the video file can be processed to obtain all the voices of the target person in the video file.
[0277] After obtaining all the voices of the target person, a series of conventional audio processing techniques can be applied to the voice, such as voice changing and reverb, and then combined with the image stream of the original video to create a new video.
[0278] Figure 11B This is another exemplary schematic diagram of the data flow for extracting the voice of a single person using the video processing method provided in the embodiments of this application.
[0279] Figure 11B The content shown is mostly related to Figure 11A The content shown is the same, so the same content will not be repeated.
[0280] Figure 11B The content shown is consistent with Figure 11A The difference is that, after extracting the voiceprint of the target person, voiceprint-based sound recognition is performed on the audio other than the audio corresponding to the target image frame to obtain the target person's voice. Then, the voice of the person separated by visual semantic features is concatenated with the voice of the target person extracted by voiceprint-based sound recognition to obtain the complete voice of the target person.
[0281] Understandably, in comparison Figure 11B The content shown is consistent with Figure 11AAs shown, after obtaining the voiceprint of the target person, whether to process all the audio data of the original video to obtain the entire voice of the target person, or to splice the separately obtained voices, depends on the quality (signal-to-noise ratio, duration) of the target person's voice separated from the audio information corresponding to the target image frame by the electronic device. If the signal-to-noise ratio of the target person's voice separated from the audio information corresponding to the target image frame by the electronic device is high and the duration is long, then the following can be chosen: Figure 11B The method shown obtains the complete voice of the target person by audio splicing.
[0282] After implementing the video processing method provided in this application embodiment, the electronic device can determine the voice of a person in the video, that is, separate the voice of a certain person from the original video. Based on the voice of that person, the electronic device can filter out the voice of that person from the audio data of the original video to obtain the audio data of the original video that does not contain the voice of that person.
[0283] There are many ways to filter out the voice of a person after it has been acquired from the audio data of the original video. For example, it can be done using various digital signal processing methods such as Least Mean Square (LMS) filtering and Wiener filtering. No specific method is specified here.
[0284] When an electronic device separates the voice of a person from the original video that has been enhanced, weakened, or voice-modified, it can enhance, weaken, or modify the voice of that person separately, and then re-overlay and splice it with the audio data of the original video that does not contain the voice of that person.
[0285] Alternatively, when the electronic device extracts the voice of a person from the original video to enhance or weaken it, it performs weakening or enhancement processing on the audio data of the original video that does not contain the person's voice, and then superimposes it with the extracted voice of the person from the original video. Alternatively, since the electronic device can determine the voiceprint of a person when extracting their voice from the original video, for example in step S804, it can process the audio data of the original video based on the person's voiceprint to directly enhance, weaken, or change the person's voice.
[0286] Next, the following provides exemplary scenarios of an electronic device executing the video processing method provided in the embodiments of this application, as well as interface diagrams presented by the electronic device in different scenarios.
[0287] The video processing method provided in this application includes various scenarios such as user shooting videos, user editing videos, user watching videos, user video calls, and audio focusing by in-vehicle voice assistants, and is not limited to these scenarios.
[0288] First, the filming location is described.
[0289] Figure 12 This is an exemplary schematic diagram of the electronic device interface during shooting, provided in an embodiment of this application.
[0290] like Figure 12 As shown, when the electronic device is recording video, it can execute the video processing method provided in this application embodiment to process the real-time video stream during the recording process. Furthermore, the electronic device can display control 1201 to inform the user that the video is being processed in real time.
[0291] When the electronic device starts recording video, it continuously acquires video image frames and audio information. Furthermore, the electronic device can perform face detection on the acquired image frames to determine if a complete human face exists. When the electronic device determines that the image frame contains a complete human face, or when the electronic device determines that the visual semantic features of the person in the image frame can be extracted, the electronic device can display control 1201 on the recording interface.
[0292] Control 1201 can be displayed permanently or flashing. The content displayed in control 1201 can be "Voice Enhancement," "Automatic Noise Reduction," etc. The control can be implemented using components such as Button, EditText, ListView, and RecyclerView (taking Android as an example).
[0293] In response to a user clicking or long-pressing control 1201, the electronic device can determine the voiceprint of the person based on their visual semantic features, and then determine the person's voice in the video in subsequent time periods based on the voiceprint. Furthermore, when other people appear in the image frames acquired by the electronic device, the voice of those other people can also be determined using the same method.
[0294] It's worth noting that electronic devices can determine the voice of any person by either extracting the voiceprint from the video based on visual semantic features, or by determining parts of the person's voice based on both visual semantic features and the voiceprint, and then splicing them together; or, when all image frames include a complete face, the person's voice can be determined based on visual semantic features. For more detailed descriptions, please refer to... Figure 7 , Figure 11A and Figure 11B The corresponding textual descriptions will not be repeated here. The following section, using the shooting scenario as an example, specifically illustrates how an electronic device implements the video processing method provided in this application's embodiments:
[0295] (1): Corresponding to steps S801 and S802, the user can configure the shooting settings before and during video recording, so that the electronic device can start implementing the video processing method provided in the embodiments of this application. Alternatively, when the electronic device determines that the image frame of the real-time video includes a complete face, it displays an interactive control 1201 and determines whether to start implementing the video processing method provided in the embodiments of this application in response to the user's interaction.
[0296] This corresponds to an electronic device continuously performing face detection on newly added image frames after acquiring real-time image and audio information, and classifying the acquired and newly added image frames into one or more groups of target image frames. Different groups of target image frames correspond to the faces of different target individuals.
[0297] (2): Corresponding to steps S803 and S804, the electronic device independently extracts multiple sets of visual semantic features from multiple sets of target image frames, and independently extracts the voice of the target person based on each set of visual semantic features. After acquiring the voice of any target person, the electronic device can extract the voiceprint of that target person.
[0298] (3): Corresponding to step S805, but with some differences, since the electronic device does not know at the current moment whether the image frames of the video at subsequent moments contain the face of the target person, after determining the voiceprint of the target person, the electronic device can extract the voice of the target person based on the voiceprint when processing the video at subsequent moments, or it can determine the visual semantic features of the target person in the image frames of the video at subsequent moments, and then extract the voice of the target person based on the visual semantic features.
[0299] (4): After the electronic device determines the voice of the target person, it can save the voice of the target person and the video taken locally independently.
[0300] Similarly, when a user acts as a broadcaster, the electronic device acquires sound including the user's voice and other sounds. The electronic device applies the video processing method provided in this application to the real-time captured video, amplifies the user's voice, and then sends the processed video data to a server for rebroadcasting. This allows viewers to hear the broadcaster's voice or the voice of a person designated by the broadcaster more clearly.
[0301] Optionally, after the electronic device implements the video processing method provided in the embodiments of this application, it determines that the processed video data may only include image data and the voice of the target person.
[0302] Secondly, it introduces the scenarios of video editing.
[0303] Figure 13A , Figure 13B Another exemplary schematic diagram of an electronic device interface provided in an embodiment of this application.
[0304] like Figure 13A As shown, the video subpage of the electronic device's gallery displays multiple videos, including a video control for "Video 1: People," "Video 2: Ferris Wheel," "Video 3: Beach," and "Video 4: Park." Each video corresponds to an interactive control used to respond to user clicks. In response to the user clicking the control corresponding to a video, the electronic device determines the video selected by the user.
[0305] For example, in response to a user clicking control 1301, the electronic device determines that the video selected by the user is "Video 1: People" and redirects to the video playback interface. The video playback interface is as follows: Figure 13B As shown. In response to the user clicking the control 1301, the electronic device can determine that "Video 1: Person" is the original video in step S801.
[0306] like Figure 13B As shown, the video playback interface displayed on the electronic device includes a video option control 1302. This video option control may include: a send control, an edit control 1303, a delete control, and more controls.
[0307] In response to the user clicking the editing control 1303, the electronic device can determine that the video is the original video from step S801 and redirect to the video editing page, as shown below. Figure 13C As shown.
[0308] Figure 13C Another exemplary schematic diagram of an electronic device interface provided in an embodiment of this application.
[0309] like Figure 13C As shown, the electronic device displays a video editing toolbar 1304 on the video editing page. The video toolbar includes multiple video editing sub-controls, such as filter template controls, clipping controls, background music controls, and voice processing controls 1305.
[0310] It is worth noting that the original video in step S801 has been determined before the electronic device displays the voice processing control 1305.
[0311] In response to the user clicking the voice processing control 1305, the electronic device first performs face recognition on all image frames of "Video 1: People" and determines one or more sets of target image frames.
[0312] If the electronic device performs face recognition on all image frames of "Video 1: People" and determines that the video does not contain a face, then the video sub-toolbar 1306 will not be displayed.
[0313] When an electronic device identifies one or more target image frames, it displays the video sub-toolbar 1306. The video sub-toolbar 1204 may include one or more sets of controls, each used to process the voices of different people in the video. That is, the number of controls included in the video sub-toolbar may be related to the number of people in the video, or to the number of separable voices of different people.
[0314] More specifically, the electronic device can traverse the image frames in the video and determine the number of people in the video based on face detection. Specifically, when an image frame in the video includes a person's face, that person is considered part of the video. In this case, the number of controls in the video sub-toolbar can be related to the number of people in the video. Alternatively, after determining the number of people in the video, the electronic device can identify target image frames for different people. After extracting visual semantic features from the target image frames of different people, it can separate the voices of those different people from the corresponding audio information of the target image frames, and determine the voiceprints of different people based on their voiceprints. Here, the number of voiceprints determined by the electronic device is consistent with the number of separable voices of people. In this case, the number of controls in the video sub-toolbar is related to the number of separable voices of people.
[0315] For example, after the electronic device performs face detection on the video, the result of clustering the target image frames is two groups of target image frames. That is, when the electronic device determines that the number of people in the video is 2, the video sub-toolbar 1204 can include 2 sets of controls. The first set of controls includes controls 1307, 1308, 1309, and 1310, and the second set of controls includes controls 1311, 1312, 1313, and 1314.
[0316] The process of the electronic device determining the target image frame can be referred to the textual description of step S802 above, and will not be repeated here.
[0317] Control 1307 can display a person's face image or other person identifiers to inform the user that controls 1308, 1309, and 1310 are functional controls that operate on the voice of the person displayed by control 1307. The content displayed by control 1307 originates from the person's face determined by the electronic device during face recognition of the video image frames, or other identifiers that can distinguish different people. For example, as mentioned above... Figure 2A , Figure 2B as well as Figure 2A , Figure 2B As the corresponding text description shows, the content displayed by control 1305 can be an image corresponding to the face region of a person.
[0318] The controls can display text to indicate their function to the user. For example, control 1308 displays the text "Enhancement", control 1309 displays the text "Voice Changer", and control 1310 displays the text "Focus".
[0319] For ease of description, the person displayed by control 1306 will be referred to as Person 1. Then, control 1308 can change only the volume of Person 1's voice in the video; control 1309 can change only the voice of Person 1 in the video; control 1310 can retain only the voice of Person 1 in the video, that is, only Person 1's voice is heard when the video is played.
[0320] Similarly, let's call the person displayed by control 1311 "person 2". Then, control 1312 can change only the volume of person 2's voice in the video; control 1313 can change only the voice of person 2 in the video; control 1314 can retain only the voice of person 2 in the video, that is, only the voice of person 2 will be heard when the modified video is played.
[0321] If the user selects controls 1310 and 1314, only the audio of characters 1 and 2 will play when the video is played. Controls 1310 and 1314 are selected when clicked by the user; they are deselected when clicked again.
[0322] Although the electronic device has already determined one or more sets of target image frames, the number of people in the video, or the separable voices of people in the video before displaying the video sub-toolbar 1306, the timing of the electronic device determining the voices of people from the video can be categorized into several situations, including:
[0323] Firstly, in response to a user clicking the voice processing control 1305, the electronic device can extract the visual semantic features of the target person corresponding to each set of target image frames based on a decoupled network. Then, based on the visual semantic features of the target person and the audio-visual fusion module, it extracts the voice of the target person from the audio information corresponding to the target image frame, thereby obtaining the voiceprint of the target person. After obtaining the voiceprint of the target person, the electronic device determines the voice of the target person based on the voiceprint and the audio information of the video.
[0324] Secondly, taking person 1 as an example, in response to the user clicking control 1308, control 1309, and control 1310, the electronic device determines the target image frame corresponding to person 1 from one or more sets of target image frames. And as described in the first part above, it determines the sound of person 1.
[0325] Among them, in response to the user clicking control 1308, the electronic device changes the voice of character 1 in the video, including:
[0326] Figure 13D , Figure 13E This is an exemplary schematic diagram illustrating how an electronic device can change a person's voice, as provided in an embodiment of this application.
[0327] The electronic device has determined the voice of character 1 before the user clicks control 1308; or, in response to the user clicking control 1308, the electronic device determines the voice of character 1.
[0328] After confirming the voice of character 1, and after the electronic device receives the user's click control 1308, such as Figure 13D As shown, when the electronic device plays a video, the audio played is the original video sound plus the voice of character 1 * coefficient. This coefficient is related to the user's settings and can be positive or negative, but cannot be 0.
[0329] Or, such as Figure 13E As shown, after determining the voice of Person 1, and after the electronic device receives the user's click control 1308, the electronic device first removes Person 1's voice from the original video audio based on Person 1's voice / voiceprint, obtaining the original video audio excluding Person 1's voice. When the electronic device plays the video, the played audio is the video audio excluding Person 1's voice and Person 1's voice * coefficient.
[0330] When the user clicks controls 1309, 1310, and 1311, the electronic device changes the voice of person 1 in the video. For details, please refer to the text description of the electronic device changing the voice of person 1 in the video after the user clicks control 1308. It will not be repeated here.
[0331] Finally, users can save their modifications to the video and generate a new video.
[0332] The following describes, in the context of video editing, a specific example of an electronic device implementing the video processing method provided in this application embodiment:
[0333] (1) Corresponding to step S801, in response to the user clicking control 1301, the electronic device determines that the video selected by the user is "Video 1: People", that is, "Video 1: People" is the original video.
[0334] (2) Corresponding to step S802, when the user is editing the video, in response to the user clicking control 1303 and before displaying control 1304, the electronic device may have already executed step S802, that is, the electronic device performs face detection on the image frames of the video and determines N sets of target image frames, that is, determines that the number of people included in the video is N, then the electronic device can display N sets of controls in the video sub-toolbar 1306.
[0335] Alternatively, in response to the user clicking the control 1305, the electronic device begins to execute step S802, which is that the electronic device needs to perform face detection on the image frames of the video and determine that the number of people included in the video is N.
[0336] For example, if N=2, the electronic device can display a person's face or a person's identifier on controls 1205 and 1209.
[0337] (3) Corresponding to steps S803, S804, and S805, there are several cases:
[0338] First, in response to a user clicking control 1301 or 1305, the electronic device independently extracts the visual semantic features of N individuals from N sets of target image frames based on a decoupled network. After obtaining the visual semantic features of the N individuals, it independently determines the voices of the N individuals from the audio information corresponding to the N sets of target image frames, and further determines the voiceprints of the N individuals.
[0339] In response to the user clicking control 1308, the electronic device determines that the user-specified person is Person 1, that is, the target person is Person 1. The electronic device determines the voiceprint of Person 1 from the voiceprints of N people, and separates the voice of Person 1 from the audio information of the original video.
[0340] After determining the voice of character 1, you can refer to... Figure 13D , Figure 13E As shown, when the electronic device plays the audio of "Video 1: Person", the sound of Person 1 becomes louder or softer compared to the original sound of "Video 1: Person".
[0341] Secondly, similar to the first method, in response to a user clicking control 1301 or 1305, the electronic device determines the voiceprints of N individuals. However, unlike the first method, the electronic device does not need to wait for a response to a user clicking control 1308, but continues to independently determine the voices of the N individuals based on their voiceprints.
[0342] In response to the user clicking control 1308, the electronic device determines that the user-specified character is Character 1, that is, the target character is Character 1. The electronic device determines the voice of Character 1 from the voices of N characters.
[0343] After determining the voice of character 1, you can refer to... Figure 13D , Figure 13E As shown, when the electronic device plays the audio of "Video 1: Person", the sound of Person 1 becomes louder or softer compared to the original sound of "Video 1: Person".
[0344] Third, in response to the user clicking control 1308, the electronic device first determines that the user-specified person is Person 1, that is, the target person is Person 1. The electronic device determines the target image frame corresponding to Person 1 from N groups of target image frames.
[0345] After identifying the target image frame corresponding to Person 1, the electronic device uses a decoupling network to determine the visual semantic features of Person 1 from the target image frame. Following the acquisition of the visual semantic features, the device determines Person 1's voice from the audio information corresponding to the target image frame and further determines Person 1's voiceprint. Based on Person 1's voiceprint, the electronic device separates Person 1's voice from the audio information of the original video.
[0346] After determining the voice of character 1, you can refer to... Figure 13D , Figure 13E As shown, when the electronic device plays the audio of "Video 1: Person", the sound of Person 1 becomes louder or softer compared to the original sound of "Video 1: Person".
[0347] It's worth noting that after an electronic device determines a person's voice from the audio information corresponding to the target image frame, there are multiple methods to extract the complete voice of that person from the audio information of the original video. For details, please refer to... Figure 11A and Figure 11B The textual descriptions in the text will not be repeated here.
[0348] Thirdly, it introduces the scenarios in which the video is played.
[0349] Figure 14A , Figure 14B This is an exemplary schematic diagram of the interface of an electronic device when playing video, as provided in an embodiment of this application.
[0350] like Figure 14A As shown, when an electronic device plays videos, such as TV series, movies, short videos, self-shot videos, online course videos, etc., the electronic device displays the video toolbar 1401 in response to various interaction methods such as clicking, long pressing, and specific gestures.
[0351] The video toolbar 1401, or its sub-toolbars, may include one or more controls. Users can click these controls through various interactive methods to personalize the voices of people in the video image, such as enhancing or changing their voices. Alternatively, the electronic device can play only the voices of user-specified individuals. The number of controls in the sub-toolbars of the video toolbar may be related to the number of separable human voices determined by the electronic device in the video, or to the number of faces detected by the electronic device from the images in the video data. The electronic device can determine the number of people / faces included in the images in the video by executing step S802.
[0352] For example, the video toolbar's sub-toolbar can include eight controls: control 1402, control 1403, control 1404, control 1405, control 1406, control 1407, control 1408, and control 1409. When the user clicks control 1402, the electronic device can amplify the voice of character 1, making it clearer for the user to hear. When the user clicks control 1403, the electronic device can change character 1's voice to other timbres, such as changing it to a cartoonish or crisp voice. When the user clicks control 1404, the electronic device will only play character 1's voice. When the user clicks control 1405, the electronic device can perform further processing on character 1's voice before playback.
[0353] Similarly, when the user clicks control 1406, the electronic device can amplify the voice of character 2, making it clearer for the user to hear. When the user clicks control 1407, the electronic device can change the voice of character 2 to other timbres, such as changing it to a cartoonish voice or a crisp voice. When the user clicks control 1408, the electronic device only plays the voice of character 1. When the user clicks control 1409, the electronic device can perform further processing on the voice of character 2 before playback.
[0354] The controls can display text to indicate their function to the user. For example, control 1402 displays the text "Enhancement," control 1403 displays the text "Voice Changer," and control 1404 displays the text "Focus," etc. Figure 14B As shown, the number of controls in the video toolbar 1401 can be independent of the number of people / faces in the video. The video toolbar 1401 includes controls 1410, 1411, 1412, and 1413.
[0355] When the user clicks control 1410, the electronic device can amplify the voices of all characters in the video, making them clearer to the user. When the user clicks control 1411, the electronic device can change the voices of all characters in the video to other timbres, such as changing them to cartoon-style voices or crisp voices. When the user clicks control 1412, the electronic device plays only the voices of all characters in the video. When the user clicks control 1413, the electronic device can perform further processing on the voices of all characters in the video before playback.
[0356] Figure 14C This is another exemplary schematic diagram of the interface for playing videos provided in the embodiments of this application.
[0357] As shown in the figure, when an electronic device plays videos, such as TV series, movies, short videos, self-shot videos, online courses, etc., the user can quickly adjust the voice of the person in the video by clicking on them.
[0358] For example, when a user clicks on the image of person 2 on the right side of the video image, the electronic device can display the controls of the video toolbar 1401, including: control 1420, control 1421, control 1422, control 1423, and control 1424. Among them, the video toolbar 1420 can float above the image of person 2.
[0359] Users can click control 1420 to increase the volume of character 2's voice; users can click control 1421 to decrease the volume of character 2's voice; users can click control 1422 to change character 2's voice to a different timbre; or users can click control 1423 to make the electronic device play only character 2's voice; or users can click control 1424 to perform more selective processing on character 2's voice.
[0360] Obviously, when a user watches an online course, and the electronic device implements the video processing method provided in this application embodiment, the user can selectively increase the volume of the instructor's voice or play only the instructor's voice to improve the user experience.
[0361] The following describes, in the context of video playback, the details of how an electronic device implements the video processing method provided in this application embodiment:
[0362] (1) Corresponding to step S801, before the electronic device plays the video, the electronic device needs to acquire the video, which is the original video.
[0363] (2) Corresponding to step S802, the electronic device can perform face detection on the image frames of the video to determine the number of people and their positions in the image frames of the video. The electronic device can determine the number of controls in the video toolbar 1401 based on the number of people.
[0364] In response to the user's operation of playing a video, the electronic device performs face detection on all image frames of the video and identifies one or more groups of target image frames.
[0365] There are some differences from step S802, in Figure 14C In the scenario shown, the target person is the person specified by the user. Since the electronic device can determine the position of the person in the video's image frames based on face detection, it can determine that the user selected the specified person through various interactive methods such as clicking. This specified person is the target person. Furthermore, the electronic device only needs to determine the set of target image frames corresponding to this target person from multiple sets of target image frames.
[0366] (3) Corresponding to steps S803, S804, and S805, which are similar to the scenarios in video editing, there are also several different situations:
[0367] First, in response to a user's video playback operation, the electronic device determines one or more sets of target image frames, directly determines the visual semantic features of people in the video, then determines the voiceprints of people in the video, and finally determines the voices of people in the video.
[0368] Secondly, in response to the user clicking the control in 1401 of the video toolbar, the electronic device first determines the person specified by the user as the target person, and then determines the corresponding target image frame of the target person from one or more sets of target image frames. The electronic device determines the visual semantic features of the target person from the target image frame corresponding to the target person, then determines the voiceprint of the target person, and finally determines the voice of the target person from the audio information of the video.
[0369] For more specific details, please refer to the textual description of the video editing scenario above, which will not be repeated here.
[0370] Fourth, it introduces the scenarios during video calls.
[0371] Figure 15 This is an exemplary schematic diagram of the electronic device interface during a video call, provided in an embodiment of this application.
[0372] like Figure 15 As shown, during a video call, a local user can use methods such as... Figure 14A , Figure 14CThe interaction method shown allows you to select a person in the image (such as the other user) and then quickly adjust that person's voice.
[0373] During a video call, if the other end of the electronic device is in a noisy environment, such as when someone is playing music, the user may not be able to hear the other end's voice clearly. Alternatively, if the other end's electronic device captures multiple people in the video, and the local user wants to hear a specific person's voice more clearly, the electronic device can implement the video processing method provided in this application to separate the voice of the person (such as the other end user) from the video, and then perform enhancement, noise reduction, and other processing to allow the user to hear the voice of the target person clearly, thereby improving the user experience. Alternatively, the user can choose to play only their own voice.
[0374] For example, in response to a user clicking on a person in a video call, the electronic device displays a video toolbar 1501. The video toolbar 1501 includes multiple controls, such as controls 1502, 1503, 1504, 1505, and 1506.
[0375] The following describes, using a video call scenario as an example, the details of how an electronic device implements the video processing method provided in this application embodiment:
[0376] (1) Corresponding to step S801, after the video call starts, the local electronic device can save the video transmitted by the other electronic device locally. This video is the original video, and the number of image frames in the video continues to increase.
[0377] (2) Corresponding to step S802, the electronic device can perform face detection on the image frames of real-time video and can determine one or more sets of target image frames.
[0378] (3) Corresponding to steps S803 and S804, at time 1, in response to the user clicking on a person in the video call, the electronic device determines the specified person selected by the user and takes the specified person as the target person, and then determines a set of target image frames corresponding to the target person from multiple sets of target image frames.
[0379] After identifying the target image frame of the target person, the electronic device determines the visual semantic features of the person, and then determines the person's voiceprint.
[0380] (4) Corresponding to step S805, but different from step S805, after obtaining the user's voiceprint, the voice of the target person in the video after time 1 can be determined based on the voiceprint, instead of determining the voice of the target person from the local video based on the voiceprint.
[0381] For more specific details, please refer to the textual description of the video editing scenario above, which will not be repeated here.
[0382] Fifth, the in-vehicle voice assistant scenario is introduced.
[0383] Figure 16A , Figure 16B This is an exemplary schematic diagram of an audio focusing scenario for an in-vehicle voice assistant provided in an embodiment of this application.
[0384] like Figure 16A As shown, the vehicle is equipped with audio input devices (such as microphones) and audio output devices (such as cameras). With multiple users inside the vehicle, when multiple users are speaking, the voice assistant may be unable to capture the target user's voice and thus fail to respond to their commands; or, when multiple users are speaking, the superposition of sound waves within the vehicle's interior can reduce the voice assistant's voice recognition accuracy, leading to incorrect responses.
[0385] In this scenario, the vehicle's infotainment system (such as HiCar) can implement the video processing method provided in this application embodiment to extract the voice of the person driving the vehicle from the video, allowing the voice assistant to respond solely to the voice of the person driving the vehicle. Alternatively, for a training vehicle, the voice of the person in the passenger seat can be extracted from the video, allowing the voice assistant to respond solely to the voice of the person in the passenger seat.
[0386] like Figure 16A As shown, during normal driving, the vehicle's infotainment system can identify the person in the driver's seat as Person 1 based on image information from video frame images. The system can then determine Person 1's voice from the audio information corresponding to the target image frame based on Person 1's visual semantic features, and obtain Person 1's voiceprint. The electronic device then determines Person 1's voice in subsequent videos based on Person 1's voiceprint and the visual semantic features of Person 1 in subsequent target image frames.
[0387] After acquiring the voice of Person 1 in real time, the vehicle's infotainment system can transmit Person 1's voice to the voice assistant in the infotainment system in real time.
[0388] Understandably, since the voice of Person 1 transmitted to the voice assistant by the vehicle's infotainment system is free from interference from other people's voices, the voice assistant can improve the accuracy of voice recognition.
[0389] The following section provides specific examples of how electronic devices implement the video processing method provided in this application embodiment, using an in-vehicle voice assistant scenario as an example:
[0390] (1) Corresponding to step S801, the electronic device obtains the real-time video data as the original video.
[0391] (2) Corresponding to step S802, but different from step S802, the electronic device can directly determine the person in the driver's seat as person 1 based on the image frames of the video, and then obtain a set of target image frames corresponding to person 1. Furthermore, for image frames acquired by the vehicle in real time, it can directly determine whether the image frame is a target image frame.
[0392] It is worth noting that, in most cases, considering the relatively fixed structure of the vehicle's interior (driver's seat, passenger's seat) and the position of the camera on the vehicle's infotainment system, all image frames can be considered as target image frames.
[0393] (3) Corresponding to steps S803, S804, and S805, the electronic device can determine the user's visual semantic features from the target image frame and determine the user's voice based on the user's visual semantic features.
[0394] Understandably, in the scenario of an in-vehicle voice assistant, in most cases, all image frames include the complete face of Person 1, so Person 1's voice can be directly determined based on the person's visual semantic features.
[0395] For more specific details, please refer to the textual description of the video editing scenario above; it will not be repeated here.
[0396] Finally, the hardware and software architectures of the electronic devices provided in the embodiments of this application are described below by way of example.
[0397] The electronic device in the embodiments of this application can be a single electronic device, such as a mobile electronic device or a PC, etc., without limitation.
[0398] Figure 17 This is an exemplary schematic diagram of the hardware structure of the electronic device 100 provided in the embodiments of this application.
[0399] The following detailed description uses electronic device 100 as an example. It should be understood that electronic device 100 may have more or fewer components than shown in the figures, may combine two or more components, or may have different component configurations. The various components shown in the figures can be implemented in hardware, software, or a combination of hardware and software, including one or more signal processing and / or application-specific integrated circuits.
[0400] Electronic device 100 may include: processor 110, external memory interface 120, internal memory 121, universal serial bus (USB) interface 130, charging management module 140, power management module 141, battery 142, antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, sensor module 180, button 190, motor 191, indicator 192, camera 193, display screen 194, and subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an accelerometer sensor 180E, a distance sensor 180F, a proximity sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.
[0401] It is understood that the structures illustrated in the embodiments of the present invention do not constitute a specific limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0402] Processor 110 may include one or more processing units, such as: application processor (AP), modem processor, graphics processing unit (GPU), image signal processor (ISP), controller, memory, video codec, digital signal processor (DSP), baseband processor, and / or neural network processing unit (NPU), etc. Different processing units may be independent devices or integrated into one or more processors.
[0403] The controller can be the nerve center and command center of the electronic device 100. The controller can generate operation control signals according to the instruction opcode and timing signals to complete the control of fetching and executing instructions.
[0404] The processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store instructions or data that the processor 110 has just used or that are used repeatedly. If the processor 110 needs to use the instruction or data again, it can retrieve it directly from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.
[0405] In some embodiments, the processor 110 may include one or more interfaces. Interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.
[0406] The I2C interface is a bidirectional synchronous serial bus, including a serial data line (SDA) and a serial clock line (SCL). In some embodiments, the processor 110 may include multiple I2C buses. The processor 110 can couple to the touch sensor 180K, charger, flash, camera 193, etc., through different I2C bus interfaces. For example, the processor 110 can couple to the touch sensor 180K through the I2C interface, enabling the processor 110 and the touch sensor 180K to communicate through the I2C bus interface, thereby realizing the touch function of the electronic device 100.
[0407] The I2S interface can be used for audio communication. In some embodiments, the processor 110 may include multiple I2S buses. The processor 110 can be coupled to the audio module 170 via the I2S bus to enable communication between the processor 110 and the audio module 170. In some embodiments, the audio module 170 can transmit audio signals to the wireless communication module 160 via the I2S interface to enable the function of answering phone calls through a Bluetooth headset.
[0408] The PCM interface can also be used for audio communication, sampling, quantizing, and encoding analog signals. In some embodiments, the audio module 170 and the wireless communication module 160 can be coupled via the PCM bus interface. In some embodiments, the audio module 170 can also transmit audio signals to the wireless communication module 160 via the PCM interface, enabling the function of answering phone calls through a Bluetooth headset. Both the I2S interface and the PCM interface can be used for audio communication.
[0409] The UART interface is a universal serial data bus used for asynchronous communication. This bus can be a bidirectional communication bus. It converts the data to be transmitted between serial and parallel communication. In some embodiments, the UART interface is typically used to connect the processor 110 and the wireless communication module 160. For example, the processor 110 communicates with the Bluetooth module in the wireless communication module 160 via the UART interface to implement Bluetooth functionality. In some embodiments, the audio module 170 can transmit audio signals to the wireless communication module 160 via the UART interface to enable music playback through Bluetooth headphones.
[0410] The MIPI interface can be used to connect the processor 110 to peripheral devices such as the display screen 194 and the camera 193. The MIPI interface includes a camera serial interface (CSI) and a display serial interface (DSI). In some embodiments, the processor 110 and the camera 193 communicate via the CSI interface to enable the electronic device 100 to capture images. The processor 110 and the display screen 194 communicate via the DSI interface to enable the electronic device 100 to display images.
[0411] The GPIO interface is configurable via software. It can be configured as a control signal or a data signal. In some embodiments, the GPIO interface can be used to connect the processor 110 to a camera 193, a display screen 194, a wireless communication module 160, an audio module 170, a sensor module 180, etc. The GPIO interface can also be configured as an I2C interface, an I2S interface, a UART interface, a MIPI interface, etc.
[0412] The SIM interface can be used to communicate with the SIM card interface 195 to transmit data to or read data from the SIM card.
[0413] USB port 130 is a USB standard compliant interface, specifically a Mini USB port, Micro USB port, USB Type-C port, etc. USB port 130 can be used to connect a charger to charge electronic device 100, and can also be used for data transfer between electronic device 100 and peripheral devices. It can also be used to connect headphones for audio playback. This interface can also be used to connect other electronic devices, such as AR devices.
[0414] It is understood that the interface connection relationships between the modules illustrated in the embodiments of the present invention are merely illustrative and do not constitute a structural limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may also employ different interface connection methods or combinations of multiple interface connection methods as described in the above embodiments.
[0415] The charging management module 140 is used to receive charging input from the charger. The charger can be a wireless charger or a wired charger.
[0416] The power management module 141 is used to connect the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140 to power the processor 110, internal memory 121, external memory, display 194, camera 193, and wireless communication module 160, etc.
[0417] The wireless communication function of electronic device 100 can be realized through antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modem processor and baseband processor, etc.
[0418] Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in electronic device 100 can be used to cover one or more communication frequency bands. Different antennas can also be multiplexed to improve antenna utilization. For example, antenna 1 can be multiplexed as a diversity antenna for a wireless local area network. In some other embodiments, the antennas can be used in conjunction with tuning switches.
[0419] The mobile communication module 150 can provide solutions for wireless communication, including 2G / 3G / 4G / 5G, applied to the electronic device 100. The mobile communication module 150 may include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves via antenna 1, and perform filtering, amplification, and other processing on the received electromagnetic waves before transmitting them to a modem processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modem processor and convert it into electromagnetic waves for radiation via antenna 1. In some embodiments, at least some functional modules of the mobile communication module 150 may be housed in the processor 110. In some embodiments, at least some functional modules of the mobile communication module 150 and at least some modules of the processor 110 may be housed in the same device.
[0420] The modem processor may include a modulator and a demodulator. The modulator modulates the low-frequency baseband signal to be transmitted into a mid-to-high frequency signal. The demodulator demodulates the received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After processing by the baseband processor, the low-frequency baseband signal is transmitted to the application processor. The application processor outputs sound signals through an audio device (not limited to speaker 170A, receiver 170B, etc.) or displays images or videos through the display screen 194. In some embodiments, the modem processor may be a separate device. In other embodiments, the modem processor may be independent of the processor 110 and may be housed in the same device as the mobile communication module 150 or other functional modules.
[0421] The wireless communication module 160 can provide solutions for wireless communication applications on the electronic device 100, including wireless local area networks (WLANs) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), and infrared (IR) technologies. The wireless communication module 160 can be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via antenna 2, performs frequency modulation and filtering of the electromagnetic wave signals, and sends the processed signal to processor 110. The wireless communication module 160 can also receive signals to be transmitted from processor 110, perform frequency modulation and amplification, and convert them into electromagnetic waves for radiation via antenna 2.
[0422] In some embodiments, antenna 1 of electronic device 100 is coupled to mobile communication module 150, and antenna 2 is coupled to wireless communication module 160, enabling electronic device 100 to communicate with networks and other devices via wireless communication technology. The wireless communication technology may include Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Time Division Code Division Multiple Access (TD-SCDMA), Long Term Evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technologies, etc. The GNSS may include the Global Positioning System (GPS), the Global Navigation Satellite System (GLONASS), the BeiDou Navigation Satellite System (BDS), the Quasi-Zenith Satellite System (QZSS), and / or satellite-based augmentation systems (SBAS).
[0423] Electronic device 100 implements display functions through a GPU, a display screen 194, and an application processor. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations and for graphics rendering. Processor 110 may include one or more GPUs, which execute program instructions to generate or modify display information.
[0424] Display screen 194 is used to display images, videos, etc. Display screen 194 includes a display panel. The display panel may be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a Mini LED, a MicroLED, a Micro-OLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, electronic device 100 may include one or N displays 194, where N is a positive integer greater than 1.
[0425] Electronic device 100 can perform shooting functions through ISP, camera 193, video codec, GPU, display screen 194 and application processor, thereby acquiring real-time video data.
[0426] The ISP (Image Signal Processor) is used to process data fed back from the camera 193. For example, when taking a picture, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, and the camera's photosensitive element transmits the electrical signal to the ISP for processing, transforming it into an image visible to the naked eye. The ISP can also perform algorithmic optimization on image noise and brightness. The ISP can also optimize parameters such as exposure and color temperature of the shooting scene. In some embodiments, the ISP can be set in the camera 193.
[0427] Camera 193 is used to capture still images or videos. An object is projected onto a photosensitive element by generating an optical image through the lens. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then passed to an ISP for conversion into a digital image signal. The ISP outputs the digital image signal to a DSP for processing. The DSP converts the digital image signal into image signals in standard RGB, YUV, or other formats. In some embodiments, the electronic device 100 may include one or N cameras 193, where N is a positive integer greater than 1.
[0428] Digital signal processors (DSPs) are used to process digital signals. Besides digital image signals, they can also process other digital signals. For example, when electronic device 100 selects a frequency, the DSP can perform Fourier transforms on the frequency energy.
[0429] Video codecs are used to compress or decompress digital video. Electronic device 100 may support one or more video codecs. Thus, electronic device 100 can play or record videos in various encoding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.
[0430] An NPU (Neural Processing Unit) is a computational processor for neural networks (NNs). By borrowing the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, it can rapidly process input information and continuously learn on its own. NPUs enable intelligent cognitive applications in electronic devices, such as image recognition, facial recognition, voice recognition, and text understanding.
[0431] In this embodiment of the application, the electronic device can determine the face region in the image data of the video based on the CPU / GPU / NPU; and determine the facial features in the face region.
[0432] In this embodiment of the application, the electronic device can extract visual semantic features from the image data of the video based on the CPU / GPU / NPU.
[0433] In this embodiment of the application, the electronic device can run code to implement the audio and video fusion module based on CPU / GPU / NPU, thereby separating the voices of different people in the video and further extracting the voiceprints of different people.
[0434] In this embodiment of the application, the electronic device can extract the voice of the target person from the video based on the voiceprint of the target person, using a CPU / GPU / NPU.
[0435] Internal memory 121 may include one or more random access memory (RAM) and one or more non-volatile memory (NVM).
[0436] Random access memory can include static random-access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM, for example, fifth generation DDR SDRAM is generally called DDR5 SDRAM), etc.
[0437] Non-volatile memory can include disk storage devices and flash memory.
[0438] In this embodiment, the non-real-time video may be located in non-volatile memory.
[0439] Flash memory can be classified according to its operating principle, including NOR FLASH, NAND FLASH, 3D NAND FLASH, etc.; according to the level of the storage cell, including single-level cell (SLC), multi-level cell (MLC), triple-level cell (TLC), quad-level cell (QLC), etc.; and according to the storage specification, including universal flash storage (UFS) and embedded multimedia card (eMMC), etc.
[0440] The random access memory can be directly read and written by the processor 110. It can be used to store executable programs (such as machine instructions) of the operating system or other running programs, as well as user and application data.
[0441] Non-volatile memory can also store executable programs and user and application data, and can be pre-loaded into random access memory for direct reading and writing by the processor 110.
[0442] The external memory interface 120 can be used to connect to external non-volatile memory, thereby expanding the storage capacity of the electronic device 100. The external non-volatile memory communicates with the processor 110 through the external memory interface 120 to perform data storage functions. For example, music, video, and other files can be stored in the external non-volatile memory.
[0443] Electronic device 100 can implement audio functions, such as music playback and recording, through audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, and application processor.
[0444] The audio module 170 is used to convert digital audio information into analog audio signals for output, and also to convert analog audio input into digital audio signals. The audio module 170 can also be used for encoding and decoding audio signals. In some embodiments, the audio module 170 may be located in the processor 110, or some functional modules of the audio module 170 may be located in the processor 110.
[0445] The speaker 170A, also known as a "loudspeaker," is used to convert audio electrical signals into sound signals. The electronic device 100 can listen to music or make hands-free calls through the speaker 170A.
[0446] The receiver 170B, also known as the "earpiece," is used to convert audio electrical signals into sound signals. When the electronic device 100 receives a telephone call or voice message, the receiver 170B can be brought close to the ear to hear the sound.
[0447] Microphone 170C, also known as a "microphone" or "voice transducer," is used to convert sound signals into electrical signals. When making a phone call or sending a voice message, the user can speak by bringing their mouth close to microphone 170C, inputting the sound signal into microphone 170C. Electronic device 100 may have at least one microphone 170C. In some embodiments, electronic device 100 may have two microphones 170C, which, in addition to collecting sound signals, can also perform noise reduction. In other embodiments, electronic device 100 may also have three, four, or more microphones 170C, which can collect sound signals, reduce noise, identify the sound source, and perform directional recording, etc.
[0448] The 170D headphone jack is used to connect wired headphones. The 170D headphone jack can be a USB 130 interface or a 3.5mm Open Mobile Terminal Platform (OMTP) standard interface, a CTIA (Cellular Telecommunications Industry Association of the USA) standard interface.
[0449] Pressure sensor 180A is used to sense pressure signals and convert them into electrical signals. In some embodiments, pressure sensor 180A can be disposed on display screen 194. There are many types of pressure sensors 180A, such as resistive pressure sensors, inductive pressure sensors, and capacitive pressure sensors. A capacitive pressure sensor may include at least two parallel plates with conductive material. When force is applied to pressure sensor 180A, the capacitance between the electrodes changes. Electronic device 100 determines the pressure intensity based on the change in capacitance. When a touch operation is applied to display screen 194, electronic device 100 detects the intensity of the touch operation based on pressure sensor 180A. Electronic device 100 can also calculate the touch position based on the detection signal from pressure sensor 180A. In some embodiments, touch operations applied to the same touch position but with different touch operation intensities can correspond to different operation commands. For example, when a touch operation with an intensity less than a first pressure threshold is applied to the SMS application icon, a command to view an SMS is executed. When a touch operation with an intensity greater than or equal to the first pressure threshold is applied to the SMS application icon, a command to create a new SMS is executed.
[0450] The gyroscope sensor 180B can be used to determine the motion attitude of the electronic device 100. In some embodiments, the gyroscope sensor 180B can determine the angular velocity of the electronic device 100 about three axes (i.e., the x, y, and z axes). The gyroscope sensor 180B can be used for image stabilization. For example, when the shutter is pressed, the gyroscope sensor 180B detects the angle of the shake of the electronic device 100, calculates the distance that the lens module needs to compensate based on the angle, and allows the lens to counteract the shake of the electronic device 100 by moving in the opposite direction, thus achieving image stabilization. The gyroscope sensor 180B can also be used in navigation and motion-sensing game scenarios.
[0451] The barometric pressure sensor 180C is used to measure air pressure. In some embodiments, the electronic device 100 calculates altitude using the air pressure value measured by the barometric pressure sensor 180C to assist in positioning and navigation.
[0452] The magnetic sensor 180D includes a Hall sensor. The electronic device 100 can use the magnetic sensor 180D to detect the opening and closing of the flip cover. In some embodiments, when the electronic device 100 is a flip phone, the electronic device 100 can detect the opening and closing of the flip cover using the magnetic sensor 180D. Then, based on the detected opening and closing state of the cover or the flip cover, features such as automatic flip unlocking can be set.
[0453] The 180E accelerometer can detect the magnitude of acceleration of electronic device 100 in various directions (typically three axes). When electronic device 100 is stationary, it can detect the magnitude and direction of gravity. It can also be used to identify the posture of electronic devices and applied to applications such as screen orientation switching and pedometers.
[0454] A distance sensor 180F is used to measure distance. Electronic device 100 can measure distance via infrared or laser. In some embodiments, during a shooting scene, electronic device 100 can utilize the distance sensor 180F to measure distance for rapid focusing.
[0455] The proximity sensor 180G may include, for example, a light-emitting diode (LED) and a light detector, such as a photodiode. The LED may be an infrared LED. The electronic device 100 emits infrared light outward through the LED. The electronic device 100 uses the photodiode to detect infrared reflected light from nearby objects. When sufficient reflected light is detected, it can be determined that there is an object near the electronic device 100. When insufficient reflected light is detected, the electronic device 100 can determine that there is no object near the electronic device 100. The electronic device 100 may use the proximity sensor 180G to detect when a user holds the electronic device 100 close to their ear for a call, so as to automatically turn off the screen to save power. The proximity sensor 180G can also be used in holster mode and pocket mode for automatic unlocking and locking of the screen.
[0456] The ambient light sensor 180L is used to sense the brightness of ambient light. The electronic device 100 can adaptively adjust the brightness of the display screen 194 based on the sensed ambient light brightness. The ambient light sensor 180L can also be used to automatically adjust the white balance when taking pictures. The ambient light sensor 180L can also work with the proximity sensor 180G to detect whether the electronic device 100 is in a pocket to prevent accidental touches.
[0457] The fingerprint sensor 180H is used to collect fingerprints. The electronic device 100 can utilize the characteristics of the collected fingerprints to achieve fingerprint unlocking, accessing application locks, taking photos with fingerprints, answering calls with fingerprints, etc.
[0458] Temperature sensor 180J is used to detect temperature. In some embodiments, electronic device 100 uses the temperature detected by temperature sensor 180J to execute a temperature handling strategy. For example, when the temperature reported by temperature sensor 180J exceeds a threshold, electronic device 100 performs thermal protection by reducing the performance of a processor located near temperature sensor 180J to reduce power consumption. In other embodiments, when the temperature is below another threshold, electronic device 100 heats battery 142 to prevent abnormal shutdown of electronic device 100 due to low temperature. In still other embodiments, when the temperature is below yet another threshold, electronic device 100 boosts the output voltage of battery 142 to prevent abnormal shutdown due to low temperature.
[0459] Touch sensor 180K, also known as a "touch panel," can be located on display screen 194. The touch sensor 180K and display screen 194 together form a touchscreen, also known as a "touch screen." Touch sensor 180K detects touch operations applied to or near it. The touch sensor can transmit the detected touch operation to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through display screen 194. In other embodiments, touch sensor 180K may also be located on the surface of electronic device 100, in a different position than display screen 194.
[0460] Buttons 190 include a power button, volume buttons, etc. Buttons 190 can be mechanical buttons or touch-sensitive buttons. Electronic device 100 can receive button input and generate key signal inputs related to user settings and function control of electronic device 100.
[0461] Motor 191 can generate vibration alerts. Motor 191 can be used for incoming call vibration alerts or for touch vibration feedback. For example, different vibration feedback effects can correspond to touch operations performed on different applications (such as taking photos, playing audio, etc.). Motor 191 can also correspond to different vibration feedback effects for touch operations performed on different areas of the display screen 194. Different application scenarios (such as time reminders, receiving messages, alarm clocks, games, etc.) can also correspond to different vibration feedback effects. The touch vibration feedback effect can also be customized.
[0462] Indicator 192 can be an indicator light, used to indicate charging status, power changes, or to indicate messages, missed calls, notifications, etc.
[0463] The SIM card interface 195 is used to connect a SIM card. The SIM card can be inserted into or removed from the SIM card interface 195 to make contact with and detach from the electronic device 100. The electronic device 100 can support one or N SIM card interfaces, where N is a positive integer greater than 1. The SIM card interface 195 can support Nano SIM cards, Micro SIM cards, and other SIM cards. Multiple cards can be inserted into the same SIM card interface 195 simultaneously. The multiple cards can be of the same or different types. The SIM card interface 195 is also compatible with different types of SIM cards. The SIM card interface 195 is also compatible with external memory cards. The electronic device 100 interacts with the network through the SIM card to realize functions such as calls and data communication.
[0464] Figure 18 This is an exemplary schematic diagram of the software structure of the electronic device 100 provided in the embodiments of this application.
[0465] A layered architecture divides software into several layers, each with a clear role and function. Layers communicate with each other through software interfaces. In some embodiments, the system is divided into four layers, from top to bottom: the application layer, the application framework layer, the system library layer, and the kernel layer.
[0466] The application layer can include a series of application packages.
[0467] like Figure 18 As shown, the application package may include applications (also known as apps) such as camera, gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, video, and SMS.
[0468] The application framework layer provides application programming interfaces (APIs) and a programming framework for applications in the application layer. The application framework layer includes some predefined functions.
[0469] like Figure 18 As shown, the application framework layer may include a window manager, content provider, view system, phone manager, resource manager, notification manager, local profile management assistant (LPA), etc.
[0470] The window manager is used to manage windowed applications. It can retrieve screen size, determine the presence of a status bar, lock the screen, and capture screenshots, among other things.
[0471] Content providers store and retrieve data, making that data accessible to applications. This data may include videos, images, audio, made and received phone calls, browsing history and bookmarks, phone books, etc.
[0472] A view system includes visual controls, such as controls for displaying text and controls for displaying images. View systems can be used to build applications. A display interface can consist of one or more views. For example, a display interface including a text notification icon could include views for displaying text and views for displaying images.
[0473] The phone manager is used to provide communication functions for electronic device 100. For example, it manages call status (including connection and disconnection).
[0474] The file explorer provides applications with various resources, such as localized strings, icons, images, layout files, video files, and more.
[0475] The notification manager allows applications to display notifications in the status bar. These notifications can be used to deliver informational messages and can disappear automatically after a short pause, requiring no user interaction. For example, the notification manager can be used to notify users of completed downloads or message alerts. The notification manager can also display notifications as icons or scrolling text in the top status bar, such as notifications from background applications, or as dialog-style notifications on the screen. Examples include displaying text messages in the status bar, emitting sounds, vibrating electronic devices, and flashing indicator lights.
[0476] The runtime includes the core libraries and the virtual machine. The runtime is responsible for the scheduling and management of the operating system.
[0477] The core library consists of two parts: one part contains the functionalities that the Java language needs to call, and the other part is the core library itself.
[0478] The application layer and application framework layer run in a virtual machine. The virtual machine executes the Java files of the application layer and application framework layer as binary files. The virtual machine is used to perform functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection.
[0479] System libraries can include multiple functional modules. For example: surface manager, media libraries, 3D graphics processing libraries (e.g., OpenGL ES), 2D graphics engines (e.g., SGL), etc.
[0480] The Surface Manager is used to manage the display subsystem and provides the fusion of two-dimensional (2D) and three-dimensional (3D) layers for multiple applications.
[0481] The media library supports playback and recording of various common audio and video formats, as well as still image files. It supports multiple audio and video encoding formats, such as MPEG4, H.264, MP3, AAC, AMR, JPG, and PNG.
[0482] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, compositing, and layer processing.
[0483] A 2D graphics engine is a graphics engine for 2D drawing.
[0484] The system library also includes a video processing library, which allows software developers to implement features such as... by calling the API interfaces provided by the video processing library. Figure 12 , Figure 13A , Figure 13B , Figure 13C , Figure 14A , Figure 14B , Figure 14C The effect shown is to process real-time video streams or non-real-time video files, separate the voice of the target person in the video, and select some adaptive processing.
[0485] The kernel layer is the layer between hardware and software. The kernel layer includes at least the display driver, camera driver, audio driver, sensor driver, and virtual card driver.
[0486] Figure 19 Another exemplary schematic diagram of the software structure of the electronic device 100 provided in the embodiments of this application.
[0487] In some embodiments, the system is divided into four layers, from top to bottom: the application layer, the framework layer, the system service library, and the kernel layer.
[0488] The application layer includes system applications and third-party non-system applications.
[0489] The framework layer provides application layer applications with user program frameworks and capability frameworks in multiple languages such as JAVA / C / C++ / JS, as well as multi-language framework APIs for various software and hardware services.
[0490] The system service layer includes: a set of basic system capabilities subsystems, a set of basic software service subsystems, a set of enhanced software service subsystems, and a set of hardware service subsystems.
[0491] The system's basic capability subsystem set supports the operating system's operation, scheduling, and migration across multiple devices. This subsystem set may include: a distributed soft bus, distributed data management, distributed task scheduling, and common infrastructure subsystems. The system service layer and framework layer jointly implement the multi-modal input subsystem and graphics subsystem.
[0492] The basic software service subsystem provides common and general software services to the operating system, and may include: event notification subsystem, multimedia subsystem, etc.
[0493] The enhanced software service subsystem set provides differentiated software services for different devices and may include: IoT proprietary business subsystems. The enhanced software service subsystem set may also include the video processing method provided in the embodiments of this application.
[0494] Software developers can achieve, for example, by calling the API interfaces provided by the Enhanced Software Services Subsystem suite. Figure 12 , Figure 13A , Figure 13B , Figure 13C , Figure 14A , Figure 14B , Figure 14C The effect shown is to process real-time video streams or non-real-time video files, separate the voice of the target person in the video, and select some adaptive processing.
[0495] The hardware service subsystem set provides hardware services to the operating system and may include: IoT proprietary hardware service subsystem.
[0496] It is worth noting that, depending on the deployment environment of different device types, the above-mentioned system basic capability subsystem set, basic software service subsystem set, enhanced software service subsystem set, and hardware service subsystem set can be re-divided according to other functional granularities.
[0497] The kernel layer comprises the kernel abstraction layer and the driver subsystem. The kernel abstraction layer includes multiple kernels and, by shielding the differences between them, provides basic kernel capabilities to higher layers, such as thread / process management, memory management, file system management, and network management. The driver subsystem provides software developers with a unified framework for peripheral access and driver development and management.
[0498] It is worth noting that, depending on the operating system and possible future upgrades, the software structure of electronic devices can be divided in other ways based on the operating system.
[0499] As used in the above embodiments, depending on the context, the term "when..." can be interpreted as meaning "if...", "after...", "in response to determining...", or "in response to detecting...". Similarly, depending on the context, the phrase "when determining..." or "if (the stated condition or event) is interpreted as meaning "if determining...", "in response to determining...", "when (the stated condition or event) is detected", or "in response to detecting (the stated condition or event)".
[0500] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. This computer program product includes one or more computer instructions. When these computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic cable, digital persona cable) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive), etc.
[0501] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This program can be stored in a computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM or random access memory (RAM), magnetic disks, or optical disks.
Claims
1. A video processing method applied to electronic devices, characterized in that, include: The electronic device acquires a saved first video, wherein at least a portion of the image frames displayed in the first video include a first object, and the first video includes a first audio. The electronic device separates a second audio and a fourth audio from the first audio; wherein the second audio is the sound of the first object in the first audio, and the fourth audio is audio data in the first audio that does not include the second audio; After the electronic device separates the second audio and the fourth audio from the first audio, it displays a first interface. The first interface includes a first control and a fourth control, wherein the first control corresponds to the fourth audio and the fourth control corresponds to the second audio. In response to the operation performed on the first control, obtain the coefficient input by the user; The electronic device enhances or weakens the fourth audio based on the coefficients to obtain processed audio data of the fourth audio. The electronic device superimposes the processed audio data of the fourth audio with the second audio in the time domain to obtain the third audio; The electronic device plays the third audio.
2. The method according to claim 1, characterized in that, Before the electronic device acquires the saved first video, the method further includes: The electronic device captures the first video; The electronic device stores the first video; The electronic device acquires the saved first video, specifically including: When the electronic device displays the editing page of the first video, it retrieves the saved first video.
3. The method according to claim 1 or 2, characterized in that, The method further includes: The first interface also includes a second control, which is used to play only the second audio. In response to an operation performed on the second control, the electronic device plays the second audio.
4. The method according to claim 1 or 2, characterized in that, The first object satisfies the first condition. The first condition includes: the pitch angle of the object's face is within a preset pitch angle range and / or the rotation of the object's face is within a preset rotation angle range and / or the tilt angle of the object's face is within a preset tilt angle range.
5. The method according to claim 1 or 2, characterized in that, The electronic device separates a second audio and a fourth audio from the first audio, specifically including: The electronic device determines the corresponding first segmented audio based on the at least a portion of the image frames; The electronic device determines the visual semantic features of the first object based on the at least a portion of the image frames, wherein the visual semantic features are features of facial morphology related to speech and sound. The electronic device determines the sound of the first object in the first truncated audio based on the visual semantic features of the first object and the first truncated audio. The electronic device determines the voiceprint of the first object based on the sound of the first object in the first captured audio; The electronic device separates the second audio and the fourth audio from the first audio based on the voiceprint of the first object.
6. The method according to claim 1 or 2, characterized in that, The electronic device separates a second audio and a fourth audio from the first audio, specifically including: If the electronic device determines that the first object is included in all image frames of the first video, the first object satisfies the first condition. The electronic device determines the visual semantic features of the first object based on all image frames of the first video, and the visual semantic features are features of facial features related to speech and sound. The electronic device separates the second audio and the fourth audio from the first audio based on the visual semantic features of the first object.
7. The method according to claim 5, characterized in that, The electronic device determines the voiceprint of the first object based on the sound of the first object in the first captured audio, specifically including: The electronic device filters out sound segments from the sound of the first object in the first captured audio, where the signal-to-noise ratio is greater than a signal-to-noise ratio threshold and the duration is greater than a duration threshold. The electronic device determines the voiceprint of the first object based on the sound segment.
8. A video processing method applied to electronic devices, characterized in that, include: The electronic device acquires a saved first video, the first video including a first audio; The electronic device determines a first portion of image frames based on the first video, the first portion of image frames including a first object, and the first object in the first portion of image frames satisfies a first condition; The electronic device determines the first segmented audio corresponding to the first portion of the image frame; The electronic device determines the visual semantic features of the first object based on the first part of the image frames, and the visual semantic features are facial morphological features related to speech and sound. The electronic device determines the sound of the first object in the first truncated audio based on the visual semantic features of the first object and the first truncated audio. The electronic device determines the voiceprint of the first object based on the sound of the first object in the first captured audio; The electronic device separates a second audio and a fourth audio from the first audio based on the voiceprint of the first object; wherein the second audio is the sound of the first object in the first audio, and the fourth audio is audio data in the first audio that does not include the second audio; After the electronic device separates the second audio and the fourth audio from the first audio, it displays a first interface. The first interface includes a first control and a fourth control, wherein the first control corresponds to the fourth audio and the fourth control corresponds to the second audio. In response to the operation performed on the first control, obtain the coefficient input by the user; The electronic device enhances or weakens the fourth audio based on the coefficient to obtain processed audio information of the fourth audio. The electronic device superimposes the processed audio data of the fourth audio with the second audio in the time domain to obtain the third audio.
9. The method according to claim 8, characterized in that, Before the electronic device acquires the saved first video, the method further includes: The electronic device captures the first video; The electronic device stores the first video; The electronic device acquires the saved first video, specifically including: When the electronic device displays the editing page of the first video, it retrieves the saved first video.
10. The method according to claim 8 or 9, characterized in that, The method further includes: The electronic device determines a second portion of image frames based on the first video, the second portion of image frames including a second object, and the second object in the second portion of image frames satisfies the first condition; The electronic device determines the second segmented audio corresponding to the second portion of the image frame; The electronic device determines the visual semantic features of the second object based on the second part of the image frames, and the visual semantic features are facial morphological features related to speech and sound. The electronic device determines the sound of the first object in the second truncated audio based on the visual semantic features of the first object and the first truncated audio. The electronic device determines the voiceprint of the second object based on the sound of the second object in the second captured audio; The electronic device determines the audio corresponding to the second object based on the voiceprint of the second object and the first audio.
11. The method according to claim 8 or 9, characterized in that, The first condition includes: the pitch angle of the object's face is within a preset pitch angle range and / or the rotation of the object's face is within a preset rotation angle range and / or the tilt angle of the object's face is within a preset tilt angle range.
12. The method according to claim 8 or 9, characterized in that, The electronic device determines the voiceprint of the first object based on the sound of the first object in the first captured audio, specifically including: The electronic device filters out sound segments from the sound of the first object in the first captured audio, where the signal-to-noise ratio is greater than a signal-to-noise ratio threshold and the duration is greater than a duration threshold. The electronic device determines the voiceprint of the first object based on the sound segment.
13. An electronic device, characterized in that, The electronic device includes: one or more processors and memory; The memory is coupled to the one or more processors, the memory being used to store computer program code, the computer program code including computer instructions, the one or more processors invoking the computer instructions to cause the electronic device to perform the method as described in any one of claims 1 to 12.
14. A chip system applied to an electronic device, characterized in that, The chip system includes one or more processors, the processors being configured to invoke computer instructions to cause the electronic device to perform the method as described in any one of claims 1 to 12.
15. A computer-readable storage medium comprising instructions, characterized in that, When the instructions are executed on an electronic device, the electronic device causes the electronic device to perform the method as described in any one of claims 1 to 12.
Citation Information
Patent Citations
Automatic dubbing method and apparatus
CN108780643A
Video processing method and electronic equipment
CN110602424A
Video speaker identification method and device, computer equipment and storage medium
CN111785279A