Speech processing method and model training method, and electronic device

By using multimodal fusion processing and model training, noise and reverberation in video and audio data are suppressed, solving the problem of reverberation enhancement caused by audio zoom, improving audio quality and user experience, and reducing hardware upgrade costs.

CN115881161BActive Publication Date: 2026-01-30HUAWEI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111148233.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-29
Publication Date
2026-01-30
Estimated Expiration
2041-09-29

AI Technical Summary

Technical Problem

During video recording, audio zoom increases reverberation, affecting audio quality and blurring the sound signal. Existing technologies have failed to effectively suppress noise and reverberation.

Method used

By performing multimodal fusion processing on video and audio data, including lip reading and frequency domain transformation, combined with preset suppression and sequence transformation models, noise and reverberation are suppressed. Then, audio zoom is performed based on zoom parameters, and finally, appropriate reverberation is added according to the scene to improve audio realism.

Benefits of technology

It improves the quality of video and audio data after zooming, enhances the user experience, reduces hardware upgrade costs, and enables audio zooming during recording and playback.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115881161B_ABST
    Figure CN115881161B_ABST
Patent Text Reader

Abstract

This application provides a speech processing method, a model training method, and an electronic device. The speech processing method includes: when it is determined that a video frame has zoomed in, acquiring zoom parameters, first video-speech data of the video, and video frame data after zooming; then, performing multimodal fusion processing on the zoomed video frame data and the first video-speech data to obtain second video-speech data; next, zooming the second video-speech data based on the zoom parameters to obtain third video-speech data, and outputting the third video-speech data. In this way, through multimodal fusion processing, noise and reverberation in the video-speech data are effectively suppressed, and zooming is performed only on the noise- and reverberation-suppressed video-speech data, thereby improving the quality of the zoomed video-speech data and enhancing the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing, and more particularly to a speech processing method, a model training method, and an electronic device. Background Technology

[0002] Currently, many users enjoy recording videos using electronic devices such as cameras and mobile phones. To improve the quality of recorded videos, some devices integrate audio zoom functionality. Audio zoom, analogous to image zoom, means that the volume of the recorded audio changes as the focal length of the camera changes.

[0003] During indoor video recording, after the subject emits a sound signal, the signal not only reaches the electronic device directly, but also is reflected by obstacles such as walls, ceilings, and floors (i.e., reverberation). This means the recorded audio contains not only the sound signal emitted by the subject but also its reverberation. Therefore, zooming in on the audio not only changes the volume of the sound signal but may also increase the reverberation, making the sound signal blurry. Summary of the Invention

[0004] To address the aforementioned technical problems, this application provides a speech processing method, a model training method, and an electronic device. In this speech processing method, reverberation and noise in the video speech data are first suppressed, and then the video speech data after noise and reverberation suppression is zoomed in. This improves the quality of the zoomed video speech data and enhances the user experience.

[0005] In a first aspect, embodiments of this application provide a voice processing method applied to an electronic device. The method includes: when it is determined that a video frame has zoomed in, acquiring zoom parameters, first video-audio data of the video, and video frame data after zooming; subsequently, performing multimodal fusion processing on the zoomed video frame data and the first video-audio data to obtain second video-audio data; then, zooming the second video-audio data based on the zoom parameters to obtain third video-audio data, and outputting the third video-audio data. Thus, through multimodal fusion processing, noise and reverberation in the video-audio data are effectively suppressed, and zooming is performed only on the noise- and reverberation-suppressed video-audio data, improving the quality of the zoomed video-audio data and enhancing the user experience.

[0006] For example, the zoom parameter is the zoom magnification.

[0007] For example, zooming the second video audio data based on zoom parameters means adjusting the audio loudness of the second video audio data.

[0008] For example, electronic devices include, but are not limited to: cameras, camcorders, mobile phones, tablets, etc.

[0009] According to the first aspect, second video-speech data is obtained by performing multimodal fusion processing on the zoomed-out video image data and the first video-speech data, including: performing lip-reading based on the zoomed-out video image data to extract phoneme features; performing frequency domain transformation based on the first video-speech data to obtain first speech features; and performing multimodal fusion of the phoneme features and the first speech features to obtain the second video-speech data. In this way, by fusing the image and sound modalities, the suppression effect on noise and reverberation in the first video-speech data can be improved.

[0010] According to the first aspect, or any implementation thereof, before performing frequency domain transformation based on the first video-speech data to obtain the first speech feature, the method further includes: suppressing noise and reverberation in the first video-speech data using a preset suppression model to obtain fourth video-speech data; performing frequency domain transformation based on the first video-speech data to obtain the first speech feature includes: performing frequency domain transformation on the fourth video-speech data to obtain the first speech feature. Thus, by initially suppressing noise and reverberation in the first video-speech data before performing multimodal fusion processing, the suppression effect on noise and reverberation in the first video-speech data can be improved.

[0011] According to the first aspect, or any implementation thereof, a second video-speech data is obtained by performing multimodal fusion processing on phoneme features and first speech features, including: concatenating phoneme features and first speech features to obtain concatenated features; converting the concatenated features into second speech features using a preset sequence conversion model, the sequence conversion model being trained based on video frame data and video-speech data; and performing speech synthesis based on the second speech features to obtain the second video-speech data. This enables multimodal fusion processing of zoomed-in video frame data and first video-speech data.

[0012] According to the first aspect, or any implementation of the first aspect above, the third video and audio data is output, including: synthesizing video data from the zoomed-out video image data and the third video and audio data, and then outputting the video data. In this way, the zoomed-out video image data and the zoomed-out video and audio data can be synthesized into video data and then output.

[0013] For example, when the current scenario is a video recording scenario, the synthesized video data can be stored in the corresponding storage area. When the current scenario is a video playback scenario, the synthesized video data can be output to the playback module for playback.

[0014] According to the first aspect, or any implementation of the first aspect above, after obtaining the third video and audio data, the method further includes: adding reverb to the third video and audio data; and outputting the third video and audio data, including: outputting the reverb-added third video and audio data. This allows the output video and audio data to have reverb, increasing the realism of the video and audio data and improving the user experience.

[0015] According to the first aspect, or any implementation of the first aspect above, reverberation is added to the third video and audio data, including: determining the distance between a target face and an electronic device in the zoomed video frame based on the zoomed video frame data, wherein the target face is the face of the user corresponding to the first video and audio data; determining the sound field impulse response based on the distance between the target face and the electronic device; and adding reverberation to the third video and audio data based on the sound field impulse response. In this way, by using zoom parameters, the reverberation recorded at a zoom distance from the sound source is simulated, improving the realism of the zoomed audio.

[0016] According to the first aspect, or any implementation of the first aspect above, reverberation is added to the third video audio data, including: scene recognition based on the zoomed video image data to determine the scene identifier corresponding to the zoomed video image data; and adding reverberation to the third video audio data using a preset audio reverberation addition model based on the scene identifier, zoom parameters, and the third video audio data. In this way, by combining the sound field and zoom parameters, a more realistic reverberation can be simulated compared to recordings at a zoom distance from the sound source, further improving the realism of the zoomed audio.

[0017] According to the first aspect, or any implementation of the first aspect above, after determining the scene identifier corresponding to the zoomed video image data, the method further includes: determining whether the scene identifier is a preset scene identifier; when the scene identifier is not a preset scene identifier, performing the step of outputting based on the third video audio data; when the scene identifier is a preset scene identifier, performing the step of adding reverb to the third video audio data based on the scene identifier, zoom parameters, and the third video audio data using a preset audio reverb addition model.

[0018] For example, the preset scene identifier is an indoor scene identifier. Therefore, when the video is recorded in an indoor scene, reverb is added to the zoomed video and audio data to increase the realism of the video and audio data. When the video is recorded in an outdoor scene, there is no need to add reverb to the zoomed video and audio data; the zoomed video and audio data can be output directly.

[0019] For example, video data is synthesized by combining zoomed-in video image data with third-party video-audio data after adding reverberation, and then output as video data. This enhances the user's audiovisual experience.

[0020] According to the first aspect, or any implementation of the first aspect above, when the focal length of the camera of the electronic device changes during video recording, it is determined that the video image has zoomed; when the size of the video image changes during video playback, it is determined that the video image has zoomed. In this way, audio zoom can be achieved when the user adjusts the camera focal length during video recording, or when the user adjusts the size of the video image during playback.

[0021] According to the first aspect, or any implementation of the first aspect above, the electronic device includes at least one microphone. This enables audio zoom in electronic devices such as cameras with only one microphone, without requiring hardware upgrades and reducing costs.

[0022] Secondly, embodiments of this application provide a model training method, which includes: firstly, collecting first video data and second video data, wherein the first video data is recorded at a first distance from the sound source, and the second video data is recorded at a second distance from the sound source, the first distance being greater than the second distance; the first video data includes first video frame data and first video speech data, and the second video data includes second video frame data and second video speech data. Subsequently, lip-reading is performed on the first video frame data and the second video frame data respectively to obtain corresponding first phoneme features and second phoneme features; and frequency domain transformation is performed on the first video speech data and the second video speech data respectively to obtain corresponding first speech features and second speech features; then, the first phoneme features, the second phoneme features, the first speech features, and the second speech features are used to train a sequence conversion model. Thus, by using multimodal data to train the sequence conversion model, the noise and reverberation suppression effect of the sequence conversion module can be improved.

[0023] For example, the first video data is video data A, the first video audio data is video audio data A, the first video frame data is video frame data A, the first phoneme feature is phoneme feature A, and the first speech feature is speech feature A. The second video data is video data B, the second video audio data is video audio data B, the second video frame data is video frame data B, the second phoneme feature is phoneme feature B, and the second speech feature is speech feature B.

[0024] For example, the first distance is LA and the second distance is LB.

[0025] According to the second aspect, the first phoneme feature, the second phoneme feature, the first speech feature, and the second speech feature all include X frames, where X is a positive integer. The sequence conversion model includes an encoder and a decoder. The sequence conversion model is trained using the first phoneme feature, the second phoneme feature, the first speech feature, and the second speech feature, including: concatenating the first speech feature and the first phoneme feature of X frames to obtain X-frame concatenated features; wherein the second phoneme feature and the second speech feature of X frames are both label data, and the X-frame concatenated features are sample data. Then, the X-frame concatenated features are input to the encoder to obtain the X-frame matrix output by the encoder; next, the X-frame matrix, the second speech feature of the i-th frame from the second speech feature of X frames, and the (i-1)-th frame speech feature output by the decoder are input to the decoder, which performs the i-th decoding to obtain the i-th frame speech feature, the i-th frame phoneme feature, and the length identifier output by the decoder, where i is an integer greater than 1. Subsequently, based on the speech features of the i-th frame and the second speech features of the i-th frame, the phoneme features of the i-th frame and the second phoneme features of the i-th frame, and the length identifier and the preset length identifier, the loss function value is determined, and the model parameters of the decoder and encoder are adjusted according to the loss function value. In this way, the sequence conversion model can converge quickly, improving the model training efficiency.

[0026] For example, the speech feature of the i-th frame is the speech feature C of the i-th frame, and the phoneme feature of the i-th frame is the phoneme feature C of the i-th frame.

[0027] Thirdly, embodiments of this application provide an electronic device, including: a memory and a processor, the memory being coupled to the processor; the memory storing program instructions, which, when executed by the processor, cause the electronic device to perform the voice processing method in the first aspect or any possible implementation thereof.

[0028] The third aspect and any implementation thereof correspond to the first aspect and any implementation thereof, respectively. The technical effects of the third aspect and any implementation thereof are similar to those of the first aspect and any implementation thereof, and will not be repeated here.

[0029] Fourthly, embodiments of this application provide an electronic device, including: a memory and a processor, the memory being coupled to the processor; the memory storing program instructions, when executed by the processor, causing the electronic device to perform the model training method in the second aspect or any possible implementation of the second aspect.

[0030] The fourth aspect and any implementation thereof correspond to the second aspect and any implementation thereof, respectively. The technical effects of the fourth aspect and any implementation thereof can be found in the technical effects of the second aspect and any implementation thereof, as described above, and will not be repeated here.

[0031] Fifthly, embodiments of this application provide a chip including one or more interface circuits and one or more processors; the interface circuits are used to receive signals from the memory of an electronic device and send signals to the processors, the signals including computer instructions stored in the memory; when the processor executes the computer instructions, it causes the electronic device to perform the voice processing method in the first aspect or any possible implementation of the first aspect.

[0032] The fifth aspect and any implementation thereof correspond to the first aspect and any implementation thereof, respectively. The technical effects of the fifth aspect and any implementation thereof are similar to those of the first aspect and any implementation thereof, and will not be repeated here.

[0033] In a sixth aspect, embodiments of this application provide a chip including one or more interface circuits and one or more processors; the interface circuits are used to receive signals from the memory of an electronic device and send signals to the processors, the signals including computer instructions stored in the memory; when the processor executes the computer instructions, it causes the electronic device to execute the model training method in the second aspect or any possible implementation of the second aspect.

[0034] The sixth aspect and any implementation thereof correspond to the second aspect and any implementation thereof, respectively. The technical effects of the sixth aspect and any implementation thereof are similar to those of the second aspect and any implementation thereof, and will not be repeated here.

[0035] In a seventh aspect, embodiments of this application provide a computer storage medium storing a computer program that, when run on a computer or processor, causes the computer or processor to execute the speech processing method in the first aspect or any possible implementation thereof.

[0036] The seventh aspect and any implementation thereof correspond to the first aspect and any implementation thereof, respectively. The technical effects of the seventh aspect and any implementation thereof are similar to those of the first aspect and any implementation thereof, and will not be repeated here.

[0037] Eighthly, embodiments of this application provide a computer storage medium storing a computer program that, when run on a computer or processor, causes the computer or processor to execute the model training method in the second aspect or any possible implementation thereof.

[0038] The eighth aspect and any implementation thereof correspond to the second aspect and any implementation thereof, respectively. The technical effects corresponding to the eighth aspect and any implementation thereof are similar to those corresponding to the second aspect and any implementation thereof, and will not be repeated here.

[0039] Ninthly, embodiments of this application provide a computer program product comprising a software program that, when executed by a computer or processor, causes the steps of the speech processing method in the first aspect or any possible implementation thereof to be performed.

[0040] The ninth aspect and any implementation thereof correspond to the first aspect and any implementation thereof, respectively. The technical effects corresponding to the ninth aspect and any implementation thereof are similar to those corresponding to the first aspect and any implementation thereof, and will not be repeated here.

[0041] In a tenth aspect, embodiments of this application provide a computer program product comprising a software program that, when executed by a computer or processor, causes the steps of the model training method in the second aspect or any possible implementation thereof to be performed.

[0042] The tenth aspect and any implementation thereof correspond to the second aspect and any implementation thereof, respectively. The technical effects corresponding to the tenth aspect and any implementation thereof are similar to those corresponding to the second aspect and any implementation thereof, and will not be repeated here.

[0043] Eleventhly, embodiments of this application provide a voice processing device, disposed in an electronic device, the device comprising:

[0044] The data acquisition module is used to acquire zoom parameters, the first video audio data of the video, and the video image data after zooming when it is determined that the video image has zoomed.

[0045] The multimodal processing module is used to perform multimodal fusion processing on the zoomed video image data and the first video audio data to obtain the second video audio data;

[0046] An audio gain adjustment module is used to zoom the second video audio data based on zoom parameters to obtain the third video audio data.

[0047] The output module is used to output third-party video and audio data.

[0048] According to aspect eleven, the multimodal processing module includes:

[0049] The preprocessing module is used to perform lip reading based on the zoomed video image data in order to extract phoneme features;

[0050] The frequency domain conversion module is used to perform frequency domain conversion based on the first video speech data to obtain the first speech features;

[0051] The feature fusion module is used to obtain the second video speech data by performing multimodal fusion of phoneme features and first speech features.

[0052] According to the eleventh aspect, or any implementation of the eleventh aspect above, the device further includes:

[0053] The first noise and reverberation suppression module is used to suppress noise and reverberation in the first video speech data using a preset suppression model to obtain the fourth video speech data.

[0054] The frequency domain conversion module is used to perform frequency domain conversion on the fourth video speech data to obtain the first speech features.

[0055] According to aspect eleven, or any implementation thereof, the feature fusion module includes: a second noise and reverberation suppression module and a speech synthesis module.

[0056] The second noise and reverberation suppression module includes:

[0057] The splicing module is used to splice the phoneme features and the first speech features to obtain the spliced ​​features;

[0058] The sequence conversion module is used to convert splicing features into second speech features using a preset sequence conversion model. The sequence conversion model is trained based on video image data and video speech data.

[0059] The speech synthesis module is used to synthesize speech based on the second speech features to obtain the second video speech data.

[0060] According to the eleventh aspect, or any of the implementations of the eleventh aspect above,

[0061] The output module is used to synthesize video data from zoomed video image data and third-party video and audio data, and output the video data.

[0062] According to the eleventh aspect, or any implementation of the eleventh aspect above, the device further includes:

[0063] The reverb addition module is used to add reverb to the third video audio data after the third video audio data is obtained;

[0064] The output module is used to output the third video audio data after adding reverb.

[0065] According to aspect eleven, or any implementation thereof, the reverb addition module includes:

[0066] The distance determination module is used to determine the distance between the target face in the zoomed video frame and the electronic device based on the zoomed video frame data. The target face is the face of the user corresponding to the first video and audio data.

[0067] The impulse response determination module is used to determine the sound field impulse response based on the distance between the target face and the electronic device;

[0068] The scene-based reverb addition module is used to add reverb to third-party video and audio data based on the sound field impulse response.

[0069] According to aspect eleven, or any implementation thereof, the reverb addition module includes:

[0070] The scene recognition module is used to perform scene recognition based on the zoomed video image data in order to determine the scene identifier corresponding to the zoomed video image data.

[0071] The scene-based reverb addition module is used to add reverb to the third video audio data based on scene identifiers, zoom parameters, and third video audio data using a preset audio reverb addition model.

[0072] According to the eleventh aspect, or any implementation of the eleventh aspect above, the device further includes:

[0073] The judgment module is used to determine whether the scene identifier is a preset scene identifier after determining the scene identifier corresponding to the zoomed video image data.

[0074] The output module is used to execute the step of outputting third video and audio data when the scene identifier is not the preset scene identifier;

[0075] The scene-based reverb addition module is used to perform the following steps when the scene identifier is a preset scene identifier: using a preset audio reverb addition model, based on the scene identifier, zoom parameters, and third video audio data, to add reverb to the third video audio data.

[0076] According to the eleventh aspect, or any of the implementations of the eleventh aspect above,

[0077] When the focal length of the camera on an electronic device changes during video recording, it is determined that the video image has zoomed in;

[0078] When the size of the video frame changes during video playback, it is determined that the video frame has zoomed in.

[0079] According to the eleventh aspect, or any of the implementations of the eleventh aspect above,

[0080] The electronic device includes at least one microphone.

[0081] The eleventh aspect and any implementation thereof correspond to the first aspect and any implementation thereof, respectively. The technical effects corresponding to the eleventh aspect and any implementation thereof can be found in the technical effects corresponding to the first aspect and any implementation thereof, as described above, and will not be repeated here.

[0082] In a twelfth aspect, embodiments of this application provide a model training apparatus, the apparatus comprising:

[0083] The data collection module is used to collect first video data and second video data. The first video data is recorded at a first distance from the sound source, and the second video data is recorded at a second distance from the sound source. The first distance is greater than the second distance. The first video data includes first video frame data and first video audio data, and the second video data includes second video frame data and second video audio data.

[0084] The forward computation module is used to perform lip reading recognition on the first video frame data and the second video frame data respectively to obtain the corresponding first phoneme features and second phoneme features; and to perform frequency domain conversion on the first video speech data and the second video speech data respectively to obtain the corresponding first speech features and second speech features.

[0085] The backpropagation module is used to train the sequence conversion model using the first phoneme feature, the second phoneme feature, the first speech feature, and the second speech feature.

[0086] According to the twelfth aspect, the first phoneme feature, the second phoneme feature, the first speech feature, and the second speech feature all include X frames, where X is a positive integer, and the sequence conversion model includes an encoder and a decoder.

[0087] The backpropagation module is used to concatenate the first speech feature and the first phoneme feature of frame X to obtain the concatenated feature of frame X. The second phoneme feature and the second speech feature of frame X are both label data, while the concatenated feature of frame X is sample data. The concatenated feature of frame X is input to the encoder to obtain the X-frame matrix output by the encoder. The X-frame matrix, the second speech feature of frame i in the second speech feature of frame X, and the speech feature of frame i-1 output by the decoder are input to the decoder. The decoder performs the i-th decoding to obtain the speech feature of frame i, the phoneme feature of frame i, and a length identifier, where i is an integer greater than 1. Based on the speech feature of frame i and the second speech feature of frame i, the phoneme feature of frame i and the second phoneme feature of frame i, and the length identifier and a preset length identifier, a loss function value is determined. The model parameters of the decoder and encoder are adjusted according to the loss function value.

[0088] The twelfth aspect and any implementation thereof correspond to the second aspect and any implementation thereof, respectively. The technical effects corresponding to the twelfth aspect and any implementation thereof are similar to those corresponding to the second aspect and any implementation thereof, and will not be repeated here. Attached Figure Description

[0089] Figure 1 This is a schematic diagram of a scene as an example.

[0090] Figure 2 This is a schematic diagram of a scene as an example.

[0091] Figure 3 This is a schematic diagram illustrating the processing procedure as an example.

[0092] Figure 4 This is a schematic diagram of the module structure as an example.

[0093] Figure 5 This is a schematic diagram of the module structure as an example.

[0094] Figure 6 This is a schematic diagram illustrating the training process as an example.

[0095] Figure 7 This is a schematic diagram illustrating the training process as an example.

[0096] Figure 8 This is a schematic diagram illustrating the training process as an example.

[0097] Figure 9 This is a schematic diagram illustrating the processing procedure as an example.

[0098] Figure 10 This is a schematic diagram illustrating the processing procedure as an example.

[0099] Figure 11 This is a schematic diagram of the module structure as an example.

[0100] Figure 12 This is a schematic diagram illustrating the processing procedure as an example.

[0101] Figure 13 This is a schematic diagram illustrating the processing procedure as an example.

[0102] Figure 14 This is a schematic diagram of the module structure as an example.

[0103] Figure 15a This is a schematic diagram of the structure of an exemplary voice processing device;

[0104] Figure 15b This is a schematic diagram of the structure of an exemplary voice processing device;

[0105] Figure 16 This is a schematic diagram of the structure of an exemplary model training device;

[0106] Figure 17 This is a schematic diagram of the structure of an exemplary device. Detailed Implementation

[0107] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0108] In this article, the term "and / or" is merely a description of the relationship between related objects, indicating that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone.

[0109] The terms "first" and "second," etc., used in the specification and claims of this application are used to distinguish different objects, not to describe a specific order of objects. For example, "first target object" and "second target object," etc., are used to distinguish different target objects, not to describe a specific order of target objects.

[0110] In the embodiments of this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design that is described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the terms "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.

[0111] In the description of the embodiments in this application, unless otherwise stated, "multiple" means two or more. For example, multiple processing units means two or more processing units; multiple systems means two or more systems.

[0112] Figure 1 This is a schematic diagram of an exemplary scenario.

[0113] Reference Figure 1 One possible scenario is when a user uses an electronic device to record a video.

[0114] Reference Figure 1 (1) For example, the video recording interface 101 includes one or more controls, including but not limited to: a stop recording option, a pause recording option, a focus adjustment option 103, etc. For example, the video recording interface 101 displays a video frame corresponding to the recorded video frame data, and the video frame includes the user 102. For example, the video recording screen 101 also displays information such as recording duration, which is not limited in this application.

[0115] For example, when a user needs to zoom in on a user in a video frame, they can increase the focal length of the electronic device's camera. (See reference...) Figure 1 (2) The user can slide the focus adjustment option 103 to the right, and the electronic device will respond to the user's operation by increasing the camera's focus. After the camera's focus is increased, the electronic device can continue recording video and display the video frame corresponding to the recorded video frame data on the video recording interface, such as... Figure 1 As shown in (3). Figure 1 (3) 104 represents the user in the video frame after the focal length is increased.

[0116] Figure 2 This is a schematic diagram of an exemplary scenario.

[0117] Reference Figure 2 One possible scenario is when a user is playing a video.

[0118] Reference Figure 2 (1) For example, the video playback interface 201 includes one or more controls, including but not limited to: an exit option, a pause playback option, etc. For example, the video playback interface 201 displays a video frame, in which the user 202 is included. For example, the video playback interface 201 also displays information such as the total video duration and the video playback duration, which are not limited in this application.

[0119] For example, the user's two fingers can be pressed respectively. Figure 2(2) The direction of the arrow indicates that when operating on the video playback interface 201, the electronic device responds to the user's operation by enlarging the video screen, such as... Figure 2 As shown in (3). Figure 2 (3) 203 refers to the user in the video frame after the video frame is magnified.

[0120] For example, regardless of whether it is by Figure 1 (1) or Figure 1 (2) to Figure 1 (3), or by Figure 2 (1) or Figure 2 (2) to Figure 2 (3) The video images have all been zoomed in. At this time, the audio in the video can also be zoomed in to improve the user's audiovisual experience.

[0121] It's worth noting that in video recording scenarios, when a user increases and then decreases the focal length, the video image will also zoom in. This can also be used to zoom in on the audio, enhancing the user's audiovisual experience. Similarly, in video playback scenarios, when a user zooms in and then zooms out, the video image will also zoom in, again enhancing the audiovisual experience.

[0122] Figure 3 This is a schematic diagram illustrating the processing procedure as an example.

[0123] Reference Figure 3 For example, an electronic device may include an audio zoom module and an output module. For example, the audio zoom module is used for audio zoom, and the output module is used to output zoomed video data, which includes zoomed video image data and zoomed video audio data. It should be understood that... Figure 3 The electronic device shown is merely one example of an electronic device, and electronic devices can have more than... Figure 3 This application does not limit the number of modules described herein, whether more or fewer.

[0124] For example, electronic devices may include, but are not limited to: cameras, camcorders, mobile phones, tablets, etc.

[0125] For example, an electronic device may have one voice acquisition module, such as a microphone, or it may have multiple voice acquisition modules; this application does not limit this.

[0126] For example, the audio zoom module may include a preprocessing module, a first noise and reverberation suppression module, a second noise and reverberation suppression module, a speech synthesis module, and an audio gain adjustment module. For example, any two or more of the preprocessing module, the first noise and reverberation suppression module, the second noise and reverberation suppression module, the speech synthesis module, and the audio gain adjustment module may be integrated into a single module; this application does not impose any limitation on this. It should be understood that... Figure 3 The audio zoom module shown is just one example of an audio zoom module, and audio zoom modules can have more than... Figure 3 This application does not limit the number of modules described herein, whether more or fewer.

[0127] For example, the preprocessing module can be used for face detection, voice detection, lip reading recognition, and distance calculation between a face and an electronic device.

[0128] For example, the first noise and reverberation suppression module is used for noise and reverberation suppression.

[0129] For example, the second noise and reverberation suppression module is used for noise and reverberation suppression.

[0130] For example, the speech synthesis module is used to synthesize video audio data.

[0131] For example, the audio gain adjustment module is used to amplify or reduce the loudness of the voice in video audio data.

[0132] Continue to refer to Figure 3 For example, the process of audio zoom in an electronic device can be as follows:

[0133] S301, input the first video and audio data to the preprocessing module.

[0134] S302 inputs zoom parameters to the preprocessing module.

[0135] S303 inputs the zoomed video image data to the preprocessing module and the output module respectively.

[0136] For example, if the current scenario is a user recording video using an electronic device, the electronic device can adjust the camera's focus in response to the user's adjustment command. Then, it uses the camera with the adjusted focus to capture video image data, thus obtaining zoomed-in video image data. Furthermore, zoom parameters can be determined based on the user's adjustment command. For example, zoom parameters can refer to the zoom magnification. The zoom parameters, the first video and audio data recorded by the electronic device, and the zoomed-in video image data (hereinafter referred to as zoomed-in video image data) are then input into the preprocessing module.

[0137] For example, if the current scenario is a user playing video on an electronic device, the system can respond to the received user's adjustment operation regarding the video screen size by adjusting the video screen size and determining the corresponding zoom parameters. Then, the zoom parameters, the first video audio data of the video data, and the zoomed video screen data can be input into the preprocessing module.

[0138] For example, while inputting the zoomed video image data into the preprocessing module, the zoomed video image data can also be input into the output module, so that the output module can generate video data and output it based on the zoomed first video audio data and the zoomed video image data.

[0139] It should be understood that S301, S302 and S303 can be executed simultaneously or sequentially, and this application does not restrict the execution order of S301, S302 and S303.

[0140] S304, the preprocessing module outputs the first video audio data to the first noise and reverberation suppression module.

[0141] S305, the preprocessing module outputs phoneme characteristics to the second noise and reverberation suppression module.

[0142] S306, the preprocessing module outputs zoom parameters to the audio gain adjustment module.

[0143] For example, after receiving the first video audio data, the preprocessing module can perform human voice detection based on the first video audio data. For instance, a VAD (Voice Activity Detection) algorithm can be used to extract features from the first video audio data, and then, based on the extracted features and the constructed Gaussian model, it can be determined whether the first video audio data contains human voice. Alternatively, the first video audio data can be input into a pre-trained deep learning model, and then, based on the output of the deep learning module, it can be determined whether the first video audio data contains human voice. The human voice detection can employ relatively mature existing technologies, which will not be elaborated upon here.

[0144] For example, after receiving the zoomed-out video image data, the preset processing module can perform face detection based on the zoomed-out video image data. For instance, the zoomed-out video image data can be input into a pre-trained deep convolutional neural network, and the output of the deep convolutional neural network can be used to determine whether the zoomed-out video image data contains a face. The face detection can employ relatively mature existing technologies, which will not be elaborated upon here.

[0145] Exemplarily, when only human voices are detected, on the one hand, S304 can be executed, that is, output the first video voice data to the first noise and reverberation suppression module. On the other hand, S306 can be executed, that is, output the zoom parameter to the audio gain adjustment module.

[0146] Exemplarily, when only faces are detected, S306 can be executed, that is, output the zoom parameter to the audio gain adjustment module.

[0147] Exemplarily, when human voices and faces are detected, on the one hand, S304 can be executed, that is, output the first video voice data to the first noise and reverberation suppression module. On the other hand, S306 can be executed, that is, output the zoom parameter to the audio gain adjustment module. On the other hand, the preprocessing module can perform lip reading based on the zoomed video frame data. Among them, faces can be continuously recognized from the zoomed video frame data, the person who is speaking can be judged, and the continuous lip movement features of this person can be extracted. Then the continuously changing features are input into the lip reading model to recognize the pronunciation corresponding to the speaker's lip shape; then according to the recognized pronunciation, the most likely natural language sentence can be calculated. Then the phoneme features can be extracted from the recognized natural language sentence, and the phoneme features are used to characterize the features of the speech content in the first video voice data. Among them, a phoneme is the smallest speech unit divided according to the natural attributes of speech. Analyzed according to the pronunciation actions in a syllable, one action constitutes one phoneme. For example, the Chinese syllable "ā" has only one phoneme, "ài" has two phonemes, "dài" has three phonemes, etc. Then S305 is executed, and the phoneme features are output to the second noise and reverberation suppression module.

[0148] Exemplarily, when no human voices and no faces are detected, then S306 can be executed, and the zoom parameter is output to the audio gain adjustment module.

[0149] It should be understood that S304 and S305 are determined whether to be executed according to the human voice detection result and the face detection result; only S304 can be executed, or S304 and S305 can be executed. Whether faces and human voices are detected or not, S306 can be executed.

[0150] S307, the first noise and reverberation suppression module outputs the fourth video voice data to the second noise and reverberation suppression module.

[0151] Exemplarily, the first noise and reverberation suppression module can be a suppression model, and the suppression model can be a CNN (Convolutional Neural Networks), whose input is video voice data and the output is also video voice data.

[0152] For example, the first noise and reverberation suppression module can be trained in advance, and then the trained first noise and reverberation suppression module can be used to process the first video speech data to output the fourth video speech data.

[0153] For example, training data can be collected. For example, two electronic devices, electronic device 1 and electronic device 2, each equipped with a voice acquisition module (such as a microphone), can be used to simultaneously acquire voice data output from the same sound source (such as electronic device 3) to collect training data. Electronic device 1 can be positioned at a first location, with a distance L1 between the first location and the location of electronic device 3, and electronic device 2 can be positioned at a second location, with a distance L2 between the second location and the location of electronic device 3. L1 is less than L2, and both L1 and L2 are positive integers, which can be set according to requirements; this application does not impose any limitations on this. Then, electronic device 3 can be controlled to play the voice data, while electronic devices 1 and 2 can be controlled to synchronously record audio. For example, after electronic device 3 plays multiple segments of voice data, electronic device 1 can acquire multiple segments of voice data 1, and electronic device 2 can acquire multiple segments of voice data 2. For example, the segments of voice data 1 and 2 acquired by electronic devices 1 and 2 from the same segment of voice data played by electronic device 3 can be used as a set of training data. For example, the speech data 1 collected by electronic device 1 can be equivalent to near-field speech data, and the speech data collected by electronic device 2 can be equivalent to far-field speech data. Near-field speech data has less reverberation and noise than far-field speech data. Therefore, speech data 1 in each training data set can be used as label data, and speech data 2 can be used as sample data.

[0154] For example, the first noise and reverberation suppression module can be trained using each set of training data. The following explanation uses a set of training data as an example. For example, the set of training data can be input into the first noise and reverberation suppression module, which performs forward computation on the sample data in the set of training data and outputs speech data. Then, based on the speech data output by the first noise and reverberation suppression module and the label data in the set of training data, a loss function value is calculated. The model parameters of the first noise and reverberation suppression module are adjusted with the goal of minimizing the loss function value; that is, backpropagation is performed. Then, this method can be used to sequentially train the first noise and reverberation suppression module model using each set of training data until the loss function value meets a first preset loss condition, or the number of training iterations meets a first preset number of training iterations, or the performance of the first noise and reverberation suppression module meets a first preset performance condition. The first preset loss condition, the first preset number of training iterations, and the first preset performance condition can all be set according to requirements, and this application does not impose any restrictions on them.

[0155] Since near-field speech data (i.e., speech data 1) contains less noise and reverberation than far-field speech data (i.e., speech data 2), training the first noise and reverberation suppression module with training data composed of near-field and far-field speech data allows the module to learn model parameters for removing noise and reverberation. Thus, using the trained first noise and reverberation suppression module, noise and reverberation can be removed from the speech data.

[0156] It should be understood that multiple sets of training data can be used to train the first noise and reverberation suppression module each time, and this application does not limit this.

[0157] It should be noted that, for example, electronic device 1 can be omitted, and only electronic device 2 can be used to record the voice data output by electronic device 3. Then, the voice data 2 collected by electronic device 2 and the voice data output by electronic device 3 can be used as training data (wherein, the voice data 2 collected by electronic device 2 is sample data, and the voice data output by electronic device 3 is label data) to train the first noise and reverberation suppression module. The training process of the first noise and reverberation suppression module described above can be referred to, and will not be repeated here.

[0158] For example, after the first noise and reverberation suppression module suppresses the noise and reverberation in the first video audio data, it can obtain the fourth video audio data; then the fourth video audio data can be output to the second noise and reverberation suppression module.

[0159] For example, assuming the first video audio data consists of N (N is a positive integer) frames, the first noise and reverberation suppression module can suppress the noise and reverberation in the first video audio data to obtain N frames of fourth video audio data.

[0160] S308, the second noise and reverberation suppression module outputs the first speech feature to the speech synthesis module.

[0161] For example, the second noise and reverberation suppression module can suppress noise and reverberation again to ensure that the subsequent video and audio data are closer to clean video and audio data. Clean video and audio data can refer to the video and audio data corresponding to the user's voice, which does not contain noise and reverberation.

[0162] Figure 4 This is a schematic diagram of a module structure as an example.

[0163] Reference Figure 4 For example, the electronic device may also include a frequency domain conversion module, and the second noise and reverberation suppression module may include a splicing module and a sequence conversion module. It should be understood that... Figure 4 The second noise and reverberation suppression module shown is merely one example of a second noise and reverberation suppression module, and the second noise and reverberation suppression module can have more than Figure 4 This application does not limit the number of modules described herein, whether more or fewer.

[0164] For example, the frequency domain conversion module can be used to convert time-domain speech data into frequency-domain speech features.

[0165] For example, the splicing module can be used to splice feature vectors.

[0166] For example, the sequence transformation module can refer to a sequence transformation model, which may include an encoder and a decoder. For example, the encoder can be used to encode the input sequence to generate a matrix. For example, the decoder can be used to decode the matrix output by the encoder to generate a sequence.

[0167] For example, the first noise and reverberation suppression module can input the fourth video-speech data to the frequency domain conversion module, which then converts the fourth video-speech data to obtain the first speech features. For example, the frequency domain conversion module can perform an STFT (Short-time Fourier transform) on the fourth video-speech data to obtain the first speech features. For example, when the fourth video-speech data includes N frames, N sets of first speech features can be obtained, with each set of first speech features corresponding to one frame of the fourth video-speech data. For example, a set of first speech features can include M1 (M1 is a positive integer) elements; that is, each set of first speech features can be an M1-dimensional sequence or vector. For example, the first speech features can be a Mel (Mel-spectrogram).

[0168] For example, assuming the first video audio data consists of N frames, and the zoomed-in video frame data also includes N frames, each zoomed-in video frame can correspond to one frame of the first or fourth video audio data. In this case, the phoneme features generated by the preprocessing module can also include N sets, with one set of phoneme features corresponding to one zoomed-in video frame. For example, one set of phoneme features can include M² (M² is a positive integer) elements; that is, each set of phoneme features can be an M²-dimensional sequence or vector.

[0169] For example, when the preprocessing module outputs phoneme features to the splicing module, the frequency domain conversion module can also output the first speech feature to the splicing module. The splicing module can then splice the first speech feature and the phoneme feature to obtain spliced ​​features, which are then output to the encoder. For example, the splicing module can splice the phoneme features obtained from lip-reading of each zoomed-out frame of video data with the first speech feature of the corresponding fourth frame of video audio data, resulting in N sets of spliced ​​features. Each set of spliced ​​features corresponds to one frame of fourth video audio data, and each set includes M3 (M3 is a positive integer) elements, where M3 = M2 + M1. That is, each set of spliced ​​features can be an M3-dimensional sequence or vector.

[0170] For example, when the preprocessing module does not output phoneme features to the splicing module, the frequency domain conversion module can output the first speech features to the splicing module, and then the splicing module can directly output the first speech features to the encoder.

[0171] For example, when the preprocessing module does not output phoneme features to the splicing module, the frequency domain conversion module can directly output the first speech features to the encoder.

[0172] For example, after receiving the splicing features, the encoder can encode the splicing features to obtain the third speech features, and then output the third speech features to the decoder. For example, the third speech features may include N groups, each group of third speech features corresponds to one frame of fourth video speech data, and each group of third speech features can be an M3*M4 matrix, where M4 is a positive integer, which can be preset according to requirements, and this application does not limit it.

[0173] For example, after receiving the third speech feature, the decoder can decode the third speech feature to obtain the second speech feature. For example, the second speech feature can include N groups, each group of second speech features corresponding to one frame of fourth video speech data. For example, the second speech feature can be a Mel spectrum.

[0174] Figure 5 This is a schematic diagram of a module structure as an example.

[0175] Reference Figure 5 (1) For example, the encoder may include a CNN and a BLTSM (Bidirectional Long Short-Term Memory) network. Data is input to the CNN and then processed sequentially by the CNN and BLSTM before being output. It should be understood that... Figure 5(1) The encoder shown is only one example of an encoder, and the encoder may have more or fewer networks than those shown in the figure, which is not limited in this application.

[0176] It should be noted that the CNN in the encoder and the CNN corresponding to the first noise and reverberation suppression module can be CNNs with different structures.

[0177] Reference Figure 5 (2) For example, the decoder may include Linear layers (i.e., fully connected layers), LTSM1 (Long Short-Term Memory), Attention layers, LSTM2, and Prenet (Progressive Recurrent Network), where Prenet may consist of multiple convolutional layers. The Linear layer may include at least one, with one output per Linear layer, and the output of the Linear layer can be fed back to the Prenet as input. It should be understood that... Figure 5 (2) The decoder shown is only one example of a decoder, and the decoder may have more or fewer networks / layers than those shown in the figure, which is not limited in this application.

[0178] For example, the inputs to the decoder training phase and the application phase are different, the number of linear layers used is different, and the outputs are also different.

[0179] The training process of the sequence conversion module is described below. For example, the encoder and decoder in the sequence conversion module can be trained jointly.

[0180] Figure 6 This is a schematic diagram illustrating the training process as an example.

[0181] Reference Figure 6 Video audio data A and video screen data A belong to the same video data, namely video data A. Video audio data B and video screen data B belong to the same video data, namely video data B.

[0182] For example, electronic device A can be positioned at location A, with a distance of LA between location A and the user's recording location, and electronic device B can be positioned at location B, with a distance of LB between location B and the user's recording location. Here, LA is greater than LB, and both LA and LB are positive integers. The specific settings can be customized as needed, and this application does not impose any restrictions on this. When the user is at the user's recording location and ready to record, electronic devices A and B can be started to record video. After a period of recording, electronic devices A and B can be stopped. Then, the video data A recorded by electronic device A and the video data B recorded by electronic device B can be aligned, and the aligned video data A and B can be used as a set of training data. Subsequently, when the user is at the user's recording location and ready to record again, electronic devices A and B can be controlled again to record video for a period of time before stopping, and the video data A and B obtained from the second recording can be aligned, using the aligned video data A and B as another set of training data for the decoder. It should be noted that this application does not limit the duration of each video recording by electronic devices A and B, nor does it limit the number of times electronic devices A and B can record video. The following example illustrates how to train a sequence conversion module using a set of training data.

[0183] For example, suppose a set of training data includes X (X is a positive integer) frames of video data A and X frames of video data B, where each frame of video data A corresponds to one frame of video data B. For example, video data A in the set of training data can be used as sample data, and video data B can be used as label data. For example, X frames of video data A include X frames of video audio data A and X frames of video frame data A, and X frames of video data B include X frames of video audio data B and X frames of video frame data B.

[0184] For example, X-frame video audio data A and X-frame video frame data A can be separated from X-frame video data A, and X-frame video audio data B and X-frame video frame data B can be separated from X-frame video data B. Then, X-frame video audio data A and X-frame video audio data B are input as two separate channels to the first noise and reverberation suppression module, and X-frame video frame data A and X-frame video frame data B are input as two separate channels to the preprocessing module.

[0185] For example, after receiving X-frame video and audio data A and X-frame video and audio data B, the first noise and reverberation suppression module performs noise and reverberation suppression on X-frame video and audio data A to obtain X-frame video and audio data A', and outputs X-frame video and audio data A' to the frequency domain conversion module. On the other hand, the first noise and reverberation suppression module performs noise and reverberation suppression on X-frame video and audio data B to obtain X-frame video and audio data B', and outputs X-frame video and audio data B' to the frequency domain conversion module. After receiving X-frame video and audio data A' and X-frame video and audio data B', the frequency domain conversion module converts X-frame video and audio data A' into corresponding X-frame audio features A, and outputs X-frame audio features A to the splicing module. On the other hand, the frequency domain conversion module converts X-frame video and audio data B' into corresponding X-frame audio features B, and outputs X-frame audio features B to the encoder's Prenet. The data processing procedures for the first noise and reverberation suppression module and the frequency domain conversion module can be referred to the description above, and will not be repeated here. For example, speech feature A and speech feature B are mel spectra.

[0186] For example, after receiving X frames of video image data A and X frames of video image data B, the preprocessing module can perform lip-reading on X frames of video image data A to obtain X-frame phoneme features A, and output X-frame phoneme features A to the splicing module. Conversely, the preprocessing module can perform lip-reading on X frames of video image data B to obtain X-frame phoneme features B, and output X-frame phoneme features B.

[0187] For example, after receiving X frames of speech features A and X frames of phoneme features A, the splicing module can splice each frame of phoneme features A and its corresponding frame of speech features A to obtain X frames of spliced ​​features. Alternatively, the splicing module can input all X frames of spliced ​​features into the encoder, which processes them to obtain an X-frame matrix. This X-frame matrix is ​​then input into the Attention layer of the decoder.

[0188] For example, the frequency domain conversion module can input X frames of speech features B into the decoder's Prenet in X parts. In other words, the splicing module inputs one frame of speech features B into the decoder's Prenet each time.

[0189] Figure 7 This is a schematic diagram illustrating the training process as an example.

[0190] Reference Figure 6 and Figure 7For example, the encoder outputs the X-frame matrix to the Attention layer of the decoder, and the frequency domain transformation module inputs the first frame speech feature B into the Prenet of the decoder, and after inputting the preset feature into the Prenet of the decoder, the decoder can begin the first decoding. For example, the preset feature can be a vector with all elements being 0, or a sequence of all zeros. For example, since the Prenet of the decoder has two inputs, the preset feature can be used as the other input when inputting the first frame speech feature B to complete the input. The first frame speech feature B, after being processed by the Prenet, yields a processing result A, which is then input into LSTM2 for processing. LSTM2 processes the input feature to obtain a processing result B, which is then input into both the Attention layer and LSTM1. The Attention layer then processes the X-frame matrix and the processing result B input from LSTM2 to obtain a processing result C, which is then output into LSTM1. LSTM1 can process the processing result C from the input of the Attention layer and the processing result B from the input of LSTM1 to obtain the corresponding processing result D. Then, LSTM1 outputs the obtained processing result D to the three fully connected layers Linear1, Linear2 and Linear3 respectively.

[0191] For example, after Linear1 receives the processing result D from the input of LSTM1, it can process the processing result D and output the phoneme features C of the first frame.

[0192] For example, after receiving the processing result D from the LSTM1 input, Linear2 can process the result D and output a length identifier of the decoded length. The decoded length can be the number of times the decoder decodes after the encoder inputs concatenated features. The length identifier can be used to characterize the length of the output speech feature C, that is, the number of frames of speech feature C, and the length identifier can be represented by a value between 0 and 1.

[0193] For example, after receiving the processing result D from the LSTM1 input, Linear3 can process the result D, output the first frame speech feature C, and input the first frame speech feature C into Prenet. For example, the speech feature C is a mel spectrum.

[0194] For example, loss function value 1 can be calculated based on the phoneme features C and B of the first frame, loss function value 2 can be calculated based on the speech features C and B of the first frame, and loss function value 3 can be calculated based on the length identifier of the first frame and a preset length identifier (the preset length identifier can be determined based on the number of frames of the video speech data B). Then, with the goal of minimizing loss function values ​​1, 2, and 3, backpropagation is performed on the encoder and decoder to adjust the model parameters of the encoder and decoder.

[0195] Continue to refer to Figure 6 and Figure 7 For example, the frequency domain transformation module inputs the speech features B of frame i (where i is an integer, ranging from 2 to X) into the Prenet of the decoder, and feeds back the speech features C of frame i-1 into the Prenet of the decoder. Then, the decoder can begin the i-th decoding operation. At this time, the Prenet processes the speech features B of frame i and the speech features C of frame i-1, obtaining a processing result A. This processing result A is then input into LSTM2 for further processing. LSTM2 processes the input features, obtaining a processing result B. This processing result B is then input into both the Attention layer and LSTM1. The Attention layer then processes the X-frame matrix and the processing result B input from LSTM2, obtaining a processing result C, which is then output into LSTM1. LSTM1 processes the processing result C input from the Attention layer and the processing result B input from LSTM1, obtaining the corresponding processing result D. Then, LSTM1 outputs the obtained processing result D to the three fully connected layers Linear1, Linear2 and Linear3 respectively.

[0196] For example, after Linear1 receives the processing result D from the input of LSTM1, it can process the processing result D and output the phoneme feature C of the i-th frame.

[0197] For example, after receiving the processing result D from the input of LSTM1, Linear2 can process the processing result D and output the length identifier of the decoded length.

[0198] For example, after receiving the processing result D from the LSTM1 input, Linear3 can process the result D to output the speech feature C of the i-th frame and input the speech feature C of the i-th frame into Prenet. For example, the speech feature C is a mel spectrum.

[0199] For example, loss function value 1 can be calculated based on the phoneme features C and B of the i-th frame, loss function value 2 can be calculated based on the speech features C and B of the i-th frame, and loss function value 3 can be calculated based on the length identifier of the i-th frame and a preset length identifier (the preset length identifier can be determined based on the number of frames of the video speech data B). Then, with the goal of minimizing loss function value 1, loss function value 2, and loss function value 3, backpropagation is performed on the encoder and decoder to adjust the model parameters of the encoder and decoder.

[0200] Continue to refer to Figure 6 and Figure 7 For example, the frequency domain conversion module inputs the speech feature B of the (i+1)th frame into the Prenet of the decoder, and feeds back the speech feature C of the ith frame into the Prenet of the decoder. Then, the decoder can begin the (i+1)th decoding. At this time, the Prenet can process the speech feature B of the (i+1)th frame and the speech feature C of the ith frame to obtain the processing result A, and then input the processing result A into LSTM2 for processing. Subsequently, the processing procedures of LSTM2, the Attention layer, LSTM1, Linear1, Linear2, and Linear3, as well as the corresponding inputs and outputs, can be referred to the description above, and will not be repeated here.

[0201] For example, training of the encoder and decoder can be stopped when the loss function value 1, loss function value 2, and loss function value 3 all satisfy the second preset loss condition, or when the number of training iterations meets the second preset training iterations, or when the performance of the second noise and reverberation suppression module meets the second preset performance condition. The second preset loss condition, the second preset number of training iterations, and the second preset performance condition can all be set as needed, and this application does not impose any restrictions on them.

[0202] In this way, compared with the prior art, this application uses the output of the previous frame as input, which enables the model to converge faster, improves training efficiency, and also improves the performance of the trained model.

[0203] Furthermore, compared to existing methods, this application employs a feature training module that combines phoneme features and speech features, increasing the variety of labels in the training module. This improves the performance of the trained model, making the speech features of the trained module closer to those of clean video speech data, and enhancing the noise and reverberation suppression effect of the trained module.

[0204] For example, in real-world scenarios, there may be situations where the preprocessing module fails to detect a face. In such cases, the preprocessing module does not need to perform lip reading detection and cannot output phoneme features to the splicing module of the second noise and reverberation suppression module. Therefore, the sequence conversion module can be trained using only video and audio data A and video and audio data B.

[0205] Figure 8 This is a schematic diagram illustrating the training process as an example.

[0206] Reference Figure 8 For example, video speech data A and video speech data B are used to train the sequence conversion module. For example, video speech data A and video speech data B can be split into two paths and input to a first noise and reverberation suppression module. Then, the first noise and reverberation suppression module performs noise and reverberation suppression on video speech data A, outputting video speech data A' to the frequency domain conversion module, and performs noise and reverberation suppression on video speech data B, outputting video speech data B' to the frequency domain conversion module. The frequency domain conversion module converts video speech data A' into speech feature A, inputting speech feature A into the CNN of the encoder; and converts video speech data B' into speech feature B, inputting speech feature B into the PreNet of the decoder. The training process of the encoder and decoder of the sequence conversion module based on speech features A and speech features B can be referred to... Figure 6 as well as Figure 7 and Figure 6 as well as Figure 7 The description will not be repeated here.

[0207] The following explains how to use the sequence conversion module.

[0208] Figure 9 This is a schematic diagram illustrating the processing procedure as an example.

[0209] For example, when the second noise and reverberation suppression module receives the phoneme features output by the preprocessing module (i.e., detects a face), the second noise and reverberation suppression module can use the following... Figure 6 After training, the sequence transformation module performs data processing, which can be referred to... Figure 9 For example, the processing procedures of the preprocessing module, the first noise and reverberation suppression module, the frequency domain conversion module, and the splicing module can be referred to the description above, and will not be repeated here.

[0210] Reference Figure 9For example, the splicing module can input N frames of spliced ​​features, obtained by splicing N frames of first speech features and N frames of phoneme features, into the CNN of the encoder. After the N frames of spliced ​​features are processed by the CNN and BLSTM in the encoder in sequence, the N frames of third speech features are output to the Attention layer of the decoder.

[0211] For example, the decoder's attention layer can begin decoding after receiving N frames of third speech features.

[0212] First decoding:

[0213] For example, during the first decoding, two preset features are input into the Prenet of the decoder. For example, the Prenet can process the preset features to obtain processing result 1, and output processing result 1 to LSTM2.

[0214] For example, LSTM2 processes result 1 to obtain result 2. Then, on the one hand, result 2 is output to the Attention layer, and on the other hand, result 2 is output to LSTM1.

[0215] For example, the Attention layer processes the processing result 2 and the third speech features of N frames to obtain the processing result 3, and then outputs the processing result 3 to LSTM1.

[0216] For example, LSTM1 processes result 3 and result 2 to obtain result 4, and then outputs result 4 to Linear2 and Linear3 respectively.

[0217] For example, after receiving processing result 4, Linear3 processes result 4 to obtain the first frame of second speech features. Then, on the one hand, it outputs the first frame of second speech features to the corresponding storage area, and on the other hand, it outputs the first frame of second speech features to Prenet. For example, since Prenet has two inputs during training, Linear3 can divide the first frame of second speech features into two paths and output them to Prenet, with each path being one frame of second speech features.

[0218] For example, after receiving processing result 4, Linear2 processes result 4 to determine the length identifier. For instance, if the current frame is the first frame and the number of frames X of the video and audio data B used during training is 100, then the length identifier is determined to be 0.01. The length identifier is then output to Prenet.

[0219] For example, after receiving the length identifier of the first frame, Prenet determines whether to proceed with the next decoding based on the length identifier. For example, it can determine whether the length identifier meets a preset stop decoding condition, where the preset stop decoding condition includes: the length identifier equals a preset length identifier, such as 1. When it is determined that the length identifier equals the preset length identifier, it can be determined that the length identifier meets the preset stop decoding condition, and decoding can be stopped. When it is determined that the length identifier is less than the preset length identifier, it can be determined that the length identifier does not meet the preset stop decoding condition, and decoding can be proceeded. For example, assuming the preset length identifier is 1 and the length identifier is 0.01, Prenet can then determine that the decoder needs to perform a second decoding.

[0220] Second decoding:

[0221] For example, Prenet can process the second speech features of the first frame to obtain processing result 1, and output processing result 1 to LSTM2.

[0222] For example, LSTM2 processes result 1 to obtain result 2. Then, on the one hand, result 2 is output to the Attention layer, and on the other hand, result 2 is output to LSTM1.

[0223] For example, the Attention layer processes the processing result 2 and the third speech features of N frames to obtain the processing result 3, and then outputs the processing result 3 to LSTM1.

[0224] For example, LSTM1 processes result 3 and result 2 to obtain result 4, and then outputs result 4 to Linear2 and Linear3 respectively.

[0225] For example, after receiving processing result 4, Linear3 processes result 4 to obtain the second frame's second speech features. Then, on one hand, it outputs the second frame's second speech features to the corresponding storage area, and on the other hand, it outputs the second frame's second speech features to Prenet. For example, Linear3 can split the second frame's second speech features into two paths and output them to Prenet.

[0226] For example, after receiving processing result 4, Linear2 processes result 4 to determine the length identifier. For instance, if the current frame is the second frame and the number of frames X of the video and audio data B used during training is 100, then the length identifier is determined to be 0.02. The length identifier is then output to Prenet.

[0227] For example, after receiving the length identifier of the second frame, Prenet determines whether to proceed with the next decoding step based on the length identifier. For example, if the length identifier equals a preset length identifier, it can be determined that the length identifier meets a preset stop decoding condition, and decoding can be stopped. If the length identifier is less than the preset length identifier, it can be determined that the length identifier does not meet the preset stop decoding condition, and decoding can proceed with the next step, at which point the decoder can perform a third decoding step. For example, assuming the preset length identifier is 1 and the length identifier is 0.02, Prenet can determine that the decoder needs to perform a third decoding step.

[0228] For example, the decoding process from the third to the Nth decoding can refer to the decoding process of the second decoding described above, and will not be repeated here. After the decoder performs the Nth decoding, Prenet can determine that the length identifier meets the stopping decoding condition. At this time, the decoder can stop decoding and then output the stored N frames of second speech features to the speech synthesis module.

[0229] Figure 10 This is a schematic diagram illustrating the processing procedure as an example.

[0230] For example, when the second noise and reverberation suppression module does not receive the output phoneme features from the preprocessing module (i.e., no face is detected), the second noise and reverberation suppression module can use... Figure 8 After training, the sequence transformation module performs data processing, which can be referred to... Figure 10 For example, the processing procedures of the preprocessing module, the first noise and reverberation suppression module, the frequency domain conversion module, and the splicing module can be referred to the description above, and will not be repeated here.

[0231] Reference Figure 10 For example, the frequency domain transformation module can input N frames of first speech features into the CNN of the encoder. After the N frames of first speech features are processed sequentially by the CNN and BLSTM in the encoder, N frames of third speech features are output to the Attention layer of the decoder. For example, after receiving the N frames of third speech features, the Attention layer of the decoder can start decoding, which can be referred to in the above description of... Figure 9 The description will not be repeated here.

[0232] S309, the speech synthesis module outputs the second video speech data to the audio gain adjustment module.

[0233] For example, after receiving N frames of second speech features, the speech synthesis module can synthesize N frames of second video speech data based on the N frames of second speech features.

[0234] Figure 11 This is a schematic diagram of a module structure as an example.

[0235] For example, the speech synthesis module includes CNN_1, a residual network, and CNN_2. For example, CNN_1 and CNN_2 may have the same network structure or different network structures, and this application does not limit this.

[0236] For example, a speech synthesis module may include at least one residual module, each residual module may include an upsampling layer and a residual network, and the residual network may include multiple convolutional layers. For example, when the speech synthesis module includes multiple residual modules, the sampling factor of the upsampling layer in each residual module may be the same or different, and this application does not impose any limitation on this. For instance, the speech synthesis module may include four residual modules: residual module 1, residual module 2, residual module 3, and residual module 4, wherein the sampling factor of the upsampling layer in residual module 1 and residual module 2 is 1 / 8, and the sampling factor of the upsampling layer in residual module 3 and residual module 4 is 1 / 2. It should be noted that this application does not limit the number of residual modules or the sampling factor of the upsampling layer in each residual module.

[0237] It should be understood that, Figure 11 The speech synthesis module shown is merely an example of a speech synthesis module, and a speech synthesis module may have more or fewer modules than those shown in the figure; this application makes no limitation in this regard.

[0238] For example, the speech synthesis module can be trained in advance, and then the second speech features can be used as input to the speech synthesis module. The speech synthesis module synthesizes the second video speech data based on the second speech features.

[0239] For example, the training process of the speech synthesis module will now be described using the aforementioned set of video data B as an example. For example, after the video data B passes through the preprocessing module, and then through the trained first noise and reverberation suppression module and the trained second noise and reverberation suppression module, the trained second noise and reverberation suppression module can output speech features Y. The specific process can be found in [reference needed]. Figure 9 And the above text is aimed at Figure 9The description will not be repeated here. Then, the X-frame speech features Y can be input into CNN_1 in the speech synthesis module, sequentially passing through CNN_1, the residual module, and CNN_2 to output X-frame video-speech data Y. The X-frame video-speech data Y is then compared with X-frame video-speech data B to determine the loss function value. The speech synthesis module is backpropagated to minimize this loss function value, adjusting the model parameters of the speech synthesis module. Furthermore, the speech synthesis module can be trained using the multiple sets of video data B obtained above until the loss function value meets the third preset loss condition, or the number of training iterations meets the third preset training iterations, or the performance of the third noise and reverberation suppression module meets the third preset performance condition. Training of the speech synthesis module can then be stopped. The third preset loss condition, the third preset number of training iterations, and the third preset performance condition can all be set according to requirements, and this application does not impose any restrictions on them.

[0240] For example, after the speech synthesis module synthesizes the second video speech data, it can output the second video speech data to the audio gain adjustment module.

[0241] S310, the audio gain adjustment module outputs the third video and audio data to the output module.

[0242] For example, the audio gain adjustment module can determine the audio loudness adjustment factor based on the zoom parameters input by the preprocessing module. For example, the audio loudness adjustment factor can be equal to the zoom factor.

[0243] For example, when the preprocessing module does not detect human voice, it can also output the first video audio data to the audio gain adjustment module, which then zooms the first video audio data. For example, the audio gain adjustment module can adjust the audio loudness of the first video audio data based on a determined audio loudness adjustment factor to obtain the zoomed-out first video audio data.

[0244] For example, when the preprocessing module detects human voice, the audio gain adjustment module can receive the second video audio data input from the speech synthesis module, and then zoom in on the second video audio data. For example, the audio gain adjustment module can adjust the audio loudness of the second video audio data based on a determined audio loudness adjustment factor to obtain the third video audio data.

[0245] For example, when the audio gain adjustment module receives the zoomed-out first video audio data, it can output the zoomed-out first video audio data to the output module. When the audio gain adjustment module receives the third video audio data, it can output the third video audio data to the output module.

[0246] S311, the output module outputs video data.

[0247] For example, the output module can align the input zoomed-out first video audio data and zoomed-out video frame data and synthesize them into video data.

[0248] For example, the output module can align the input third video audio data and the zoomed video image data and synthesize the video data.

[0249] For example, if the current scenario is a video recording scenario, the output module can store the obtained video data in the corresponding storage area. If the current scenario is a video playback scenario, the output module can output the obtained video data to the playback module, which will then play the video data.

[0250] In summary, employing a second noise and reverberation suppression module with superior noise and reverberation suppression capabilities allows for cleaner removal of noise and reverberation from video and audio data, resulting in cleaner zoomed-in video and audio data. This ensures that the audio loudness of reverberation and noise in the recorded audio data remains unchanged after zooming. This improves the quality of the zoomed-in video and audio data, enhancing the user's listening experience and ultimately improving the user experience.

[0251] In one possible scenario, the reverberation recorded at a zoom distance from the sound source can be simulated, and then the reverberation can be added to the third video audio data to increase the realism of the output audio and further improve the user experience.

[0252] One possible approach is to simulate reverberation based on the distance between the target face and the electronic device; where the target face is the face of the user corresponding to the first video and audio data.

[0253] Figure 12 This is a schematic diagram illustrating the processing procedure as an example.

[0254] Reference Figure 12 For example, an electronic device may include an audio zoom module and an output module. It should be understood that... Figure 12 The electronic device shown is merely an example of an electronic device, and electronic devices may have more or fewer modules than those shown in the figure; this application does not limit this.

[0255] For example, the audio zoom module may include a preprocessing module, a first noise and reverberation suppression module, a second noise and reverberation suppression module, a speech synthesis module, an audio gain adjustment module, and a reverberation addition module. For example, any two or more of the preprocessing module, the first noise and reverberation suppression module, the second noise and reverberation suppression module, the speech synthesis module, the audio gain adjustment module, and the reverberation addition module can be integrated into a single module; this application does not impose any limitation on this. It should be understood that... Figure 12The audio zoom module shown is merely an example of an audio zoom module, and an audio zoom module may have more or fewer modules than those shown in the figure; this application makes no limitation in this regard.

[0256] For example, the functions of the preprocessing module, the first noise and reverberation suppression module, the second noise and reverberation suppression module, the speech synthesis module, and the audio gain adjustment module can be referred to the description above, and will not be repeated here. For example, the preprocessing module is also used to calculate the distance between the face and the electronic device.

[0257] For example, the reverb adding module is used to add reverb.

[0258] Reference Figure 12 For example, the process of audio zoom in an electronic device can be as follows:

[0259] S1201, input the first video and audio data to the preprocessing module.

[0260] S1202 inputs zoom parameters to the preprocessing module.

[0261] S1203 inputs the zoomed video image data to the preprocessing module and the output module respectively.

[0262] S1204, the preprocessing module outputs the first video audio data to the first noise and reverberation suppression module.

[0263] S1205, the preprocessing module outputs phoneme characteristics to the second noise and reverberation suppression module.

[0264] S1206, the preprocessing module outputs zoom parameters to the audio gain adjustment module.

[0265] For example, S1201 to S1206 can be referred to the description of S301 to S306 above, and will not be repeated here.

[0266] S1207, the preprocessing module outputs the distance between the target face and the electronic device to the reverb addition module.

[0267] For example, when the preprocessing module detects a face, it can calculate the distance between the target face and the electronic device.

[0268] One possible approach is to calculate the target face area based on the zoomed-in video footage, and then calculate the target face's area relative to the electronic device screen, i.e., the target face's screen-to-body ratio. Then, based on this ratio, the distance between the target face and the electronic device can be calculated. Specifically, empirical values ​​for the face's screen-to-body ratio and empirical values ​​for the distance between the face and the electronic device can be derived through historical statistics or experiments. Then, based on the calculated target face screen-to-body ratio, this correlation can be used to determine the distance between the target face and the electronic device.

[0269] In one possible approach, the target face area can be calculated based on the zoomed video image data, and the distance between the target face and the electronic device can be obtained based on the functional relationship between the target face area and the distance between the face and the electronic device.

[0270] One possible approach is to use dual cameras with two inputs to perform dual-range measurement and calculate the distance between the target face and the electronic device.

[0271] One possible approach is to use depth devices such as structured light in electronic devices to measure the distance between the target face and the electronic device.

[0272] It should be noted that this application does not limit the method by which the preprocessing module determines the distance between the face and the electronic device.

[0273] S1208, the first noise and reverberation suppression module outputs the fourth video and audio data to the second reverberation suppression module.

[0274] S1209, the second noise and reverberation suppression module outputs the second speech feature to the speech synthesis module.

[0275] S1210, the speech synthesis module outputs the second video speech data to the audio gain adjustment module.

[0276] S1211, the audio gain adjustment module outputs the third video audio data to the reverb addition module.

[0277] The exemplary steps S1208 to S1211 can be referred to the description of S127 to S110 above, and will not be repeated here.

[0278] S1212, the reverb addition module outputs the fifth video and audio data to the output module.

[0279] For example, after receiving the distance between the target face and the electronic device, the reverberation addition module can determine the sound field impulse response based on this distance. The sound field impulse response refers to the sequence of signals radiated by a pulse sound source received at the receiving position in the sound field.

[0280] For example, when the sound field is a room, the sound field impulse response is the room impulse response. For instance, the length, width, and height of the room, the location of the sound source (i.e., the user corresponding to the first video and audio data), and the distance between the sound source and the electronic device (i.e., the distance between the target face and the electronic device) can be estimated; then, based on the length, width, and height of the room, the location of the sound source, and the distance between the sound source and the electronic device, the room impulse response can be calculated.

[0281] For example, assuming the sound source is located at the center of the room, and the room's length, width, and height are all twice the distance between the sound source and the electronic device, the location of the sound source, as well as the room's length, width, and height, can be determined based on this distance. It should be understood that other methods can also be used to determine the location of the sound source, and the room's length, width, and height; this application does not limit the methods used in this regard.

[0282] For example, after determining the sound field impulse response, the reverb addition module can perform calculations based on the sound field impulse response and the third video audio data to add reverb to the third video audio data, thus obtaining the fifth video audio data. This allows the addition of reverb recorded at the zoom distance from the sound source to the video audio data after noise suppression and reverb reduction, enhancing the listening experience of the video audio data.

[0283] In other words, the fifth video audio data is video audio data with reverb added. Then the reverb adding module can output the fifth video audio data to the output module.

[0284] S1213, the output module outputs video data.

[0285] For example, after receiving the fifth video audio data and the zoomed video frame data, the output module can align the fifth video audio data and the zoomed video frame data and synthesize them into video data, then output the synthesized video data. In this way, the audio in the output video data is reverberant, making the audio data heard by the user after playing the video data more realistic and improving the user experience.

[0286] Figure 13 This is a schematic diagram illustrating the processing procedure as an example.

[0287] Reference Figure 13 For example, an electronic device may include an audio zoom module and an output module. It should be understood that... Figure 13 The electronic device shown is merely an example of an electronic device, and electronic devices may have more or fewer modules than those shown in the figure; this application does not limit this.

[0288] For example, the audio zoom module may include a preprocessing module, a first noise and reverberation suppression module, a second noise and reverberation suppression module, a speech synthesis module, an audio gain adjustment module, and a reverberation addition module. For example, any two or more of the preprocessing module, the first noise and reverberation suppression module, the second noise and reverberation suppression module, the speech synthesis module, the audio gain adjustment module, and the reverberation addition module can be integrated into a single module; this application does not impose any limitation on this. It should be understood that... Figure 13 The audio zoom module shown is merely an example of an audio zoom module, and an audio zoom module may have more or fewer modules than those shown in the figure; this application makes no limitation in this regard.

[0289] For example, the functions of the preprocessing module, the first noise and reverberation suppression module, the second noise and reverberation suppression module, the speech synthesis module, and the audio gain adjustment module can be referred to the description above, and will not be repeated here.

[0290] For example, the reverb adding module is used to add reverb.

[0291] For example, the process of audio zoom in an electronic device can be as follows:

[0292] S1301, input the first video and audio data to the preprocessing module.

[0293] S1302, input zoom parameters to the preprocessing module.

[0294] For example, S1301 to S1302 can be referred to the description of S301 to S302 above.

[0295] S1303 inputs the zoomed video image data to the preprocessing module, reverb addition module, and output module respectively.

[0296] For example, in addition to inputting the zoomed video image data into the preprocessing module and the output module, the zoomed video image data can also be input into the reverb adding module, so that the subsequent reverb adding module can add reverb to the third video and audio data based on the zoomed video image data.

[0297] S1304, the preprocessing module outputs the first video audio data to the first noise and reverberation suppression module.

[0298] S1305, the preprocessing module outputs phoneme characteristics to the second noise and reverberation suppression module.

[0299] S1306, the preprocessing module outputs zoom parameters to the audio gain adjustment module.

[0300] S1307, the first noise and reverberation suppression module outputs the fourth video and audio data to the second reverberation suppression module.

[0301] S1308, the second noise and reverberation suppression module outputs the second speech feature to the speech synthesis module.

[0302] S1309, the speech synthesis module outputs the second video speech data to the audio gain adjustment module.

[0303] S1310, the audio gain adjustment module outputs the third video audio data to the reverb addition module.

[0304] The exemplary steps S1007 to S1010 can be referred to the description of S307 to S310 above, and will not be repeated here.

[0305] S1311, the reverb addition module outputs the fifth video audio data to the output module.

[0306] Figure 14 This is a schematic diagram of a module structure as an example.

[0307] Reference Figure 14 For example, the reverb addition module may include a scene recognition module and a scene-based reverb addition module. The scene recognition module is used for scene identification, and the scene-based reverb addition module is used to add reverb to the video and audio data; it can be an audio reverb addition model.

[0308] Continue to refer to Figure 14 For example, in S1303, the zoomed video image data is input to the scene recognition module. The scene recognition module can identify the scene in the zoomed video image data and determine the corresponding scene identifier. The scene identifier includes, but is not limited to: outdoor label, bedroom label, living room label, office label, vehicle label, restaurant label, venue label, outdoor label, etc. This application embodiment does not limit this, and one scene identifier corresponds to one scene.

[0309] For example, the scene recognition module can be a scene recognition model, which can be a neural network. For example, the scene recognition module can be pre-trained, and then used to recognize zoomed-in video footage. For example, training video footage data corresponding to various scenes (such as outdoors, bedroom, living room, office, car interior, restaurant, stadium, etc.) can be pre-collected, and a corresponding reference scene identifier can be set for each training video footage data according to the scene it corresponds to, with one reference scene identifier corresponding to one scene. For example, for each scene, the collected training video footage data can include multiple images, and then one training video footage data and its corresponding reference scene identifier can be used as a set of training data, thus obtaining multiple sets of training data.

[0310] For example, the training of a scene recognition module using a set of training data will be used as an example. This set of training data can be input into the scene recognition module, which can perform forward computation on the training video frame data in the set of training data and output scene identifiers. Then, the scene identifiers output by the scene recognition module can be compared with reference scene identifiers in the set of training data to calculate the corresponding loss function value. The scene recognition module is then trained with the goal of minimizing this loss function value. The scene recognition module can then be trained using multiple sets of training data in the same way until the loss function value meets the fourth preset loss condition, or the number of training iterations meets the fourth preset number of training iterations, or the performance of the scene recognition module meets the fourth preset performance condition. The fourth preset loss condition, the fourth preset number of training iterations, and the fourth preset performance condition can all be set according to requirements, and this application does not impose any restrictions on them.

[0311] It should be understood that multiple sets of training data can be used to train the scene recognition module each time, and this application does not impose any restrictions on this.

[0312] For example, after the zoomed video image data is input into the trained scene recognition module, the trained scene recognition module can perform scene recognition on the zoomed video image data and output the scene identifier for the zoomed video image data to the scene-based reverb addition module.

[0313] For example, all indoor tags (such as outdoor tags, bedroom tags, living room tags, office tags, car tags, restaurant tags, and venue tags) can be set as preset tags. For example, when the scene identifier determined by the scene recognition module is a preset tag, the scene identifier can be output to the scene-based reverb addition module. When the scene identifier determined by the scene recognition module is not a preset tag, the audio gain adjustment module can be instructed to output the third video and audio data to the output module.

[0314] Continue to refer to Figure 14 For example, the scene recognition module outputs a scene identifier to the scene-based reverb addition module, the audio gain adjustment module outputs the fourth video-audio data to the scene-based reverb addition module, and the preprocessing module outputs zoom parameters to the scene-based reverb addition module. The scene-based reverb addition module can process the received third video-audio data, scene identifier, and zoom parameters to add reverb to the third video-audio data, thus obtaining the fifth video-audio data.

[0315] For example, the scene-based reverb addition module can be an audio reverb addition model, which can also be a neural network. For example, the scene-based reverb addition module can be pre-trained, and then the trained scene-based reverb addition module can be used to add reverb to the third video audio data. For example, training data can be collected. For example, electronic devices D, E, and F can be deployed in scene 1 (e.g., a bedroom). Electronic device F can be used as the sound source, and its position can be fixed. Electronic device D is set at position D, with a distance of LD between position D and electronic device F, and electronic device E is set at position E, with a distance of LE between position E and electronic device F. Here, LD is greater than LE, and both LD and LE are positive integers, which can be set according to requirements; this application does not impose any restrictions on this. When electronic device F starts playing audio data, electronic devices D and E can be started to record audio. After recording for a period of time, electronic devices D and E can be controlled to stop recording. For example, the audio data D that electronic device D can record, and the audio data E that electronic device E can record, can be used to record audio data. For example, zoom parameters can be calculated based on the camera parameters of electronic device D and the distance between electronic device F and electronic device D. Then, the zoom parameters corresponding to electronic device D at position D1, the voice data E recorded by electronic device E, the measured voice data D recorded by electronic device D, and the scene identifier corresponding to scene 1 are used as a set of training data. For example, when electronic device D is deployed at position D and electronic device E is deployed at position E, controlling electronic device F to play various different voice data and controlling electronic devices D and E to record audio can yield multiple sets of training data. For example, in scene 1, electronic device D can also be deployed to other positions, and then controlling electronic device F to play various different voice data and controlling electronic devices D and E to record audio can also yield multiple sets of training data. For example, in scene 2 (such as a venue), electronic device F can be fixed, and electronic device E can be fixed at position E, and then electronic device D can be deployed sequentially at different positions. When electronic device D is deployed at a certain position, controlling electronic device F to play voice data and controlling electronic devices D and E to record audio can also yield multiple sets of training data. This application does not limit the number of scenarios for collecting training data, the location of electronic device D in different scenarios, or the duration and quantity of voice data played by electronic device F.

[0316] For example, for each set of training data, the scene identifier, speech data E, and zoom parameters in that set of training data can be used as sample data, and the speech data D can be used as label data to train the scene-based reverb addition module. For example, the training of the scene-based reverb addition module using a set of training data will be illustrated below. For example, a set of training data can be input into the scene-based reverb addition module, which performs forward calculations based on the scene identifier, speech data E, and zoom parameters in that set of training data, outputting reverb-enhanced video and speech data. Then, the reverb-enhanced video and speech data output by the scene-based reverb addition module is compared with the label data in that set of training data, a loss function value is calculated, and the scene-based reverb addition module is trained with the goal of minimizing this loss function value. Then, the scene-based reverb addition module can be trained using multiple sets of training data in the above manner until the loss function value meets the fifth preset loss condition, or the number of training iterations meets the fifth preset number of training iterations, or the performance of the scene-based reverb addition module meets the fifth preset performance condition. The fifth preset loss condition, the fifth preset number of training iterations, and the fifth preset performance condition can all be set as needed, and this application does not impose any restrictions on them.

[0317] It should be noted that, for example, electronic device E can also be omitted. Instead, the voice data played by electronic device F, scene identifiers, zoom parameters of electronic device D, and voice data D collected by electronic device D can be collected as a set of training data (wherein, the voice data played by electronic device F, scene identifiers, and zoom parameters of electronic device D are sample data, and the voice data D collected by electronic device D is label data) to train the scene-based reverb addition module. Refer to the description above, which will not be repeated here.

[0318] It should be understood that multiple sets of training data can be used to train the scene-based reverberation addition module each time, and this application does not limit this.

[0319] S1312, the output module outputs video data.

[0320] For example, S1312 can be referred to the description of S1213 above, and will not be repeated here.

[0321] In this way, by combining sound field and zoom parameters, it is possible to simulate a more realistic reverberation recorded at a zoom distance from the sound source, thereby further improving the realism of the zoomed audio.

[0322] Figure 15a This is a schematic diagram of the structure of an exemplary voice processing device.

[0323] Reference Figure 15aFor example, the voice processing device 1500 is disposed in an electronic device.

[0324] Reference Figure 15a For example, the voice processing device 1500 includes:

[0325] The data acquisition module 1501 is used to acquire zoom parameters, the first video audio data of the video, and the video image data after zooming when it is determined that the video image has zoomed.

[0326] The multimodal processing module 1502 is used to perform multimodal fusion processing on the zoomed video image data and the first video audio data to obtain the second video audio data.

[0327] The audio gain adjustment module 1503 is used to zoom the second video audio data based on the zoom parameters to obtain the third video audio data.

[0328] Output module 1504 is used to output third video and audio data.

[0329] For example, when the data acquisition module 1501 determines that the video frame has zoomed, it can acquire zoom parameters, first video audio data, and video frame data after zooming. Then, on one hand, the data acquisition module 1501 inputs the first video audio data and the video frame data after zooming to the multimodal processing module 1502; on the other hand, it inputs the zoom parameters to the audio gain adjustment module 1503. After receiving the first video audio data and the video frame data after zooming, the multimodal processing module 1502 can perform multimodal fusion processing on the zoomed video frame data and the first video audio data to obtain second video audio data; then, it inputs the second video audio data to the audio gain adjustment module 1503. After receiving the zoom parameters and the second video audio data, the audio gain adjustment module 1503 can zoom the second video audio data based on the zoom parameters to obtain third video audio data, and then inputs the third video audio data to the output module 1504, which outputs the third video audio data. Thus, through the multimodal fusion processing of the multimodal processing module 1502, noise and reverberation in the video and audio data are effectively suppressed. Consequently, the audio gain adjustment module 1503 only zooms on the video and audio data after noise and reverberation suppression, improving the quality of the zoomed video and audio data and enhancing the user experience.

[0330] Figure 15b This is a schematic diagram of the structure of an exemplary voice processing device.

[0331] Reference Figure 15b For example, the multimodal processing module 1502 includes:

[0332] The preprocessing module 15021 is used to perform lip reading based on the zoomed video image data in order to extract phoneme features;

[0333] The frequency domain conversion module 15022 is used to perform frequency domain conversion based on the first video speech data to obtain the first speech features;

[0334] The feature fusion module 15023 is used to obtain second video speech data by performing multimodal fusion of phoneme features and first speech features.

[0335] Reference Figure 15b For example, the voice processing device 1500 also includes:

[0336] The first noise and reverberation suppression module 1505 is used to suppress noise and reverberation in the first video speech data using a preset suppression model to obtain the fourth video speech data.

[0337] The frequency domain conversion module 15022 is used to perform frequency domain conversion on the fourth video speech data to obtain the first speech feature.

[0338] Reference Figure 15b For example, the feature fusion module 15023 includes: a second noise and reverberation suppression module 150231 and a speech synthesis module 150232.

[0339] The second noise and reverberation suppression module 150231 includes:

[0340] The splicing module 1502311 is used to splice phoneme features and first speech features to obtain spliced ​​features;

[0341] The sequence conversion module 1502312 is used to convert splicing features into second speech features using a preset sequence conversion model. The sequence conversion model is trained based on video image data and video speech data.

[0342] The speech synthesis module 150232 is used to perform speech synthesis based on the second speech features to obtain the second video speech data.

[0343] For example, the output module 1504 is used to synthesize video data using zoomed video image data and third video audio data, and output video data.

[0344] Reference Figure 15b For example, the voice processing device 1500 also includes:

[0345] The reverb addition module 1506 is used to add reverb to the third video audio data after the third video audio data is obtained;

[0346] Output module 1504 is used to output the third video audio data after adding reverb.

[0347] Reference Figure 15b For example, the reverb adding module 1506 includes:

[0348] The distance determination module 15061 is used to determine the distance between the target face in the zoomed video frame and the electronic device based on the zoomed video frame data. The target face is the face of the user corresponding to the first video and audio data.

[0349] The impulse response determination module 15062 is used to determine the sound field impulse response based on the distance between the target face and the electronic device;

[0350] The impulse response-based reverb addition module 15063 is used to add reverb to third-party video audio data based on the impulse response of the sound field.

[0351] Reference Figure 15b For example, the reverb adding module 1506 includes:

[0352] The scene recognition module 15064 is used to perform scene recognition based on the zoomed video image data in order to determine the scene identifier corresponding to the zoomed video image data.

[0353] The scene-based reverb addition module 15065 is used to add reverb to the third video audio data based on scene identifiers, zoom parameters, and third video audio data using a preset audio reverb addition model.

[0354] Reference Figure 15b For example, the voice processing device 1500 also includes:

[0355] The judgment module 1507 is used to determine whether the scene identifier is a preset scene identifier after determining the scene identifier corresponding to the zoomed video image data.

[0356] Output module 1504 is used to execute the step of outputting third video and audio data when the scene identifier is not the preset scene identifier;

[0357] The scene-based reverb addition module 15065 is used to perform the following steps when the scene identifier is a preset scene identifier: using a preset audio reverb addition model, and adding reverb to the third video audio data based on the scene identifier, zoom parameters, and third video audio data.

[0358] For example, when the focal length of the camera of an electronic device changes during video recording, it is determined that the video image has zoomed; when the size of the video image changes during video playback, it is determined that the video image has zoomed.

[0359] For example, the electronic device includes at least one microphone.

[0360] Figure 16 This is a schematic diagram of the structure of a model training device as an example.

[0361] Reference Figure 16 For example, the model training device 1600 includes:

[0362] The data collection module 1601 is used to collect first video data and second video data. The first video data is recorded at a first distance from the sound source, and the second video data is recorded at a second distance from the sound source. The first distance is greater than the second distance. The first video data includes first video frame data and first video audio data, and the second video data includes second video frame data and second video audio data.

[0363] The forward computing module 1602 is used to perform lip reading recognition on the first video frame data and the second video frame data respectively to obtain the corresponding first phoneme features and second phoneme features; and to perform frequency domain conversion on the first video speech data and the second video speech data respectively to obtain the corresponding first speech features and second speech features.

[0364] The backpropagation module 1603 is used to train the sequence conversion model using the first phoneme feature, the second phoneme feature, the first speech feature, and the second speech feature.

[0365] For example, the data collection module 1601 can collect first video data and second video data, and then input the first video data and second video data into the forward computing module 1602. After receiving the first video data and second video data, the forward computing module 1602 can perform lip-reading recognition on the first video data and the second video data respectively to obtain the corresponding first phoneme features and second phoneme features; and perform frequency domain transformation on the first video speech data and the second video speech data respectively to obtain the corresponding first speech features and second speech features; then input the first phoneme features, second phoneme features, first speech features and second speech features into the backpropagation module 1603. After receiving the first phoneme features, second phoneme features, first speech features and second speech features, the backpropagation module 1603 can use the first phoneme features, second phoneme features, first speech features and second speech features to train the sequence conversion model. In this way, by using multimodal data to train the sequence conversion model, the noise and reverberation suppression effect of the sequence conversion module can be improved.

[0366] For example, the first phoneme feature, the second phoneme feature, the first speech feature, and the second speech feature all include X frames, where X is a positive integer, and the sequence conversion model includes an encoder and a decoder.

[0367] The backpropagation module 1603 is used to concatenate the first speech feature and the first phoneme feature of frame X to obtain the concatenated feature of frame X; wherein, the second phoneme feature and the second speech feature of frame X are both label data, and the concatenated feature of frame X are sample data; the concatenated feature of frame X is input to the encoder to obtain the X-frame matrix output by the encoder; the X-frame matrix, the second speech feature of frame i in the second speech feature of frame X, and the speech feature of frame i-1 output by the decoder are input to the decoder, and the decoder performs the i-th decoding to obtain the speech feature of frame i, the phoneme feature of frame i, and the length identifier output by the decoder, where i is an integer greater than 1; the loss function value is determined based on the speech feature of frame i and the second speech feature of frame i, the phoneme feature of frame i and the second phoneme feature of frame i, and the length identifier and the preset length identifier, and the model parameters of the decoder and the encoder are adjusted according to the loss function value.

[0368] In one example, Figure 17 A schematic block diagram illustrating an embodiment of the present application shows an apparatus 1700. The apparatus 1700 may include a processor 1701 and a transceiver / transceiver pin 1702, and optionally, a memory 1703.

[0369] The various components of device 1700 are coupled together via bus 1704, which includes a data bus, a power bus, a control bus, and a status signal bus. However, for clarity, all buses are referred to as bus 1704 in the figure.

[0370] Optionally, the memory 1703 can be used for the instructions in the foregoing method embodiments. The processor 1701 can be used to execute the instructions in the memory 1703, control the receive pin to receive signals, and control the transmit pin to transmit signals.

[0371] Device 1700 may be an electronic device or a chip of an electronic device in the above method embodiments.

[0372] All relevant content of each step involved in the above method embodiments can be referenced from the functional description of the corresponding functional module, and will not be repeated here.

[0373] This embodiment also provides a computer storage medium storing computer instructions. When the computer instructions are executed on an electronic device, the electronic device performs the aforementioned method steps to implement the speech processing method and model training method in the above embodiment.

[0374] This embodiment also provides a computer program product that, when run on a computer, causes the computer to perform the aforementioned steps to implement the speech processing method and model training method in the above embodiments.

[0375] In addition, embodiments of this application also provide an apparatus, which may specifically be a chip, component or module. The apparatus may include a connected processor and a memory. The memory is used to store computer execution instructions. When the apparatus is running, the processor can execute the computer execution instructions stored in the memory to cause the chip to execute the speech processing method and model training method in the above-described method embodiments.

[0376] In this embodiment, the electronic device, computer storage medium, computer program product or chip are all used to execute the corresponding method provided above. Therefore, the beneficial effects that can be achieved can be referred to the beneficial effects of the corresponding method provided above, and will not be repeated here.

[0377] Through the above description of the embodiments, those skilled in the art will understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.

[0378] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another apparatus, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0379] The units described as separate components may or may not be physically separate. A component shown as a unit can be one or more physical units; that is, it can be located in one place or distributed in multiple different locations. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0380] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0381] Any content in the various embodiments of this application, as well as any content in the same embodiment, can be freely combined. Any combination of the above content is within the scope of this application.

[0382] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, in essence, or the parts that contribute to the prior art, or all or part of the technical solutions, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0383] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

[0384] The steps of the methods or algorithms described in conjunction with the embodiments of this application can be implemented in hardware or by a processor executing software instructions. The software instructions can consist of corresponding software modules, which can be stored in random access memory (RAM), flash memory, read-only memory (ROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, portable hard disks, CD-ROMs, or any other form of storage medium well known in the art. An exemplary storage medium is coupled to a processor, enabling the processor to read information from and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can reside in an ASIC.

[0385] Those skilled in the art will recognize that the functions described in the embodiments of this application in one or more of the above examples can be implemented using hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or code on a computer-readable medium. Computer-readable media include computer storage media and communication media, wherein communication media include any medium that facilitates the transfer of a computer program from one place to another. Storage media can be any available medium that can be accessed by a general-purpose or special-purpose computer.

[0386] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. A voice processing method, characterized by, The method is applied to an electronic device, and the method comprises: When it is determined that zooming occurs in a video picture, a zooming parameter, first video voice data of the video, and video picture data after zooming are obtained; Second video voice data is obtained by performing multi-modal fusion processing on the video picture data after zooming and the first video voice data; Third video voice data is obtained by zooming the second video voice data based on the zooming parameter; and the third video voice data is outputted; The second video voice data is obtained by performing multi-modal fusion processing on the video picture data after zooming and the first video voice data, and the method comprises: Lip speech recognition is performed based on the video picture data after zooming to extract phoneme features; Frequency domain conversion is performed based on the first video voice data to obtain first voice features; The second video voice data is obtained by performing multi-modal fusion on the phoneme features and the first voice features.

2. The method of claim 1, wherein, Before the frequency domain conversion based on the first video voice data is performed to obtain the first voice features, the method further comprises: Noise and reverberation in the first video voice data are suppressed by using a preset suppression model to obtain fourth video voice data; The frequency domain conversion based on the first video voice data is performed to obtain the first voice features, and the method comprises: The fourth video voice data is subjected to frequency domain conversion to obtain the first voice features.

3. The method according to claim 1 or 2, characterized in that, The second video voice data is obtained by performing multi-modal fusion processing on the phoneme features and the first voice features, and the method comprises: The phoneme features and the first voice features are spliced to obtain spliced features; The spliced features are converted into second voice features by using a preset sequence conversion model, and the sequence conversion model is obtained based on video picture data and video voice data; Voice synthesis is performed based on the second voice features to obtain the second video voice data.

4. The method according to claim 1 or 2, characterized in that, The third video voice data is outputted, and the method comprises: Video data is synthesized by using the video picture data after zooming and the third video voice data, and the video data is outputted.

5. The method according to claim 1 or 2, characterized in that, After the third video voice data is obtained, the method further comprises: Reverberation is added to the third video voice data; The third video voice data is outputted, and the method comprises: The third video voice data to which reverberation is added is outputted.

6. The method of claim 5, wherein, The third video voice data is added with reverberation, and the method comprises: A distance between a target face in the video picture data after zooming and the electronic device is determined based on the video picture data after zooming, and the target face is a face of a user corresponding to the first video voice data; A sound field impulse response is determined based on the distance between the target face and the electronic device; Reverberation is added to the third video voice data based on the sound field impulse response.

7. The method of claim 5, wherein, The third video voice data is added with reverberation, and the method comprises: Scene recognition is performed based on the video picture data after zooming to determine a scene identifier corresponding to the video picture data after zooming; The preset audio reverberation adding model is used to add reverberation to the third video voice data based on the scene identifier, the zoom parameter and the third video voice data.

8. The method of claim 7, wherein, After the scene identifier corresponding to the zoomed video picture data is determined, the method further includes: determining whether the scene identifier is a preset scene identifier; when the scene identifier is not the preset scene identifier, performing the step of outputting the third video voice data; when the scene identifier is the preset scene identifier, performing the step of using the preset audio reverberation adding model to add reverberation to the third video voice data based on the scene identifier, the zoom parameter and the third video voice data.

9. The method of any one of claims 1, 2, 6-8, wherein: when the focal length of the camera of the electronic device changes during the recording of the video, it is determined that the video picture zooms; when the size of the video picture changes during the playing of the video, it is determined that the video picture zooms.

10. The method of any one of claims 1, 2, 6-8, wherein: the electronic device includes at least one microphone.

11. An electronic device, comprising: including: a memory and a processor, the memory being coupled to the processor; the memory stores program instructions, when the program instructions are executed by the processor, the electronic device executes the voice processing method of any one of claims 1 to 10.

12. A chip, characterized by one or more interface circuits and one or more processors; the interface circuit is used to receive signals from the memory of the electronic device and send the signals to the processor, the signals include computer instructions stored in the memory; when the processor executes the computer instructions, the electronic device executes the voice processing method of any one of claims 1 to 10.

13. A computer storage medium, characterized in that The computer readable storage medium stores a computer program, when the computer program runs on the computer or the processor, the computer or the processor executes the method as claimed in any one of claims 1 to 10.

14. A computer program product, characterised in that, The computer program product includes a software program, when the software program is executed by a computer or a processor, the steps of the method of any one of claims 1 to 10 are executed. The computer program product includes a software program, when the software program is executed by a computer or a processor, the steps of the method of any one of claims 1 to 10 are executed.

Citation Information

Patent Citations

  • Video processing method and electronic equipment

    CN110602424A

  • Hearing aid systems and methods

    CN113196803A