Audio processing method, device, electronic device and storage medium

Through the merger and reconstruction technology of audio features and tone characteristics, the problem of difficulty and high cost of dubbing actors in overseas markets is solved, and the tone of multiple characters is automatically converted, meeting the dialogue requirements of film and television dramas, reducing costs and improving dubbing efficiency.

CN114842858BActive Publication Date: 2025-08-22CHENGDU IQIYI INTELLIGENT INNOVATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210457487.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-27
Publication Date
2025-08-22
Estimated Expiration
2042-04-27

AI Technical Summary

Technical Problem

In overseas markets, it is difficult and expensive to find dubbing actors who meet quality requirements, and it is difficult for the existing technology to use manual dubbing on a large scale, especially due to the problem of character tone type requirements and dubbing schedule matching.

Method used

By extracting the dubbing features and tone features in the target video audio file, combining them for audio reconstruction, and generating a new audio file with the target tone, realizing the effect of automatically converting the tone of one dubbing object into multiple tone, while retaining the dubbing content and emotions.

Benefits of technology

The conversion of multiple roles can be achieved without additional voice actors, meeting the requirements of dialogue between film and television dramas, reducing costs and improving dubbing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114842858B_ABST
    Figure CN114842858B_ABST
Patent Text Reader

Abstract

The present invention relates to an audio processing method, device, electronic device and storage medium, wherein the audio processing method includes: obtaining a first audio file corresponding to a target video, extracting dubbing features corresponding to the dubbing content of a first dubbing object in the first audio file, wherein the first language in the first audio file is different from the second language in the original audio file of the target video; obtaining timbre features corresponding to a second dubbing object, wherein the first dubbing object and the second dubbing object have different timbres; merging the dubbing features and the timbre features to obtain an audio spectrum; and performing audio reconstruction based on the audio spectrum to obtain a second audio file corresponding to the target video. The embodiment of the present application can automatically convert the timbre of the first dubbing object into the timbre of the second dubbing object while retaining the content and emotion of the dubbing of the first dubbing object.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to an audio processing method, device, electronic device, and storage medium. Background Art

[0002] As my country's culture goes global at an increasingly faster pace, a large number of domestic film and television dramas go overseas every year, while a large number of foreign-language film and television dramas are also introduced into China. Localized dubbing of film and television dramas has become an important constraint for the landing of domestic dramas overseas or for overseas dramas in China.

[0003] However, finding sufficient voice actors overseas with the right quality for the roles is far more difficult than finding Chinese voice actors, and the time and cost of finding dubbing resources are even higher. Furthermore, very few Chinese dubbing professionals are currently employed, while the existing stock of foreign-language films and TV series is even larger and the cost of dubbing a single episode is very high, making manual dubbing even more unaffordable. Furthermore, the characters in the dubbed films and TV series generally have specific requirements for their voices and variety, and the availability and compatibility of voice actors are major factors hindering the widespread use of manual dubbing. Summary of the Invention

[0004] In order to solve the above technical problems or at least partially solve the above technical problems, the present application provides an audio processing method, device, electronic device and storage medium.

[0005] In a first aspect, the present application provides an audio processing method, comprising:

[0006] Obtaining a first audio file corresponding to a target video, and extracting dubbing features corresponding to dubbing content of a first dubbing object in the first audio file, wherein a first language in the first audio file is different from a second language in the original audio file of the target video;

[0007] Acquiring a timbre feature corresponding to a second dubbing object, wherein the first dubbing object and the second dubbing object have different timbres;

[0008] Combining the dubbing feature and the timbre feature to obtain an audio spectrum;

[0009] Audio reconstruction is performed based on the audio spectrum to obtain a second audio file corresponding to the target video.

[0010] Optionally, obtaining a first audio file corresponding to the target video includes:

[0011] Obtaining an original audio file, a dubbed audio file, a first dialogue text, and a second dialogue text corresponding to a target video, wherein the first dialogue text is obtained by performing speech recognition on the dubbed audio file and does not include character information; the dubbed audio file is dubbed in a first language different from a second language of the original audio file of the target video; and the second dialogue text corresponds to the original audio file and includes character information;

[0012] Determining, based on the target video, the first dialogue text, and the original audio file, a speaking time period in which a face belonging to the same character speaks and the dialogue content corresponding to the speaking time period;

[0013] Separating the dubbing audio file into audio tracks according to the speaking time period, the dialogue content corresponding to the speaking time period, and the second dialogue text to obtain the speaking time period of each character and the audio file corresponding to the time period;

[0014] The audio file corresponding to the time period when any character speaks is determined as the first audio file corresponding to the target video.

[0015] Optionally, determining, based on the target video, the first dialogue text, and the original audio file, a speech time period during which faces belonging to the same character speak and dialogue content corresponding to the speech time period includes:

[0016] Extracting a timestamp of a face appearance in the target video;

[0017] Extract the timestamp of the voiceprint segment in the original audio file;

[0018] extracting the timestamp of the first language dialogue segment from the first dialogue text;

[0019] Matching the timestamp of the voiceprint segment with the timestamp of the face appearance to obtain the speaking time period of the faces belonging to the same role;

[0020] The time period in which the face belonging to the same character speaks is matched with the timestamp of the first language dialogue segment to obtain the dialogue content corresponding to the speaking time period.

[0021] Optionally, the dubbing audio file is divided into audio tracks according to the speaking time period, the dialogue content corresponding to the speaking time period, and the second dialogue text, to obtain the speaking time period of each character and the audio file corresponding to the time period, including:

[0022] Matching the speech time period, the dialogue content corresponding to the speech time period, and the second dialogue text to obtain the speech time period of each character;

[0023] The dubbing audio file is divided into audio tracks according to the time period of each character's speech to obtain audio files corresponding to the time period.

[0024] Optionally, the dubbing feature includes: a content feature, and extracting the dubbing feature corresponding to the dubbing content of the first dubbing object in the first audio file includes:

[0025] Inputting the first audio file into a preset speech recognition encoder to obtain recognition content;

[0026] The identified content is input into a preset content encoder to obtain the content features.

[0027] Optionally, the dubbing feature includes a rhythm feature, and extracting the dubbing feature corresponding to the dubbing content of the first dubbing object in the first audio file includes:

[0028] Inputting the first audio file into a preset speech self-supervised learning pre-training model to obtain output data;

[0029] The output data is input into a preset prosody encoder to obtain the prosody feature.

[0030] Optionally, obtaining a second timbre feature corresponding to the second dubbing object includes:

[0031] Obtaining the original audio file of the target video;

[0032] Extracting the original voiceprint features of the original dubbing object from the original audio file;

[0033] Searching for a voiceprint identifier corresponding to the original voiceprint feature in a preset voiceprint library;

[0034] The timbre feature of the dubbing object corresponding to the voiceprint identifier is determined as the second timbre feature of the second dubbing object.

[0035] Optionally, after performing audio reconstruction based on the audio spectrum to obtain a second audio file corresponding to the target video, the method further includes:

[0036] Obtaining the original audio file of the target video;

[0037] Performing volume detection on the original audio file to obtain a plurality of first volume values ​​corresponding to the first timestamps;

[0038] searching the second audio file for a second volume value corresponding to each of the first timestamps;

[0039] If the difference between the first volume value and the second volume value is greater than a preset threshold, the second volume value is adjusted to the first volume value to obtain an adjusted second audio file.

[0040] Optionally, after performing audio reconstruction based on the audio spectrum to obtain a second audio file corresponding to the target video, the method further includes:

[0041] Performing sound effect detection on the original audio file to obtain a plurality of sound effect types corresponding to the second timestamps;

[0042] Sound effects are added to the second audio file according to the sound effect types corresponding to the multiple second timestamps to obtain an adjusted second audio file.

[0043] In a second aspect, the present application provides an audio processing device, comprising:

[0044] a first acquisition module configured to acquire a first audio file corresponding to a target video and extract content features and rhythmic features of a dubbing content of a first dubbing object in the first audio file, wherein a first language in the first audio file is different from a second language in the original audio file of the target video;

[0045] A second acquisition module is used to acquire a timbre feature corresponding to a second dubbing object, wherein the first dubbing object and the second dubbing object have different timbres;

[0046] a merging module, configured to merge the content feature, the rhythm feature, and the timbre feature to obtain an audio spectrum;

[0047] A reconstruction module is used to perform audio reconstruction based on the audio spectrum to obtain a second audio file corresponding to the target video.

[0048] Optionally, the first acquisition module includes:

[0049] a first acquisition unit configured to acquire an original audio file, a dubbing audio file, a first dialogue text, and a second dialogue text corresponding to a target video, wherein the first dialogue text is obtained by performing speech recognition on the dubbing audio file and does not include character information; the dubbing audio file is dubbed in a first language different from a second language of the original audio file of the target video; and the second dialogue text corresponds to the original audio file and includes character information;

[0050] A first determining unit is configured to determine, based on the target video, the first dialogue text, and the original audio file, a speech time period during which a face belonging to the same character speaks and the dialogue content corresponding to the speech time period;

[0051] A track splitting unit is configured to split the dubbing audio file into audio tracks according to the speech time period, the dialogue content corresponding to the speech time period, and the second dialogue text, to obtain the speech time period of each character and the audio file corresponding to the time period;

[0052] The second determining unit is configured to determine an audio file corresponding to a time period in which any character speaks as a first audio file corresponding to the target video.

[0053] Optionally, the first determining unit includes:

[0054] A first extraction subunit is used to extract a timestamp of a face appearance in the target video;

[0055] The second extraction subunit is used to extract the timestamp of the voiceprint appearance segment from the original audio file;

[0056] a third extraction subunit, configured to extract the timestamp of the first language speech segment from the first speech text;

[0057] A first matching subunit is configured to match the timestamp of the voiceprint segment with the timestamp of the face appearance to obtain a speech time period in which the faces belonging to the same character spoke;

[0058] The second matching subunit is used to match the time period when the face of the same character speaks with the timestamp of the first language dialogue segment to obtain the dialogue content corresponding to the speaking time period.

[0059] Optionally, the track division unit includes:

[0060] A third matching subunit is configured to match the speech time period, the dialogue content corresponding to the speech time period, and the second dialogue text to obtain the speech time period of each character;

[0061] The track division sub-unit is used to divide the dubbing audio file into audio tracks according to the time period of each character's speech, and obtain audio files corresponding to the time period.

[0062] Optionally, the dubbing feature includes: content feature, and the first acquisition module includes:

[0063] A first input unit is used to input the first audio file into a preset speech recognition encoder to obtain recognition content;

[0064] The second input unit is configured to input the identified content into a preset content encoder to obtain the content feature.

[0065] Optionally, the dubbing feature includes a rhythm feature, and the first acquisition module includes:

[0066] A third input unit is used to input the first audio file into a preset speech self-supervised learning pre-training model to obtain output data;

[0067] The fourth input unit is configured to input the output data into a preset prosody encoder to obtain the prosody feature.

[0068] Optionally, the second acquisition module includes:

[0069] A second acquiring unit, configured to acquire the original audio file of the target video;

[0070] An extraction unit, configured to extract the original voiceprint features of the original dubbing object from the original audio file;

[0071] A first searching unit is configured to search a preset voiceprint library for a voiceprint identifier corresponding to the original voiceprint feature;

[0072] The third determining unit is configured to determine the timbre feature of the dubbing object corresponding to the voiceprint identifier as the second timbre feature of the second dubbing object.

[0073] Optionally, after the reconstruction unit, the device further includes:

[0074] A third acquisition module is used to obtain the original audio file of the target video;

[0075] a volume detection module, configured to perform volume detection on the original audio file to obtain first volume values ​​corresponding to a plurality of first timestamps;

[0076] a first search module, configured to search the second audio file for a second volume value corresponding to each of the first timestamps;

[0077] The volume adjustment module is configured to adjust the second volume value to the first volume value if the difference between the first volume value and the second volume value is greater than a preset threshold, thereby obtaining an adjusted second audio file.

[0078] Optionally, after the consumption unit, the device further includes:

[0079] a sound effect detection module, configured to perform sound effect detection on the original audio file to obtain a plurality of sound effect types corresponding to the second timestamps;

[0080] The sound effect adjustment module is used to add sound effects to the second audio file according to the sound effect types corresponding to the multiple second time stamps to obtain an adjusted second audio file.

[0081] In a third aspect, the present application provides an electronic device, comprising a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other via the communication bus;

[0082] Memory for storing computer programs;

[0083] The processor is configured to implement any of the audio processing methods described in the first aspect when executing a program stored in the memory.

[0084] In a fourth aspect, the present application provides a computer-readable storage medium, on which a program of an audio processing method is stored. When the program of the audio processing method is executed by a processor, the steps of any of the audio processing methods described in the first aspect are implemented.

[0085] The above technical solution provided by the embodiment of the present application has the following advantages compared with the prior art:

[0086] The present application only retains the dubbing features of the first dubbing object in the first audio file, does not use the timbre features of the first dubbing object, merges the dubbing features of the first dubbing object with the timbre features of the second dubbing object, so that the second audio file reconstructed based on the merged audio spectrum can have the timbre of the second dubbing object and retain the dubbing features, thereby automatically converting the timbre of the first dubbing object into the timbre of the second dubbing object while retaining the content and emotion of the dubbing of the first dubbing object. Furthermore, it is convenient to convert the timbre of the first dubbing object in all first audio files corresponding to the target video file into the timbre of the corresponding second dubbing object respectively, without the need for other dubbing actors, to achieve the effect of having one first dubbing object match the timbre of multiple second dubbing objects, and at the same time retain the rich emotions of the first dubbing object's speech, thereby meeting the requirements of film and television drama scenes for dialogue. BRIEF DESCRIPTION OF THE DRAWINGS

[0087] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.

[0088] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0089] Figure 1 A flowchart of an audio processing method provided in an embodiment of the present application;

[0090] Figure 2 Provided in the embodiments of this application Figure 1 A flow chart of step S101;

[0091] Figure 3 Provided in the embodiments of this application Figure 1 A flow chart of step S102;

[0092] Figure 4 Another flow chart of an audio processing method provided in an embodiment of the present application;

[0093] Figure 5 Another flow chart of an audio processing method provided in an embodiment of the present application;

[0094] Figure 6 A schematic diagram illustrating the principle of an audio processing method in practical application provided by an embodiment of the present application;

[0095] Figure 7 A schematic diagram of the principle of a sound conversion model in practical application provided by an embodiment of the present application;

[0096] Figure 8 A schematic diagram illustrating the principles of another audio processing method in practical applications provided by an embodiment of the present application;

[0097] Figure 9 A schematic diagram of the principle of another sound conversion model in practical application provided by an embodiment of the present application;

[0098] Figure 10 A structural diagram of an audio processing device provided in an embodiment of the present application;

[0099] Figure 11 A structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0100] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0101] Finding sufficient voice actors overseas with the right quality for the roles is far more difficult than finding Chinese voice actors, and the time and expense involved in finding dubbing resources are even higher. Furthermore, very few Chinese dubbing professionals are currently employed, while the existing stock of foreign-language films and TV series is enormous and the cost of dubbing a single episode is very high. Using only manual dubbing is even more unaffordable. Furthermore, the characters in the dubbed films and TV series generally have specific requirements for their voices and variety, and the availability and compatibility of voice actors are major factors hindering the widespread use of manual dubbing.

[0102] To this end, the embodiments of the present application provide an audio processing method, device, electronic device, storage and medium, so that in the scenario of video export, that is, when a video originally dubbed in Chinese is exported overseas or a video originally dubbed in a non-Chinese manner is imported into China, the method automatically retains only the dubbing features of the first dubbing object in the first audio file, does not use the timbre features of the first dubbing object, and merges the dubbing features of the first dubbing object with the timbre features of the second dubbing object, so that the second audio file reconstructed based on the merged audio spectrum can have the timbre of the second dubbing object and retain the dubbing features, thereby automatically converting the timbre of the first dubbing object into the timbre of the second dubbing object while retaining the content and emotion of the dubbing of the first dubbing object. Furthermore, it is convenient to convert the timbre of the first dubbing object in all first audio files corresponding to the target video file into the timbre of the corresponding second dubbing object, without the need for other dubbing actors, to achieve the effect of having one first dubbing object match the timbre of multiple second dubbing objects, while retaining the rich emotions of the first dubbing object's speech, thereby meeting the requirements of film and television drama scenes for dialogue.

[0103] like Figure 1 As shown, the embodiment of the present application provides an audio processing method, which may include the following steps:

[0104] Step S101: Acquire a first audio file corresponding to a target video, and extract dubbing features corresponding to the dubbing content of a first dubbing object in the first audio file.

[0105] In one embodiment of the present application, the target video refers to a video to be exported overseas. The target video may correspond to an original audio file and a dubbing audio file. The original audio file refers to an audio file dubbed in Chinese, that is, the second language may exemplarily refer to Chinese. In actual applications, in order to export the target video overseas so that overseas people can understand it without watching subtitles, the target video needs to be dubbed into the corresponding language. Therefore, the dubbing audio file refers to an audio file dubbed in a language other than Chinese for the target video, that is, the first language may exemplarily refer to a non-Chinese language. In actual applications, the dubbing audio file is generally dubbed by a first dubbing object. In order to complete the dubbing quickly, it can also be completed by two or more first dubbing objects. The first dubbing object refers to a dubbing person who can dub the target video into languages ​​other than Chinese. Other languages ​​may exemplarily be small languages ​​in Southeast Asia, etc.

[0106] In another embodiment of the present application, the target video refers to a video to be imported into the country. The target video may correspond to an original audio file and a dubbing audio file. The original audio file refers to an audio file using non-Chinese dubbing, that is, the second language can exemplarily refer to a non-Chinese language. In actual applications, in order to import the target video into the country so that domestic people can understand it without watching subtitles, it is necessary to dub the target video from other languages ​​into Chinese. The other languages ​​can exemplarily be small languages ​​in Southeast Asia, etc. Therefore, the dubbing audio file refers to an audio file using Chinese dubbing for the target video, that is, the first language can exemplarily refer to Chinese. In actual applications, the dubbing audio file is generally completed by a first dubbing object. In order to complete the dubbing quickly, it can also be completed by two or more first dubbing objects. The first dubbing object refers to a dubbing person who can dub the target video into Chinese.

[0107] The first audio file corresponds to a character with lines in the target video. Each character may correspond to at least one audio file. The first audio file may be part of a dubbing audio file. Therefore, the first language in the first audio file is different from the second language in the original audio file of the target video.

[0108] The dubbing features are used to characterize the content, emotional state, and rhythm of the first dubbing object. For example, the dubbing features may include content features and rhythm features.

[0109] In this step, the first audio files corresponding to the target video can be obtained one by one in a preset order, and then the dubbing features of the dubbing content of the first dubbing object can be extracted from the first audio files. In actual applications, a pre-trained encoder can be used to extract the dubbing features.

[0110] Step S102: Acquire the timbre feature corresponding to the second dubbing object.

[0111] Since the target video generally has multiple characters with lines, the timbres of these characters are generally different, and each character generally corresponds to a timbre. In order to dub the voices of different characters in the target video using different timbres, it is necessary to select a dubbing person with a different timbre from the first dubbing object, that is, the second dubbing object.

[0112] In an embodiment of the present application, a timbre feature library may be pre-constructed, which stores timbre features of multiple second dubbing objects. In this step, the timbre features corresponding to the second dubbing objects may be obtained from the timbre feature library.

[0113] Step S103, combining the dubbing feature and the timbre feature to obtain an audio spectrum;

[0114] In this step, a decoder may be used to combine and decode the dubbing features and the timbre features to obtain an audio spectrum. For example, the audio spectrum may refer to a Mel spectrum.

[0115] Step S104: reconstructing audio based on the audio spectrum to obtain a second audio file corresponding to the target video.

[0116] A vocoder may be used to perform audio reconstruction on the audio spectrum to reconstruct the audio into a playable waveform file to obtain a second audio file.

[0117] Based on the above steps, since only the dubbing features in the first audio file are retained and the dubbing features are merged with the timbre features of the second dubbing object, it is equivalent to replacing the timbre features of the first dubbing object in the first audio file with the timbre features of the second dubbing object to obtain the second audio file. In actual applications, this process can be repeated for each first audio file in the dubbing audio file in this way, and the timbre features of the first dubbing object in different first audio files are replaced with the timbre features of corresponding different second dubbing objects. When multiple second audio files are combined to obtain a complete target dubbing file, the dubbing audio file obtained by dubbing one first dubbing object is converted into a target audio file dubbed by multiple second dubbing objects.

[0118] In actual applications, the second audio file generated by audio reconstruction may have certain mechanical sounds, current sounds and a certain degree of noise, so this application uses DSP technology to denoise and repair the sound to ensure the playback effect of the second audio file.

[0119] The present application only retains the dubbing features of the first dubbing object in the first audio file, does not use the timbre features of the first dubbing object, merges the dubbing features of the first dubbing object with the timbre features of the second dubbing object, so that the second audio file reconstructed based on the merged audio spectrum can have the timbre of the second dubbing object and retain the dubbing features, thereby automatically converting the timbre of the first dubbing object into the timbre of the second dubbing object while retaining the content and emotion of the dubbing of the first dubbing object. Furthermore, it is convenient to convert the timbre of the first dubbing object in all first audio files corresponding to the target video file into the timbre of the corresponding second dubbing object respectively, without the need for other dubbing actors, to achieve the effect of having one first dubbing object match the timbre of multiple second dubbing objects, and at the same time retain the rich emotions of the first dubbing object's speech, thereby meeting the requirements of film and television drama scenes for dialogue.

[0120] In another embodiment of the present application, the first audio file corresponding to the target video is obtained in step S101, such as Figure 2 Shown, including:

[0121] Step S201: Obtain the original audio file, dubbing audio file, first dialogue text, and second dialogue text corresponding to the target video.

[0122] In the embodiment of the present application, the first dialogue text is obtained by performing speech recognition on the dubbing audio file, and does not contain character information. The dubbing audio file is obtained by dubbing in a first language different from the second language of the original audio file of the target video. The second dialogue text corresponds to the original audio file and contains character information. For example, the second dialogue text can be obtained by translating a dialogue script in the second language into the first language. The dialogue script in the second language contains the character information, so the second dialogue text contains the character information.

[0123] Step S202, determining a speech time period during which a face belonging to the same character speaks and the speech content corresponding to the speech time period based on the target video, the first dialogue text, and the original audio file;

[0124] In this step, a face appearance timestamp may be first extracted from the target video. Specifically, face recognition may be performed on each image frame of the target video, and the moment of the current image frame is recorded when a face is recognized to obtain a face appearance timestamp.

[0125] Then, extract the timestamp of the voiceprint appearance segment from the original audio file. Specifically, perform voiceprint recognition in the original audio file. When a voiceprint is recognized, record the time when the voiceprint is recognized to obtain the timestamp of the voiceprint appearance segment.

[0126] Then, extracting the timestamp of the first language speech segment from the first speech text. Specifically, character recognition can be performed on the first speech text. When a speech segment is recognized, the current time is recorded to obtain the timestamp of the first language speech segment.

[0127] Then, the timestamp of the voiceprint segment and the timestamp of the face appearance are matched, and faces belonging to the same character are determined, and the speaking time period of the faces belonging to the same character is obtained, that is, the time period in which each face belonging to the same character speaks;

[0128] Finally, the time period when the faces belonging to the same character speak can be matched with the timestamps of the first language dialogue segments to obtain the dialogue content corresponding to the speaking time period, that is, what lines each face belonging to the same character said during its speaking time period.

[0129] Step S203, dividing the dubbing audio file into audio tracks according to the speaking time period, the dialogue content corresponding to the speaking time period, and the second dialogue text, to obtain the speaking time period of each character and the audio file corresponding to the time period;

[0130] In this step, the speaking time period, the dialogue content corresponding to the speaking time period and the second dialogue text are matched to obtain the speaking time period of each character; the dubbing audio file is divided into audio tracks according to the speaking time period of each character to obtain the audio file corresponding to the time period.

[0131] For example, the first language lines corresponding to a certain face within a period of time are:

[0132] "Face No. 3 01:10:01 Mom";

[0133] "Face No. 3 01:10:02 I'm going";

[0134] "Face No. 3 goes to school at 01:10:03";

[0135] "Face No. 2 01:10:06 OK";

[0136] "Face No. 2 01:10:10 on the way";

[0137] "Face No. 2 01:10:11 Be careful";

[0138] The corresponding second language lines in the second line text are:

[0139] "Xiaohong 01:10:01 Mom";

[0140] "Xiaohong 01:10:02 I'll go";

[0141] "Xiaohong goes to school at 01:10:03";

[0142] "Mom 01:10:06 OK";

[0143] "Mom 01:10:10 on the way";

[0144] "Mom 01:10:11 Be careful";

[0145] By matching the two, we can get that the character corresponding to face number 3 is Xiaohong. The time period when Xiaohong said the line "Mom, I'm going to school" is from 01:10:01 to 01:10:03, and the time period when Mom said the line "Be careful on the road" is from 01:10:06 to 01:10:11. Therefore, the dubbing audio file can be split into different audio tracks at 01:10:01 to 01:10:03 and 01:10:06 to 01:10:11 to obtain audio file A corresponding to Xiaohong in the time period, and audio file B corresponding to Mom in the time period from 01:10:06 to 01:10:11.

[0146] In actual applications, before splitting the tracks, you can use a speech recognition algorithm to compare the dialogue text with the recognized text to detect errors such as missing tracks. When missing tracks are detected, manual error correction is performed to improve the accuracy of the audio file after splitting the tracks.

[0147] Step S204: determine the audio file corresponding to the time period when any character speaks as the first audio file corresponding to the target video.

[0148] In this step, each audio file can be determined as the first audio file corresponding to the target video one by one in a certain order. For example, audio file A is determined as the first audio file. After audio reconstruction is performed based on the first audio file to obtain the second audio file, audio file B is determined as the first audio file, and so on.

[0149] The embodiment of the present application can automatically divide a complete dubbing audio file into multiple first audio files, so as to replace the timbre of the first dubbing object in the first audio file corresponding to each character with the timbre of the second dubbing object to obtain a second audio file.

[0150] In another embodiment of the present application, the dubbing features include content features and rhythm features. Step S101 extracts the content features and rhythm features corresponding to the dubbing content of the first dubbing object in the first audio file, including:

[0151] The first audio file is input into a preset speech recognition encoder to obtain recognition content, and the recognition content is input into a preset content encoder to obtain the content features.

[0152] For example, the first audio file (Source Audio) can be input into the end-to-end automatic speech recognition encoder (E2EASR Encoder) to obtain the recognized content (BN), and the recognized content (BN) can be input into the content encoder (content encoder) to obtain the content feature (content vector).

[0153] The first audio file is input into a preset speech self-supervised learning pre-training model to obtain output data, and the output data is input into a preset prosody encoder to obtain the prosody feature.

[0154] For example, the first audio (Source Audio) can be input into a speech self-supervised learning pre-trained model (VQ-wav2vec pre-trained model) to obtain output data (VQW2V), and then the output data (VQW2V) can be input into a prosody encoder to obtain a prosody feature (prosody vector).

[0155] The embodiment of the present application can automatically extract content features and rhythm features respectively by model, so as to retain only the content features and rhythm features of the first audio file, and then merge them with the timbre features of the second dubbing object to obtain the second audio file after audio reconstruction.

[0156] In another embodiment of the present application, step S102 obtains the second timbre feature corresponding to the second dubbing object, such as Figure 3 Shown, including:

[0157] Step S301, obtaining the original audio file of the target video;

[0158] Step S302, extracting the original voiceprint features of the original dubbing object from the original audio file;

[0159] In the embodiment of the present application, the original dubbing object corresponds to the first audio file, that is, the original dubbing object refers to the dubbing actor used to dub the character corresponding to the first audio file.

[0160] In the original audio file, different characters will be dubbed by different voice actors. In order to reflect the voice characteristics of the characters and be closer to the original audio file, when replacing the timbre, a voice actor whose voice print is closer to the original voice actor can be selected. Therefore, in this step, the original voice print characteristics of the original voice actor can be extracted.

[0161] Step S303, searching a preset voiceprint library for a voiceprint identifier corresponding to the original voiceprint feature;

[0162] In order to facilitate the storage of different voiceprints and the timbre features corresponding to the voiceprints, a voiceprint database can be pre-built. The voiceprint database stores the correspondence between multiple groups of voiceprint identifiers, voiceprint features and timbre features. In actual applications, the voice audio of multiple second dubbing objects can be collected, the voiceprint is calculated (VoicePrint database calculate) on the voice audio, the calculated voiceprint features are stored in the voiceprint database, and the voice audio is input into the speaker encoder (speaker encoder) to extract the feature vector of the short speech of the specified speaker to obtain the timbre features of the second dubbing object.

[0163] In this step, the original voiceprint features extracted in step S302 can be used to query each voiceprint feature in the voiceprint database (VoicePrint query) to obtain the voiceprint identifier corresponding to the successfully matched voiceprint feature.

[0164] Step S304: Determine the timbre feature of the dubbing object corresponding to the voiceprint identifier as the second timbre feature of the second dubbing object.

[0165] The embodiment of the present application can automatically obtain the timbre characteristics of the original dubbing object in the original audio file, and when the timbre characteristics are replaced, the replaced timbre characteristics are closer to the timbre characteristics of the dubbing actor in the original audio file and more in line with the plot.

[0166] In another embodiment of the present application, after performing audio reconstruction based on the audio spectrum in step S104 to obtain a second audio file corresponding to the target video, as shown in FIG. Figure 4 As shown, the method further includes:

[0167] Step S401, obtaining the original audio file of the target video;

[0168] Step S402: performing volume detection on the original audio file to obtain a plurality of first volume values ​​corresponding to first timestamps;

[0169] The original audio file in the embodiment of the present application can correspond to the same time period as the first audio file. For example, if the time period corresponding to the first audio file in the entire dubbing audio file is 00:05:20-00:05:30, the original audio file should also select the audio segment from 00:05:20-00:05:30.

[0170] In this step, a sliding window mechanism can be used to detect the volume of the original audio file. The time slice length of the sliding window (TimeSlide) can be selected according to actual needs, such as 0.5 seconds, to obtain multiple groups of volume values ​​with first timestamps, that is, multiple first volume values ​​corresponding to the first timestamps.

[0171] Step S403: searching the second audio file for a second volume value corresponding to each of the first timestamps;

[0172] In this step, the second volume values ​​corresponding to the first timestamps may be searched for one by one in the second audio file to obtain a plurality of second volume values ​​corresponding to the first timestamps.

[0173] Step S404: If the difference between the first volume value and the second volume value is greater than a preset threshold, the second volume value is adjusted to the first volume value to obtain an adjusted second audio file.

[0174] The difference between the first volume value and the second volume value can be calculated and compared with a preset threshold. If the difference is greater than the preset threshold, it indicates that the difference between the two is too large and the second volume value needs to be adjusted to make the second volume value closer to or equal to the first volume value.

[0175] The embodiment of the present application can ensure that the volume value corresponding to each time point in the second audio file is the same as or similar to the volume value corresponding to the corresponding time point in the original audio file, maintain the volume stability, and avoid the volume of the second audio file from fluctuating.

[0176] In another embodiment of the present application, after performing audio reconstruction based on the audio spectrum in step S104 to obtain a second audio file corresponding to the target video, as shown in FIG. Figure 5 As shown, the method further includes:

[0177] Step S501: Perform sound effect detection on the original audio file to obtain a plurality of sound effect types corresponding to second timestamps;

[0178] In the embodiments of the present application, sound effects refer to effects created by sound, and refer to special effects added to the soundtrack to enhance the realism, atmosphere or dramatic message of a scene, such as the effect of a human voice on the phone, the effect of a human voice in a cave, etc.

[0179] In actual applications, an end-to-end model can be used to perform sound effect detection and sampling on the original audio file to obtain a sound effect type with a timestamp. The end-to-end model of the embodiment of the present application can support 9 types of sound effect types: low-pass sound effect type, high-pass sound effect type, band-pass sound effect type, band-pass (with gain) sound effect type, full-pass sound effect type, peak sound effect type, lowshelf sound effect type, highshelf sound effect type and notch sound effect type, etc.

[0180] Step S502: Add sound effects to the second audio file according to the sound effect types corresponding to the multiple second timestamps to obtain an adjusted second audio file.

[0181] Since different sound effect types are added to the original dubbing file to suit the plot requirements, in order to make the second audio file more restored to the original dubbing file, it has the same sound effect type. Therefore, according to the sound effect type detected at the second timestamp, the same sound effect can be added to the position of the second timestamp in the second audio file to achieve a better fit with the original dubbing file.

[0182] The embodiment of the present application can ensure that the sound effect type corresponding to each time point in the second audio file is the same as the sound effect type corresponding to the corresponding time point in the original audio file, thereby avoiding the situation where the second audio file does not add the corresponding sound effect, resulting in the user having difficulty understanding the plot, and improving the user's sense of immersion in watching the video.

[0183] For ease of understanding, this application also provides an embodiment of an audio processing method in a practical application scenario, as follows:

[0184] like Figure 6 As shown, after the target video whose original audio file is in Chinese is translated into non-Chinese, an amateur non-Chinese dubbing artist A performs the overall dubbing to obtain a dubbing audio file.

[0185] Manual dubbing materials often have errors such as missing tracks, so the dubbing audio file can be detected for missing tracks and corrected to obtain a corrected dubbing audio file.

[0186] The subtitle file translated from Chinese into non-Chinese does not contain character information, and character information is necessary for selecting the voice conversion model (VC) of the character and the second dubbing object. Therefore, the subtitle character needs to be split to add character information to the subtitle file. Specifically: the face appearance timestamp is extracted from the target video, the voiceprint appearance segment timestamp is extracted from the original audio file, and the non-Chinese dialogue segment appearance timestamp is extracted from the first dialogue text. The voiceprint appearance segment timestamp, the face appearance timestamp, and the non-Chinese dialogue segment appearance timestamp are merged to obtain the corresponding dialogue content of each face in different time periods. The corresponding dialogue content of each face in different time periods is matched with the second dialogue text to obtain the audio time period corresponding to each character. Based on the audio time period corresponding to each character, intelligent track division is performed to obtain the audio files corresponding to each character in different time periods, namely: character track 1, character track 2...character track N. Each character track is equivalent to any first audio file in the aforementioned embodiment.

[0187] Each character's track is input into the voice conversion model separately. The role of the voice conversion model is to retain the emotion, rhythm and content of amateur non-Chinese voice actor A, and only replace his timbre characteristics with the timbre characteristics of non-Chinese voice actor B, non-Chinese voice actor C or non-Chinese voice actor D. Specifically, Figure 7 As shown, a character track (i.e., a first audio file) voiced by an amateur non-Chinese dubbing artist A is input into an encoder. The encoder extracts content A and rhythm A from the character track. In practice, timbre A may not be extracted. Voiceprint casting can be performed as needed, whereby a voice suitable for a film or TV drama dubbing scene is selected from non-Chinese voices B, C, and D. In actual operation, a suitable voice ID (i.e., the voiceprint identifier in the aforementioned embodiment) is selected, and the timbre features corresponding to the voice ID are obtained. Content A, rhythm A, and the obtained timbre features are then input into a decoder. The decoder combines these three to obtain an audio spectrum, which is then reconstructed to produce a second audio file dubbed by the non-Chinese dubbing artist with voices B, C, or D.

[0188] The voice reconstructed based on the audio spectrum generated by the sound conversion model may have certain mechanical sounds, current sounds, and a certain degree of noise, so DSP technology can be used to restore the sound quality of the sound in the second audio file to remove noise and restore the sound quality.

[0189] Since the volume of the reconstructed voice may be different at different times, the volume of the original audio file and the corresponding timestamp can be detected, and the volume of the voice in the second audio file can be repaired accordingly. In addition, the sound effect of the original audio file and the corresponding timestamp can be detected, and the corresponding sound effect can be added to the second audio file accordingly.

[0190] For ease of understanding, this application also provides an embodiment of an audio processing method in a practical application scenario, as follows:

[0191] like Figure 8 As shown, the original audio file is a non-Chinese target video that is translated into Chinese and then dubbed by an amateur Chinese dubbing artist A to obtain a dubbing audio file.

[0192] Manual dubbing materials often have errors such as missing tracks, so the dubbing audio file can be detected for missing tracks and corrected to obtain a corrected dubbing audio file.

[0193] The subtitle file translated from non-Chinese into Chinese does not contain character information, and character information is necessary for selecting the voice conversion model (VC) of the character and the second dubbing object. Therefore, the subtitle character needs to be split to add character information to the subtitle file. Specifically: the face appearance timestamp is extracted from the target video, the voiceprint appearance segment timestamp is extracted from the original audio file, and the Chinese line segment appearance timestamp is extracted from the first line text. The voiceprint appearance segment timestamp, the face appearance timestamp, and the Chinese line segment appearance timestamp are merged to obtain the corresponding line content of each face in different time periods. The corresponding line content of each face in different time periods is matched with the second line text to obtain the audio time period corresponding to each character. Based on the audio time period corresponding to each character, intelligent track division is performed to obtain the audio files corresponding to each character in different time periods, namely: character track 1, character track 2...character track N. Each character track is equivalent to any first audio file in the aforementioned embodiment.

[0194] Each character's track is input into the voice conversion model separately. The role of the voice conversion model is to retain the emotion, rhythm and content of amateur Chinese voice actor A, and only replace his timbre characteristics with the timbre characteristics of Chinese voice actor B, Chinese voice actor C or Chinese voice actor D. Specifically, Figure 9 As shown, a character track (i.e., a first audio file) of an amateur Chinese voice actor A is input into an encoder. The encoder extracts content A and rhythm A from the character track. In practice, timbre A may not be extracted. Voiceprint casting can be performed as needed, whereby a dubbing timbre suitable for a film or TV drama dubbing scene is selected from Chinese timbres B, C, and D. In actual operation, a suitable timbre ID (i.e., the voiceprint identifier in the aforementioned embodiment) is selected, and the timbre features corresponding to the timbre ID are obtained. Content A, rhythm A, and the obtained timbre features are then input into a decoder. The decoder combines these three to obtain an audio spectrum, which is then reconstructed to produce a second audio file dubbed by the Chinese voice actor with timbre B, C, or D.

[0195] The voice reconstructed based on the audio spectrum generated by the sound conversion model may have certain mechanical sounds, current sounds, and a certain degree of noise, so DSP technology can be used to restore the sound quality of the sound in the second audio file to remove noise and restore the sound quality.

[0196] Since the volume of the reconstructed voice may be different at different times, the volume of the original audio file and the corresponding timestamp can be detected, and the volume of the voice in the second audio file can be repaired accordingly. In addition, the sound effect of the original audio file and the corresponding timestamp can be detected, and the corresponding sound effect can be added to the second audio file accordingly.

[0197] In another embodiment of the present application, an audio processing device is also provided. Figure 10 As shown, the audio processing device includes:

[0198] A first acquisition module 11 is configured to acquire a first audio file corresponding to a target video and extract content features and rhythmic features of a dubbing content of a first dubbing object in the first audio file, wherein the language of the first audio file is a first language different from the second language of the original audio file of the target video;

[0199] A second acquisition module 12 is configured to acquire a timbre feature corresponding to a second dubbing object, wherein the first dubbing object and the second dubbing object have different timbres;

[0200] a merging module 13, configured to merge the content feature, the rhythm feature, and the timbre feature to obtain an audio spectrum;

[0201] The reconstruction module 14 is configured to perform audio reconstruction based on the audio spectrum to obtain a second audio file corresponding to the target video.

[0202] Optionally, the first acquisition module includes:

[0203] a first acquisition unit configured to acquire an original audio file, a dubbing audio file, a first dialogue text, and a second dialogue text corresponding to a target video, wherein the first dialogue text is obtained by performing speech recognition on the dubbing audio file and does not include character information; the dubbing audio file is dubbed in a first language different from a second language of the original audio file of the target video; and the second dialogue text corresponds to the original audio file and includes character information;

[0204] a first determining unit, configured to determine, based on the target video, the first dialogue text, and the original audio file, a speech time period during which a face of the same character speaks and dialogue content corresponding to the speech time period;

[0205] A track splitting unit is configured to split the dubbing audio file into audio tracks according to the speech time period, the dialogue content corresponding to the speech time period, and the second dialogue text, to obtain the speech time period of each character and the audio file corresponding to the time period;

[0206] The second determining unit is configured to determine an audio file corresponding to a time period in which any character speaks as a first audio file corresponding to the target video.

[0207] Optionally, the first determining unit includes:

[0208] A first extraction subunit is used to extract a timestamp of a face appearance in the target video;

[0209] The second extraction subunit is used to extract the timestamp of the voiceprint appearance segment from the original audio file;

[0210] a third extraction subunit, configured to extract the timestamp of the first language speech segment from the first speech text;

[0211] A first matching subunit is configured to match the timestamp of the voiceprint segment with the timestamp of the face appearance to obtain a speech time period in which the faces belonging to the same character spoke;

[0212] The second matching subunit is used to match the time period when the face of the same character speaks with the timestamp of the first language dialogue segment to obtain the dialogue content corresponding to the speaking time period.

[0213] Optionally, the track division unit includes:

[0214] A third matching subunit is configured to match the speech time period, the dialogue content corresponding to the speech time period, and the second dialogue text to obtain the speech time period of each character;

[0215] The track division sub-unit is used to divide the dubbing audio file into audio tracks according to the time period of each character's speech, and obtain audio files corresponding to the time period.

[0216] Optionally, the dubbing feature includes: content feature, and the first acquisition module includes:

[0217] A first input unit is used to input the first audio file into a preset speech recognition encoder to obtain recognition content;

[0218] The second input unit is configured to input the identified content into a preset content encoder to obtain the content feature.

[0219] Optionally, the dubbing feature includes a rhythm feature, and the first acquisition module includes:

[0220] A third input unit is used to input the first audio file into a preset speech self-supervised learning pre-training model to obtain output data;

[0221] The fourth input unit is configured to input the output data into a preset prosody encoder to obtain the prosody feature.

[0222] Optionally, the second acquisition module includes:

[0223] A second acquiring unit, configured to acquire the original audio file of the target video;

[0224] An extraction unit, configured to extract the original voiceprint features of the original dubbing object from the original audio file;

[0225] A first searching unit is configured to search a preset voiceprint library for a voiceprint identifier corresponding to the original voiceprint feature;

[0226] The third determining unit is configured to determine the timbre feature of the dubbing object corresponding to the voiceprint identifier as the second timbre feature of the second dubbing object.

[0227] Optionally, after the reconstruction unit, the device further includes:

[0228] A third acquisition module is used to obtain the original audio file of the target video;

[0229] a volume detection module, configured to perform volume detection on the original audio file to obtain first volume values ​​corresponding to a plurality of first timestamps;

[0230] a first search module, configured to search the second audio file for a second volume value corresponding to each of the first timestamps;

[0231] The volume adjustment module is configured to adjust the second volume value to the first volume value if the difference between the first volume value and the second volume value is greater than a preset threshold, thereby obtaining an adjusted second audio file.

[0232] Optionally, after the consumption unit, the device further includes:

[0233] a sound effect detection module, configured to perform sound effect detection on the original audio file to obtain a plurality of sound effect types corresponding to the second timestamps;

[0234] The sound effect adjustment module is used to add sound effects to the second audio file according to the sound effect types corresponding to the multiple second time stamps to obtain an adjusted second audio file.

[0235] In another embodiment of the present application, an electronic device is provided, including a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other via the communication bus;

[0236] Memory for storing computer programs;

[0237] The processor is configured to implement the audio processing method described in any of the aforementioned method embodiments when executing the program stored in the memory.

[0238] In an electronic device provided by an embodiment of the present invention, a processor executes a program stored in a memory to achieve the following: retaining only the dubbing features of a first dubbing object in a first audio file, not using the timbre features of the first dubbing object, merging the dubbing features of the first dubbing object with the timbre features of a second dubbing object, so that a second audio file reconstructed based on the merged audio spectrum can have the timbre of the second dubbing object and retain the dubbing features, thereby automatically converting the timbre of the first dubbing object into the timbre of the second dubbing object while retaining the content and emotion of the dubbing of the first dubbing object. Furthermore, the timbre of the first dubbing object in all first audio files corresponding to the target video file can be easily converted into the timbre of the corresponding second dubbing object. Without the need for other dubbing actors, the effect of having one first dubbing object match the timbre of multiple second dubbing objects can be achieved, and the rich emotions of the first dubbing object's speech can be retained, thereby meeting the requirements of film and television drama scenes for dialogue.

[0239] The communication bus 1140 mentioned in the above electronic device can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. The communication bus 1140 can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 11 Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.

[0240] The communication interface 1120 is used for communication between the electronic device and other devices.

[0241] The memory 1130 may include a random access memory (RAM) or a non-volatile memory (non-volatile memory), such as at least one disk storage. Alternatively, the memory may be at least one storage device located away from the processor.

[0242] The above-mentioned processor 1110 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components.

[0243] In another embodiment of the present application, a computer-readable storage medium is provided, on which a program of an audio processing method is stored. When the program of the audio processing method is executed by a processor, the steps of the audio processing method described in any of the aforementioned method embodiments are implemented.

[0244] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or device comprising the element.

[0245] The foregoing description is intended only to provide specific embodiments of the present invention, which will enable those skilled in the art to understand and implement the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not intended to be limited to the embodiments shown herein, but is intended to be accorded the widest scope consistent with the principles and novel features claimed herein.

Claims

1. An audio processing method, characterized in that: include: Obtaining a first audio file corresponding to a target video, and extracting dubbing features corresponding to dubbing content of a first dubbing object in the first audio file, wherein a first language in the first audio file is different from a second language in the original audio file of the target video; Acquiring a timbre feature corresponding to a second dubbing object, wherein the first dubbing object and the second dubbing object have different timbres; Combining the dubbing feature and the timbre feature to obtain an audio spectrum; Performing audio reconstruction based on the audio spectrum to obtain a second audio file corresponding to the target video; Wherein, the obtaining of the first audio file corresponding to the target video includes: obtaining the original audio file, dubbing audio file, first dialogue text and second dialogue text corresponding to the target video, wherein the first dialogue text is obtained by performing speech recognition on the dubbing audio file and does not contain character information, the dubbing audio file is obtained by dubbing in a first language different from the second language of the original audio file of the target video, and the second dialogue text corresponds to the original audio file and contains character information; determining the speaking time period of a face belonging to the same character and the dialogue content corresponding to the speaking time period according to the target video, the first dialogue text and the original audio file; performing audio track division on the dubbing audio file according to the speaking time period, the dialogue content corresponding to the speaking time period and the second dialogue text to obtain the speaking time period of each character and the audio file corresponding to the time period; determining the audio file corresponding to the speaking time period of any character as the first audio file corresponding to the target video; The method of determining the speaking time period of faces speaking belonging to the same character and the speech content corresponding to the speaking time period based on the target video, the first dialogue text, and the original audio file includes: extracting the timestamp of face appearance in the target video; extracting the timestamp of voiceprint appearance in the original audio file; extracting the timestamp of first language dialogue appearance in the first dialogue text; matching the timestamp of voiceprint appearance with the timestamp of face appearance to obtain the speaking time period of faces speaking belonging to the same character; matching the timestamp of face appearance with the timestamp of first language dialogue appearance to obtain the speech content corresponding to the speaking time period.

2. The audio processing method according to claim 1, wherein: The dubbing audio file is divided into audio tracks according to the speaking time period, the dialogue content corresponding to the speaking time period, and the second dialogue text, to obtain the speaking time period of each character and the audio file corresponding to the time period, including: Matching the speech time period, the dialogue content corresponding to the speech time period, and the second dialogue text to obtain the speech time period of each character; The dubbing audio file is divided into audio tracks according to the time period of each character's speech to obtain audio files corresponding to the time period.

3. The audio processing method according to claim 1, wherein: Obtaining a second timbre feature corresponding to the second dubbing object includes: Obtaining the original audio file of the target video; Extracting the original voiceprint features of the original dubbing object from the original audio file; Searching for a voiceprint identifier corresponding to the original voiceprint feature in a preset voiceprint library; The timbre feature of the dubbing object corresponding to the voiceprint identifier is determined as the second timbre feature of the second dubbing object.

4. The audio processing method according to claim 1, wherein: After performing audio reconstruction based on the audio spectrum to obtain a second audio file corresponding to the target video, the method further includes: Obtaining the original audio file of the target video; Performing volume detection on the original audio file to obtain a plurality of first volume values ​​corresponding to the first timestamps; searching the second audio file for a second volume value corresponding to each of the first timestamps; If the difference between the first volume value and the second volume value is greater than a preset threshold, the second volume value is adjusted to the first volume value to obtain an adjusted second audio file.

5. The audio processing method according to claim 1, wherein: After performing audio reconstruction based on the audio spectrum to obtain a second audio file corresponding to the target video, the method further includes: Performing sound effect detection on the original audio file to obtain a plurality of sound effect types corresponding to the second timestamps; Sound effects are added to the second audio file according to the sound effect types corresponding to the multiple second timestamps to obtain an adjusted second audio file.

6. An audio processing device, characterized in that: include: a first acquisition module configured to acquire a first audio file corresponding to a target video and extract content features and rhythmic features of a dubbing content of a first dubbing object in the first audio file, wherein a first language in the first audio file is different from a second language in the original audio file of the target video; A second acquisition module is configured to acquire a timbre feature corresponding to a second dubbing object, wherein the first dubbing object and the second dubbing object have different timbres; a merging module, configured to merge the content feature, the rhythm feature, and the timbre feature to obtain an audio spectrum; a reconstruction module, configured to perform audio reconstruction based on the audio spectrum to obtain a second audio file corresponding to the target video; Wherein, the obtaining of the first audio file corresponding to the target video includes: obtaining the original audio file, dubbing audio file, first dialogue text and second dialogue text corresponding to the target video, wherein the first dialogue text is obtained by performing speech recognition on the dubbing audio file and does not contain character information, the dubbing audio file is obtained by dubbing in a first language different from the second language of the original audio file of the target video, and the second dialogue text corresponds to the original audio file and contains character information; determining the speaking time period of a face belonging to the same character and the dialogue content corresponding to the speaking time period according to the target video, the first dialogue text and the original audio file; performing audio track division on the dubbing audio file according to the speaking time period, the dialogue content corresponding to the speaking time period and the second dialogue text to obtain the speaking time period of each character and the audio file corresponding to the time period; determining the audio file corresponding to the speaking time period of any character as the first audio file corresponding to the target video; The method of determining the speaking time period of faces speaking belonging to the same character and the speech content corresponding to the speaking time period based on the target video, the first dialogue text, and the original audio file includes: extracting the timestamp of face appearance in the target video; extracting the timestamp of voiceprint appearance in the original audio file; extracting the timestamp of first language dialogue appearance in the first dialogue text; matching the timestamp of voiceprint appearance with the timestamp of face appearance to obtain the speaking time period of faces speaking belonging to the same character; matching the timestamp of face appearance with the timestamp of first language dialogue appearance to obtain the speech content corresponding to the speaking time period.

7. An electronic device, characterized in that: It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus; Memory for storing computer programs; A processor, configured to implement the audio processing method according to any one of claims 1 to 5 when executing a program stored in a memory.

8. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a program of the audio processing method, and when the program of the audio processing method is executed by a processor, the steps of the audio processing method according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Voice conversion method and device, corresponding model training method and device, equipment and storage medium

    CN112466275A

  • Voice conversion method and device and computer system

    CN113808576A