Audiobook production method, production device and storage medium

By combining character and scene sound effects in audio book production, the problem of monotonous AI reading sound is solved and the auditory effect of audio books is improved.

CN116403561BActive Publication Date: 2025-08-08TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202310312863.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-28
Publication Date
2025-08-08
Estimated Expiration
2043-03-28

AI Technical Summary

Technical Problem

Among the existing audio book production methods, the reading sound generated by AI reading is relatively monotonous, resulting in poor auditory effects of audio books.

Method used

By obtaining the text of the audio book, target sentences related to the characters and scenes are determined, and character reading sounds matching the audio characteristics are generated based on the character information, and scene sound effects matching the scene information are generated based on the scene information, and added to the appropriate position of the character reading sounds to form the target audio.

Benefits of technology

The audio auditory effect of audio books is improved, so that the audio of the target sentence not only contains the sound of character reading, but also scene sound effects, enhancing the auditory experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116403561B_ABST
    Figure CN116403561B_ABST
Patent Text Reader

Abstract

The embodiment of the present application discloses a method for producing an audio book, a production device, and a storage medium for use in the field of audio technology. The method of the embodiment of the present application includes: obtaining the text corresponding to the audio book; determining the target sentences related to the character and the scene in the text; performing voice processing on the target sentence according to the audio features corresponding to the character information of the target sentence, and obtaining the character reading sound that matches the audio features; obtaining the scene sound effect that matches the scene information according to the scene information of the target sentence; determining the sentence position of the scene information in the target sentence, and adding the scene sound effect to the audio segment of the character reading sound corresponding to the sentence position, and obtaining the target audio corresponding to the target sentence. By adding the corresponding scene sound effect to the character reading sound of the target sentence, the audio of the target sentence contains not only the character reading sound but also the scene sound effect, thereby improving the audio hearing effect of the audio book.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of audio technology, and in particular to a method for producing an audio book, a production device, and a storage medium. Background Art

[0002] Existing audiobooks generally read text aloud into audio, allowing users to listen to the text, such as audiobooks of novels. The reading method can be manual or generated by technology.

[0003] Existing audiobooks produced through manual reading require the text to be read aloud one by one. When the text contains a lot of words, the production cost and time are huge, and the generation efficiency is low. With the development of deep neural network technology, AI audiobooks directly generated by AI reading technology can quickly synthesize reading sounds; AI reading is a technology that converts text into sound through deep neural networks. Existing AI reading is synthesized through text-to-speech speech synthesis technology. When converting text to speech, AI reading generally distinguishes the different characters corresponding to each sentence in the text, converts the sentences of different characters into sound waveforms with the audio characteristics corresponding to the characters, and obtains the sound of the character reading the sentence.

[0004] However, in the existing method of producing audio books through AI reading, the sentences corresponding to the characters in the text only contain the voice of the characters reading, and the resulting reading voice is relatively monotonous, and the auditory effect of the audio book is poor. Summary of the Invention

[0005] The embodiments of the present application provide a method for producing an audio book, a production device, and a storage medium, which can effectively improve the audio listening effect of the audio book.

[0006] The present application provides a method for producing an audio book, including:

[0007] Get the text corresponding to the audiobook;

[0008] Identifying target sentences in the text that are relevant to the character and the scene;

[0009] Performing voice processing on the target sentence according to the audio features corresponding to the character information of the target sentence to obtain a character reading voice that matches the audio features;

[0010] Obtaining a scene sound effect that matches the scene information according to the scene information of the target sentence;

[0011] The sentence position of the scene information in the target sentence is determined, and the scene sound effect is added to the audio segment of the character's reading sound corresponding to the sentence position to obtain the target audio corresponding to the target sentence.

[0012] Furthermore, the determining the sentence position of the scene information in the target sentence includes:

[0013] Obtaining a first phoneme sequence of scene content text corresponding to the scene information;

[0014] Determine the sequence position of the first phoneme sequence in the phoneme sequence corresponding to the target sentence.

[0015] Furthermore, adding the scene sound effect to the audio segment corresponding to the position of the sentence in the character's reading sound includes:

[0016] Determine, according to the sequence position, an audio frame sequence corresponding to the first phoneme sequence in the character's reading voice;

[0017] Using the preset audio frame corresponding to the audio frame sequence as the sound effect start frame of the scene sound effect;

[0018] Determining the duration of the scene sound effect according to the sound source in the scene information;

[0019] The scene sound effect is added to the character's reading voice according to the sound effect start frame and the duration of the sound effect.

[0020] Furthermore, determining the duration of the scene sound effect according to the sound source in the scene information includes:

[0021] If the volume change value of the sound source within the preset time period is greater than the preset volume threshold, the scene sound effect is determined to be a trigger sound, and the duration of the trigger sound is less than the preset duration;

[0022] If the volume change value of the sound source within the preset time length is less than the preset volume threshold, it is determined that the scene sound effect is environmental background sound, and the duration of the sound effect of the environmental background sound is greater than the preset time length.

[0023] Furthermore, the adding of the scene sound effect to the character's reading voice includes:

[0024] If there are multiple scene sound effects in the same audio frame, and the trigger sound and the ambient background sound are included in the multiple scene sound effects, the volume of the trigger sound is increased based on a preset volume balancing technology so that the volume of the trigger sound is higher than the volume of the ambient background sound.

[0025] Furthermore, adding the scene sound effect to the character's reading voice according to the sound effect start frame and the duration of the sound effect includes:

[0026] Fading the scene sound effect into the character's reading voice at the sound effect start frame, and continuously increasing the volume of the scene sound effect until it reaches a preset volume;

[0027] The volume of the scene sound effect is reduced during the preset end period of the sound effect duration, and the scene sound effect is faded out from the character's reading voice.

[0028] Furthermore, determining target sentences in the text that are related to the character and the scene includes:

[0029] Determining whether the semantic information of the sentence in the text contains preset character dialogue information, or determining whether the semantic information of the sentence in the text matches the semantic information in the preset character sentence set;

[0030] If the preset character dialogue information exists, or matches the semantic information in the preset character sentence set, then determining that the sentence is a character sentence related to the character;

[0031] If the role sentence contains the preset scene semantics, the role sentence is determined to be the target sentence.

[0032] Furthermore, the step of performing voice processing on the target sentence according to the audio features corresponding to the character information of the target sentence to obtain a character reading voice matching the audio features includes:

[0033] Inputting the target sentence into a preset reading model, and determining the timbre characteristics corresponding to the character information of the target sentence in the preset reading model;

[0034] Converting the phoneme sequence corresponding to the target sentence into a target audio feature based on the timbre feature corresponding to the character information;

[0035] The target audio feature is converted into a sound waveform to obtain a character reading voice that matches the audio feature.

[0036] Furthermore, obtaining a scene sound effect that matches the scene information according to the scene information of the target sentence includes:

[0037] According to the preset scene semantics in the scene information, a scene sound effect matching the scene information is determined from a preset sound effect library, wherein the preset sound effect library contains a plurality of scene sound effects corresponding to the scene semantics.

[0038] The present application also provides an audio book production device, including:

[0039] CPU, memory and input / output interfaces;

[0040] The memory is a short-term storage memory or a persistent storage memory;

[0041] The central processing unit is configured to communicate with the memory and execute instructions in the memory to perform the above method.

[0042] An embodiment of the present application further provides a computer-readable storage medium, comprising instructions, which, when executed on a computer, enable the computer to execute the above method.

[0043] It can be seen from the above technical solutions that the embodiments of the present application have the following advantages:

[0044] The method of the embodiment of the present application includes: obtaining the text corresponding to the audio book; determining the target sentences related to the characters and the scenes in the text; performing voice processing on the target sentences according to the audio features corresponding to the character information of the target sentences, and obtaining the character reading sound that matches the audio features; obtaining the scene sound effects that match the scene information according to the scene information of the target sentences; determining the sentence position of the scene information in the target sentences, and adding the scene sound effects to the audio segment of the character reading sound corresponding to the sentence position, and obtaining the target audio corresponding to the target sentence. By adding the corresponding scene sound effects to the character reading sound of the target sentence, the character reading sound corresponding to the target sentence is combined with the scene sound effects, so that the audio of the target sentence contains not only the character reading sound but also the scene sound effects, thereby improving the audio hearing effect of the audio book. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments recorded in the present application. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.

[0046] Figure 1 A communication architecture diagram for producing an audio book disclosed in an embodiment of the present application;

[0047] Figure 2 A flowchart for producing an audio book disclosed in an embodiment of the present application;

[0048] Figure 3 A flowchart of another audio book production process disclosed in an embodiment of the present application;

[0049] Figure 4 A logical structure diagram of the production of an audio book disclosed in an embodiment of the present application;

[0050] Figure 5 A schematic diagram of the structure of an acoustic model disclosed in an embodiment of the present application;

[0051] Figure 6 A diagram of an audio book production device disclosed in an embodiment of the present application;

[0052] Figure 7 This is a diagram of another audio book production device disclosed in an embodiment of the present application. DETAILED DESCRIPTION

[0053] In order to make the purpose, technical solutions and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limiting this application. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0054] In the following description, references to "one embodiment" or "an example" and similar expressions describe a subset of all possible embodiments. However, it is understood that "one embodiment" or "an example" may refer to the same subset or different subsets of all possible embodiments, and may be combined with each other without conflict. In the following description, the term "plurality" refers to at least two.

[0055] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.

[0056] The existing audio book production process is to obtain the corresponding text and then convert the text into the corresponding audio. Users can get the information corresponding to the text by listening to the corresponding audio. Figure 1As shown, the existing audio book production device 101 is connected to the audio book reader 102, and the connection can be a wired or wireless network connection, which is not limited here. It is understandable that the audio book reader 102 can be a mobile phone, a handheld e-book or other device, and the text can be a novel or other written work, which is not limited here. The audio book production device 101 can obtain the text for which the audio book needs to be produced from the audio book reader 102, convert the text into audio, and then transmit the audio to the audio book reader 102. The audio book reader 102 can play the corresponding audio when displaying the text. The audio book production device 101 can be connected to one or more audio book readers 102, which is not limited here. The audio book production device 101 can be set inside the audio book reader 102, or it can be set outside the audio book reader 102, which is not limited here. Existing AI reading is synthesized through text-to-speech speech synthesis technology. When converting text into speech, AI reading generally distinguishes the different characters corresponding to each sentence in the text, converts the sentences of different characters into sound waveforms with the audio characteristics corresponding to the characters, and obtains the voice of the character reading the sentence. However, in the existing method of producing audio books through AI reading, the sentences corresponding to the characters in the text only contain the voice of the characters reading, and the obtained reading voice is relatively monotonous, and the auditory effect of the audio book is poor. Therefore, the embodiment of the present application provides a method for producing an audio book, which can improve the audio listening experience of the audio book, such as Figure 2 As shown, the details are as follows:

[0057] 201. Obtain the text corresponding to the audio book.

[0058] In an embodiment of the present application, an audio book production device can obtain the text corresponding to the audio book. Specifically, the text to be produced into the audio book can be obtained from an audio book reader connected to the audio book production device, the text can be converted into audio, and then the audio can be transmitted back to the audio book reader. It is understood that the audio book production device can obtain the text corresponding to the audio book from the connected audio book reader or from a storage library of the audio book production device itself, and the specific method is not limited here.

[0059] 202. Identify target sentences in the text that are relevant to the character and the scene.

[0060] After obtaining the corresponding text of the audiobook, the audiobook production device can determine target sentences in the text that are relevant to the characters and the scene. It is understood that "relevant to the characters" refers to sentences that are spoken and described by a preset character in the text, and "relevant to the scene" refers to sentences that contain corresponding scene information. The target sentence is relevant to both the character and the scene, meaning that the target sentence spoken and described by the preset character contains corresponding scene information. It is understood that the target sentence refers to a specific sentence in the text, which typically contains multiple target sentences. Specifically, whether the sentence is relevant to the characters or the scene can be determined based on the semantic information of the sentence in the text. The role information in the sentence's semantic information can be used to determine whether the sentence is relevant to the characters. The role information refers to the character corresponding to the reading of the sentence. For example, if the sentence is "Character A said to Character B," the character corresponding to the sentence is Character A. The role information can be understood as the character's name, which can be an elderly person, a young person, a student, a teacher, etc., without limitation here. After determining that the sentence is a sentence related to the character, it is possible to further determine whether the sentence is a sentence related to the scene. This can be done by determining whether the semantic information of the sentence contains scene information. The scene information refers to the scene that appears in the sentence. For example, when the sentence includes "the sound of the car starting up," the scene corresponding to the sentence is the scene of the engine starting up. If it is further determined that the sentence is a sentence related to the scene, then the sentence is determined to be a target sentence related to the character and the scene. It is understandable that it is also possible to first determine whether the sentence in the text is a sentence related to the scene, and then determine whether the sentence is related to the character. This is not limited to the specifics here.

[0061] 203. Perform voice processing on the target sentence according to the audio features corresponding to the character information of the target sentence to obtain a character reading voice that matches the audio features.

[0062] After determining the target sentences in the text that are relevant to the role and the scene, the target sentences can be voiced according to the audio features corresponding to the role information of the target sentences, and the character reading sound that matches the audio features can be obtained. Among them, the audio features corresponding to the role information of the target sentence refer to the audio features of the pronunciation (speech) of the character, and the audio features include timbre, loudness, etc., which are not specifically limited here. For example, when the role of the target sentence is an old man, the corresponding audio features are relatively turbid audio features. When the role of the target sentence is a student, the corresponding audio features are relatively brisk audio features. The target sentence is voiced according to the audio features corresponding to the role information of the target sentence, and a reading sound that matches the audio features is obtained, and the reading sound is consistent with the audio features. For example, when the role of the target sentence is a student, the reading sound of the target sentence is a relatively brisk sound.

[0063] 204. Obtain a scene sound effect that matches the scene information according to the scene information of the target sentence.

[0064] After determining the target sentence in the text that is relevant to the character and the scene, a scene sound effect that matches the scene information can be obtained based on the scene information of the target sentence. Specifically, a scene sound effect that matches the scene information can be obtained based on the scene semantics in the scene information. For example, if the scene semantics is "a train is coming," the scene sound effect can be determined to be the roar of a train. In this case, the corresponding roar of a train can be downloaded from the Internet or searched for the roar of a train in the sound effect library of the audio book production device. The specifics are not limited here.

[0065] It is understandable that there is no limitation on the execution order of step 203 and step 204 .

[0066] 205. Add scene sound effects to the audio segment of the character's reading sound corresponding to the scene information to obtain target audio corresponding to the target sentence.

[0067] After obtaining the character's reading voice and the scene sound effects that match the scene information, the scene sound effects can be added to the audio segment corresponding to the scene information in the character's reading voice to obtain the target audio corresponding to the target sentence. Specifically, the sentence position of the scene information in the target sentence can be determined, and the scene sound effects can be added to the audio segment corresponding to the sentence position of the character's reading voice to obtain the target audio corresponding to the target sentence. It is understood that the scene content text describing the scene information occupies a certain sentence position in the target sentence. For example, if the target sentence is "Zhang San said: I heard a knock on the door," the scene content text of the scene information is "knock on the door," and the sentence position of "knock on the door" in the target sentence can be determined. Adding the scene sound effects to the audio segment corresponding to the sentence position of the character's reading voice can be understood as adding the scene sound effects to the audio segment reading the scene content text when the character's reading voice reads the scene content text of the scene information; for example, the scene sound effects can be added to the audio segment when the character's reading voice reads the scene content text of the scene information. The scene sound effects can be added to the first audio frame of the reading scene content text or to the first few frames of the reading scene content text, although the specific details are not limited here. For example, for the sound of “knock on the door”, the scene sound effect can be added to the first audio frame of the word “knock”, or the scene sound effect can be added to the first few frames of the word “knock”.

[0068] It is understandable that the semantic information of different target sentences in the text is different, and the obtained character information and scene sound effects are also different. When mixing the character reading sounds and scene sound effects of multiple target sentences, a target audio containing the characteristics of multiple characters and multiple scene sound effects can be obtained. When playing the target audio, the audio of the current sentence played has the reading sound that matches the audio features corresponding to the character information and the scene sound effects of the current sentence.

[0069] It can be seen that the method of the embodiment of the present application includes: obtaining the text corresponding to the audio book; determining the target sentences related to the character and the scene in the text; performing voice processing on the target sentence according to the audio features corresponding to the character information of the target sentence, and obtaining the character reading sound that matches the audio features; obtaining the scene sound effect that matches the scene information according to the scene information of the target sentence; determining the sentence position of the scene information in the target sentence, and adding the scene sound effect to the audio segment of the character reading sound corresponding to the sentence position, and obtaining the target audio corresponding to the target sentence. By adding the corresponding scene sound effect to the character reading sound of the target sentence, the character reading sound corresponding to the target sentence is combined with the scene sound effect, so that the audio of the target sentence contains not only the character reading sound, but also the scene sound effect, thereby improving the audio hearing effect of the audio book.

[0070] The above describes the production process of audio books. Figure 3 The audiobook production process is described in detail. The specific steps are as follows:

[0071] 301. Obtain the text corresponding to the audio book.

[0072] It should be noted that step 301 is similar to the above step 201 and will not be described in detail here.

[0073] 302. Determine target sentences related to the role and the scene based on the semantic information of the sentences in the text.

[0074] In the embodiment of the present application, the audiobook production device can determine target sentences related to characters and scenes based on the semantic information of sentences in the text. Specifically, it can determine whether the semantic information of the sentence in the text contains preset character dialogue information, or whether the semantic information of the sentence in the text matches the semantic information in a preset character sentence set. If the preset character dialogue information exists, or if the semantic information matches the semantic information in the preset character sentence set, the sentence is determined to be a character sentence related to the character. If the character sentence contains preset scene semantics, the character sentence is determined to be a target sentence.

[0075] It is understandable that the audiobook production device can analyze the semantic information of the sentence in the text to determine whether the role information exists, so as to determine whether the sentence is a role sentence related to the role. Determining the role information includes classifying the sentences in the text into roles or not. Specifically, Figure 4As shown, the inclusion of multiple roles can be determined based on the text type. For example, news stories are typically presented by a single person, so role classification is unnecessary. However, novels or plays typically include multiple characters, each representing different personality traits, so role classification is required. When performing role classification, natural language processing techniques can be used to process text types such as novels and plays that contain multiple characters, resulting in character segmentation results. For example, the current sentence is spoken by character Zhang San, the next sentence is spoken by character Li Si, and the previous sentence is narration. Determining whether the semantic information of a sentence in the text matches the semantic information in a set of preset character sentences can be performed using a natural language processing model trained on big data. Determining whether the semantic information of a sentence in the text contains dialogue information for the preset characters can be performed using a rule-based post-processing method. Determining whether the semantic information of a sentence in the text matches the semantic information in the set of preset character sentences is specifically determined by method A, while determining whether the semantic information of a sentence in the text contains dialogue information for the preset characters is determined by method B.

[0076] Method A: Determine whether the semantic information of the sentence in the text matches the semantic information in the preset role sentence set.

[0077] When the audiobook production device determines whether the semantic information of a sentence in the text matches the semantic information in the preset role sentence set, it can compare the semantic information of the sentence in the text with the semantic information in the sentence training set to determine whether the sentence has role information. Specifically, the preset language processing model can be trained based on the sentence training set, and each sentence in the sentence training set has a corresponding role label; the sentence training set can be Chinese text big data, and the preset language processing model can be a Bert model, which can extract the text features of the paragraph where the current sentence is located. The text feature refers to the representation vector extracted based on the natural language processing model. The representation vector is a set of data used to represent the position of the current text in the feature space. Texts with different characteristics will be in different feature spaces. For example, the three sentences before and after the current sentence are regarded as a paragraph. During the training stage, supervised training can be performed through the preset language processing model by marking the roles of the sentences in the text, such as the name of the character of each sentence in the novel. When determining whether a sentence has role information, the sentence to be classified can be input. If the predicted role name can be obtained, it can be determined that the sentence has role information. Specifically, the sentence can be input into a preset language processing model. In the preset language processing model, the semantic information of the sentence is compared with the semantic information in the sentence training set. If there is a first sentence in the sentence training set that matches the sentence, it is determined that the sentence has role information. After determining that the sentence has role information, the role label corresponding to the first sentence can be used as the role information corresponding to the sentence. It can be understood that during the training process, the preset language processing model can learn the sentences in the sentence training set and the corresponding role labels by itself. Specifically, the training process can be to select the same specific field or a specific field that appears with a certain probability from the sentences belonging to the same preset role label. After that, it can be determined whether the sentence has the specific field. If so, it is determined that the sentence has role information and it is determined that the semantic information of the sentence in the text matches the semantic information in the preset role sentence set.

[0078] Method B: Determine whether there is preset character dialogue information in the semantic information of the sentence in the text.

[0079] Determining whether the semantic information of a sentence in a text contains pre-set character dialogue information is a rule-based post-processing method. This involves analyzing the role fields contained in the current sentence. When pre-set character dialogue information is present, the presence of pre-set character dialogue information is determined. Upon determining the presence of pre-set character dialogue information, the sentence's character information can also be determined. Specifically, the presence of pre-set character dialogue information is determined within the sentence's semantic information. If so, the name of the speaking character is determined from the pre-set character dialogue information, and the speaking character name is used as the corresponding character information for the sentence. For example, if the sentence is "Zhang San said to Li Si, 'The weather is so nice today!'," then the sentence contains the character names "Zhang San" and "Li Si," and someone is clearly speaking to someone else. Therefore, the corresponding character for the sentence is "Zhang San." If the sentence is a continuous dialogue and no specific speaker is explicitly labeled, the reader can only estimate the speaker's identity through semantic comprehension. For such sentences where the roles are not obvious, the prediction results of a natural language processing model (method A) can be used to determine the sentence's character information.

[0080] It is understandable that only one of Method A and Method B can be executed, or Method B can be executed after Method A. After determining that the sentence is a sentence related to a role, the role information obtained by comparing the two methods can be more accurately determined. For example, if the first role information is obtained after executing Method A and the second role information is obtained after executing Method B, when the first role information and the second role information are the same, the role information of the sentence is determined to be the first role information (i.e., the second role information); when the first role information and the second role information are different, the second role information can be used as the role information of the target sentence, and the second role information of Method B can be used as the basis.

[0081] 303. Perform voice-activated processing on the target sentence in a preset reading model according to the audio features corresponding to the character information of the target sentence, and obtain a character reading voice that matches the audio features.

[0082] After determining the target sentence related to the character and the scene, the target sentence can be voiced in a preset reading model based on the audio features corresponding to the character information of the target sentence, thereby obtaining a character reading voice that matches the audio features. The target sentence can be input into the preset reading model, and the timbre features corresponding to the character information of the target sentence can be determined in the preset reading model; based on the timbre features corresponding to the character information, the phoneme sequence corresponding to the target sentence can be converted into target audio features; and the target audio features can be converted into sound waveforms to obtain a character reading voice that matches the audio features.

[0083] Specifically, in the preset reading model, the target sentence is voiced into a sound waveform based on the audio features corresponding to the character information of the target sentence. The preset reading model can be an AI reading model. AI reading is generally generated by technology. Through speech synthesis technology, sentences of different characters can generate sound content corresponding to different timbres. For example, the sentence content of Zhang San, input "Zhang San" and the sentence content into the AI reading model, and through technology generation, you can get the sound of the first timbre; the sentence content of Li Si, input "Li Si" and the sentence content into the AI reading model, and through technology generation, you can get the sound of the second timbre, and the first timbre and the second timbre are different.

[0084] Specifically, the audio features corresponding to the character information include timbre features. The target sentence is input into a preset reading model, which includes: an acoustic front end, an acoustic model, and a vocoder. The acoustic front end includes text-to-phoneme conversion (such as Chinese G2P), word segmentation (such as Jieba word segmentation), text regularization, and other parts, which convert the input text into a phoneme sequence, that is, the target sentence is converted into a phoneme sequence in the acoustic front end. Specifically, the text can be converted into a phoneme sequence according to the pronunciation order of the characters in the target sentence. For example, if the input target sentence is "the weather is good", the corresponding pinyin sequence is "tian1qi4hao3". The pinyin sequence is then further decomposed according to the initials and finals, and information such as word segmentation and prosody is added as the phoneme sequence. The acoustic model can be based on a non-autoregressive solution (such as Fastspeech) or an autoregressive solution (such as Tacotron), which converts the phoneme sequence into audio features (such as Mel spectrum). Specifically, the timbre features corresponding to the character information of the target sentence are determined in the acoustic model, and the phoneme sequence is converted into the target audio features based on the timbre features corresponding to the character information, that is, the phoneme sequence is converted into the audio features corresponding to each audio frame. The audio features include timbre, energy, fundamental frequency and other information. The vocoder can be a solution based on an adversarial network (such as Hifigan), a solution based on a recurrent neural network (such as wavernn), etc., which converts the audio features into sound waveforms, that is, the target audio features are converted into sound waveforms in the vocoder to obtain a reading sound that matches the audio features. It is understandable that the AI reading model can be pre-trained to learn the timbre features corresponding to each character, and then select different timbre features according to different character names when in use, and then convert the phoneme sequence into audio features based on the corresponding timbre features, and then convert the audio features into sound waveforms. It is understandable that the character reading voice of the character dialogue generally does not include the name of the speaking character, but rather the text of the character's voice is audio-converted into the character reading voice. For example, if the target sentence is "Zhang San said: I heard a knock on the door", then Zhang San's character reading voice is: I heard a knock on the door.

[0085] Furthermore, in order to provide the AI reading model with more timbre options as much as possible, the acoustic model can be improved, such as Figure 5 As shown, the role information of the target sentence includes: a role identifier for the target sentence; multiple role identifiers and corresponding timbre features can be added to the audio coding network of the acoustic model; a target role identifier that matches the role identifier of the target sentence is determined from the multiple role identifiers, and the timbre features corresponding to the target role identifier are used as the timbre features corresponding to the role identifier of the target sentence. It is understood that the ID encoding information of different speakers (roles) and the audio feature encoding information of the input speech can be added to the audio coding network of the input speech. The ID encoding information can be obtained through a convolutional network, and the audio feature encoding information can be extracted using a pre-trained timbre extraction model, such as wav2vec, to extract feature vectors containing timbre information and content information unrelated to timbre, thereby decoupling timbre and content from the input speech. Then, a feature classification network is added after the audio features predicted by the acoustic model to analyze whether the audio features predicted by the acoustic model correspond to the expected timbre. This post-processing helps the acoustic model better model the timbre of inputs with different timbres. In the figure, Encoder represents the encoder of the acoustic model, VarianceAdapter represents the variable adapter, and Decoder represents the decoder. At the same time, for the vocoder, speech of various timbres can be used for training. The resulting vocoder can learn the pronunciation characteristics of different timbres and can more accurately synthesize the corresponding sound waveform when used.

[0086] Furthermore, the voice processing can also be performed by artificial reading, and the target sentence can be artificially read aloud as a sound waveform based on the timbre characteristics corresponding to the role information of the target sentence. Specifically, artificial reading (real person reading) is to artificially read aloud the target sentence as a sound waveform based on the timbre characteristics corresponding to the role information of the target sentence. Among them, the audio features include timbre features; the target role that needs to be artificially read aloud is selected from the multiple role information contained in the text; it is understandable that after the text is classified by role, multiple sentences belonging to the same role can be grouped together; the sentence corresponding to the target role is transmitted to a real person anchor whose pronunciation conforms to the timbre characteristics corresponding to the target role for real person voice recording, and the sound waveform returned by the real person anchor is received, and the recorded sound is used as the reading sound that matches the audio features. It is understandable that for real person reading, the recording party can select the role sentence to be read aloud by the real person, send it to the designated real person anchor via the network for sound recording, and then transmit the recorded sound to the recording party via the network. Taking a novel as an example, it may contain dozens of characters, involving age changes from old to middle-aged, young to young, and gender differences between men and women. Live reading can provide live reading audio content for some of the characters.

[0087] 304. Determine, based on a preset sound effect library, a scene sound effect that matches the scene information of the target sentence.

[0088] After determining a target sentence related to the character and the scene, a scene sound effect that matches the scene information of the target sentence is determined from a preset sound effect library. The scene sound effect that matches the scene information can be determined from the preset sound effect library based on the preset scene semantics within the semantic information of the target sentence. The preset sound effect library contains multiple scene sound effects corresponding to the scene semantics. A determination can be made as to whether the semantic information of the target sentence contains the preset scene semantics; if so, the scene sound effect that matches the preset scene semantics is retrieved from the preset sound effect library. It is understood that the preset sound effect library can pre-store multiple scene sound effects, such as a car horn or the sound of reading a book, or more scene sound effects can be acquired in real time via a network connection. The preset sound effect library is typically stored in the audiobook production device. By analyzing the semantic information of the input text, if a semantics matching the specified scene is found, the corresponding sound from the sound effect library is used. Natural language processing techniques are used to analyze whether the current sentence contains the specified scene semantics. For example, if the sentence "a train is coming," a train whistle sound file is retrieved from the sound effect library. For example, "walking over here in high heels" means finding the sound file of high heels stepping on the ground from the sound effect library.

[0089] It is understandable that the execution order of step 303 and step 304 is not limited here.

[0090] 305. Determine the sentence position of the scene information in the target sentence.

[0091] In an embodiment of the present application, when mixing the character's reading voice and the scene sound effects, it is necessary to determine the sentence position of the scene information in the target sentence. Among them, the first phoneme sequence of the scene content text corresponding to the scene information can be obtained, that is, the scene content text is converted into the first phoneme sequence. It can be understood that phonemes are the smallest speech units divided according to the natural properties of the language. Each text has a corresponding phoneme, and the phoneme can be represented by pinyin; if the target sentence is "Zhang San said: I heard a knock on the door", at this time, the scene content text is: knock on the door, and the corresponding first phoneme sequence is "qi, ao, me, n". After obtaining the first phoneme sequence corresponding to the scene information, it can be determined that the first phoneme sequence is located in the sequence position of the phoneme sequence corresponding to the target sentence, and the sequence position is used as the sentence position. It can be understood that the target sentence can be converted into the corresponding phoneme sequence, and then the sequence position of the first phoneme sequence in the phoneme sequence corresponding to the target sentence can be determined.

[0092] 306. Add scene sound effects to the audio segment corresponding to the sentence position of the character's reading sound to obtain target audio corresponding to the target sentence.

[0093] After determining the sentence position of the scene information in the target sentence, the scene sound effect can be added to the audio segment of the character's reading sound corresponding to the sentence position to obtain the target audio corresponding to the target sentence; wherein, the scene sound effect can be added to the audio segment of the character's reading sound according to the sequence position of the first phoneme sequence in the phoneme sequence corresponding to the target sentence, specifically including the following steps 3061 to 3064.

[0094] 3061. Determine an audio frame sequence corresponding to a first phoneme sequence of scene content text corresponding to the scene information in the character's reading voice.

[0095] In an embodiment of the present application, the audio frame sequence corresponding to the first phoneme sequence of the scene content text corresponding to the scene information in the character's reading voice can be determined. Specifically, after determining the sequence position of the first phoneme sequence in the phoneme sequence corresponding to the target sentence, the audio frame sequence corresponding to the first phoneme sequence in the character's reading voice is determined according to the sequence position. For example, if the target sentence is: "Zhang San said: I heard a knock on the door", the first phoneme sequence is "qi, ao, me, n", and the character's reading voice is Zhang San reading: I heard a knock on the door; the character's reading voice can be converted into a corresponding reading phoneme sequence based on the acoustic characteristics of the character's reading voice, such as converting the character's reading voice into a corresponding reading phoneme sequence through RNNLSTM speech recognition technology; it can be understood that the character's reading voice is obtained by vocalizing the target sentence according to the audio characteristics, and the reading phoneme sequence obtained by converting the character's reading voice is the same as the phoneme sequence corresponding to the target sentence. In the reading phoneme sequence, each pronunciation phoneme corresponds to multiple audio frames. The first phoneme sequence can be compared with the reading phoneme sequence, that is, the phoneme sequence "qi, ao, me, n" can be found from the reading phoneme sequence, and the multiple audio frames corresponding to each pronunciation phoneme in the phoneme sequence can be determined, and then the audio frame sequence corresponding to the first phoneme sequence in the character's reading voice can be determined, that is, the audio segment corresponding to the phoneme sequence "qi, ao, me, n" in the character's reading voice can be determined.

[0096] 3062. Use the preset audio frame corresponding to the audio frame sequence as the sound effect start frame of the scene sound effect.

[0097] In an embodiment of the present application, it is necessary to determine which audio frame in the character's reading voice begins to add the scene sound effect, that is, to determine the starting frame of the scene sound effect in the character's reading voice. After determining the audio frame sequence corresponding to the first phoneme sequence in the character's reading voice, the preset audio frame corresponding to the audio frame sequence can be used as the sound effect starting frame of the scene sound effect. Specifically, the audio frame ranked first in the audio frame sequence can be used as the sound effect starting frame of the scene sound effect, or the audio frame preceding the audio frame sequence can be used as the sound effect starting frame of the scene sound effect, or the last audio frame in the audio frame sequence can be used as the sound effect starting frame of the scene sound effect, and the specific details are not limited here. For example, if the first phoneme sequence is "qi, ao, me, n", the audio frame ranked first in the pronunciation order of the multiple audio frames corresponding to the pronunciation phoneme "qi" can be used as the sound effect starting frame of the scene sound effect, or the audio frame preceding the first audio frame can be used as the sound effect starting frame of the scene sound effect. It is understandable that the sound effect starting frame is an audio frame in the character's reading voice, which is used to indicate that the scene sound effect starts to be added at this sound effect starting frame.

[0098] 3063. Determine the duration of the scene sound effect based on the sound source in the scene information.

[0099] In an embodiment of the present application, when adding scene sound effects to the voice of the character reading aloud, it is also necessary to determine the end time of the scene sound effects. The duration of the scene sound effects can be determined based on the sound source in the scene information; specifically, the type of scene sound effects can be determined based on the volume change of the sound source in the actual scene; specifically, in the actual scene, if the volume change value of the sound source within the preset duration is greater than the preset volume threshold, the scene sound effect is determined to be a trigger sound, and the duration of the trigger sound is less than the preset duration; the preset duration can be 10 seconds or 12 seconds, which is not specifically limited here; the preset volume threshold can be 50dB or 60dB, which is not specifically limited here; the preset duration can be 2 seconds or 3 seconds, which is not specifically limited here. It can be understood that if the volume change value of the sound source within the preset time is greater than the preset volume threshold, it can be determined that the volume change of the sound emitted by the sound source is relatively abrupt, such as knocking on the door, objects falling, etc., and the sound duration of the sound source is generally short. At this time, it can be determined that the scene sound effect corresponding to the sound source is the trigger sound; the sound effect duration of the trigger sound is generally short, and the sound effect duration of the trigger sound can be set to 1 second or 2 seconds, that is, the sound effect duration of the scene sound effect is 1 second or 2 seconds.

[0100] If the volume change value of the sound source within the preset duration is less than the preset volume threshold, the scene sound effect is determined to be ambient background sound, and the duration of the ambient background sound effect is greater than the preset duration. Among them, the duration of the ambient background sound effect is greater than the duration of the trigger sound effect. It is understandable that if the volume change value of the sound source within the preset duration is less than the preset volume threshold, the volume change of the sound emitted by the sound source is relatively gentle, such as the sound of a bell, the sound of traffic, and other continuous surrounding sounds, and the duration of the sound of the sound source is generally longer. At this time, it can be determined that the scene sound effect corresponding to the sound source is ambient background sound, and the duration of the sound effect of the ambient background sound is generally longer. The duration of the sound effect of the ambient background sound can be set to 10 seconds or 8 seconds, that is, the duration of the sound effect of the scene sound effect is 10 seconds or 8 seconds. It is understandable that the sound effect of the ambient background sound can also be set to last until the end of the target sentence read by the character, so that the character reading sound of the target sentence has the corresponding ambient background sound from the sound effect start frame to the end of the character reading sound.

[0101] 3064. Add scene sound effects to the character's reading voice based on the sound effect start frame and sound effect duration.

[0102] After determining the sound effect start frame and the duration of the scene sound effect, the scene sound effect can be added to the character's reading sound according to the sound effect start frame and the sound effect duration. Specifically, the scene sound effect can be faded into the character's reading sound at the sound effect start frame corresponding to the character's reading sound, and the volume of the scene sound effect can be increased until it reaches a preset volume. Since the added scene sound effect serves as background sound, the preset volume is generally lower than the volume of the character's reading sound; that is, a low-volume scene sound effect is added to the character's reading sound at the sound effect start frame, and then the volume of the scene sound effect is continuously increased until it reaches the preset volume. The volume of the scene sound effect is reduced at the preset end period of the sound effect duration, and the scene sound effect is faded out of the character's reading sound, that is, when the sound effect duration is about to end, the volume of the scene sound effect is gradually reduced from the preset volume until the volume of the scene sound effect is reduced to zero or the scene sound effect can no longer be heard. It is understandable that the volume of the scene sound effect is continuously increased when fading in. When it is increased to a preset volume, the preset volume is maintained until the character reading sound is faded out. The character reading sound can also be directly faded out without being increased to the preset volume.

[0103] It is understandable that the purpose of adding scene sound effects to the character's reading voice with fade-in and fade-out is to avoid the sense of lag caused by abrupt splicing when the two sounds (the character's reading voice and the scene sound effects) overlap in time. For example, if the target sentence read by the character is: Zhang San said: "I heard a knock on the door", then the audio frame that ranks first in the pronunciation order among the multiple audio frames corresponding to the pronunciation phoneme "qi" can be used as the sound effect starting frame of the scene sound effect, and the scene sound effect can be faded into the character's reading voice. At this time, the sound effect of the scene sound effect lasts for 2 seconds, and the scene sound effect can be faded out of the character's reading voice starting from 1.5 seconds.

[0104] Furthermore, when adding scene sound effects to the character's reading, if multiple scene sound effects overlap within the same audio frame—that is, if multiple scene sound effects are located within the same audio frame, and if a trigger sound and ambient background sound are present within these multiple scene sound effects—then the ambient background sound can be reduced and the trigger sound can be enhanced to emphasize the trigger sound. It's understandable that the trigger sound changes abruptly, and its volume fluctuates very rapidly. For example, if the target sentence is "The sound of traffic outside disturbed me, and then I heard someone knocking on the door," the ambient background sound corresponding to the "traffic" may persist until the trigger sound corresponding to the "knock" begins. To enhance the user's auditory experience, the trigger sound needs to be emphasized. The volume of the trigger sound can be increased based on a preset volume balancing technique, so that it is higher than the ambient background sound. Specifically, the ambient background sound and trigger sound within the same audio frame can be volume-balanced and normalized to a preset volume. Volume balancing involves calculating the average power of adjacent sound files. This average power is calculated by taking the logarithm of the squared average of the speech sampling points over time, expressed in decibels. Then, the audio is normalized to the same or specified volume. The volume of the trigger sound in the audio frame is then increased, while the volume of the ambient background sound is reduced, so that the volume of the trigger sound is higher than the volume of the ambient background sound. It is understood that when multiple scene sound effects overlap, if you want to highlight a particular scene sound effect, you can also increase (enhance) the volume of that scene sound effect using the preset volume balancing technology.

[0105] In one feasible method, after obtaining the target audio corresponding to the target sentence, the character reading voice of the target sentence can be aligned with the corresponding text input preset voice alignment model, and subtitles with timestamps can be output, that is, through voice alignment technology, the reading voice without scene sound effects and the corresponding text are used as input to obtain output subtitles with timestamps. Specifically, the Kaldi-based voice-to-text alignment technology can be used to obtain the start and end time of each text in the voice, or other technical solutions can be used. It should be noted that in the case of long texts, such as novels, the start and end time of paragraphs or single sentences are generally used as the time finally presented to the user, rather than the word-by-word time. This can avoid visual fatigue of the user while knowing the corresponding moment of the current voice reading.

[0106] The present application embodiment provides a device for producing an audio book, such as Figure 6 Shown, including:

[0107] An acquisition unit 601 is used to acquire the text corresponding to the audio book;

[0108] A determination unit 602 is configured to determine target sentences in the text that are relevant to the character and the scene;

[0109] The processing unit 603 is configured to perform voice processing on the target sentence according to the audio features corresponding to the character information of the target sentence, and obtain a character reading voice that matches the audio features;

[0110] An execution unit 604 is configured to obtain a scene sound effect that matches the scene information according to the scene information of the target sentence;

[0111] The audio mixing unit 605 is configured to determine the sentence position of the scene information in the target sentence, and add the scene sound effect to the audio segment of the character's reading sound corresponding to the sentence position to obtain the target audio corresponding to the target sentence.

[0112] The embodiment of the present application provides an audio book production device 700, such as Figure 7 Shown, including:

[0113] CPU 701, memory 702 and input / output interface 703;

[0114] The memory 702 is a temporary storage memory or a permanent storage memory;

[0115] The central processing unit 701 is configured to communicate with the memory 702 and execute instructions in the memory 702 to perform the above-mentioned audio book production method.

[0116] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0117] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0118] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0119] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0120] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

Claims

1. A method for producing an audio book, characterized in that: include: Get the text corresponding to the audiobook; Identifying target sentences in the text that are relevant to the character and the scene; Performing voice processing on the target sentence according to the audio features corresponding to the character information of the target sentence to obtain a character reading voice that matches the audio features; Obtaining a scene sound effect that matches the scene information according to the scene information of the target sentence; Determining a sentence position of the scene information in the target sentence, and adding the scene sound effect to an audio segment of the character's reading sound corresponding to the sentence position, to obtain a target audio corresponding to the target sentence; The sentence position includes: the sequence position of the first phoneme sequence corresponding to the scene information in the phoneme sequence corresponding to the target sentence; The adding of the scene sound effect into the audio segment corresponding to the position of the sentence in the character's reading sound comprises: Determine, according to the sequence position, an audio frame sequence corresponding to the first phoneme sequence in the character's reading voice; Using the preset audio frame corresponding to the audio frame sequence as the sound effect start frame of the scene sound effect; Determining the duration of the scene sound effect according to the sound source in the scene information; The scene sound effect is added to the character's reading voice according to the sound effect start frame and the duration of the sound effect.

2. The production method according to claim 1, characterized in that Determining the sentence position of the scene information in the target sentence includes: Obtaining a first phoneme sequence of scene content text corresponding to the scene information; The sequence position of the first phoneme sequence in the phoneme sequence corresponding to the target sentence is determined, and the sequence position is used as the sentence position.

3. The production method according to claim 1, characterized in that The determining of the duration of the scene sound effect according to the sound source in the scene information includes: If the volume change value of the sound source within the preset time period is greater than the preset volume threshold, the scene sound effect is determined to be a trigger sound, and the duration of the trigger sound is less than the preset duration; If the volume change value of the sound source within the preset duration is less than the preset volume threshold, the scene sound effect is determined to be ambient background sound, and the duration of the sound effect of the ambient background sound is greater than the preset duration.

4. The production method according to claim 3, characterized in that: Adding the scene sound effect to the character's reading voice includes: If there are multiple scene sound effects in the same audio frame, and the trigger sound and the ambient background sound are included in the multiple scene sound effects, the volume of the trigger sound is increased based on a preset volume balancing technology so that the volume of the trigger sound is higher than the volume of the ambient background sound.

5. The production method according to claim 1, characterized in that: Adding the scene sound effect to the character's reading voice according to the sound effect start frame and the duration of the sound effect includes: Fading the scene sound effect into the character's reading voice at the sound effect start frame, and increasing the volume of the scene sound effect until it reaches a preset volume; The volume of the scene sound effect is reduced during the preset end period of the sound effect duration, and the scene sound effect is faded out from the character's reading voice.

6. The manufacturing method according to claim 1, characterized in that Determining target sentences in the text that are relevant to the role and the scene includes: Determining whether the semantic information of the sentence in the text contains preset character dialogue information, or determining whether the semantic information of the sentence in the text matches the semantic information in the preset character sentence set; If the preset character dialogue information exists, or matches the semantic information in the preset character sentence set, then determining that the sentence is a character sentence related to the character; If the role sentence contains the preset scene semantics, the role sentence is determined to be the target sentence.

7. The production method according to claim 1, characterized in that: The step of performing voice processing on the target sentence according to the audio features corresponding to the character information of the target sentence to obtain a character reading voice matching the audio features comprises: Inputting the target sentence into a preset reading model, and determining the timbre characteristics corresponding to the character information of the target sentence in the preset reading model; Converting the phoneme sequence corresponding to the target sentence into a target audio feature based on the timbre feature corresponding to the character information; The target audio feature is converted into a sound waveform to obtain a character reading voice that matches the audio feature.

8. The production method according to claim 1, characterized in that: The step of obtaining a scene sound effect that matches the scene information according to the scene information of the target sentence includes: According to the preset scene semantics in the scene information, a scene sound effect matching the scene information is determined from a preset sound effect library, wherein the preset sound effect library contains a plurality of scene sound effects corresponding to the scene semantics.

9. An audio book production device, characterized in that: include: CPU, memory and input / output interfaces; The memory is a transient storage memory or a persistent storage memory; The central processing unit is configured to communicate with the memory and execute instructions in the memory to perform the method according to any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that The method comprises instructions which, when executed on a computer, cause the computer to execute the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Method and device for generating audio for plain text document

    CN110491365A