Audio book video generation method and device, electronic equipment, medium and program product
By generating audio book videos and combining multi-frame images and audio data, the problem of audio book ignoring visual feelings is solved, a better auditory and visual experience is achieved, and users' understanding and attraction to the story are enhanced.
Patent Information
- Application Number
- CN202510510921.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-07-22
AI Technical Summary
Existing audio books ignore the user's visual feelings, resulting in poor auditory experience.
By obtaining multi-frame images and audio data of story text, audio book videos are generated, and combined with biographical models and audio animation resources, enhance visual and auditory experiences.
It improves users' understanding of story content, enhances auditory and visual experience, and enhances the attractiveness of audio books.
Smart Images

Figure CN120353967A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technologies, and in particular, to a method, apparatus, electronic device, medium, and program product for generating an audiobook video. Background Art
[0002] With the improvement of people's quality of life, people increasingly attach importance to enriching their spiritual world through reading. Whether it is a paper book or an e-book, it can provide users with rich knowledge and spiritual nourishment. With the development of technology, audiobooks have emerged, enabling people to obtain book content more conveniently and quickly. Summary of the Invention
[0003] According to some embodiments of the present disclosure, there is provided a method for generating an audiobook video, including: obtaining multiple frames of images corresponding to a story text, where the multiple frames of images are used to represent key plots of the story text; determining target audio data in the audio data corresponding to the story text; and generating an audiobook video corresponding to the story text based on the target audio data and the multiple frames of images.
[0004] According to some other embodiments of the present disclosure, there is provided an apparatus for generating an audiobook video, including: an image acquisition module configured to obtain multiple frames of images corresponding to a story text, where the multiple frames of images are used to represent key plots of the story text; an audio acquisition module configured to determine target audio data in the audio data corresponding to the story text; and a video generation module configured to generate an audiobook video corresponding to the story text based on the target audio data and the multiple frames of images.
[0005] According to some embodiments of the present disclosure, there is provided an electronic device, including: a processor; and a memory coupled to the processor for storing instructions, where when the instructions are executed by the processor, the processor executes the method for generating an audiobook video according to any one of the embodiments in the present disclosure.
[0006] According to some embodiments of the present disclosure, there is provided a computer-readable storage medium having computer instructions stored thereon, where when the computer instructions are executed by a processor, the method for generating an audiobook video according to any one of the embodiments in the present disclosure is executed.
[0007] According to some embodiments of the present disclosure, there is provided a computer program product, including: computer instructions, where when the computer instructions are executed by a processor, the method for generating an audiobook video according to any one of the embodiments in the present disclosure is implemented.
[0008] Other features, aspects, and advantages of the present disclosure will become clear through the following detailed description of the exemplary embodiments of the present disclosure with reference to the accompanying drawings. Brief Description of the Drawings
[0009] Embodiments of the present disclosure will be described below with reference to the accompanying drawings. It should be understood that the drawings in the following description only relate to some embodiments of the present disclosure and do not constitute a limitation to the present disclosure. In the drawings:
[0010] Figure 1 A flowchart showing a method for generating an audiobook video according to some embodiments of the present disclosure;
[0011] Figure 2 A flowchart showing a process for obtaining multiple frames of images corresponding to a story text according to some embodiments of the present disclosure;
[0012] Figure 3 A flowchart showing a process for obtaining multiple frames of images corresponding to a story text according to some other embodiments of the present disclosure;
[0013] Figure 4 A flowchart showing a process for generating subtitle text according to some embodiments of the present disclosure;
[0014] Figure 5 A flowchart showing a process for adding audio effect resources according to some embodiments of the present disclosure;
[0015] Figure 6 A flowchart showing a method for generating an audiobook video according to some other embodiments of the present disclosure;
[0016] Figure 7 A block diagram showing an audiobook video generation device according to some embodiments of the present disclosure;
[0017] Figure 8 A block diagram showing an electronic device according to some embodiments of the present disclosure;
[0018] Figure 9 A block diagram showing an electronic device according to some other embodiments of the present disclosure.
[0019] It should be understood that for ease of description, the dimensions of the various parts shown in the drawings are not necessarily drawn to actual scale. The same or similar reference numerals are used in the various drawings to represent the same or similar components. Therefore, once an item is defined in one drawing, it may not be further discussed in subsequent drawings. Detailed Embodiments
[0020] The technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. It should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein.
[0021] It should be understood that the various steps described in the method embodiments of the present disclosure can be executed in different orders and / or executed in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard. Unless otherwise specifically stated, the relative arrangements, numerical expressions, and numerical values of the components and steps set forth in these embodiments should be construed as merely exemplary and not limiting the scope of the present disclosure.
[0022] The term "comprising" and its variants used in the present disclosure mean open terms that include at least the subsequent elements / features but do not exclude other elements / features, that is, "including but not limited to". The term "based on" means "at least partially based on".
[0023] It should be noted that the concepts such as "first", "second", etc. mentioned in the present disclosure are only used to distinguish different devices, modules, or units, and are not used to limit the order of the functions performed by these devices, modules, or units or their interdependencies. Unless otherwise specified, the concepts such as "first", "second", etc. are not intended to imply that the objects so described must be in a given order in terms of time, space, ranking, or any other way.
[0024] It should be noted that the modifications of "one" and "plural" mentioned in the present disclosure are illustrative rather than restrictive. Those skilled in the art should understand that, unless otherwise clearly indicated in the context, it should be understood as "one or more".
[0025] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.
[0026] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present disclosure are all information and data authorized by the user or fully authorized by all parties. And the collection, use, and processing of the relevant data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation entrances are provided for the user to choose to authorize or refuse.
[0027] The embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings, but the present disclosure is not limited to these specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. In addition, in one or more embodiments, specific features, structures, or characteristics can be combined in any suitable manner that will be clear to those of ordinary skill in the art from the present disclosure.
[0028] In the context of the present disclosure, an image may refer to any one of a variety of images, such as a color image, a grayscale image, etc. It should be noted that, in the context of this specification, the type of the image is not specifically limited. In addition, the image can be any suitable image, such as an original image, or an image that has been subjected to specific processing on the original image, such as preliminary filtering, anti-aliasing, color adjustment, contrast adjustment, normalization, and so on. It should be noted that the preprocessing operations may also include other types of preprocessing operations known in the art, which will not be described in detail here.
[0029] In the related art, although audiobooks can enable people to obtain book content more conveniently and quickly, they ignore people's visual experience. The present disclosure provides a scheme for generating an audiobook video, which can enhance the user's auditory experience and visual experience while providing book content conveniently and quickly.
[0030] Figure 1 A flowchart showing a method for generating an audiobook video according to some embodiments of the present disclosure is shown.
[0031] As Figure 1 shown, the method for generating an audiobook video includes: step S1, obtaining multiple frames of images corresponding to the story text, where the multiple frames of images are used to represent the key plots of the story text; step S2, determining target audio data in the audio data corresponding to the story text; step S3, generating an audiobook video corresponding to the story text based on the target audio data and the multiple frames of images.
[0032] The story text can be book text, such as novel text, or can be script text adapted from the book text, such as containing lines applied to fields such as film and television, drama, etc. The multiple frames of images are a visual presentation of the story text. By playing the multiple frames of images, it is beneficial to help users better understand the story content.
[0033] The story text can correspond to multiple audio data, and the target audio data can be selected according to requirements. The story text can also correspond to one audio data, and this audio data is the target audio data.
[0034] The target audio data and the multiple frames of images are synthesized, and each audio segment in the target audio data corresponds to the corresponding image.
[0035] In the above embodiments, an audiobook video corresponding to the story text is generated based on the target audio data and the multiple frames of images, which improves the charm of the story, can meet both the user's auditory requirements and visual requirements, enhances the user's ability to understand the story content, and thus enhances the user attraction.
[0036] Next, the method for generating an audiobook video will be further introduced in conjunction with Figures 2 to 6 Hereinafter.
[0037] Figure 2 A schematic flowchart showing the process of obtaining multiple frames of images corresponding to a story text according to some embodiments of the present disclosure.
[0038] As Figure 2 shown, obtaining multiple frames of images corresponding to a story text includes: Step S101, dividing the story text into multiple storyboard texts according to the storyboard granularity; Step 102, generating corresponding images for each storyboard text in the multiple storyboard texts based on an image generation model.
[0039] For example, the story text is divided according to a specified word count range to obtain multiple storyboard texts. In some embodiments, if the word count of a text segment in the story text is within the specified word count range, the text segment is divided into a storyboard text. If the sum of the word counts of multiple text segments in the story text is within the specified word count range, these multiple text segments are divided into a storyboard text. In this way, when playing multiple frames of images, the images do not change too fast or too slow.
[0040] The image generation model can generate corresponding images according to the input content. The specific image generation model used is not limited in the present disclosure.
[0041] In this embodiment, by generating corresponding images for each storyboard text and finally splicing the multiple frames of images, a smooth picture can be provided for the audiobook video.
[0042] Figure 3 A schematic flowchart showing the process of obtaining multiple frames of images corresponding to a story text according to other embodiments of the present disclosure.
[0043] As Figure 3 shown, obtaining multiple frames of images corresponding to a story text includes: Step S111, intercepting multiple images from the audiovisual media content corresponding to the story text according to the storyboard granularity, where the audiovisual media content includes comics or videos; Step S112, using the multiple images as the multiple frames of images.
[0044] For example, if there are already corresponding comics or short videos for the story text, some frames in the comics or short videos can be used as the multiple frames of images corresponding to the story text, reducing the image generation process and improving the generation efficiency of the audiobook video.
[0045] Before generating the audiobook video, it is necessary to select corresponding target audio data for the story text. In some embodiments, at least one of the first audio data and the second audio data is selected as the target audio data from the audio data corresponding to the story text. The first audio data is generated based on a text-to-speech model, and the second audio data is generated by a user's performance based on the story text.
[0046] The first audio data is TTS (Text-To-Speech) audio data, that is, a voice audio file generated by converting text content through TTS technology. TTS audio data can be generated quickly, saving time and cost, and TTS audio data has high consistency in terms of pronunciation, intonation, speech rate, etc., making the audio quality stable. By adjusting parameters, different timbres, intonations, speech rates, etc. can be obtained.
[0047] The second audio data is audiobook audio data, specifically, for example, live audiobook audio data, that is, a book form in which a real user reads a book aloud and generates audio. Through the voice performance of the real user, including intonation, speech rate, tone, etc., characters are shaped and the atmosphere is created. The live audiobook audio data has better sound quality and supports multi-person broadcasting, being able to vividly display the plot in the book and enabling listeners to obtain a better auditory experience.
[0048] The target audio data can include only the first audio data, or only the second audio data, or both the first audio data and the second audio data. For example, for each character in the story text, an audio clip from the first audio data can be selected, and for the narration in the story text, an audio clip from the second audio data can be selected. Or, for each character in the story text, an audio clip from the second audio data can be selected, and for the narration in the story text, an audio clip from the first audio data can be selected. Or, some characters in the story text select audio clips from the first audio data, some characters select audio clips from the second audio data, and the narration selects the corresponding audio clip according to the actual situation.
[0049] In this embodiment, by selecting the corresponding target audio data for the story text, a rich auditory experience can be brought to the user.
[0050] In some embodiments, if the target audio data includes the first audio data, in the audio data corresponding to the story text, determining the target audio data includes: in the audio data corresponding to the story text, selecting the first audio data corresponding to the target timbre.
[0051] For example, if the first audio data is multi-timbre single-cast audio data, then the first audio data of the target timbre can be selected for the entire story text. For example, the target timbre can be selected according to the overall atmosphere of the story text. For example, for a warm love story, a gentle and soft timbre can be selected, which helps to create a romantic and sweet atmosphere and enables listeners to better immerse themselves in the story; for a horror novel, a gloomy and deep timbre can be selected to enhance the horror atmosphere.
[0052] In this embodiment, by selecting a rich variety of timbres, the story becomes more vivid, thereby making the audiobook video more attractive and improving the quality of the audiobook video.
[0053] In some embodiments, if the target audio data includes the first audio data, determining the target audio data in the audio data corresponding to the story text includes: for each character and the narrator in the story text, selecting a first audio segment with a corresponding target timbre, and different first audio segments come from one or more of the first audio data.
[0054] For example, the first audio data can be multi-timbre multi-cast audio data, and an audio frequency band with a corresponding target timbre is selected for each character and the narrator. For example, based on the description information of each character, an audio segment with a corresponding target timbre is selected for each character. These audio segments can come from the same first audio data or from different first audio data. For example, a story may include children, good people, bad people, etc. An audio segment with a lively and cute timbre can be selected for children, an audio segment with a warm and gentle timbre can be selected for good people, an audio segment with a deep and hoarse timbre can be selected for bad people, etc., and an audio segment with a steady timbre can be selected for the narrator.
[0055] The respective audio segments are spliced in the order of the story text, so as to better lead the listener into the story context and improve the attractiveness of the audiobook video.
[0056] In some embodiments, if the target audio data includes the second audio data, determining the target audio data in the audio data corresponding to the story text includes: for each character and the narrator in the story text, selecting a corresponding second audio segment, and different second audio segments come from one or more of the second audio data.
[0057] For example, an audio frequency band with a corresponding target timbre is selected for each character and the narrator. These audio segments can come from the same second audio data or from different second audio data. Since the second audio data is generated by real users reading books, if a second audio data is generated by a user reading a book, different second audio data have different timbres. Therefore, audio segments from different second audio data can be selected for each character and the narrator in the story. Of course, the timbres of each character can also not be distinguished, and a second audio data can be directly selected as the target audio data for the story text. If a second audio data is generated by multiple users reading books according to different characters, then this second audio data can also be directly selected as the target audio data for the story text.
[0058] In this embodiment, by selecting the audio frequency bands corresponding to the target timbres for the characters and the narrator in the story text, the listener can be more easily immersed in the story.
[0059] When generating an audiobook video, in addition to multiple frames of images and target audio data, subtitle text can also be included.
[0060] In some embodiments, if the first audio data is selected for the story text, then generating the audiobook video corresponding to the story text based on the target audio data and the multiple frames of images includes: using the story text as the subtitle text; generating the audiobook video corresponding to the story text based on the first audio data, the multiple frames of images, and the subtitle text.
[0061] Since the first audio data is audio data generated based on the TTS technology, this audio data is basically consistent with the story text. Therefore, the story text can be used as the subtitle text. The generated audiobook video based on the first audio data, multiple frames of images, and subtitle text can satisfy both the user's auditory experience and visual experience. For some users with visual and auditory impairments, by displaying the subtitle text, it can better help the users understand the story content.
[0062] In some embodiments, if the second audio data is selected for the story text, since the user may add self-interpreted elements when reading the story aloud, such as adding or deleting words or adjusting the sentence order, etc., the audio data and the story text may not be completely consistent. Therefore, it is necessary to generate subtitle text that matches the second audio data.
[0063] Figure 4 The flowchart showing the process of generating subtitle text according to some embodiments of the present disclosure is shown.
[0064] As Figure 4 shown, in step S41, convert the second audio data into a performance text, and the performance text includes multiple text segments; in step S42, based on the story text, correct at least one of the multiple text segments in terms of words and sentence breaks to generate subtitle text.
[0065] Generating the audiobook video corresponding to the story text based on the target audio data and the multiple frames of images includes: generating the audiobook video corresponding to the story text based on the second audio data, the multiple frames of images, and the subtitle text.
[0066] For example, the ASR text corresponding to the second audio data is obtained by using ASR (Automatic Speech Recognition) technology, and the ASR text is the performance text. The performance text may contain many homophonic typos and so the performance text cannot be used directly as a subtitle text. In this embodiment, the story text is used to identify homophonic typos and some sentence segmentation problems in each text segment, while respecting the reasonable interpretation in the oral broadcast, which can improve the accuracy of the subtitle text, that is, obtain a subtitle text that corresponds to the second audio data and is accurate. Then, based on the second audio data, multiple frames of images and subtitle text, a more vivid and touching audio book video is generated to meet the user's auditory and visual experience.
[0067] Next, we will continue to introduce how to correct the text, sentence breaks, etc. of the broadcast text.
[0068] In some embodiments, based on the story text, correcting homophonic typos in each of the plurality of text segments is performed.
[0069] Compare the differences between the story text and the performance text, and identify homophonic typos in the performance text. Refer to the content of the story text, replace the homophonic typos in the performance text, but do not modify the non-typos, for example, do not delete or add text, and do not interpret it by itself, so that the generated subtitle text can remain basically consistent with the second audio data. For example, "eh" in the story text corresponds to "I" in the performance text. Since "I" is not a homophonic character of "eh", "I" in the performance text is not corrected. For another example, "I didn't expect" in the story text corresponds to "I didn't expect it" in the performance text. "I didn't expect it" is not a homophonic character of "I didn't expect it", so "I didn't expect it" in the performance text is not corrected. For another example, the Chinese pinyin of "food" in the story text is "shiwu", while the Chinese pinyin of "fifteen" in the performance text is also "shiwu". Since "fifteen" is a homophonic character of "food", "fifteen" in the performance text needs to be corrected.
[0070] In some embodiments, based on the story text, adjacent text segments with incorrect sentence breaks among the multiple text segments are merged or split.
[0071] For example, the text segment in the story text is "I also have a day to win an award", but the performance text is divided into two text segments, "I also have" and "a day to win an award", so the two text segments need to be merged into one text segment. For another example, if the text segments in the story text are "In the kitchen, my mother is cutting vegetables" and "The water in the pot is boiling, and steam is coming out", but the performance text is "In the kitchen, my mother is cutting vegetables, and the water in the pot is boiling, and steam is coming out", then one text segment in the performance text needs to be segmented.
[0072] By merging or splitting adjacent text segments with incorrect sentence breaks among multiple text segments, the presentation text expression can be made more accurate and complete.
[0073] In some embodiments, based on the story text, at least one of text and sentence segmentation correction is performed on the contents in the multiple text segments, which are included in both the performance text and the story text.
[0074] For example, the content in the story text may be more than that in the performance text, and the final subtitle text does not need to add content that is not in the performance text; or, the content in the performance text does not appear in the story text, in this case, there is no need to delete the content that does not appear in the story text, and respect the reasonable interpretation in the oral broadcast. Only the content included in both the performance text and the story text is corrected to improve the accuracy of the correspondence between the subtitle text and the second audio data. By flexibly processing the text interpretation, the subtitle text can be made natural and smooth.
[0075] In the case where the audio in the audio book video is the second audio data, the method of obtaining multiple frames of images corresponding to the story text includes: dividing the subtitle text into multiple storyboard texts according to the storyboard granularity; and generating a corresponding image for each of the multiple storyboard texts based on the raw image model.
[0076] For example, a corresponding image is generated based on the subtitle text corrected by the performance text, so that the generated image has a higher degree of match with the subtitle text and audio data, thereby improving the consistency of the image, audio and subtitle in the audio book video.
[0077] In order to further enhance the overall experience of audiobook videos for users, you can also add audio and video effects to audiobook videos. Figure 5 This article introduces how to add audio and animation resources to audiobook videos.
[0078] Figure 5 A schematic diagram of a process for adding audio motion effect resources according to some embodiments of the present disclosure is shown.
[0079] like Figure 5As shown, in step S51, provide corresponding audio sound effect resources for at least one plot in the story text; in step S52, add the audio sound effect resources to the audiobook video.
[0080] The plot corresponding to the story text is also the plot corresponding to the subtitle text.
[0081] By understanding the content in the story text, the plots that need sound effects can be extracted, and appropriate audio sound effect resources can be matched for each plot. For example, by analyzing the semantics, emotions, and scene changes of the story text, the corresponding audio sound effect resources can be accurately matched. Such as sound effects of fireworks, heavy rain, falling leaves, lightning, etc., so that the audiobook video can increase the user's immersion, improve the user's attention, and enhance the rich auditory and visual experience.
[0082] In some embodiments, providing corresponding audio sound effect resources for at least one plot in the story text includes: extracting at least one of the weather keyword and the atmosphere signal word corresponding to the at least one plot; based on at least one of the weather keyword and the atmosphere signal word, providing the audio sound effect resources for the at least one plot.
[0083] For example, by extracting semantic features, the weather keyword or the atmosphere signal word can be identified. Weather keywords include, for example, rain, snow, lightning, etc. Atmosphere signal words can also be called emotion signal words, such as sadness, intense, romantic, surprise, etc. Add light rain special effects for the plot environment with light rain weather and some sad scenes; add heavy rain special effects for the plot environment with heavy rain or rainstorm weather and some disaster scenes; add light snow special effects for the plot environment with light snow weather and some aesthetic and romantic scenes; add heavy snow special effects for the plot environment with heavy snow scenes; add lightning special effects for the plot environment with lightning weather and some thriller twists, sudden events, etc.; add fireworks special effects for New Year and some celebration, surprise events, etc.; add falling leaf special effects for autumn falling leaves and some parting, memory events, etc.
[0084] Adding audio sound effect resources through weather keywords or atmosphere signal words can make the added audio sound effect resources more in line with the story scene, not appear obtrusive, and enhance the story's attractiveness.
[0085] Based on at least one of the weather keyword and the atmosphere signal word, providing the audio sound effect resources for the at least one plot includes: in response to both identifying the weather keyword and the atmosphere signal word, providing the audio sound effect resources for the at least one plot based on the weather keyword.
[0086] For example, set the priority of the weather keyword to be greater than the priority of the atmosphere signal word, so as to achieve that the special effects corresponding to natural disasters take precedence over the special effects corresponding to the environmental atmosphere.
[0087] Of course, those skilled in the art can set the priorities of weather keywords and atmosphere signal words according to the actual situation. For example, the priority of the atmosphere signal word can also be set to be greater than that of the weather keyword.
[0088] In some embodiments, providing the audio dynamic effect resources for the at least one plot includes: identifying a plurality of key plots in the story text; and providing corresponding audio dynamic effect resources for each of the plurality of key plots.
[0089] In this embodiment, only matching the corresponding audio dynamic effect resources for the key plots can save resources and improve the generation efficiency of the audiobook video.
[0090] In some embodiments, providing the audio dynamic effect resources for the at least one plot includes: determining time stamps corresponding to the at least one plot; and determining start and end times of the audio dynamic effect resources matched by the at least one plot based on the time stamps corresponding to the at least one plot.
[0091] For example, determine the time stamps of each plot in the audio data or the story text, and determine the timing of adding special effects according to the start and end times of the audio dynamic effect resources matched by the plot, so as to improve the consistency of the audio dynamic effect resources with the audio data and the story text.
[0092] Determining the start and end times of the audio dynamic effect resources matched by the at least one plot based on the time stamps corresponding to the at least one plot includes: advancing the start time in the time stamp corresponding to the target plot in the at least one plot by a specified time as the start time of the audio dynamic effect resources matched by the target plot; and using the end time in the time stamp corresponding to the target plot as the end time of the audio dynamic effect resources matched by the target plot.
[0093] For example, triggering the audio dynamic effect resources 0.1 - 0.5 seconds in advance, and the audio dynamic effect resources are extended until the corresponding subtitle text ends, which can achieve a more fascinating effect.
[0094] For the audio dynamic effect resources matched with the first sentence in the story text, in order to reduce the situation of only having special effects for the video picture, the audio dynamic effect resources may not be triggered in advance.
[0095] Taking the generation of an audiobook video from an audiobook audio as an example, the generation process of the audiobook video of the present disclosure will be introduced below. Figure 6 The flowchart showing the method for generating an audiobook video according to some other embodiments of the present disclosure is shown.
[0096] As Figure 6 shown, in step S61, the audiobook audio is converted into a performing text.
[0097] In step S62, the broadcast text is corrected based on the story text to obtain the subtitle text.
[0098] In step S63, the subtitle text is segmented to obtain the segmented text.
[0099] In step S64, corresponding images are generated based on the segmented text.
[0100] In step S65, audio effect resources are matched for the subtitle text.
[0101] In step S66, the audiobook audio, audio effect resources, images, and subtitle text are edited to generate an audiobook video.
[0102] For example, using some cloud editing API (Application Programming Interface) interfaces, the audiobook audio, audio effect resources, images, and subtitle text are constructed into a complete live-action audiobook video, thereby improving the user's visual and auditory experience as a whole.
[0103] The above is the method for generating an audiobook video provided by the present disclosure. The following will be combined with Figure 7 Describe an audiobook video generation device according to an embodiment of the present disclosure, which is used to execute any one of the above-mentioned methods for generating an audiobook video.
[0104] Figure 7 A block diagram of an audiobook video generation device according to some embodiments of the present disclosure is shown.
[0105] As Figure 7 shown, the audiobook video generation device 7 includes an image acquisition module 71, an audio acquisition module 72, and a video generation module 73.
[0106] The image acquisition module 71 is configured to acquire multiple frames of images corresponding to the story text, and the multiple frames of images are used to represent the key plots of the story text; the audio acquisition module 72 is configured to determine target audio data in the audio data corresponding to the story text; the video generation module 73 is configured to generate an audiobook video corresponding to the story text based on the target audio data and the multiple frames of images.
[0107] In the above embodiment, an audiobook video corresponding to the story text is generated based on the target audio data and multiple frames of images, which improves the charm of the story, can meet both the user's auditory requirements and visual requirements, and enhances the user's attractiveness.
[0108] In some embodiments, the image acquisition module 71 is configured to divide the story text into multiple shot texts according to shot granularity; and generate corresponding images for each of the multiple shot texts based on an image generation model.
[0109] Alternatively, the image acquisition module 71 is configured to intercept multiple images from the corresponding audio-visual media content of the story text according to shot granularity, where the audio-visual media content includes comics or videos; and use the multiple images as the multiple frames of images.
[0110] In some embodiments, the audio acquisition module 72 is configured to select at least one of first audio data and second audio data from the audio data corresponding to the story text as the target audio data, where the first audio data is generated based on a text-to-speech model, and the second audio data is generated by a user's performance based on the story text.
[0111] In some embodiments, when the target audio data includes the first audio data, the audio acquisition module 72 is configured to select the first audio data corresponding to a target voice from the audio data corresponding to the story text; or select first audio segments with corresponding target voices for each character and the narrator in the story text, and different first audio segments come from one or more of the first audio data.
[0112] In some embodiments, the video generation module 73 is configured to use the story text as subtitle text; and generate the audiobook video corresponding to the story text based on the first audio data, the multiple frames of images, and the subtitle text.
[0113] In some embodiments, when the target audio data includes the second audio data, the audio acquisition module 72 is configured to select corresponding second audio segments for each character and the narrator in the story text, and different second audio segments come from one or more of the second audio data.
[0114] In some embodiments, the audiobook video generation device further includes a text correction module (not shown), which is configured to convert the second audio data into performance text, where the performance text includes multiple text segments; and correct at least one of the words and sentence breaks of the multiple text segments based on the story text to generate subtitle text, where the video generation module 73 is configured to generate the audiobook video corresponding to the story text based on the second audio data, the multiple frames of images, and the subtitle text.
[0115] In some embodiments, the text correction module is configured to correct the homophonic misspelled words in each of the multiple text segments based on the story text.
[0116] In some embodiments, the text correction module is configured to merge or split adjacent text segments with incorrect sentence breaks among the multiple text segments based on the story text.
[0117] In some embodiments, the text correction module is configured to correct at least one of the text and sentence breaks for the content that is included in both the broadcast text and the story text among the multiple text segments based on the story text.
[0118] In some embodiments, the image acquisition module 71 is configured to divide the subtitle text into multiple storyboard texts according to the storyboard granularity; and generate corresponding images for each storyboard text among the multiple storyboard texts based on the image generation model.
[0119] In some embodiments, the video generation module 73 is further configured to provide corresponding audio effect resources for at least one plot in the story text; and add the audio effect resources to the audiobook video.
[0120] In some embodiments, providing corresponding audio effect resources for at least one plot in the story text includes: extracting at least one of the weather keyword and the atmosphere signal word corresponding to the at least one plot; and providing the audio effect resources for the at least one plot based on at least one of the weather keyword and the atmosphere signal word.
[0121] In some embodiments, providing the audio effect resources for the at least one plot based on at least one of the weather keyword and the atmosphere signal word includes: in response to both the weather keyword and the atmosphere signal word being recognized, providing the audio effect resources for the at least one plot based on the weather keyword.
[0122] In some embodiments, providing the audio effect resources for the at least one plot includes: identifying multiple key plots in the story text; and providing corresponding audio effect resources for each key plot among the multiple key plots.
[0123] In some embodiments, providing the audio effect resources for the at least one plot includes: determining the time stamp corresponding to the at least one plot; and determining the start and end times of the audio effect resources matched by the at least one plot based on the time stamp corresponding to the at least one plot.
[0124] In some embodiments, determining the start and end times of the audio effect resources matched to the at least one scenario based on the timestamps corresponding to the at least one scenario includes: advancing the start time in the timestamps corresponding to a target scenario in the at least one scenario by a specified time as the start time of the audio effect resources matched to the target scenario; and using the end time in the timestamps corresponding to the target scenario as the end time of the audio effect resources matched to the target scenario.
[0125] It should be noted that the above-mentioned individual units are only logical modules divided according to their specific implemented functions, rather than being used to limit the specific implementation manners. For example, they can be implemented in software, hardware, or a combination of software and hardware. In actual implementation, the above-mentioned individual units can be implemented as independent physical entities, or can also be implemented by a single entity (for example, a processor (such as a CPU or DSP, etc.), an integrated circuit, etc.). In addition, the above-mentioned individual units are shown by dashed lines in the drawings to indicate that these units may not actually exist, and the operations / functions they implement can be implemented by the processing circuit itself.
[0126] Figure 8 The block diagram of an electronic device according to some embodiments of the present disclosure is shown.
[0127] As Figure 8 shown, the electronic device 8 includes a processor 82; and a memory 81 coupled to the processor for storing instructions, which when executed by the processor, cause the processor to execute the audiobook video generation method as described above.
[0128] The memory 81 is used to store one or more computer-readable instructions. The memory 81 may include any combination of various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory, including but not limited to random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), read-only memory (ROM), flash memory. The memory 81 can store, for example, an operating system, application programs, a boot loader (BootLoader), a database, and other programs, and can also store various application programs and various data, etc.
[0129] The processor 82 is used to run computer-readable instructions to implement the audiobook video generation method described in any of the foregoing embodiments. The specific implementation of each step of the method can be referred to the above embodiments, and the repeated parts are not described herein again.
[0130] The processor 82 can be configured to execute Figures 1 to 6The steps therein. The processor 82 can be embodied as various processing devices, such as a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The central processing unit (CPU) can be of the X86 or ARM architecture, etc.
[0131] The processor 82 and the memory 81 can communicate with each other directly or indirectly. For example, the processor 82 and the memory 81 can communicate through a network. The network can include a wireless network, a wired network, and / or any combination of a wireless network and a wired network. The processor 82 and the memory 81 can also communicate with each other through a system bus, and the present disclosure does not limit this.
[0132] It should be noted that Figure 8 The components of the electronic device 8 shown are merely exemplary and not restrictive. According to actual application requirements, the electronic device 8 can also have other components. The processor 82 can control other components in the electronic device 8 to perform desired functions.
[0133] The electronic device 8 can be implemented in a software, firmware, and / or hardware manner and can be integrated in a device installed with relevant application programs.
[0134] In the above embodiments, by storing data instructions in the memory and then processing the above instructions through the processor, the overall visual and auditory experience of the user can be improved.
[0135] Figure 9 A block diagram of an electronic device according to some other embodiments of the present disclosure is shown. In some embodiments, the audiobook video generation device is presented in the form of an electronic device.
[0136] Figure 9 The electronic device 9 shown can be a computer system with a dedicated hardware structure and can perform corresponding functions when installed with relevant application programs.
[0137] The electronic device includes but is not limited to mobile terminals such as smartphones, laptop computers, personal digital assistants (PDAs), tablet personal computers (Tablet PCs), portable multimedia players (PMPs), vehicle terminals (such as vehicle navigation terminals), wearable devices, etc., and fixed terminals such as digital TVs, desktop computers, etc.
[0138] Such as Figure 9As shown, a central processing unit (CPU) 91 performs various processes according to a program stored in a read-only memory (ROM) 92 or a program loaded from a storage section 98 into a random access memory (RAM) 93. In the RAM 93, data required when the CPU 91 performs various processes and the like is stored as needed. The central processing unit is merely exemplary and may also be other types of processors, such as the various processors described above. The ROM 92, the RAM 93, and the storage section 98 may be various forms of computer-readable storage media. It should be noted that although Figure 9 the ROM 92, the RAM 93, and the storage section 98 are shown separately, one or more of them may be combined or located in the same or different memories or storage modules.
[0139] The CPU 91, the ROM 92, and the RAM 93 are connected to each other via a bus 94. An input / output interface 95 is also connected to the bus 94.
[0140] The following components are connected to the input / output interface 95: an input section 96, such as a touch screen, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output section 97, including a display, such as a cathode ray tube (CRT), a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage section 98, including a hard disk, a magnetic tape, etc.; and a communication section 99, including a network interface card such as a LAN card, a modem, etc. The communication section 99 allows communication processing to be performed via a network such as the Internet. It is easily understood that although Figure 9 parts in the electronic device 9 are shown to communicate via the bus 94, they may also communicate via a network or other means, where the network may include a wireless network, a wired network, and / or any combination of a wireless network and a wired network.
[0141] As needed, a drive 910 is also connected to the input / output interface 95. A removable medium 911, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is mounted on the drive 910 as needed, so that a computer program read therefrom is installed in the storage section 98 as needed.
[0142] In the case where the above series of processes are implemented by software, a program constituting the software can be installed from a network such as the Internet or a storage medium such as the removable medium 911.
[0143] According to an aspect of the present disclosure, there is provided a computer-readable storage medium having computer instructions stored thereon, wherein when the computer instructions are executed by a processor, the above-described audiobook video generation method is implemented.
[0144] The computer-readable storage medium can enhance the charm of the story.
[0145] According to one aspect of the present disclosure, there is provided a computer program product, including: computer instructions, which when executed by a processor, implement the above-mentioned audiobook video generation method.
[0146] This computer program product can enhance the charm of the story.
[0147] According to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, some embodiments of the present disclosure include a computer program product that, when run on a computer, causes the computer to implement the interaction method described in any of the foregoing embodiments. The computer program product includes computer instructions carried on a computer-readable medium, including program code for executing the interaction method shown in the flowchart. In such an embodiment, the computer instructions can be downloaded and installed from the network through the communication part 99, or installed from the storage part 98, or installed from the ROM 92. When the computer program is executed by the CPU 91, the interaction method of the embodiments of the present disclosure is executed.
[0148] It should be noted that in the context of the present disclosure, a computer-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device.
[0149] A computer-readable medium can be a computer-readable storage medium, a computer-readable signal medium, or any combination of the two.
[0150] Computer-readable storage media include, but are not limited to: electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or components, or any combination of the above. More specific examples of computer-readable storage media can include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus, or device. Computer instructions are stored on the computer-readable storage medium, and when executed by the processor, implement the interaction method described in any of the foregoing embodiments.
[0151] A computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. The computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the foregoing.
[0152] The foregoing computer-readable medium may be included in the aforementioned electronic device; or may exist separately without being assembled into the electronic device.
[0153] In some embodiments, a computer program is also provided, including: instructions that, when executed by a processor, cause the processor to execute the audiobook video generation method described in any of the foregoing embodiments. For example, the instructions may be embodied as computer program code.
[0154] In embodiments of the present disclosure, computer program code for performing the operations of the present disclosure may be written in one or more programming languages or combinations thereof. The foregoing programming languages include but are not limited to object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network (including a local area network (LAN) or a wide area network (WAN)), or may be connected to an external computer (e.g., by using an Internet service provider to connect through the Internet).
[0155] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, interactive methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0156] The functions described above can be performed, at least in part, by one or more hardware logic components. By way of example, and without limitation, exemplary hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), and the like.
[0157] Although some specific embodiments of the present disclosure have been described in detail by way of example, those skilled in the art should understand that the above examples are for illustrative purposes only and are not intended to limit the scope of the present disclosure. Those skilled in the art should understand that the above embodiments can be modified without departing from the scope and spirit of the present disclosure. The scope of the present disclosure is defined by the appended claims.
Claims
1. A method for generating an audiobook video, comprising: Obtaining multiple frames of images corresponding to a story text, where the multiple frames of images are used to represent key plots of the story text; Determining target audio data from the audio data corresponding to the story text; Generating an audiobook video corresponding to the story text based on the target audio data and the multiple frames of images.
2. The method for generating an audiobook video according to claim 1, wherein, The obtaining of multiple frames of images corresponding to the story text includes: Dividing the story text into multiple storyboard texts according to the storyboard granularity; Generating corresponding images for each of the multiple storyboard texts based on an image generation model.
3. The method for generating an audiobook video according to claim 1, wherein, The obtaining of multiple frames of images corresponding to the story text includes: Intercepting multiple images from the audiovisual media content corresponding to the story text according to the storyboard granularity, where the audiovisual media content includes comics or videos; Using the multiple images as the multiple frames of images.
4. The method for generating an audiobook video according to claim 1, wherein, The determining of target audio data from the audio data corresponding to the story text includes: Selecting at least one of first audio data and second audio data from the audio data corresponding to the story text as the target audio data, where the first audio data is generated based on a text-to-speech model, and the second audio data is generated by a user's performance based on the story text.
5. The method for generating an audiobook video according to claim 4, wherein, The target audio data includes the first audio data. The determining of target audio data from the audio data corresponding to the story text includes: Selecting the first audio data corresponding to a target voice in the audio data corresponding to the story text; or Selecting first audio segments with corresponding target voices for each character and the narrator in the story text, where different first audio segments come from one or more of the first audio data.
6. The method for generating an audiobook video according to claim 5, wherein, The generating of an audiobook video corresponding to the story text based on the target audio data and the multiple frames of images includes: Using the story text as subtitle text; Generating the audiobook video corresponding to the story text based on the first audio data, the multiple frames of images, and the subtitle text.
7. The method for generating an audiobook video according to claim 4, wherein, The target audio data includes the second audio data. The determining of target audio data from the audio data corresponding to the story text includes: Selecting corresponding second audio segments for each character and the narrator in the story text, where different second audio segments come from one or more of the second audio data.
8. The method for generating an audiobook video according to claim 7, further comprising: Converting the second audio data into performance text, where the performance text includes multiple text segments; Based on the story text, correcting at least one of the multiple text segments in terms of text or sentence segmentation to generate subtitle text. The generating of an audiobook video corresponding to the story text based on the target audio data and the multiple frames of images includes: Generating the audiobook video corresponding to the story text based on the second audio data, the multiple frames of images, and the subtitle text.
9. The audiobook video generation method according to claim 8, wherein, The correcting of at least one of the multiple text segments in terms of text or sentence segmentation based on the story text includes at least one of the following: Based on the story text, correct the homophonic typos in each of the multiple text segments; Based on the story text, merge or split adjacent text segments with incorrect sentence breaks among the multiple text segments; Based on the story text, correct at least one of the text and sentence breaks for the content that is included in both the broadcast text and the story text among the multiple text segments.
10. The method for generating an audiobook video according to claim 8, wherein, The obtaining of multiple frames of images corresponding to the story text includes: Dividing the subtitle text into multiple sub-shot texts according to the sub-shot granularity; Based on the image generation model, generating corresponding images for each of the multiple sub-shot texts.
11. The method for generating an audiobook video according to any one of claims 1 to 10 further includes: Providing corresponding audio effect resources for at least one plot in the story text; Adding the audio effect resources to the audiobook video.
12. The method for generating an audiobook video according to claim 11, wherein, The providing of corresponding audio effect resources for at least one plot in the story text includes: Extracting at least one of the weather keyword and the atmosphere signal word corresponding to the at least one plot; Based on at least one of the weather keyword and the atmosphere signal word, providing the audio effect resources for the at least one plot.
13. The method for generating an audiobook video according to claim 12, wherein, The providing of the audio effect resources for the at least one plot based on at least one of the weather keyword and the atmosphere signal word includes: In response to both the weather keyword and the atmosphere signal word being recognized, providing the audio effect resources for the at least one plot based on the weather keyword.
14. The method for generating an audiobook video according to claim 12, wherein, The providing of the audio effect resources for the at least one plot includes: Identifying multiple key plots in the story text; Providing corresponding audio effect resources for each of the multiple key plots.
15. The method for generating an audiobook video according to claim 12, wherein, The providing of the audio effect resources for the at least one plot includes: Determining the time stamps corresponding to the at least one plot; Based on the time stamps corresponding to the at least one plot, determining the start and end times of the audio effect resources matched by the at least one plot.
16. The method for generating an audiobook video according to claim 15, wherein, The determining of the start and end times of the audio effect resources matched by the at least one plot based on the time stamps corresponding to the at least one plot includes: Advancing the start time in the time stamp of the target plot in the at least one plot by a specified time as the start time of the audio effect resources matched by the target plot; Taking the end time in the time stamp of the target plot as the end time of the audio effect resources matched by the target plot.
17. An audiobook video generation device includes: An image acquisition module configured to acquire multiple frames of images corresponding to the story text, the multiple frames of images being used to represent the key plots of the story text; An audio acquisition module configured to determine target audio data in the audio data corresponding to the story text; A video generation module configured to generate an audiobook video corresponding to the story text based on the target audio data and the multiple frames of images.
18. An electronic device includes: A processor; And A memory coupled to the processor for storing instructions which, when executed by the processor, cause the processor to execute the audiobook video generation method according to any one of claims 1 to 16.
19. A computer-readable storage medium having computer instructions stored thereon, wherein, When the computer instructions are executed by a processor, the audiobook video generation method according to any one of claims 1 to 16 is implemented.
20. A computer program product, comprising: It includes computer instructions which, when executed by a processor, implement the audiobook video generation method according to any one of claims 1 to 16.
Citation Information
Cited By
Children story video generation method and system based on AI
CN120812370A
AI-based children's story video generation method and system
CN120812370B
Short article information storage method, device and equipment based on artificial intelligence
CN121144536A