Image processing method and electronic equipment

CN121909491APending Publication Date: 2026-04-21DOUYIN VISION CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
DOUYIN VISION CO LTD
Filing Date
2024-08-20
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing technologies require multiple steps, such as scriptwriting, voice-over, and special effects processing, to generate dynamic comics, resulting in low processing efficiency and making it difficult to generate dynamic comics efficiently.

Method used

By identifying the characters and dialogue in the comic, an AI model is used to generate voices with distinctive timbre, which are then combined with dynamic images to automatically synthesize an animated comic.

Benefits of technology

It enables efficient generation of dynamic comics, reduces labor costs, improves processing efficiency, and provides users with a more vivid and engaging dynamic comic experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121909491A_ABST
    Figure CN121909491A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an image processing method and electronic equipment. The method comprises the following steps: determining at least one role in the cartoon and dialogue content of the at least one role; generating a sound with at least one tone for the at least one role based on the dialogue content; generating a dynamic image by determining an image of at least one role and a corresponding background; and generating a dynamic cartoon by combining the sound with the dynamic image. Therefore, a dynamic voiced cartoon can be generated. Therefore, the generation cost of the dynamic audio cartoon can be reduced, and the production speed and yield can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Method of image processing and electronic device TECHNICAL FIELD

[0001] Embodiments of the present disclosure mainly relate to the field of computer, and more particularly, to a method of image processing and an electronic device. BACKGROUND

[0002] A comic is a two-dimensional visual static picture, which has a simple composition and generally contains texts such as dialogues and side notes. A traditional comic is generally issued in the form of serialization or a single volume. With the rise of the Internet, a comic can also be read on a terminal such as a smart phone.

[0003] The development of the Internet and the popularity of smart phones have brought great changes to the form of comics. Currently, there are animated comics using methods such as videos or animation programs (Flash). With the demand for dynamic comics, a solution for efficiently generating such comics is needed.

[0004] SUMMARY

[0005] Embodiments of the present disclosure provide a method of image processing, which generates dynamic comics based on static comics and / or novels, provides users with more abundant content, and improves user experience.

[0006] In a first aspect of the present disclosure, a method of image processing is provided, including: determining at least one character in a comic and dialogue content of the at least one character; generating a sound with at least one tone for the at least one character based on the dialogue content; generating a dynamic image by determining an image of the at least one character and a corresponding background; and generating a dynamic comic by combining the sound with the dynamic image.

[0007] In a second aspect of the present disclosure, an electronic device is provided, including: at least one processing unit; at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, which, when executed by the at least one processing unit, causes the electronic device to perform actions, the actions including: determining at least one character in a comic and dialogue content of the at least one character; generating a sound with at least one tone for the at least one character based on the dialogue content; generating a dynamic image by determining an image of the at least one character and a corresponding background; and generating a dynamic comic by combining the sound with the dynamic image.

[0008] In a third aspect of the present disclosure, an apparatus is provided, including modules or components for performing the method described in the first aspect of the present disclosure.

[0009] In a fourth aspect of the disclosure, there is provided a computer-readable storage medium having stored thereon machine executable instructions, which when executed by a device, cause the device to perform the method described according to the first aspect of the disclosure.

[0010] In a fifth aspect of the disclosure, there is provided a computer program product comprising computer executable instructions, which when executed by a processor, implement the method described according to the first aspect of the disclosure.

[0011] In a sixth aspect of the disclosure, there is provided an electronic device comprising: processing circuitry configured to perform the method described according to the first aspect of the disclosure.

[0012] The summary is provided to introduce a selection of concepts that are further described below in the of the Invention. This summary is not intended to identify key or essential features of the disclosure, nor is it intended to limit the scope of the disclosure. Other features of the disclosure will be apparent from review of the disclosure, which is set forth below. of the Invention BRIEF DESCRIPTION OF DRAWINGS

[0013] The above and other features, aspects and advantages of embodiments of the present disclosure will become more apparent from the following detailed description when taken in conjunction with the accompanying drawings. In the drawings:

[0014] FIG. 1 shows a schematic view of a comic according to some embodiments of the disclosure;

[0015] FIG. 2 shows a flowchart of a method of image processing according to some embodiments of the disclosure;

[0016] FIG. 3 shows an example processing flow according to some embodiments of the disclosure;

[0017] FIG. 4 shows a schematic view of a dynamic comic according to some embodiments of the disclosure;

[0018] FIG. 5 shows a schematic view of a large language model based process according to some embodiments of the disclosure;

[0019] FIG. 6 shows a block diagram of an example apparatus according to some embodiments of the disclosure; and

[0020] FIG. 7 shows a block diagram of an example electronic device that can be used to implement embodiments of the disclosure. DETAILED DESCRIPTION

[0021] Embodiments of the present disclosure will be described below in greater detail with reference to the accompanying drawings. While certain embodiments of the present disclosure are shown in the drawings, it is understood that the present disclosure can be embodied in various forms and should not be construed as being limited to the embodiments set forth herein, but rather, the embodiments are provided so that the present disclosure can be more thoroughly and completely understood. It is understood that the drawings of the present disclosure and the embodiments are for exemplary purposes only and are not intended to limit the scope of protection of the present disclosure.

[0022] In the description of embodiments of the present disclosure, the term "includes" and its derivatives, such as "including" should be understood in an open-ended way, i.e., "including, but not limited to". The term "based on" should be understood as "based at least in part on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The terms "first", "second", and the like can refer to different or the same objects. The term "and / or" means at least one of the two terms connected by it. For example, "A and / or B" means A, B, or A and B. Other explicit and implicit definitions can also be included below.

[0023] The term "comic" in the embodiments of the present disclosure can mean a static comic image, for example, a comic work can include multiple comics (or multiple images). The term "novel" in the embodiments of the present disclosure can mean text content corresponding to a static comic, for example, which can also be referred to as a novel text, etc. In some examples, the novel can also be referred to as a script of the comic, or the novel can be processed to obtain a script of the comic, for example, the script can represent key content associated with characters, etc. displayed in the comic in the novel, such as dialogue, emotional description of the character, action description of the character, etc. In some examples, the script can also be referred to as a dialogue content, a dialogue script, or a script, etc.

[0024] The term "dynamic comic" in the embodiments of the present disclosure can also be referred to as "sound comic", "dynamic sound comic", "streaming comic", "dynamically presented comic", "comic in video form", "comic in animation type", etc., which means a streaming content of a certain time length with sound, which can have multiple image frames and associated audio content.

[0025] Embodiments of the present disclosure relate to various models, which can be Artificial Intelligence (AI) or Machine Learning (ML) models, and can be trained in various manners such as self-supervised, semi-supervised, transfer learning, etc. It can be understood that the use of the term “model” is merely illustrative and is not intended to limit embodiments of the present disclosure, and in some scenarios, the term “model” can also be replaced by any one or a combination of the following: network model, neural network model, neural network, AI model, ML model, Deep Learning (DL) model, deep network, large language model, etc. It can be understood that embodiments of the present disclosure do not limit the network structure of the models involved, which can be, for example, a convolutional neural network, a recurrent neural network, a teacher-student network based on knowledge distillation, etc.

[0026] With the development of the Internet and the popularity of mobile terminals such as smart phones and the rise of various types of multimedia such as short videos, people also hope to watch comics in a dynamic form. For example, a comic can be adapted into a dubbing script, and then the comic and the dubbing can be synthesized and added with special effects to give the comic multimedia properties. However, this approach requires separate processing of various personnel such as screenwriters, dubbers, and special effect processors, which is inefficient. Therefore, how to efficiently and accurately generate dynamic comics is one of the problems to be solved.

[0027] In view of this, embodiments of the present disclosure provide a scheme for generating dynamic comics. In this scheme, at least one character in a comic and the dialogue content of the at least one character can be determined; based on the dialogue content, a sound with at least one tone is generated for the at least one character; by determining the image of the at least one character and the corresponding background, a dynamic image is generated; and by combining the sound with the dynamic image, a dynamic comic is generated.

[0028] FIG. 1 shows a schematic diagram of a comic 100 according to some embodiments of the present disclosure. The comic 100 can be one of the comics of a comic work, where the comic work can include multiple comics. As shown in FIG. 1, the comic 100 can include multiple different elements: a background 110, at least one character 120-1 to 120-3 (collectively referred to as characters 120), and text information 130 (e.g., the dialogue bubbles 130-1 to 130-4 shown). For the convenience of explaining the embodiments in the present disclosure, the following more detailed embodiments will be described by taking the comic 100 in FIG. 1 as an example, but it should be understood that the comic 100 in FIG. 1 is only an example for convenience of description and should not be interpreted as a limitation of embodiments of the present disclosure. In actual scenarios, a comic can include more or fewer characters and can include more or fewer text information.

[0029] FIG. 2 illustrates a flowchart of a method 200 of image processing, according to some embodiments of the present disclosure. The method illustrated in FIG. 2 can be performed by an electronic device, for example, the electronic device can be a computer, including but not limited to a desktop computer, a tablet computer, a laptop computer, a smart phone, etc. It should be understood that the image in the embodiments of the present disclosure can refer to a comic, such as the comic 100 illustrated in FIG. 1, etc.

[0030] At block 210, at least one character in the comic and a dialogue content of the at least one character are determined. At block 220, based on the dialogue content, a sound with at least one tone is generated for the at least one character. At block 230, a dynamic image is generated by determining a figure of the at least one character and a corresponding background. At block 240, a dynamic comic is generated by combining the sound with the dynamic image.

[0031] Based on the method 200, a dynamic comic can be generated based on a comic and / or a novel. In this way, a more vivid and lively comic can be provided for a user, so that the user can more conveniently and efficiently understand a summary of the comic and / or the novel to make a decision on whether to continue reading.

[0032] As described above in connection with FIG. 1, a comic can include a plurality of different elements. In some embodiments, each element in the comic can be determined by analyzing the comic. Illustratively, a comic can be elementized by utilizing text and image recognition techniques. For example, based on an Optical Character Recognition (OCR) technique and a visual recognition technique, a background picture, at least one character, and text information in the comic can be determined.

[0033] For example, based on a visual recognition technique, such as an AI character recognition model or an AI face recognition model, at least one character in the comic is determined. Illustratively, a character feature of each character can be further determined, including but not limited to, a clothing color, a hair color, a hair length, a facial expression, a mouth shape, an eye expression, a body feature, etc.

[0034] For example, based on an OCR, text information in the comic is determined. Illustratively, a region where the text information is located can be further determined by utilizing, for example, an image segmentation model.

[0035] For example, a remaining part except the at least one character and the text information can be determined as a background picture.

[0036] In some embodiments, at least one timbre can be matched to at least one character. Illustratively, a character library (e.g., also referred to as a character asset library) can be constructed according to multiple comics of a comic work. For example, based on a large model picture understanding technology, a background environment and a character appearing in each comic can be identified to form a character library. Illustratively, based on a character feature of each of the multiple characters, a corresponding timbre can be assigned. It can be understood that different characters have different timbres. For example, based on different timbres, a listener can distinguish different sources.

[0037] In some implementations of the present disclosure, there is a novel corresponding to a comic. In some embodiments, the novel corresponding to the comic can be obtained. For example, the novel corresponding to the comic can be a specific chapter in the novel. Illustratively, the novel can be parsed to obtain dialogue content and non-dialogue content. For example, the non-dialogue content can include environmental description information and / or character feature description information, etc. For the convenience of description, it can be assumed that the novel corresponding to the comic 100 in FIG. 1 is:

[0038] Table 1

[0039] In some embodiments, at least one character in the comic and dialogue content of at least one character can be determined by content parsing of the novel corresponding to the comic. For example, by parsing the example of the novel described above, it can be determined that there are three characters “Ye Chen”, “Drunkard” and “Old Lady”, and the dialogue content of each character can be determined accordingly.

[0040] In other embodiments, the novel can be further processed to obtain conversational content. Illustratively, the novel can be parsed and conversational content can be obtained according to the result of the content parsing. Illustratively, the dialogue content of each character can be determined from the conversational content.

[0041] Optionally, the conversational content can be referred to as a multi-character conversational push book script, or simply referred to as a script. For example, the push book script can highlight the character features of each character, and can represent the core scenes and / or summary information in the comic through dialogue. As an example, the push book script can weaken or remove part of the description information about the environment.

[0042] Illustratively, based on a text-to-speech (TTS) technology, for example, using a multicast speech TTS model, the novel (or the conversational content obtained based on the novel) can be converted into audio. Illustratively, the dialogue content can be matched to the character, so that the audio corresponding to the dialogue content has a timbre corresponding to the character.

[0043] For example, for at least one character involved in the novel, at least one timbre of the at least one character can be determined based on the character library and the timbre corresponding to each character in the character library. For example, the at least one timbre can be utilized to generate audio of the corresponding dialogue content.

[0044] For example, the audio converted from the novel can include multiple parts, for example, the multiple parts include: a first part corresponding to first dialogue content of a first character (using a first timbre), a second part corresponding to second dialogue content of a second character (using a second timbre), a third part corresponding to third dialogue content of a third character (using a third timbre), and so on. Optionally, the multiple parts can also include parts corresponding to environmental description information in non-dialogue content (for example, using a default setting of a voice-over timbre), and the like.

[0045] In some embodiments, a dynamic image can be generated based on a comic. Illustratively, the comic can be processed to obtain an initial image. For example, the processing of the comic can include removing text information in the comic, for example, the area where the text information is located can be removed by image cutting or the like. For example, the processing of the comic can include performing a continuous processing on the area where the text information is located, for example, the area where the text information is located can be filled as part of the pixels of the background and / or the character by referring to the pixel information around the area where the text information is located. Optionally, when performing the continuous processing, the remaining comic pictures in the comic work can also be referred to, for example, the remaining comics with similar or the same background as the comic 100.

[0046] Exemplarily, the dynamic image can be generated based on the initial image by adding dynamic effects. For example, dynamic effects can be added to specific parts of at least one character, which can include one or more of the following: head, mouth, frown, hair, arm, or leg, etc. In some examples, dynamic effects such as opening the mouth can be added to the character with the dialogue content. In some examples, dynamic effects such as flowing hair, blinking, frowning, etc. can also be added to the character. In some examples, the dynamic effects of the character can be generated based on the description of the character in the novel, for example, if the text in the novel records that a certain character has made a certain action, the corresponding dynamic effect of the character can be generated based on the text record. For example, the novel describes that the emotion of the character "Ye Chen" when speaking the dialogue content is "angry", and the dynamic effect can be added to reflect this emotion. As an example, the dynamic effect can be waving the arm, etc. In some examples, the dynamic effects can be added based on the description of the environment in the novel, for example, the novel can describe content such as rain, wind, fast driving vehicles, etc., and the corresponding dynamic effects can be added. In the embodiments of the present disclosure, the dynamic effects can be realized by at least two different images, and at least a part of the initial image can be changed based on the dynamic effects to reflect the dynamic effects on the basis of the initial image.

[0047] In some implementations of the present disclosure, there is no novel corresponding to the comic. In some embodiments, the content of the text information of the comic can be parsed, and the content of the text information can be matched with at least one character to obtain the dialogue content of the at least one character.

[0048] In combination with FIG. 1, the text content can be parsed to determine that the text 130-1 and the text 130-2 correspond to the character 120-1, the text 130-3 corresponds to the character 120-2, and the text 130-4 corresponds to the character 120-3. Exemplarily, the dialog content can also be generated by analyzing the content of each text, such as determining the logical relationship, the sequence relationship, etc. between each text. Exemplarily, the dialog content can include the dialogue content of each character.

[0049] Exemplarily, the dialog content can be converted into audio based on text-to-speech (TTS) technology, for example, using a multicast speech TTS model. Exemplarily, the dialogue content can be matched with the character, so that the audio corresponding to the dialogue content has a tone corresponding to the character.

[0050] For example, for at least one character in the comic, at least one tone of the at least one character can be determined based on the character library and the tone corresponding to each character in the character library. For example, the audio of the corresponding dialogue content can be generated using the at least one tone.

[0051] For example, the converted audio can include multiple parts, e.g., the multiple parts include: a first part corresponding to a first dialogue content of a first character (using a first timbre), a second part corresponding to a second dialogue content of a second character (using a second timbre), a third part corresponding to a third dialogue content of a third character (using a third timbre), and so on.

[0052] In some embodiments, a dynamic image can be generated based on a comic. Illustratively, the comic can be processed to obtain an initial image. For example, the processing of the comic can include removing text information in the comic, e.g., the area where the text information is located can be removed by image cutting, etc. For example, the processing of the comic can include performing a continuity processing on the area where the text information is located, e.g., the area where the text information is located can be filled as a part of the pixels of the background and / or the character by referring to the pixel information around the area where the text information is located. Optionally, when performing the continuity processing, the remaining comic pictures in the comic work can also be referred to, e.g., the remaining comics with similar or the same background as the comic 100, etc.

[0053] Illustratively, a dynamic image can be generated by adding dynamic effects based on the initial image. For example, dynamic effects can be added to specific parts of at least one character. In some examples, dynamic effects such as opening the mouth can be added to the character with dialogue content. In some examples, dynamic effects such as flowing hair, blinking, frowning, etc. can also be added to the character. In some examples, the corresponding dynamic effects can be generated based on picture understanding by analyzing the image of the comic. For example, short lines near the wheels in the comic representing rapid driving of the vehicle, short lines in the background representing rain or wind, etc. In the embodiments of the present disclosure, the addition of dynamic effects can be realized by at least two different images, at least a part of the initial image can be changed based on the dynamic effects to embody the dynamic effects on the basis of the initial image.

[0054] In some implementations of the present disclosure, a dynamic comic can be generated based on a novel. In some embodiments, the novel can be obtained, e.g., as an input for generating a dynamic comic. At least one character and dialogue content of the at least one character can be determined by performing content analysis on the novel.

[0055] Illustratively, a specific chapter in the novel can be determined. For example, a text quality model can be used to determine the specific chapter based on the novel. For example, the specific chapter can be determined based on the novel using reading data of the user, or interaction data with the user, etc. As an example, it is assumed that the specific chapter includes the content shown in Table 1 as described above. The content of the specific chapter can be further analyzed to determine at least one character and dialogue content of the at least one character.

[0056] Exemplarily, at least one role can be acquired in advance, and scene content corresponding to the at least one role can be determined from the novel. For example, the scene content can include the aforementioned content as shown in Table 1. Based on the scene content, dialogue content of the at least one role can be further determined.

[0057] For example, the at least one role can be determined in advance based on the name of the role, etc. Optionally, the novel can include a list of characters appearing in the novel, etc., and the at least one role can be determined from the list.

[0058] In some examples, at least one timbre corresponding to the at least one role can be determined based on the description of the at least one role in the novel. For example, the description of a role includes feature description information of the role, such as body shape, personality, etc. Exemplarily, the corresponding timbre can be determined based on the feature description information of the role.

[0059] Exemplarily, the novel (or the dialogue content obtained based on the novel) can be converted into audio based on TTS technology, for example, by using a multi-modal speech TTS model. Exemplarily, the dialogue content can be matched with the role, so that the audio corresponding to the dialogue content has the timbre corresponding to the role.

[0060] In some examples, the image of the at least one role can be determined based on the description of the at least one role in the novel. The corresponding background can be determined based on the description of the scene and / or the environment in the novel.

[0061] The dynamic element, such as a specific part of the role or a specific object in the background, can be determined. Then, the dynamic image can be generated based on the image of the at least one role, the corresponding background, and the dynamic element.

[0062] Optionally, the initial image can be generated based on the image of the at least one role and the corresponding background, for example, the initial image can be an initial comic image, i.e., an initial image. Then, the dynamic image can be generated by adding dynamic effects on the basis of the initial image. For brevity, the specific implementation will not be repeated here, and the reader can refer to the more detailed description described above.

[0063] Further, in embodiments of the present disclosure, the sound can be synthesized with the dynamic image to obtain a dynamic comic with sound. In some embodiments, the audio can be generated based on the sound of each character with a corresponding voice, further adding background music (BGM) and / or specific sound effects. For example, the background music can include sounds such as the sound of rain, the sound of wind, etc. For example, the specific sound effects can include sounds such as waving fists, panting, falling to the ground, etc.

[0064] In some embodiments, the audio can be matched with the dynamic image (such as a dynamic video clip) based on the time length of the audio, and a dynamic comic with sound corresponding to the time length can be generated. For example, during the time when a specific character in the audio speaks a dialogue, the mouth of the specific character can be dynamically opened and closed.

[0065] Through the exemplary embodiments of the present disclosure, a dynamic comic with sound can be generated based at least on a comic. The scheme can be implemented by using a model with the aid of AI, which is more efficient. The involvement of scriptwriters and voice actors is avoided, greatly reducing the labor cost. Moreover, the scheme can process a large number of comics, improving the efficiency and output of processing. Moreover, the scheme of the present disclosure can provide users with more vivid and lifelike dynamic comics, so that users can more conveniently and efficiently understand the summary of the comic and / or novel to make a decision on whether to continue reading.

[0066] FIG. 3 shows an example processing flow 300 according to some embodiments of the present disclosure. In the flow 300, a dynamic comic with sound 310 can be generated based on at least the comic frames 302 of a comic work 301.

[0067] At 303, it can be determined whether there is a novel corresponding to the comic frames 302. If so, at 306, dialogue scripts are generated by analyzing the novel. If it is determined at 303 that there is no novel, at 304, the comic is elementally disassembled to identify the background, characters, and dialogue bubbles; then at 307, dialogue scripts can be generated by matching the background and characters. Further, at 308, voices can be generated by matching the voice of each character.

[0068] At 305, the comic can be dynamicized based on the disassembly of the comic elements. Exemplarily, the static comic can be dynamicized based on the image-to-video capability in the video generation large model. It can be understood that although not explicitly shown in FIG. 3, the process of the comic dynamicization 305 can also further incorporate the content in the novel (if any).

[0069] On the basis of the generated voice at 308, matching BGM and / or sound effects, etc. can also be added at 309, and then synthesized with the dynamicized comic to generate a dynamic comic strip at 310. For example, the sound can be automatically spliced with the dynamicized image to obtain a complete dynamic comic strip.

[0070] It should be noted that the process flow 300 shown in FIG. 3 is only for illustratively representing a possible process of generating a dynamic comic strip, but should not be interpreted as a limitation on the embodiments of the present disclosure. As an example, the voice can be generated in combination with the dialogue bubbles, etc. in the novel and comic. As an example, the comic can be dynamicized in combination with the text in the novel. As an example, the process of generating voice and dynamicizing comic can be coupled and associated with each other, rather than independent of each other. Although FIG. 3 shows a process of generating a dynamic comic based on a comic or a comic and a novel, in other examples, a dynamic comic can be generated based only on a novel.

[0071] FIG. 4 shows a schematic diagram of a dynamic comic 400 according to some embodiments of the present disclosure. Illustratively, the dynamic comic 400 can be generated based on the comic 100 in FIG. 1. It can be understood that the dynamic comic 400 is a streaming media content with a time length. At different times, the displayed pictures and / or audio are not the same. As shown in FIG. 4, as an example, multiple pictures 401-403 at different times are presented. For example, the picture 401 can be an initial image obtained by processing the comic 100. For example, the picture 402 and the picture 403 can be associated with certain dynamic effects, such as a turn of the character 120-1, a waving of the arm, a wide-open mouth; a bared teeth of the character 120-2, a clenched fist; a wide-open mouth of the character 120-3, etc.

[0072] It can be understood that the embodiments of the present disclosure involve multiple models, such as an OCR model, a visual recognition model, a picture understanding model, an image segmentation model, a TTS model, a picture-to-video model, etc., but the embodiments of the present disclosure do not limit the implementation of each model, for example, part or all of the models can be integrated into a large model. FIG. 5 shows a schematic diagram of a process 500 based on a large language model according to some embodiments of the present disclosure. As shown in FIG. 5, each model involved in the dynamic comic strip generation process can be integrated into a large language model 550. Illustratively, the input 501 of the large language model 550 can include a comic, or include a novel, or include a comic and a corresponding novel. Illustratively, the output of the large language model 550 includes a dynamic comic strip 502.

[0073] It can be understood that the above embodiments are mainly described with respect to a comic. For a comic work with multiple comics, similar processing can be performed based on each comic to obtain a dynamic sound comic work. It should be understood that the time length of the generated dynamic sound comic can be equal or unequal for different comics in the comic work, such as comic one and comic two in the comic work.

[0074] Through the above-described embodiments in combination with FIGS. 1-5, the present disclosure provides a scheme for generating a dynamic sound comic. The scheme utilizes an AI model, such as the large language model 550 shown in FIG. 5, which can fully understand the content of a comic by analyzing a novel or parsing a comic; and can generate a corresponding dynamic sound comic from a static comic picture based on TTS and video generation technology, etc. In this way, the generation cost of the dynamic sound comic can be reduced, and the production speed and yield can be improved.

[0075] It should be understood that in the embodiments of the present disclosure, "first", "second", "third", etc. are only to indicate that the plurality of objects can be different, but at the same time do not exclude that the two objects are the same, and should not be interpreted as any limitation on the embodiments of the present disclosure.

[0076] It should also be understood that the manners, cases, categories and divisions of embodiments in the embodiments of the present disclosure are only for the convenience of description and should not constitute a special limitation. The features in various manners, categories, cases and embodiments can be combined with each other as long as they are logically consistent.

[0077] It should also be understood that the above content is only to help those skilled in the art better understand the embodiments of the present disclosure, and is not intended to limit the scope of the embodiments of the present disclosure. Those skilled in the art can make various modifications or changes or combinations, etc. according to the above content. The schemes after such modifications, changes or combinations are also within the scope of the embodiments of the present disclosure.

[0078] It should also be understood that the above description focuses on the differences between the various embodiments, and the same or similar parts can be referred to or learned from each other. For brevity, they will not be repeated here.

[0079] FIG. 6 shows a schematic block diagram of an example apparatus 600 according to some embodiments of the present disclosure. The apparatus 600 can be implemented by software, hardware, or a combination of both. In some embodiments, the apparatus 600 can be implemented as an electronic device. In embodiments of the present disclosure, the electronic device can be a desktop computer, a tablet computer, etc., which is not limited by the present disclosure. As shown in FIG. 6, the apparatus 600 includes a dialogue content determination module 610, a sound generation module 620, a dynamic image generation module 630, and a dynamic comic generation module 640.

[0080] The dialogue content determination module 610 can be configured to determine at least one character in the comic and dialogue content of the at least one character. The sound generation module 620 can be configured to generate a sound with at least one tone for the at least one character based on the dialogue content. The dynamic image generation module 630 can be configured to generate a dynamic image by determining an image of the at least one character and a corresponding background. The dynamic comic generation module 640 can be configured to generate a dynamic comic by combining the sound with the dynamic image.

[0081] In some embodiments, the dialogue content determination module 610 can be configured to obtain a novel; and determine at least one character in the novel and dialogue content of the at least one character by content parsing of the novel.

[0082] In some embodiments, the dialogue content determination module 610 can be configured to determine a specific chapter in the novel by content parsing of the novel; and determine at least one character in the specific chapter and dialogue content of the at least one character based on the specific chapter.

[0083] In some embodiments, the dialogue content determination module 610 can be configured to obtain at least one predetermined character; determine scene content corresponding to the at least one character from the novel; and determine dialogue content of the at least one character based on the scene content corresponding to the at least one character. Optionally, the at least one character can be predetermined based on a comic or based on a character name in a novel.

[0084] In some embodiments, the dialogue content determination module 610 can be configured to generate dialogue content based on a content parsing result of the novel; and determine the dialogue content of the at least one character in the dialogue content.

[0085] In some embodiments, the apparatus 600 can further include an identification module configured to determine a background picture, at least one character, and text information in the comic using an OCR technique and a visual recognition technique.

[0086] Exemplarily, the dialogue content determination module 610 can be configured to match content in the text information with the at least one character to obtain dialogue content of the at least one character.

[0087] In some embodiments, the sound generation module 620 can be configured to determine at least one tone corresponding to the at least one character; and convert the dialogue content into a sound with the at least one tone using a TTS technique.

[0088] In some embodiments, the sound generation module 620 can be configured to determine at least one tone corresponding to the at least one character based on a description of the at least one character in the novel.

[0089] In some embodiments, the dynamic image generation module 630 can be configured to generate the dynamic image based on the comic by at least one of the following operations: removing text information in the comic, performing a continuous processing on an area where the text information is located, or adding a dynamic effect for a specific part of the at least one character.

[0090] Illustratively, the apparatus 600 further includes a dynamic effect determination module configured to: obtain a novel corresponding to the comic; determine characteristic description information of the at least one character by performing content analysis on the novel; and determine a specific part of the at least one character and a corresponding dynamic effect based on the characteristic description information of the at least one character. Optionally, the specific part includes at least one of the following: a head, a mouth, a frown, hair, an arm, or a leg.

[0091] Illustratively, the apparatus 600 further includes a tone determination module configured to, in response to the comic including a plurality of characters, determine a tone for each character in the plurality of characters, wherein different characters have different tones.

[0092] In some embodiments, the dynamic comic generation module 640 can be configured to match the dynamic effect in the dynamic image with a sound; add a background sound effect and / or a sound effect associated with a part of the dynamic effect; and generate a dynamic comic by synthesis.

[0093] The apparatus 600 of FIG. 6 can be used to implement the method 200 described above in connection with FIG. 2 or the flow 300 described above in connection with FIG. 3, and thus, for brevity, will not be described again.

[0094] The division of modules or units in the embodiments of the present disclosure is illustrative, and is merely a logical function division. In actual implementation, another division manner can be used, and each functional unit in the disclosed embodiments can be integrated into one unit, or can be physically separated, or two or more units can be integrated into one unit. The integrated unit can be implemented in the form of hardware, or in the form of a software functional unit.

[0095] FIG. 7 shows a block diagram of an example electronic device 700 that can be used to implement embodiments of the present disclosure. It should be understood that the device 700 shown in FIG. 7 is merely an example and should not be construed as limiting the functionality and scope of the implementations described herein. For example, the electronic device 700 can include the apparatus 600. For example, the device 700 can be used to perform the method 200 or the flow 300 described above.

[0096] As shown in FIG. 7, device 700 is in the form of a general-purpose computing device. Components of the computing device 700 can include, but are not limited to, one or more processors or processing units 710, a memory 720, a storage device 730, one or more communication units 740, one or more input devices 750, and one or more output devices 760. The processing unit 710 can be a real or virtual processor and capable of executing a program stored in the memory 720 to perform various processes. In a multi-processing system, multiple processing units execute computer-executable instructions in parallel to improve the processing power of the computing device 700.

[0097] The computing device 700 typically includes a plurality of computer storage media. Such media can be volatile, nonvolatile, removable, and / or non-removable media implemented in any method or technology for storage of information such as computer readable instructions, data structures, program modules, and other data. The memory 720 can be volatile memory (such as registers, cache, random access memory (RAM)), non-volatile memory (such as read only memory (ROM), electrically erasable programmable read only memory (EEPROM), flash memory), or some combination thereof. The storage device 730 can be a removable or non-removable media and can include machine readable media such as flash drives, disks, or any other media capable of storing information and / or data (e.g., training data for training) and accessible by the computing device 700.

[0098] The computing device 700 can further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 7, a disk drive and a disk drive interface can be provided for reading from or writing to a removable, non- volatile magnetic disk (e.g., a "floppy disk"), and an optical disk drive and disk interface can be provided for reading from or writing to a removable, non-volatile optical disk (such as a CD-ROM). In such instances, each drive can be connected to the bus by one or more data media interfaces. The memory 720 can include a computer program product 725 having one or more program modules configured to carry out the various methods or actions of the various implementations of the present disclosure.

[0099] The communication units 740 enable communications with other computing devices over a communication medium. Additionally, the functionality of the components of the computing device 700 can be implemented in a single computing cluster or a plurality of computer machines that are capable of communicating over a communication connection. Thus, the computing device 700 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network nodes in the networking environment.

[0100] The input device 750 can be one or more input devices, such as a mouse, a keyboard, a trackball, etc. The output device 760 can be one or more output devices, such as a display, a speaker, a printer, etc. The computing device 700 can also communicate with one or more external devices (not shown) such as a storage device, a display device, etc. through the communication unit 740, as needed, one or more devices that enable a user to interact with the computing device 700, or any devices (e.g., a network card, a modem, etc.) that enable the computing device 700 to communicate with one or more other computing devices. Such communication can be carried out through an Input / Output (I / O) interface (not shown).

[0101] According to an example implementation of the present disclosure, there is provided a computer-readable storage medium, such as a non-transitory computer-readable storage medium, having stored thereon computer-executable instructions, wherein the computer-executable instructions are executed by a processor to implement the method described above. According to an example implementation of the present disclosure, there is also provided a computer program product tangibly stored on a non-transitory computer-readable medium and comprising computer-executable instructions, wherein the computer-executable instructions are executed by a processor to implement the method described above. According to an example implementation of the present disclosure, there is provided a computer program product having stored thereon a computer program which, when executed by a processor, implements the method described above.

[0102] Various aspects of the disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, and computer program products according to this disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0103] The computer readable program instructions can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions can also be stored in a computer readable storage medium that can include a non-transitory computer readable storage medium that can be a computer- readable storage medium having no data, programs, program modules, e.g., instructions for operation, or digital content stored thereon or therein for a short time or not at all. The computer readable storage medium can also have other meanings inhered thereby, which can include a computer- readable storage medium encoding computation- directed instructions; i.e., installable program modules / objects; i.e., where said medium is a manufacture (Mfct). The instructions

[0104] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0105] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0106] implementations of the present disclosure have been described above, the description is illustrative only and not restrictive ones, and is not exhaustive nor limited to the disclosed implementations. Many modifications and changes can occur to one of ordinary skill in the art to which the disclosure pertains without departing from the scope and spirit of the described implementations. The choice of words in this document is intended to best describe the principles of the implementations, practical application, or improvement to the technology in the market, or to enable other ordinary skilled persons in the art to understand the various implementations disclosed herein.

Claims

1.A method of image processing, comprising: determining at least one character in a comic and a dialogue content of the at least one character; generating a sound with at least one tone for the at least one character based on the dialogue content; generating a dynamic image by determining a figure of the at least one character and a corresponding background; and generating a dynamic comic by combining the sound with the dynamic image. 2.The method of claim 1, wherein determining at least one character in a comic and a dialogue content of the at least one character comprises: obtaining a novel; and determining the at least one character in the novel and the dialogue content of the at least one character by content parsing the novel. 3.The method of claim 2, wherein determining the at least one character in the novel and the dialogue content of the at least one character by content parsing the novel comprises: determining a specific chapter in the novel by content parsing the novel; and determining the at least one character and the dialogue content of the at least one character based on the specific chapter. 4.The method of claim 2, wherein determining the at least one character in the novel and the dialogue content of the at least one character by content parsing the novel comprises: obtaining a predetermined at least one character; determining a scene content corresponding to the at least one character from the novel; and determining the dialogue content of the at least one character based on the scene content corresponding to the at least one character. 5.The method of claim 2, further comprising: generating a dialogue content based on a result of content parsing the novel; and determining the dialogue content of the at least one character in the dialogue content. 6.The method of claim 1, further comprising: obtaining an input comic; and determining a background picture, the at least one character, and text information in the comic by using an optical character recognition (OCR) technology and a visual recognition technology. 7.The method of claim 6, wherein determining at least one character in a comic and a dialogue content of the at least one character comprises: matching content in the text information with the at least one character to obtain the dialogue content of the at least one character. 8.The method of claim 1, wherein generating a sound with at least one tone for the at least one character based on the dialogue content comprises: determining at least one tone corresponding to the at least one character; and converting the dialogue content into a sound with the at least one tone by using a text-to-speech (TTS) technology. 9.The method of claim 8, wherein determining at least one tone corresponding to the at least one character comprises: determining the at least one tone corresponding to the at least one character based on a description of the at least one character in a novel. 10.The method of claim 6, wherein generating a dynamic image comprises: generating the dynamic image based on the comic by at least one of: ​ ​ ​ ​ ​ ​ ​ remove text information in the comic, continuously process a region where the text information is located, or add a dynamic effect for a specific part of the at least one character. 11.The method of claim 10, further comprising: obtaining a novel corresponding to the comic; determining characteristic description information of the at least one character by content analysis on the novel; and determining the specific part of the at least one character and a corresponding dynamic effect based on the characteristic description information of the at least one character. 12.The method of claim 10, wherein the specific part comprises at least one of a head, a mouth, a frown, hair, an arm, or a leg. 13.The method of claim 1, further comprising: in response to the comic including a plurality of characters, determining a voice tone for each of the plurality of characters, wherein different characters have different voice tones. 14.The method of claim 1, wherein generating a dynamic comic by combining the sound with the animated image comprises: matching a dynamic effect in the animated image with the sound; adding background sound effects and / or sound effects associated with partial dynamic effects; and generating the dynamic comic by composition. 15.An electronic device comprising: at least one processor; and at least one memory having computer-executable instructions stored thereon that, when executed by the at least one processor, cause the electronic device to implement the method of any one of claims 1-14. 16.A non-transitory computer-readable storage medium having computer-executable instructions stored thereon that, when executed by a processor, implement the method of any one of claims 1-14. 17.A computer program product having computer-executable instructions embodied thereon that, when executed, implement the method of any one of claims 1-14. ​ ​ ​