Video processing method and device, electronic equipment, storage medium and program product
By generating videos that match the character attributes within e-books, the problem of insufficient vividness of character images for users is solved, achieving efficient video generation and enhancing the reading experience.
Patent Information
- Application Number
- CN202511783615.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-28
- Publication Date
- 2026-03-03
AI Technical Summary
Users find the visual experience of characters in e-books lacking in vividness, and existing technologies struggle to automatically generate videos that match the characters' appearances, resulting in an unsatisfactory reading experience.
By acquiring reference images and target action templates of characters from the story text, videos that match the character attributes are generated. The model is then used for semantic understanding and action template matching to achieve automatic generation of character videos.
It improves the efficiency of character video generation and the matching with character images, thereby enhancing the user's reading and viewing experience.
Smart Images

Figure CN121603740A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, and in particular to a video processing method, apparatus, electronic device, storage medium, and program product. Background Technology
[0002] With the development of internet technology, users' reading habits have changed significantly. Today, users can read e-books anytime, anywhere via the internet.
[0003] Many e-books are presented in text format, and even if there are a few illustrations, they are manually created by the author or some users. Summary of the Invention
[0004] According to some embodiments of this disclosure, a video processing method is provided, comprising: acquiring a video generated for a character in a story text, wherein the video is generated based on a reference image of the character and a target action template, the target action template being matched with attribute information of the character, and the reference image of the character being generated based on relevant text of the character in the story text; and displaying the video.
[0005] According to some other embodiments of this disclosure, a video processing apparatus is provided, comprising: an acquisition module configured to acquire a video generated for a character in a story text, wherein the video is generated based on a reference image of the character and a target action template, the target action template being matched with attribute information of the character, and the reference image of the character being generated based on relevant text of the character in the story text; and a display module configured to display the video.
[0006] According to further embodiments of the present disclosure, an electronic device is provided, including: a processor; and a memory coupled to the processor for storing instructions, which, when executed by the processor, cause the processor to perform a video processing method according to any embodiment of the present disclosure.
[0007] According to further embodiments of the present disclosure, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the video processing method of any embodiment of the present disclosure.
[0008] According to further embodiments of the present disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the video processing method of any embodiment of the present disclosure.
[0009] Other features, aspects, and advantages of this disclosure will become clear from the following detailed description of exemplary embodiments with reference to the accompanying drawings. Attached Figure Description
[0010] Embodiments of this disclosure are described below with reference to the accompanying drawings. It should be understood that the drawings described below are merely illustrative of some embodiments of this disclosure and are not intended to limit the scope of this disclosure. In the drawings:
[0011] Figure 1 A flowchart illustrating some embodiments of the video processing method of this disclosure is shown;
[0012] Figures 2-5 A schematic diagram illustrating the display interface of some embodiments of this disclosure;
[0013] Figure 6 This invention provides a schematic diagram of the structure of a video processing apparatus according to some embodiments of the present disclosure.
[0014] Figure 7 This invention discloses schematic diagrams of the structure of electronic devices according to some embodiments of the present disclosure;
[0015] Figure 8 A schematic diagram of the structure of an electronic device according to other embodiments of the present disclosure is shown. Detailed Implementation
[0016] The technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. It should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein.
[0017] It should be understood that the various steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect. Unless otherwise specifically stated, the relative arrangement of steps set forth in these embodiments should be interpreted as merely exemplary and does not limit the scope of this disclosure.
[0018] As used in this disclosure, the term "comprising" and its variations are open-ended terms that include at least the following elements / features but do not exclude other elements / features, i.e., "including but not limited to". The term "based on" means "at least partially based on".
[0019] It should be noted that the concepts of "first," "second," etc., used in this disclosure are used only to distinguish different devices, modules, or units, and are not intended to define the order of functions performed by these devices, modules, or units or their interdependencies. Unless otherwise specified, the concepts of "first," "second," etc., are not intended to imply that the objects described herein must be in a given temporal, spatial, rank, or any other given order.
[0020] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0021] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0022] The embodiments of this disclosure are described in detail below with reference to the accompanying drawings; however, this disclosure is not limited to these specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. Furthermore, in one or more embodiments, specific features, structures, or characteristics can be combined in any suitable manner that will be apparent to those skilled in the art from this disclosure.
[0023] With the widespread dissemination of e-books, many characters in the stories have become beloved by users. Users are no longer satisfied with the textual descriptions of the characters or the images of the characters in their minds; they hope to see more vivid and concrete character images.
[0024] For the reasons mentioned above, this disclosure proposes a video processing method that can acquire and display videos generated from characters in a story text. The video is generated based on a reference image of the character and a target action template. The reference image is generated based on relevant text about the character in the story text. The reference image matches the character's appearance described in the story text, and the video generated based on the reference image better maintains the consistency of the character's appearance. The target action template matches the character's attribute information; different target action templates can be used for characters with different attributes, improving the richness of video types and their adaptability to character attributes. This video processing method allows users to transform the descriptions of characters in a story text into more vivid and concrete visual images without requiring manual video creation, thus improving video generation efficiency and enhancing the user's reading and viewing experience.
[0025] The following is combined Figures 1-5 This disclosure describes a video processing method. The video processing method of this disclosure can be executed by a video processing apparatus or an electronic device. The video processing apparatus can be implemented by software and / or hardware. For example, the video processing apparatus can be an e-book reading platform, an e-book reading application, etc., and is not limited to the examples given.
[0026] Figure 1 Flowcharts illustrating some embodiments of the video processing method disclosed herein. Figure 1 As shown, the video processing method of this embodiment includes steps S102 to S104.
[0027] In step S102, the video generated for the characters in the story text is obtained.
[0028] The story text is the story text from an e-book. For example, the video is generated based on a character's reference image and a target action template, which is matched with the character's attribute information. The character's reference image is generated based on the character's relevant text in the story text.
[0029] For example, a character's attribute information includes character setting information, such as the character's physical characteristics, gender, age, occupation, etc., and is not limited to the examples given. The character's related text can include descriptive text about the character's attributes or descriptive text about the character's appearance. The character's reference image can be generated in advance based on the character's related text, or it can be generated in real-time during the video generation process.
[0030] Multiple action templates can be configured, including a target action template. Different action templates can be used to make the character perform different actions. Action templates are not limited to being generated based on the character's action descriptions in the story text. In many cases, users prefer that the character can break free from the limitations of the scene and actions in the story text, generating different action videos, increasing the richness of the generated videos, and making the character more vivid, natural, and an independent individual.
[0031] For example, each action template includes configured action parameters, facial expression parameters, appearance parameters, scene parameters, etc., and is not limited to the examples given. The target action template matches the character's attribute information. For example, if the character is male, determining the target action template requires selecting a template of male actions. The selection of the target action template is not only based on one dimension of attribute information, but can be based on multiple dimensions of attribute information, making the generated video smoother and more in line with the character.
[0032] For example, a model can be used to generate a reference image of a character based on relevant text, and then generate a video of the character based on the reference image and a target motion template.
[0033] In step S104, the video is displayed.
[0034] For example, videos can be displayed in the story text reading interface, in the upper-level visualization component of the reading interface, in the video interface after jumping from the reading interface, or in the e-book recommendation interface. Videos can be displayed in various distribution scenarios and in various interfaces, not limited to the examples given.
[0035] The video processing method described in the above embodiments can acquire and display videos generated for characters in a story text. The videos are generated based on reference images of the characters and target action templates. The reference images are generated based on relevant text about the characters in the story text. The reference images match the characters' appearances described in the story text, and videos generated based on these reference images better maintain consistency in the characters' appearances. The target action templates match the characters' attribute information; different target action templates can be used for characters with different attributes, increasing the richness of video types and their adaptability to the characters' attributes. This video processing method allows users to transform descriptions of characters in story texts into more vivid and concrete visual images without requiring manual video creation, thus improving video generation efficiency and enhancing the user's reading and viewing experience.
[0036] The following describes how to obtain videos generated for characters in the story text.
[0037] In some embodiments, acquiring a video generated for a character in the story text includes: in response to a user selecting relevant text for a character and triggering a video generation function, acquiring a video, wherein the character's attribute information is determined based on an understanding of the story text.
[0038] Users can select relevant text about a character while reading the story text of an e-book and trigger a video generation function to generate a video for that character. The system uses a model to perform semantic understanding on the selected character's relevant text to identify the character. It then uses the model to perform semantic understanding on the story text to determine the character's attribute information. If the relevant text does not contain the attribute information needed to generate a reference image of the character, the system combines this attribute information to generate a reference image of the character. Based on the character's attribute information, a target action template is determined. Finally, based on the reference image and the target action template, a video of the character is generated.
[0039] like Figure 2 As shown, the story text is displayed in the reading interface 200. For example, a user can long-press anywhere within a text paragraph, bring up a selection handle, drag to adjust the selection range, and release to select the relevant text for a character. Alternatively, a user can select relevant text for a character by double-clicking, triple-clicking, etc. Another example is that a selection control can be displayed next to each text paragraph, allowing the user to select that paragraph as the relevant text for a character. The methods users use to select relevant text for a character are not limited to the examples given. Typically, users will select relevant text that includes the character's visual characteristics to generate a video of the character, rather than selecting text unrelated to the character.
[0040] like Figure 2As shown, in response to the user selecting relevant text for a character, the relevant text can be displayed as selected (e.g., highlighted, backgrounded, etc.), and one or more controls, including video generation control 202, can be displayed. In response to the user triggering video generation control 202, a video of the character is generated and displayed.
[0041] The method described in the above embodiments allows users to freely and flexibly select relevant text for any character to generate videos, improving the convenience and flexibility of user operation and meeting users' needs for generating videos for any character, thus enhancing the user experience.
[0042] In some embodiments, in response to a user selecting relevant text about a character and triggering a video generation function, obtaining a video includes: in response to a user selecting relevant text about a character and triggering a video generation function, displaying one or more action templates, wherein the one or more action templates include a target action template; and in response to a user selecting a target action template, displaying a video.
[0043] Users can choose a target action template from one or more action templates, or they can choose other templates. If other templates are selected, a video is generated based on the character's reference image and other templates, and then displayed. This can meet different user needs and generate more flexible and richer videos.
[0044] For example, in response to user triggers such as Figure 2 The video generation control 202 shown can display, as follows: Figure 3 The interface shown. (As shown) Figure 3 As shown, the video editing interface 300 can display multiple action template options 301, 302, and 303. For each action template, the options can display its corresponding name (identifier), sample video, etc. In response to the user selecting option 301 of the target action template, option 301 can be displayed as selected (e.g., displaying a checkmark or other identifier). The video editing interface 300 can also display a confirmation control 304. In response to the user triggering the confirmation control 304, a video is generated based on the target action template and the character's related text.
[0045] When displaying multiple action templates, the action template that matches the character's attribute information can be recommended as the type. For example, the target action template can be displayed in a specified position (first position, single line, etc.) or the target action template can be displayed as the default selected state.
[0046] The method described in the above embodiments provides one or more action templates for users to choose from, allowing users to more flexibly and freely select the videos they want to generate. This improves the flexibility and convenience of video generation, enabling users to obtain videos that better meet their needs and enhancing the user experience.
[0047] In some embodiments, in response to a user selecting relevant text for a character and triggering a video generation function, one or more action templates and a content editing area are displayed. The one or more action templates include a target action template, and the content editing area displays relevant text for the character or prompts generated based on the relevant text. In response to a user modifying the relevant text or prompts for the character, the modified relevant text or prompts for the character are displayed. In response to a user selecting a target action template, a video generated based on the relevant text or prompts for the character is obtained. The video is generated based on the relevant text or prompts for the character and the target action template.
[0048] like Figure 3 The video editing interface 300 may also include a content editing area 305, in which relevant text about the character or prompts generated based on the relevant text can be displayed. In response to the user's modification of the relevant text or prompts about the character, the modified relevant text or prompts about the character are displayed in the content editing area 305. In response to the user selecting option 301 of the target action template and triggering the confirmation control 304, a video is generated based on the target action template and the modified relevant text or prompts about the character.
[0049] The method described in the above embodiments allows users to modify the relevant text or prompts for the character, making the generated video more in line with the user's needs, improving the flexibility of video generation and the convenience of user operation, and enhancing the user experience.
[0050] In some embodiments, in response to a user selecting relevant text about a character and triggering a video generation function, acquiring a video includes: in response to a user selecting relevant text about a character and triggering an image generation function, displaying a reference image of the character; in response to a user triggering a video generation function, displaying one or more action templates, wherein the one or more action templates include a target action template; and in response to a user selecting a target action template, acquiring a video.
[0051] Users can choose to first generate a reference image of the character. If satisfied with the reference image, they can then generate a video. For example, in response to the user triggering the image generation function, multiple candidate images of the character are displayed; in response to the user selecting a candidate image as the reference image of the character and triggering the video generation function, one or more action templates are displayed; in response to the user selecting a target action template, the video is displayed.
[0052] Multiple candidate images can be generated based on the relevant text of the character selected by the user, from which the user can choose a reference image for the character used to generate the video. Figure 2 The medium video generation control 202 can be replaced with an image generation control. In response to user triggering of the image generation control, the following will be displayed: Figure 4The interface shown is as follows. Figure 4 As shown, the generated candidate images can be displayed in the image display area 401 within the image editing interface 400. Multiple candidate images can be generated; for example, thumbnails 408, 409, and 410 of multiple candidate images can be displayed. Users can trigger the display of candidate images in the image display area 401 by activating the thumbnails, or switch between displayed candidate images through specified operations (such as swiping). The currently displayed candidate image can serve as a reference image for the character selected by the user.
[0053] like Figure 4 As shown, the image editing interface 400 displays multiple action template options 402, 403, and 404. In response to the user selecting option 402 of the target action template and triggering the confirmation control 406, a video is generated based on the character's reference image and the target action template. The image editing interface can also display a modification control 405. In response to the user triggering the modification control 405, an image modification interface can be displayed. This interface can include a content editing area, which can display relevant text about the character or prompts generated based on that text. In response to the user modifying the relevant text or prompts, the modified text or prompts are displayed. In response to the user triggering the generation control, the image is regenerated based on the modified text or prompts.
[0054] like Figure 4 As shown, the image editing interface 400 can also display a regeneration control 407, which, in response to the user triggering the regeneration control 407, regenerates one or more candidate images based on the relevant text of the character.
[0055] The method described in the above embodiment first generates images corresponding to the relevant text of the character. The user can initially confirm the images to see if they meet the requirements. If they do, video generation continues. The user can also adjust the images to make the generated video even more suitable. Furthermore, using user-confirmed images to continue video generation better maintains the consistency of elements such as characters and scenes in the images, further improving the quality of the generated video and producing a video that better meets the requirements.
[0056] In some embodiments, obtaining a video generated for a character in a story text includes: in response to a user selecting a character in the story text, displaying one or more generated candidate images corresponding to that character; and in response to a user selecting a reference image of the character from one or more candidate images and triggering a video generation function, obtaining a video of the character.
[0057] One or more candidate images can be pre-generated, generated by the user performing actions such as selecting character-related text, automatically generated using a model, or generated by other users and then shared and authorized for use. For example, information about multiple candidate characters in the story text can be displayed at a specified location within the story text. This information may include at least one of the following: name and description, which can be determined based on a semantic understanding of the story text. In response to the user selecting a character from the multiple candidate characters, one or more generated candidate images corresponding to that character are displayed. For example, a viewing control can be displayed associated with the text paragraph in the story text where the character appears. In response to the user triggering the viewing control, one or more generated candidate images corresponding to the character are displayed.
[0058] Users can use candidate images of characters as reference images to generate videos, which improves the ease of operation for users, increases the utilization rate of created images, and allows users to select candidate images as reference images for characters, making the generated videos more in line with user needs and improving the user experience.
[0059] In the above embodiments, videos can be generated from texts of characters selected by the user, or videos can be automatically generated for characters in the story text.
[0060] In some embodiments, obtaining a video generated for a character in a story text includes: performing semantic understanding on the story text to determine the character, related text of the character, and attribute information of the character, wherein the related text of the character includes a description of the character's appearance; generating a reference image of the character based on the related text of the character; determining a target action template based on the attribute information of the character; and generating a video based on the reference image of the character and the target action template.
[0061] For example, semantic understanding can be performed using model story text to identify multiple characters, determine relevant text and attribute information for each character, and then generate a reference image for each character; based on the attribute information of each character, determine the target action template corresponding to each character; and based on the reference image and target action template of each character, generate a video for each character.
[0062] The method described above automatically performs semantic understanding of the story text, identifies characters, and generates character videos, improving the efficiency of character video generation. These videos can be distributed in various scenarios, such as in e-book recommendation interfaces, e-book comment areas, and around character-related text paragraphs within e-books. The generated character videos can meet the needs of different users, allowing them to select videos of interest and enhancing their reading experience.
[0063] In some embodiments, one or more characters are selected based on the interaction information corresponding to multiple characters in the story text, reference images of one or more characters are generated based on the relevant text of one or more characters and the attribute information of the characters, one or more target action templates are determined based on the attribute information of one or more characters, and videos of one or more characters are generated based on the reference images of one or more characters and the one or more target action templates.
[0064] For example, the interactive information for each character includes one or more of the following: character comments, feedback information, and related text tagging information. Comment information may include the number of comments and positive reviews; feedback information may include the number of likes and ratings; and tagging information may include the number of times a tag (such as an underline) has been added. Characters with more interactive information are those that users pay more attention to, and videos can be generated for these characters to meet the needs of more users.
[0065] The following describes the method for determining the target action template.
[0066] In some embodiments, the target action template is determined based on the type of the character, which is determined based on semantic understanding of the character's attribute information.
[0067] For example, a model can be used to perform semantic understanding of a character's attribute information to determine the character's type. Multiple types can be pre-defined, and the character type can be determined from among these. For instance, the model can include a first module, which can be trained using training samples. These training samples include attribute information of multiple sample characters and the type of attribute information annotation for each sample character. The attribute information of the multiple sample characters is input into the first module to obtain the predicted type of each of the multiple sample characters' attribute information. Based on the predicted type and the annotation type of each sample character's attribute information, a loss function is determined, and the parameters of the first module are adjusted according to the loss function until training is complete. As another example, the first module can employ a large language model. First prompt information and the character's attribute information are input into the first module to determine the character's type. The first prompt information can instruct the first module to determine the character's type. The first prompt information can also include multiple types, examples of determining the character's type, etc.
[0068] It allows configuring the correspondence between multiple character types and multiple candidate action templates, and determining the target action template from the multiple candidate action templates based on the correspondence and the determined type.
[0069] The method described in the above embodiments determines the type of a character based on semantic understanding of the character's attribute information, and determines the target action template based on the type of the character, thereby making the generated video more compatible with the character, better representing the character's attributes, and improving the user's reading and viewing experience.
[0070] In some embodiments, the target action template is selected from multiple candidate action templates, which correspond to multiple types. The multiple candidate action templates are generated in the following manner: semantic understanding is performed on multiple story texts to determine the types of multiple characters in the multiple story texts and the action description information corresponding to the multiple characters; based on the types of multiple characters, the action description information corresponding to the multiple characters is summarized to determine the reference action description information corresponding to each type; and based on the reference action description information corresponding to each type, a candidate action template for each type is generated.
[0071] Based on the interaction information of multiple story texts, multiple target story texts can be selected from multiple story texts. Based on the semantic understanding of multiple target story texts, the types of multiple characters in multiple target story texts and the corresponding action description information of multiple characters can be determined. Based on the types of multiple characters, the corresponding action description information of multiple characters can be summarized to determine the reference action description information corresponding to each type. Based on the reference action description information corresponding to each type, candidate action templates for each type can be generated.
[0072] For example, the interactive information for each story text includes one or more of the following: reading information, comment information, feedback information, and tagging information. Reading information may include the number of reads; comment information may include the number of comments and positive reviews; feedback information may include the number of likes and ratings; and tagging information may include the number of tags (underlining, etc.). Story texts with a high amount of interactive information can be selected as target story texts, and multiple candidate action templates can be generated using these target story texts.
[0073] For example, the action descriptions of multiple characters of the same type may contain identical or similar action descriptions. These action descriptions can be summarized to obtain reference action descriptions for that type. For instance, for the type "career woman," reference action descriptions might include "sitting upright, quickly signing a document with a pen in her right hand, and then closing it." Each type can correspond to one or more reference action descriptions, generating one or more candidate action templates.
[0074] The method described in the above embodiments can automatically generate candidate action templates for multiple types, which are then used to generate videos of the characters. The candidate action templates for each type can reflect the classic and iconic actions of characters of that type, resulting in videos that are more closely matched to the characters and improve the user experience.
[0075] In some embodiments, in response to the relevant text of a character, including descriptive text of at least one of a scene or plot, the target action template is determined based on at least one of the scene or plot and the type of the character, wherein at least one of the scene or plot is determined based on semantic understanding of the relevant text and the context of the relevant text.
[0076] Semantic understanding can be performed on the character's relevant text and its context to determine whether it includes descriptive information about the scene or plot. In some cases, the relevant text of the character selected by the user only includes visual descriptions; in other cases, it includes both visual descriptions and descriptions of at least one aspect of the scene or plot. The selected target action template not only matches the character type but also matches at least one aspect of the scene or plot.
[0077] For example, if the relevant text description of a character is set in an office, and the character type is a career woman, this type corresponds to multiple candidate action templates, including candidate action templates for office scenes, candidate action templates for banquet scenes, etc., then the candidate action template is selected as the target action template based on the office scene.
[0078] The method described above determines the target action template not only based on the type of the character, but also based on at least one of the scenes and plots in which the character is located, so that the generated video is more closely matched with the character and the character's related text, thereby improving the user's viewing experience.
[0079] In some embodiments, the target action template is determined from one or more candidate action templates based on at least one of the scene and plot, and an understanding of one or more candidate action templates. The one or more candidate action templates are determined from multiple candidate action templates based on the type of the character and the type corresponding to the multiple candidate action templates.
[0080] For example, based on the character type and the types corresponding to multiple candidate action templates, one or more candidate action templates corresponding to the character type are selected. The model is then used to identify at least one of the scenes and plots within these candidate action templates. From these candidate action templates, the one that matches at least one of the scenes and plots in the relevant text is selected as the target action template. If no candidate action template matches at least one of the scenes and plots in the relevant text, the target action template can be selected based on the interaction information of one or more candidate action templates. For example, the interaction information for each candidate action template includes at least one of the following: usage count, number of comments, number of positive reviews, number of likes, and rating.
[0081] The method described above can select a target action template that is more compatible with the type of the character and the related text of the character, thereby improving the quality of the generated video and enhancing the user experience.
[0082] It can generate videos for a single character or for multiple characters.
[0083] In some embodiments, there are multiple characters, and the target action template is matched with the attribute information of the multiple characters and the relationship between the multiple characters. The relationship between the multiple characters is determined based on semantic understanding of the story text.
[0084] For example, if the text related to the character selected by the user describes multiple characters, videos can be generated for multiple characters. The specific methods for the user to select multiple characters and trigger video generation can be found in the methods described in the foregoing embodiments, and will not be repeated here.
[0085] For example, semantic understanding of story text can be performed to identify multiple character groups with specified relationships, related text of character groups, attribute information of each character in each character group, and relationships between multiple characters in each character group. Based on the related text of multiple characters, reference images of multiple characters can be generated. Based on the attribute information of multiple characters and the relationships between multiple characters, target action templates can be determined. Based on the reference images of multiple characters and target action templates, videos can be generated.
[0086] For multiple characters, the relationships between them need to be determined based on the story text. For example, the characters might be mother and child, friends, or lovers. Different relationships between characters can correspond to different classic, iconic actions. Therefore, in addition to considering the attribute information of the characters, it is also necessary to determine the target action template based on the relationships between them.
[0087] The method described in the above embodiments can also generate videos for multiple characters, and the target action template is determined based on the attribute information of multiple characters and the relationship between multiple characters, making the generated video more compatible with multiple characters and improving the user's viewing experience.
[0088] In some embodiments, the target action template is selected from multiple candidate action templates, which are generated as follows: semantic understanding is performed on multiple story texts to determine multiple character groups in the multiple story texts, the relationships between characters in the multiple character groups, the types of characters in the multiple character groups, and the action description information corresponding to the multiple character groups; based on the relationships between characters in the multiple character groups and the types of characters in the multiple character groups, the action description information corresponding to the multiple characters is summarized to determine the reference action description information corresponding to each combination of relationships and types; based on the reference action description information corresponding to each combination of relationships and types, a candidate action template for each combination of relationships and types is generated, wherein each character group includes multiple characters, and each combination of relationships and types includes one relationship and the types of multiple characters.
[0089] Multiple target story texts can be selected from multiple story texts based on their interaction information. The above method is then executed on these target story texts, and will not be elaborated further here. For example, the interaction information of each story text includes one or more of the following: reading information, comment information, feedback information, and tagging information. Reading information may include the number of reads, comment information may include the number of comments and positive reviews, feedback information may include the number of likes and ratings, and tagging information may include the number of tags (underlining, etc.). Story texts with more interaction information can be selected as target story texts, and multiple candidate action templates can be generated using these target story texts.
[0090] For example, multiple role groups with the same relationship and type may have identical or similar action descriptions. These action descriptions can be summarized to obtain reference action descriptions for that relationship and type combination. For instance, if the relationship is "friends" and both roles are career-oriented women, the reference action description might include: "Both are standing side-by-side at a desk, staring intently at a computer screen; one person taps the key point on the screen, and the other nods in agreement." Each relationship and type combination can correspond to one or more reference action descriptions, generating one or more candidate action templates.
[0091] The method described in the above embodiments can automatically generate candidate action templates for combinations of multiple relationships and types, which can then be used to generate videos for multiple characters. The candidate action templates corresponding to each combination of relationships and types can reflect the classic and iconic actions of the characters associated with that combination, resulting in videos that are more closely matched to the multiple characters and improve the user experience.
[0092] In some embodiments, the relevant text for a character includes descriptive text of at least one of a scene or plot, and the target action template is determined based on at least one of the scene or plot, the types of multiple characters, and the relationships between the multiple characters. The scene or plot is determined based on semantic understanding of the relevant text and the context of the relevant text, and the types of multiple characters are determined based on attribute information of multiple characters.
[0093] It can perform semantic understanding on the relevant texts of multiple characters and their context to determine whether they include descriptive information about scenes and plots. For example, if the relevant texts of multiple characters describe an office scene, the selected target action template could be an office scene.
[0094] The method described above determines the target action template not only based on the types of multiple characters and the relationships between them, but also based on at least one of the scenes and plots in which the multiple characters are located, making the generated video more compatible with the multiple characters and their related text, thereby improving the user's viewing experience.
[0095] In some embodiments, the target action template is determined from one or more candidate action templates based on at least one of the scene and plot, and an understanding of one or more candidate action templates. The one or more candidate action templates are determined from multiple candidate action templates based on the types of multiple characters, the relationships between multiple characters, and the types and relationships corresponding to multiple candidate action templates.
[0096] For example, based on the types of multiple characters, the relationships between them, and the types corresponding to multiple candidate action templates, one or more candidate action templates corresponding to the types of multiple characters and the relationships between them are selected. The model is then used to identify at least one of the scenes and plots within these candidate action templates. From these candidate action templates, the one that matches at least one of the scenes and plots in the relevant text is selected as the target action template. If no candidate action template matches at least one of the scenes and plots in the relevant text, the target action template can be selected based on the interaction information of one or more candidate action templates. For example, the interaction information for each candidate action template includes at least one of the following: usage count, number of comments, number of positive reviews, number of likes, and rating.
[0097] The method described above can select target action templates that better match the types of multiple characters, the relationships between multiple characters, and the relevant text of the characters, thereby improving the quality of the generated video and enhancing the user experience.
[0098] The foregoing embodiments described how to determine the target motion template. Once the target motion template is determined, a video of the character can be generated based on the character's reference image and the target motion template.
[0099] In some embodiments, the actions of the characters in the video are generated based on actions in a target action template, and the background of the video is generated based on scene description text related to the characters in the story text.
[0100] Scene description text related to a character can be determined from the character's relevant text or from other text paragraphs in the story text. For example, if the character's relevant text includes scene description text, the background in the video can be generated based on that scene description text. For instance, if the scene description text describes a "palace" scene, the background in the generated video will be an image of a palace scene. The background in the video must conform to the setting of the character's relevant scene; for example, if the story text describes an ancient scene, the background in the video cannot be an image of a modern scene. Alternatively, if the character's relevant text does not include scene description text, description text for one or more scenes can be determined from the text paragraphs in the story text where the character appears, and the background in the video can then be generated.
[0101] In some embodiments, one or more action templates and one or more scenes may be displayed. In response to a user selecting a target action template and a target scene, a video is generated based on a reference image of the character, the target action template, and the target scene. Users can freely choose the scenes in the video to meet the needs of different users.
[0102] The method described above makes the generated video more closely match the scenes in the characters and story text, thereby improving the quality and effect of the generated video and enhancing the user experience.
[0103] In some embodiments, the target action template includes action configuration information and audio configuration information, and the video includes multiple images of the character. The multiple images of the character are generated based on the character's reference image and action configuration information, and the multiple images of the character include multiple dynamic images and / or static images.
[0104] The video can be a continuous sequence of images. Motion configuration information may include at least one of motion parameters or motion images, used to generate the character's actions. Audio configuration information may include background music audio or audio configuration information for character voices. For example, audio configuration information for a monologue, a dialogue between multiple characters, or a character singing. If the audio configuration information includes audio configuration information for character voices, the character's timbre is generated based on the character's attribute information, and the audio in the video is generated based on the character's timbre and the audio configuration information.
[0105] The method described in the above embodiments can generate dynamic audiovisual content corresponding to a character based on a target action template, such as a video of a character dancing, thereby enhancing the user's audiovisual experience.
[0106] In some embodiments, the switching rhythm of multiple images of a character matches the rhythm of the audio.
[0107] The audio includes background music. Multiple images are played alternately within the video, with the switching rhythm matching the audio rhythm, making the video more engaging and enhancing its audiovisual experience.
[0108] In some embodiments, in response to a user modifying the target action template, a regenerated video is displayed, wherein the regenerated video is generated based on the modified target action template and a reference image of the character.
[0109] Users can modify the video's motion template, the character's reference image, and the related text or prompts used to generate the video. For example, after the video is generated, it can be displayed in a preview interface, which can include video editing controls. In response to the user triggering these controls, the following can be displayed: Figure 3The interface shown allows for re-editing of the video, which will not be elaborated upon here. The preview interface can display a regenerate control; in response to the user triggering the regenerate control, the video is regenerated and displayed.
[0110] The preview interface can also display a publish control. In response to the user triggering the publish control, it can display... Figure 5 The comment editing interface 500 is shown. The comment editing interface 500 can display text paragraphs 501, a thumbnail 502 of the generated video, and the video thumbnail 502 can be displayed in the editing area 503. In response to the user entering information through the input area 505, the user-input information can be displayed in the editing area 503 as caption for the video. The comment editing interface 500 can also display a publish confirmation control 504. In response to the user triggering the publish confirmation control 504, the video and the user-input information can be published to the comment area of the e-book. Of course, it can also be distributed to other specified areas.
[0111] The method described in the above embodiments allows users to modify and publish the generated video, improving the flexibility of video processing, meeting various user needs, and enhancing the user experience.
[0112] This disclosure also provides a video processing apparatus, which is described below in conjunction with... Figure 6 Describe it.
[0113] Figure 6 These are structural diagrams of some embodiments of the video processing apparatus disclosed herein. Figure 6 As shown, the video processing apparatus 60 of this embodiment includes: an acquisition module 610 and a display module 620.
[0114] The acquisition module 610 is configured to acquire a video generated for a character in the story text, wherein the video is generated based on a reference image of the character and a target action template, the target action template being matched with the character's attribute information, and the reference image of the character being generated based on the relevant text of the character in the story text.
[0115] Display module 620 is configured to display video.
[0116] The video processing apparatus described in the above embodiment can acquire and display videos generated for characters in a story text. The videos are generated based on reference images of the characters and target action templates. The reference images are generated based on relevant text about the characters in the story text. The reference images match the characters' appearance as described in the story text, and videos generated based on these reference images better maintain consistency in the characters' appearance. The target action templates match the characters' attribute information; different target action templates can be used for characters with different attributes, increasing the richness of video types and their adaptability to the characters' attributes. This video processing apparatus allows users to transform descriptions of characters in story texts into more vivid and concrete visual images without requiring manual video creation, thus improving video generation efficiency and enhancing the user's reading and viewing experience.
[0117] In some embodiments, the acquisition module 610 is configured to acquire a video in response to a user selecting relevant text about a character and triggering a video generation function, wherein the character's attribute information is determined based on an understanding of the story text.
[0118] In some embodiments, the acquisition module 610 is configured to, in response to a user selecting relevant text of a character and triggering an image generation function, display a reference image of the character; in response to a user triggering a video generation function, display one or more action templates, wherein the one or more action templates include a target action template; and in response to a user selecting a target action template, acquire a video.
[0119] In some embodiments, the acquisition module 610 is configured to perform semantic understanding on the story text, determine the character, the character's related text, and the character's attribute information, wherein the character's related text includes the character's image description text; generate a reference image of the character based on the character's related text; determine a target action template based on the character's attribute information; and generate a video based on the character's reference image and the target action template.
[0120] In some embodiments, the target action template is determined based on the type of the character, which is determined based on semantic understanding of the character's attribute information.
[0121] In some embodiments, in response to the relevant text of a character, including descriptive text of at least one of a scene or plot, the target action template is determined based on at least one of the scene or plot and the type of the character, wherein at least one of the scene or plot is determined based on semantic understanding of the relevant text and the context of the relevant text.
[0122] In some embodiments, the target action template is determined from one or more candidate action templates based on at least one of the scene and plot, and an understanding of one or more candidate action templates. The one or more candidate action templates are determined from multiple candidate action templates based on the type of the character and the type corresponding to the multiple candidate action templates.
[0123] In some embodiments, there are multiple characters, and the target action template is matched with the attribute information of the multiple characters and the relationship between the multiple characters. The relationship between the multiple characters is determined based on semantic understanding of the story text.
[0124] In some embodiments, the relevant text for a character includes descriptive text of at least one of a scene or plot, and the target action template is determined based on at least one of the scene or plot, the types of multiple characters, and the relationships between the multiple characters. The scene or plot is determined based on semantic understanding of the relevant text and the context of the relevant text, and the types of multiple characters are determined based on attribute information of multiple characters.
[0125] In some embodiments, the target action template is determined from one or more candidate action templates based on at least one of the scene and plot, and an understanding of one or more candidate action templates. The one or more candidate action templates are determined from multiple candidate action templates based on the types of multiple characters, the relationships between multiple characters, and the types and relationships corresponding to multiple candidate action templates.
[0126] In some embodiments, the actions of the characters in the video are generated based on actions in a target action template, and the background of the video is generated based on scene description text related to the characters in the story text.
[0127] In some embodiments, the target action template includes action configuration information and audio configuration information, and the video includes multiple images of the character. The multiple images of the character are generated based on the character's reference image and action configuration information, and the multiple images of the character include multiple dynamic images and / or static images.
[0128] In some embodiments, the switching rhythm of multiple images of a character matches the rhythm of the audio.
[0129] In some embodiments, the target action template is selected from multiple candidate action templates, which correspond to multiple types. The multiple candidate action templates are generated in the following manner: semantic understanding is performed on multiple story texts to determine the types of multiple characters in the multiple story texts and the action description information corresponding to the multiple characters; based on the types of multiple characters, the action description information corresponding to the multiple characters is summarized to determine the reference action description information corresponding to each type; and based on the reference action description information corresponding to each type, a candidate action template for each type is generated.
[0130] This disclosure also provides an electronic device, which is described below in conjunction with... Figure 7 and 8 Describe it. Figure 7 A block diagram of an electronic device according to some embodiments of the present disclosure is shown.
[0131] like Figure 7 As shown, the electronic device 7 includes a processor 72; and a memory 71 coupled to the processor 72 for storing instructions, which, when executed by the processor 72, cause the processor 72 to perform a video processing method as described in any embodiment of this disclosure.
[0132] The electronic device described in the above embodiment can acquire and display videos generated for characters in a story text. The videos are generated based on reference images of the characters and target action templates. The reference images are generated based on relevant text about the characters in the story text. The reference images match the characters' appearances described in the story text, and videos generated based on these reference images better maintain consistency in the characters' appearances. The target action templates match the characters' attribute information; different target action templates can be used for characters with different attributes, increasing the richness of video types and their adaptability to the characters' attributes. This electronic device allows users to transform descriptions of characters in story texts into more vivid and concrete visual images without requiring manual video creation, thus improving video generation efficiency and enhancing the user's reading and viewing experience.
[0133] Memory 71 is used to store one or more computer-readable instructions. Memory 71 may include any combination of various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory, including but not limited to random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), read-only memory (ROM), and flash memory. Memory 71 may, for example, store operating systems, application programs, bootloaders, databases, and other programs, as well as various application programs and various data.
[0134] The processor 72 is configured to execute computer-readable instructions to implement the video processing method of any of the foregoing embodiments. Specific implementations of each step of the method can be found in the above embodiments; repeated details will not be elaborated upon here.
[0135] The processor 72 can be various processing devices, such as a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The central processing unit (CPU) can be based on x86 or ARM architectures, etc.
[0136] The processor 72 and the memory 71 can communicate with each other directly or indirectly. For example, the processor 72 and the memory 71 can communicate via a network. The network can include a wireless network, a wired network, and / or any combination of wireless and wired networks. The processor 72 and the memory 71 can also communicate with each other via a system bus, which is not limited in this disclosure.
[0137] It should be noted that Figure 7 The components of the electronic device 7 shown are merely exemplary and not limiting. The electronic device 7 may have other components depending on the specific application requirements. The processor 72 can control other components in the electronic device 7 to perform desired functions.
[0138] Electronic device 7 can be implemented by software, firmware and / or hardware, and can be integrated into a device with the relevant application installed.
[0139] Figure 8 Block diagrams of electronic devices according to other embodiments of the present disclosure are shown.
[0140] Figure 8 The electronic device 8 shown can be a computer system with a dedicated hardware structure, which can perform corresponding functions when the relevant application is installed.
[0141] Electronic devices include, but are not limited to, mobile terminals such as smartphones, laptops, personal digital assistants (PDAs), tablet computers (PCs), PMPs (portable multimedia players), in-vehicle terminals (such as in-vehicle navigation terminals), wearable devices, and fixed terminals such as digital televisions and desktop computers.
[0142] like Figure 8As shown, the Central Processing Unit (CPU) 81 performs various processes based on a program stored in the Read-Only Memory (ROM) 82 or a program loaded from the storage section 88 into the Random Access Memory (RAM) 83. The RAM 83 stores data required as needed when the CPU 81 performs various processes, etc. The CPU is merely exemplary; it could also be other types of processors, such as the various processors described above. The ROM 82, RAM 83, and storage section 88 can be various forms of computer-readable storage media. It should be noted that although... Figure 8 The diagram shows ROM 82, RAM 83 and storage section 88, but one or more of them may be combined or located in the same or different memory or storage modules.
[0143] CPU 81, ROM 82 and RAM 83 are interconnected via bus 84. Input / output interface 85 is also connected to bus 84.
[0144] The following components are connected to the input / output interface 85: input section 86, such as a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output section 87, including displays such as cathode ray tube (CRT), liquid crystal display (LCD), speakers, vibrators, etc.; storage section 88, including hard disk, magnetic tape, etc.; and communication section 89, including network interface cards such as LAN cards, modems, etc. The communication section 89 allows communication processing to be performed via a network such as the Internet. It is easy to understand that, although... Figure 8 The portion of the electronic device 8 shown communicates via bus 84, but it may also communicate via a network or other means, wherein the network may include a wireless network, a wired network, and / or any combination of wireless and wired networks.
[0145] As needed, drive 810 is also connected to input / output interface 85. Removable media 811, such as disks, optical disks, magneto-optical disks, semiconductor memories, etc., are installed on drive 810 as needed, so that computer programs read from them can be installed into storage section 88 as needed.
[0146] When the above series of processes are implemented through software, the program constituting the software can be installed from a network such as the Internet or a storage medium such as a removable medium 811.
[0147] According to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. This disclosure also provides a computer program product, including a computer program that, when executed by a processor, implements the video processing method of any embodiment of this disclosure.
[0148] The computer program product described in the above embodiments can acquire and display videos generated for characters in a story text. The videos are generated based on reference images of the characters and target action templates. The reference images are generated based on relevant text about the characters in the story text. The reference images match the characters' appearance as described in the story text, and videos generated based on these reference images better maintain consistency in the characters' appearance. The target action templates match the characters' attribute information; different target action templates can be used for characters with different attributes, increasing the richness of video types and their adaptability to the characters' attributes. This computer program product allows users to transform descriptions of characters in story texts into more vivid and concrete visual images without requiring manual video creation, thus improving video generation efficiency and enhancing the user's reading and viewing experience.
[0149] For example, some embodiments of this disclosure include a computer program product that, when run on a computer, causes the computer to implement the methods of any of the foregoing embodiments. The computer program product includes computer instructions carried on a computer-readable medium, comprising program code for performing the methods shown in the flowchart. In such embodiments, the computer instructions can be downloaded and installed from a network via communication section 89, or installed from storage section 88, or installed from ROM 82. When the computer program is executed by CPU 81, the video processing method of embodiments of this disclosure is performed.
[0150] This disclosure also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the video processing method of any embodiment of this disclosure.
[0151] The computer-readable storage medium described in the above embodiments can acquire and display videos generated for characters in a story text. The videos are generated based on reference images of the characters and target action templates. The reference images are generated based on relevant text about the characters in the story text. The reference images match the characters' appearances described in the story text, and videos generated based on these reference images better maintain consistency in the characters' appearances. The target action templates match the characters' attribute information; different target action templates can be used for characters with different attributes, increasing the richness of video types and their adaptability to the characters' attributes. This computer-readable storage medium allows users to transform descriptions of characters in story texts into more vivid and concrete visual images without requiring manual video creation, thus improving video generation efficiency and enhancing the user's reading and viewing experience.
[0152] It should be noted that, in the context of this disclosure, a computer-readable medium can be a tangible medium that may contain or store programs for use by or in conjunction with an instruction execution system, apparatus, or device.
[0153] A computer-readable medium may be a computer-readable storage medium, a computer-readable signal medium, or any combination thereof.
[0154] Computer-readable storage media include, but are not limited to, systems, apparatuses, or devices that are electrical, magnetic, optical, electromagnetic, infrared, or semiconductor, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. The computer-readable storage medium stores computer instructions that, when executed by a processor, implement the methods of any of the foregoing embodiments.
[0155] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, capable of sending, propagating, or transmitting programs for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0156] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0157] In some embodiments, a computer program is also provided, comprising: instructions that, when executed by a processor, cause the processor to perform the methods of any of the foregoing embodiments. For example, the instructions may be embodied in computer program code.
[0158] In embodiments of this disclosure, computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof. These programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network (including a local area network (LAN) or a wide area network (WAN)), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0159] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0160] The functions described above can be performed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary hardware logic components that can be used include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0161] While specific embodiments of this disclosure have been described in detail by way of example, those skilled in the art should understand that the examples are for illustrative purposes only and not intended to limit the scope of this disclosure. Those skilled in the art should understand that modifications can be made to the above embodiments without departing from the scope and spirit of this disclosure. The scope of this disclosure is defined by the appended claims.
Claims
1. A video processing method, comprising: Obtain a video generated for a character in a story text, wherein the video is generated based on a reference image of the character and a target action template, the target action template being matched with the character's attribute information, and the reference image of the character being generated based on relevant text of the character in the story text; The video is displayed.
2. The video processing method according to claim 1, wherein, The acquisition of videos generated for characters in the story text includes: In response to a user selecting relevant text about the character and triggering a video generation function, the video is obtained, wherein the character's attribute information is determined based on an understanding of the story text.
3. The video processing method according to claim 2, wherein, The process of responding to the user selecting relevant text for the role and triggering the video generation function to obtain the video includes: In response to the user selecting relevant text about the character and triggering the image generation function, a reference image of the character is displayed; In response to the user triggering the video generation function, one or more action templates are displayed, wherein the one or more action templates include the target action template; In response to the user selecting the target action template, the video is acquired.
4. The video processing method according to claim 1, wherein, The acquisition of videos generated for characters in the story text includes: Semantic understanding is performed on the story text to determine the character, the character's related text, and the character's attribute information, wherein the character's related text includes the character's image description text; Generate a reference image of the character based on the relevant text of the character; Based on the attribute information of the character, determine the target action template; The video is generated based on the reference image of the character and the target action template.
5. The video processing method according to claim 1, wherein, The target action template is determined based on the type of the character, and the type of the character is determined based on semantic understanding of the character's attribute information.
6. The video processing method according to claim 5, wherein, The relevant text for the character includes descriptive text of at least one of the scenes and plots, and the target action template is determined based on at least one of the scenes and plots and the type of the character, wherein at least one of the scenes and plots is determined based on semantic understanding of the relevant text and the context of the relevant text.
7. The video processing method according to claim 6, wherein, The target action template is determined from the one or more candidate action templates based on at least one of the scene and plot, and an understanding of one or more candidate action templates. The one or more candidate action templates are determined from the multiple candidate action templates based on the type of the character and the type corresponding to the multiple candidate action templates.
8. The video processing method according to claim 1, wherein, The character refers to multiple characters, and the target action template is matched with the attribute information of the multiple characters and the relationship between the multiple characters. The relationship between the multiple characters is determined based on semantic understanding of the story text.
9. The video processing method according to claim 8, wherein, The relevant text for the character includes descriptive text of at least one of the scenes and plots. The target action template is determined based on at least one of the scenes and plots, the types of the multiple characters, and the relationships between the multiple characters. The scenes and plots are determined based on semantic understanding of the relevant text and its context. The types of the multiple characters are determined based on attribute information of the multiple characters.
10. The video processing method according to claim 9, wherein, The target action template is determined from the one or more candidate action templates based on at least one of the scenarios and plots, as well as an understanding of one or more candidate action templates. The one or more candidate action templates are determined from the multiple candidate action templates based on the types of the multiple characters, the relationships between the multiple characters, and the types and relationships corresponding to the multiple candidate action templates.
11. The video processing method according to any one of claims 1-10, wherein, The actions of the character in the video are generated based on the actions in the target action template, and the background in the video is generated based on the scene description text related to the character in the story text.
12. The video processing method according to any one of claims 1-10, wherein, The target action template includes action configuration information and audio configuration information. The video includes multiple images of the character. The multiple images of the character are generated based on the reference image of the character and the action configuration information. The multiple images of the character include multiple dynamic images and / or static images.
13. The video processing method according to claim 12, wherein, The switching rhythm of the multiple images of the character matches the rhythm of the audio.
14. The video processing method according to any one of claims 1-10, wherein, The target action template is selected from multiple candidate action templates, which correspond to multiple types. The multiple candidate action templates are generated in the following way: Perform semantic understanding on multiple story texts to determine the types of multiple characters in the multiple story texts and the corresponding action description information of the multiple characters; Based on the types of the multiple roles, the action description information corresponding to the multiple roles is summarized, and the reference action description information corresponding to each type is determined. Based on the reference action description information corresponding to each type, a candidate action template for each type is generated.
15. A video processing apparatus, comprising: The acquisition module is configured to acquire a video generated for a character in the story text, wherein the video is generated based on a reference image of the character and a target action template, the target action template being matched with the attribute information of the character, and the reference image of the character being generated based on the relevant text of the character in the story text; The display module is configured to display the video.
16. An electronic device comprising: processor; as well as A memory coupled to the processor is used to store instructions that, when executed by the processor, cause the processor to perform the video processing method according to any one of claims 1 to 14.
17. A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the video processing method of any one of claims 1 to 14.
18. A computer program product comprising a computer program that, when executed by a processor, implements the video processing method according to any one of claims 1 to 14.