Content generation methods, apparatuses, electronic devices, storage media, and program products

CN120602745BActive Publication Date: 2026-09-01DOUYIN VISION CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510865574.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-25
Publication Date
2026-09-01
Estimated Expiration
2045-06-25

AI Technical Summary

Technical Problem

相关技术中,需要通过第三方修图软件进行图像编辑,或者通过PS(Photoshop,图像处理软件)等绘画工具进行二创,使得制作难度和成本较高

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120602745B_ABST
    Figure CN120602745B_ABST
Patent Text Reader

Abstract

This disclosure provides a content generation method, apparatus, electronic device, storage medium, and program product, relating to the field of computer technology. The content generation method includes: in response to a user's triggering operation, acquiring first visual content from a video and displaying an editing interface; in the editing interface, recommending at least one candidate content related to the first visual content, the at least one candidate content being determined based on at least one of a plot dimension, a character dimension, and a scene dimension in the first visual content, the at least one candidate content being derived from the video and one or more interactive content related to the video; in response to a target content determined by the user from the at least one candidate content, generating second visual content based on the first visual content and the target content.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, and in particular to a content generation method, apparatus, electronic device, storage medium, and program product. Background Technology

[0002] With the development of the film and television industry and the increase in community activity, more and more people are sharing and discussing derivative works from film and television works on social media. The related techniques often require image editing using third-party image editing software or drawing tools such as Photoshop (PS) for derivative creation, making the production process difficult and costly. Summary of the Invention

[0003] According to some embodiments of this disclosure, a content generation method is provided, comprising: in response to a user's triggering operation, acquiring first visual content in a video and displaying an editing interface; in the editing interface, recommending at least one candidate content related to the first visual content, the at least one candidate content being determined based on at least one of a plot dimension, a character dimension, and a scene dimension in the first visual content, the at least one candidate content being derived from one or more of the video and interactive content related to the video; in response to a target content determined by the user from the at least one candidate content, generating second visual content based on the first visual content and the target content.

[0004] According to other embodiments of this disclosure, a content generation apparatus is provided, comprising: a processing module configured to, in response to a user's triggering operation, acquire first visual content in a video and display an editing interface; a recommendation module configured to, on the editing interface, recommend at least one candidate content related to the first visual content, the at least one candidate content being determined based on at least one of a plot dimension, a character dimension, and a scene dimension in the first visual content, the at least one candidate content being derived from one or more of the video and interactive content related to the video; and a content generation module configured to, in response to a target content determined by the user from the at least one candidate content, generate second visual content based on the first visual content and the target content.

[0005] According to some embodiments of this disclosure, an electronic device is provided, including: a processor; and a memory coupled to the processor for storing instructions, which, when executed by the processor, cause the processor to perform a content generation method according to any embodiment of this disclosure.

[0006] According to some embodiments of the present disclosure, a computer-readable storage medium is provided that stores computer instructions thereon, wherein the computer instructions, when executed by a processor, implement the content generation method of any embodiment of the present disclosure.

[0007] According to some embodiments of this disclosure, a computer program product is provided, comprising: computer instructions that, when executed by a processor, implement the content generation method of any embodiment described in this disclosure.

[0008] Other features, aspects, and advantages of this disclosure will become clear from the following detailed description of exemplary embodiments with reference to the accompanying drawings. Attached Figure Description

[0009] Embodiments of this disclosure are described below with reference to the accompanying drawings. It should be understood that the drawings described below are merely illustrative of some embodiments of this disclosure and are not intended to limit the scope of this disclosure. In the drawings:

[0010] Figure 1 A flowchart illustrating a content generation method according to some embodiments of the present disclosure is shown;

[0011] Figure 2 A flowchart illustrating recommended candidate content according to some embodiments of the present disclosure is shown;

[0012] Figure 3 A flowchart illustrating recommended candidate content according to other embodiments of this disclosure is shown;

[0013] Figure 4 A flowchart illustrating recommended candidate content according to some embodiments of the present disclosure is shown;

[0014] Figure 5 A flowchart illustrating a content generation method according to other embodiments of the present disclosure is shown;

[0015] Figure 6 Showing video screenshots according to some embodiments of this disclosure;

[0016] Figure 7 A schematic diagram illustrating an editing interface according to some embodiments of the present disclosure;

[0017] Figure 8 A schematic diagram illustrating an editing interface according to other embodiments of the present disclosure;

[0018] Figure 9 A schematic diagram showing an editing interface according to some embodiments of the present disclosure;

[0019] Figure 10 A schematic diagram showing an editing interface according to some embodiments of the present disclosure;

[0020] Figure 11 A schematic diagram showing an editing interface according to some embodiments of the present disclosure;

[0021] Figure 12 A schematic diagram showing an editing interface according to some embodiments of the present disclosure;

[0022] Figure 13 A schematic diagram showing an editing interface according to some embodiments of the present disclosure;

[0023] Figure 14 A schematic diagram showing an editing interface according to some embodiments of the present disclosure;

[0024] Figure 15 A block diagram of a content generation apparatus according to some embodiments of the present disclosure is shown;

[0025] Figure 16 A block diagram of an electronic device according to some embodiments of the present disclosure is shown;

[0026] Figure 17 Block diagrams of electronic devices according to other embodiments of the present disclosure are shown.

[0027] It should be understood that, for ease of description, the dimensions of the various parts shown in the accompanying drawings are not necessarily drawn to actual scale. The same or similar reference numerals are used in the various drawings to denote the same or similar parts. Therefore, once an item is defined in one drawing, it may not be discussed further in subsequent drawings. Detailed Implementation

[0028] The technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. It should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein.

[0029] It should be understood that the various steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect. Unless otherwise specifically stated, the relative arrangement, numerical expressions, and values ​​of components and steps set forth in these embodiments should be interpreted as merely exemplary and do not limit the scope of this disclosure.

[0030] As used in this disclosure, the term "comprising" and its variations are open-ended terms that include at least the following elements / features but do not exclude other elements / features, i.e., "including but not limited to". The term "based on" means "at least partially based on".

[0031] It should be noted that the concepts of "first," "second," etc., used in this disclosure are used only to distinguish different devices, modules, or units, and are not intended to define the order of functions performed by these devices, modules, or units or their interdependencies. Unless otherwise specified, the concepts of "first," "second," etc., are not intended to imply that the objects described herein must be in a given temporal, spatial, rank, or any other given order.

[0032] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0033] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0034] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data shall comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation portals shall be provided for users to choose to authorize or refuse.

[0035] The embodiments of this disclosure are described in detail below with reference to the accompanying drawings; however, this disclosure is not limited to these specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. Furthermore, in one or more embodiments, specific features, structures, or characteristics can be combined in any suitable manner that will be apparent to those skilled in the art from this disclosure.

[0036] In related technologies, when creating derivative works, users need to spend considerable effort adding decorative elements to images to suit the storyline, and they also need to edit images using third-party image editing software, which increases the difficulty and cost of creation and reduces efficiency. This disclosure provides a content generation method that can reduce the difficulty and cost of creation, allowing users to create beautiful derivative content even without professional image editing skills. The solution of this disclosure will be described below with reference to specific embodiments.

[0037] Figure 1 A schematic flowchart of a content generation method according to some embodiments of the present disclosure is shown.

[0038] like Figure 1As shown, the content generation method includes: Step S1, in response to a user's trigger operation, acquiring first visual content from a video and displaying an editing interface; Step S2, in the editing interface, recommending at least one candidate content related to the first visual content, wherein the at least one candidate content is determined based on at least one of the plot dimension, character dimension, and scene dimension in the first visual content, and the at least one candidate content comes from one or more of the video and interactive content related to the video; Step S3, in response to a target content determined by the user from the at least one candidate content, generating second visual content based on the first visual content and the target content.

[0039] For example, when a user is watching video content, after capturing a frame or a segment of video, they are automatically taken to an editing interface. This interface displays one or more candidate content items. Each candidate can be a single frame or multiple frames from the video, or a portion of interactive content. Interactive content can include comments, related posts, PGC (Professionally-generated Content), and UGC (User-generated Content). Comments can be text, emojis, images, GIFs, or videos.

[0040] Among the recommended candidate content, users can select one or more target content. The first visual content and the target content are combined to generate the second visual content, which can be an image, GIF, video, etc.

[0041] In the above embodiments, users do not need to spend much effort searching for resources to edit images or videos. Instead, after the user triggers an operation, relevant candidate content is automatically recommended in the editing interface. After the user selects the target content from the candidate content, the created content is generated. Users do not need to have professional image editing skills, nor do they need to jump to a third-party platform for image retouching. This reduces the editing steps, thereby reducing the difficulty and cost of creation and improving the motivation and efficiency of creation.

[0042] Below, we will combine Figures 2 to 14 The content generation method disclosed herein will be further described.

[0043] Figure 2 A flowchart illustrating recommended candidate content according to some embodiments of this disclosure is shown.

[0044] like Figure 2As shown, recommending at least one candidate content related to the first visual content includes: step S21, performing content understanding on the first visual content, and determining at least one of the first key content in the plot dimension, the second key content in the character dimension, and the third key content in the scene dimension; step S22, determining the at least one candidate content in the video and / or the interactive content based on at least one of the first key content, the second key content, and the third key content.

[0045] For example, in the plot dimension, the first key element of the first visual content can be determined. This first key element may include, for example, the plot itself. In the character dimension, the second key element of the first visual content can be determined. This second key element may include, for example, character identity, character actions, character dialogue, and character clothing. In the scene dimension, the third key element of the first visual content can be determined. This third key element may include, for example, environmental features, object features, prop features, lighting, and color features.

[0046] Based on the first, second, and third key content, multiple candidate content items are identified. These candidate items can come from videos or interactive content. This recommendation process can be implemented using a large model.

[0047] In this embodiment, by performing content understanding on the first visual content, key content in the first visual content can be identified, and then candidate content can be obtained based on the key content, thus realizing automatic recommendation of candidate content, which facilitates users to select target content in the future and realize the secondary creation of content.

[0048] In the above embodiments, the first visual content is a frame or a video clip from the video. The final generated second visual content can also be a frame or a video clip, thus enriching the content creation.

[0049] The primary visual content includes image information but excludes non-image information. For example, after extracting images and video clips from film and television content, interactive functions, navigation bars, episode information, and other content are automatically removed. This allows users to obtain relatively complete and unobstructed image or video resources, facilitating clear and complete secondary visual content generated during subsequent editing and improving the user's visual experience.

[0050] During the process of removing non-image information, images and videos extracted from the film and television content can be compared with the plot stored in the background to obtain complete and unobstructed image or video resources. Alternatively, the client can hide interactive functions, navigation bars, episode information, etc., and then extract image or video clips to obtain relatively complete and unobstructed image or video resources.

[0051] In some embodiments, the step of performing content understanding on the first visual content and determining at least one of the first key content in the plot dimension, the second key content in the character dimension, and the third key content in the scene dimension includes: performing content understanding on the first visual content to obtain a content focus; determining the dimension related to the first visual content as the plot dimension in response to the content focus being related to event development or plot logic; and determining the first key content of the first visual content in the plot dimension.

[0052] For example, by understanding the content of the first visual element, we can identify the content focus, which reflects which dimension is emphasized in that first visual element. If the content focus can showcase story events, plot developments, or key turning points, emphasizing "what happened," that is, if the content focus is related to the development of events or plot logic, then the dimension related to that first visual element can be identified as the plot dimension.

[0053] For example, through large-scale model analysis, if the images captured by the user contain interactive behaviors between characters, such as dialogue, conflict, and cooperation; or contain key actions that drive the storyline, such as decryption, combat, and escape; or contain important moments, such as flashbacks, then the dimension related to the first visual content is the plot dimension. In this case, it is necessary to determine the first key content of the first visual content within the plot dimension.

[0054] In the above embodiments, based on the content focus corresponding to the first visual content, the dimension most relevant to the first visual content is determined as the plot dimension, and then the key content of the plot dimension can be extracted to achieve the recommendation of candidate content.

[0055] In some embodiments, the step of performing content understanding on the first visual content and determining at least one of the first key content in the plot dimension, the second key content in the character dimension, and the third key content in the scene dimension includes: performing content understanding on the first visual content to obtain a content focus; in response to the content focus being related to the character's image or personality psychology, determining the dimension related to the first visual content as the character dimension; and determining the second key content of the first visual content in the character dimension.

[0056] For example, if the focus of the content is on shaping the character's image, personality, or emotions, emphasizing "who the character is," that is, if the focus of the content is related to the character's image or personality and psychology, then the dimension related to the first visual content can be identified as the character dimension.

[0057] For example, through large-scale model analysis, if the images captured by users contain physical characteristics of a character, such as expressions, clothing, and signature props; or contain details that reflect the character's background or psychology; or contain information that reflects the character's growth stage or identity transformation, then the dimension related to the first visual content is the character dimension. In this case, it is necessary to determine the primary key content of the first visual content within the character dimension.

[0058] In the above embodiments, based on the content focus corresponding to the first visual content, the dimension most relevant to the first visual content is determined as the role dimension, and then the key content of the role dimension can be extracted to achieve the recommendation of candidate content.

[0059] In some embodiments, the step of performing content understanding on the first visual content and determining at least one of the first key content in the plot dimension, the second key content in the character dimension, and the third key content in the scene dimension includes: performing content understanding on the first visual content to obtain a content focus; determining the dimension related to the first visual content as the scene dimension in response to the content focus being related to the environmental atmosphere or spatial setting; and determining the third key content of the first visual content in the scene dimension.

[0060] For example, if the focus of the content is on building the environment, worldview, or atmosphere in which the story takes place, emphasizing "where it happens," then the dimension related to this first visual content can be identified as the scene dimension.

[0061] For example, through large-scale model analysis, the images captured by users may contain spatial structures of the environment, such as architectural layouts and natural landscapes; or details of historical context, such as the furnishings of ancient streets or technological elements of futuristic cities; or atmosphere, such as the oppressive feeling of rainy weather. This indicates that the dimension related to the first visual content is the scene dimension. In this case, it is necessary to determine the primary key content of the first visual content within the scene dimension.

[0062] In the above embodiments, based on the content focus corresponding to the first visual content, the dimension most relevant to the first visual content is determined as the scene dimension, and then the key content of the scene dimension can be extracted to achieve the recommendation of candidate content.

[0063] Those skilled in the art should understand that the focus of content in a single image frame or video clip may not be a single dimension, but rather a combination of multiple dimensions. For example, the focus of content can be used to determine whether the dimensions related to the first visual content are plot and character, plot and scene, character and scene, or plot, scene, and character. Therefore, it is necessary to determine the key content across multiple dimensions.

[0064] By analyzing the proportions of visual elements, such as character actions, environment, and facial expressions, we can determine the dimensions most relevant to the primary visual content.

[0065] If the candidate content comes from a video, then the candidate content is determined as follows: Figure 3 As shown. Figure 3 A flowchart illustrating recommended candidate content according to other embodiments of this disclosure is shown.

[0066] like Figure 3 As shown, determining at least one candidate content in the video based on at least one of the first key content, the second key content, and the third key content includes: step S2211, determining key plot points related to the first visual content based on at least one of the first key content, the second key content, and the third key content; step S2212, determining a target set of the key plot points in the video, the target set including one or more episodes; and step S2213, recommending at least one candidate content based on keyframes in the target set.

[0067] For example, if the key content of a frame in a video is a school sports meet, then the key storylines related to that frame could be athlete selection, awards ceremony, etc. This allows us to determine the target set of these key storylines within the video. This target set could be a single episode or multiple episodes. Keyframes from this target set that are related to these key storylines are then recommended as candidate content. These keyframes can be one frame or multiple frames.

[0068] In this embodiment, the location of key plot points related to the first visual content within the video can be quickly identified, thereby recommending relevant candidate content. Furthermore, since key plot points may exist across multiple episodes of the video, the cost for users to manually take screenshots across different episodes is reduced.

[0069] In some embodiments, determining the key plot related to the first visual content based on at least one of the first key content, the second key content, and the third key content includes: determining the current plot based on at least one of the character identity, character actions, character dialogue, and character clothing in the second key content, and using the key plot related to the current plot as the key plot related to the first visual content.

[0070] For example, identifying basic character information, such as identity and clothing, can lay the foundation for plot analysis; analyzing character actions or interactions can capture plot conflicts and developments; and character dialogue can link character motivations to plot logic. In this embodiment, based on character identity, actions, dialogue, and clothing, the current plot of the first-person view content can be determined, and then key plot points related to this current plot can be identified. These key plot points might be similar to the current plot, flashback scenes, future developments, or plots with turning points or conflicts. Thus, from a character perspective, candidate content can be recommended to the user.

[0071] In some embodiments, determining the key plot related to the first visual content based on at least one of the first key content, the second key content, and the third key content includes: determining the current plot based on at least one of environmental features, object features, prop features, and light and color features in the third key content, and using the key plot related to the current plot as the key plot related to the first visual content.

[0072] For example, environmental features, object features, prop features, and lighting and color features can reflect the spatiotemporal coordinates of the plot, plot clues, plot logic, and the driving force of the plot. Therefore, based on environmental features, object features, prop features, and lighting and color features, the current plot can be determined, and then the key plot points related to that current plot can be identified. In this way, from a scene perspective, candidate content can be recommended to the user.

[0073] In some embodiments, determining the key plot related to the first visual content based on at least one of the first key content, the second key content, and the third key content includes: using the key plot related to the first key content as the key plot related to the first visual content.

[0074] Since the first key content reflects the current plot, key plot points related to the current plot can be directly identified. In this way, from the perspective of plot, candidate content can be recommended to the user.

[0075] If the candidate content comes from interactive content, then the candidate content is determined as follows: Figure 4 As shown. Figure 4 A flowchart illustrating recommended candidate content according to some embodiments of the present disclosure is shown.

[0076] like Figure 4As shown, the step of determining at least one candidate content in the interactive content based on at least one of the first key content, the second key content, and the third key content is as follows: Step S2221, determining a key plot related to the first visual content based on at least one of the first key content, the second key content, and the third key content; Step S2222, determining at least one candidate content related to the key plot in the interactive content.

[0077] The specific method for determining key plot points related to the first visual content based on at least one of the first, second, and third key content has been described in the above embodiments and will not be elaborated further here. Since the key plot points are determined, candidate content can be identified from comments related to those key plot points, or from images, videos, etc., created by other users based on those key plot points. This reduces the time users spend searching for relevant content, thereby improving creation efficiency.

[0078] After the large model analyzes and identifies candidate content related to the first visual content, it can be displayed to the user. At this point, there is a large number of candidate contents displayed, so the user can further filter them, thereby achieving key recommendations.

[0079] In some embodiments, each of the at least one candidate content includes at least one of a key person and a key location, and generating second visual content based on the first visual content and the target content in response to the user determining a target content among the at least one candidate content includes: displaying reference visual content related to at least one of the key person and the key location in response to the user selecting at least one of the key person and the key location; and generating second visual content based on the first visual content and the target content in response to the user determining the target content among the reference visual content, wherein the target content includes one or more of the reference visual content.

[0080] For example, each candidate item displays either a key person or a key location. If a user selects a key person, all filtered reference visuals will include that key person. If a user selects a key location, all filtered reference visuals will include locations related to that key location. If a user selects both a key person and a key location, all filtered reference visuals will include locations related to that key location and will also include the key person, making it convenient for users to filter.

[0081] In some embodiments, in response to a user selecting one candidate content from at least one candidate content, the candidate content is spliced ​​with a first visual content; in response to a user selecting at least one of a key person and a key location related to the candidate content, reference visual content related to at least one of the key person and key location is displayed; in response to a user determining a target content in the reference visual content, the candidate content, the first visual content, and the target content are spliced ​​together to obtain a second visual content.

[0082] For example, after the system recommends multiple candidate items, the user clicks on a candidate item, which is then stitched together with the original screenshot into a single image. After the user clicks on information such as the person, they are redirected to a secondary page that displays recommended frames related to that person, allowing the user to select again, thus facilitating the user's filtering process.

[0083] Each of the at least one candidate content includes time information. After the target content is determined, the first visual content and the target content are synthesized according to the time information to generate the second visual content.

[0084] For example, the episode number and time of the recommended keyframe can be displayed. This allows multiple frames to be synthesized according to time, making the synthesized second visual content chronological and more logical, thus promoting narrative and story comprehension.

[0085] In some embodiments, the at least one candidate content includes at least one candidate text content, the at least one candidate text content being one or more of dialogue in the video and interactive content related to the video, and generating second visual content based on the first visual content and the target content in response to the target content determined by the user in the at least one candidate content includes: generating the second visual content based on the first visual content and the target text content in response to the user selecting target text content in the at least one candidate text content.

[0086] Candidate text content can be, for example, lines from dialogue or popular internet slang. Users can select target text content, which is then added to the first visual content to form the second visual content. This results in the final generated content containing both image and text content, making the content vivid and engaging while reducing the user's editing costs.

[0087] After generating the second visual content, the system can also recommend editing tools to the user to facilitate further editing of the content. For example, the system can recommend the editing tools to the user based on at least one of the image style, image color, and image content of the first visual content.

[0088] Editing tools include filters and art styles. For example, if the primary visual content includes landscapes, filters that enhance color saturation and contrast can be recommended, or an artistic art style can be suggested to give the image a unique artistic texture. If the primary visual content includes portraits, filters that beautify skin tone and soften skin texture can be recommended, or an art style that can transform the portrait into a cartoon or comic book style. If the primary visual content has a retro style, filters that add a retro atmosphere or a retro art style can be recommended. If the primary visual content has a fresh style, elegant filters can be recommended, or an art style that fills the image with a fresh and artistic feel. If the primary visual content is a warm-toned image, warm-toned filters can be recommended to enhance the warm atmosphere of the image. If the primary visual content is a cool-toned image, cool-toned filters can be recommended to give the image a cool aesthetic, and so on.

[0089] In the above embodiments, by adding editing tools to visual content, users can choose suitable editing tools to process images, thereby improving the emotional expression of the content and enhancing artistic creation.

[0090] After the system generates the second visual content, it shares the second visual content in response to the user's trigger operation on the sharing control of the editing interface.

[0091] In related technologies, users need to edit images using third-party image editing software before sharing them on other social media platforms. In this embodiment, however, users can directly share to social media platforms via a sharing control after completing content editing on the editing interface, reducing cross-platform editing and sharing costs.

[0092] The content generation method of this disclosure will now be further described with reference to a specific embodiment.

[0093] Figure 5 A flowchart illustrating a content generation method according to other embodiments of this disclosure is shown.

[0094] like Figure 5 As shown, in step 51, a screenshot of the video content is taken and the screen is automatically cleared.

[0095] For example, when a user takes a screenshot while watching a short drama, such as... Figure 6 As shown, the screenshot automatically removes interactive features 61, navigation bar 62, episode information 63, and other information, allowing users to obtain relatively complete and unobstructed image resources.

[0096] In step S52, the editing interface is displayed.

[0097] For example, such as Figure 7 As shown, the image editing process begins. Figure 7 A schematic diagram of an editing interface according to some embodiments of the present disclosure is shown. This editing interface can display relevant stills and add text. Recommended relevant stills are shown in Figure 71.

[0098] The above steps allow users to quickly and easily enter the image creation process, reducing the cumbersome process of users having to go to a third-party image creator and then come back to share the image.

[0099] In step S53, keyframes related to the screenshot are recommended.

[0100] For example, based on screenshots, the system intelligently identifies image elements and recommends relevant keyframes based on plot, characters, and scenes. The system not only selects keyframes near the current episode but can also recommend related frames across episodes. For instance, if the screenshot shows a scene of plum blossoms in episode 3 of the series, the system can identify subsequent key plot points such as red plum blossoms and characters, and then select and recommend corresponding frames. Users can also add stills to stitch and combine images.

[0101] Each keyframe displays the episode number and time of the episode, as well as key characters or locations in the image. Clicking on a key character or location leads to a secondary page displaying recommended frames featuring that character or scene, allowing users to easily filter the content.

[0102] like Figure 8 As shown, after a user clicks on a recommended still image (image 81), the two images are stitched together. After the user clicks on information such as a person, the system can identify stills related to that content. Figure 9 As shown, 91 stills related to this character are recommended. (For example...) Figure 10 As shown, after the user clicks on the recommended still from show 101, the image is further composited.

[0103] In this step, relevant stills are intelligently recommended, reducing the cost for users to manually take screenshots from different episodes.

[0104] In step S54, text related to the screenshot is recommended.

[0105] For example, based on screenshots, the system intelligently identifies image elements and recommends relevant highlight lines and trending internet terms based on plot, characters, and scenes. It also supports adding corresponding text and further image editing.

[0106] like Figure 11 As shown, when a user clicks to add text, the system displays relevant quotes, trending memes, and other related information. For example, when a user clicks to add text, the system displays relevant quotes, trending memes, and other related information. Figure 12 As shown, the text 121 is added to the image.

[0107] In this step, relevant trending texts are intelligently recommended, reducing the editing cost for users.

[0108] In step S55, share with one click.

[0109] After editing the image, you can share it directly to the community. If the user clicks... Figure 12 The completion control in the middle, such as Figure 13 As shown, users can post the generated recommended image, and can also include some comments.131 Figure 14 As shown, after the recommended image is published in the community, other users can make further comments and interactions on the image, etc.

[0110] In the above embodiments, the link of image capture, image creation, intelligent recommendation of relevant frames and text, and one-click sharing is closed within a single product, without the need to switch between devices. This reduces the cost of cross-device editing and sharing. On the one hand, even if users do not have professional image editing skills, they can still create exquisite derivative content, which improves creative motivation and efficiency. On the other hand, it allows users to complete more production and publishing within the device, enriching the content to be shared, enhancing user stickiness with the platform, and reducing user churn.

[0111] The above is the content generation method provided by the embodiments of this disclosure. Below, refer to... Figures 15 to 17 Describes a content generation apparatus according to embodiments of the present disclosure.

[0112] Figure 15 A block diagram of a content generation apparatus according to some embodiments of the present disclosure is shown.

[0113] like Figure 15 As shown, the content generation device 15 includes a processing module 151, a recommendation module 152, and a content generation module 153. The processing module 151 is configured to, in response to a user's trigger operation, acquire first visual content from a video and display an editing interface; the recommendation module 152 is configured to, on the editing interface, recommend at least one candidate content related to the first visual content, wherein the at least one candidate content is determined based on at least one of a plot dimension, a character dimension, and a scene dimension in the first visual content, and the at least one candidate content comes from one or more of the video and interactive content related to the video; the content generation module 153 is configured to, in response to a target content determined by the user from the at least one candidate content, generate second visual content based on the first visual content and the target content.

[0114] In this embodiment, after the user triggers the operation, relevant candidate content is automatically recommended in the editing interface. After the user selects the target content from the candidate content, the created content is generated. The user does not need to have professional image editing skills and does not need to jump to a third-party platform for image retouching, which reduces the editing steps. Therefore, it reduces the difficulty and cost of creation and improves the motivation and efficiency of creation.

[0115] In some embodiments, the recommendation module 152 is configured to perform content understanding on the first visual content, determine at least one of a first key content in the plot dimension, a second key content in the character dimension, and a third key content in the scene dimension; and determine at least one candidate content in the video and / or the interactive content based on at least one of the first key content, the second key content, and the third key content.

[0116] In some embodiments, the recommendation module 152 is configured to perform content understanding on the first visual content to obtain content focus; in response to the content focus being related to event development or plot logic, determine the dimension related to the first visual content as the plot dimension; and determine the first key content of the first visual content in the plot dimension.

[0117] In some embodiments, the recommendation module 152 is configured to perform content understanding on the first visual content to obtain content focus; in response to the content focus being related to a character's image or personality psychology, determine the dimension related to the first visual content as the character dimension; and determine the second key content of the first visual content in the character dimension.

[0118] In some embodiments, the recommendation module 152 is configured to perform content understanding on the first visual content to obtain content focus; in response to the content focus being related to the environmental atmosphere or spatial setting, determine the dimension related to the first visual content as the scene dimension; and determine the third key content of the first visual content in the scene dimension.

[0119] In some embodiments, the recommendation module 152 is configured to determine key plot points related to the first visual content based on at least one of the first key content, the second key content, and the third key content; determine a target set of the key plot points in the video, the target set including one or more episodes; and recommend the at least one candidate content based on keyframes in the target set.

[0120] In some embodiments, the recommendation module 152 is configured to determine key plot points related to the first visual content based on at least one of the first key content, the second key content, and the third key content; and to determine at least one candidate content related to the key plot points in the interactive content.

[0121] In some embodiments, the recommendation module 152 is configured to determine the current storyline based on at least one of the character identity, character actions, character dialogue, and character clothing in the second key content, and to use the key storyline related to the current storyline as the key storyline related to the first visual content.

[0122] In some embodiments, the recommendation module 152 is configured to determine the current storyline based on at least one of environmental features, item features, prop features, and light and color features in the third key content, and to use the key storyline related to the current storyline as the key storyline related to the first visual content.

[0123] In some embodiments, the recommendation module 152 is configured to use key storylines related to the first key content as key storylines related to the first visual content.

[0124] In some embodiments, each of the at least one candidate content includes at least one of a key person and a key location, and the content generation module 153 is configured to, in response to the user selecting at least one of the key person and the key location, display reference visual content related to at least one of the key person and the key location; and in response to the user determining the target content in the reference visual content, generate second visual content based on the first visual content and the target content, wherein the target content includes one or more of the reference visual content.

[0125] In some embodiments, the content generation module 153 is configured to synthesize the first visual content and the target content according to the time information to generate the second visual content.

[0126] In some embodiments, the at least one candidate content includes at least one candidate text content, which is derived from one or more of dialogue in the video and interactive content related to the video. The content generation module 153 is configured to generate the second visual content based on the first visual content and the target text content in response to the user selecting target text content from the at least one candidate text content.

[0127] In some embodiments, the content generation apparatus further includes a tool recommendation module (not shown) configured to recommend the editing tools to the user based on at least one of the image style, image color, and image content of the first visual content.

[0128] In some embodiments, the content generation apparatus further includes a sharing module configured to share the second visual content in response to a user's triggering operation on a sharing control of the editing interface.

[0129] In some embodiments, the first visual content includes image information but excludes non-image information; and / or the first visual content is a frame of an image or a video clip from the video.

[0130] It should be noted that the above-described units are merely logical modules divided according to their specific functions, and are not intended to limit the specific implementation method. For example, they can be implemented in software, hardware, or a combination of both. In actual implementation, the above-described units can be implemented as independent physical entities, or they can be implemented by a single entity (e.g., a processor (CPU or DSP, etc.), integrated circuit, etc.). Furthermore, the units shown in the accompanying drawings with dashed lines indicate that these units may not actually exist, and the operations / functions they perform can be implemented by the processing circuitry itself.

[0131] The system configuration can be achieved through electronic devices. Figure 16 A block diagram of an electronic device according to some embodiments of the present disclosure is shown.

[0132] like Figure 16 As shown, the electronic device 16 includes: a processor 162; and a memory 161 coupled to the processor for storing instructions, which, when executed by the processor, cause the processor to perform the content generation method as described above.

[0133] Memory 161 is used to store one or more computer-readable instructions. Memory 161 may include any combination of various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory, including but not limited to random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), read-only memory (ROM), and flash memory. Memory 161 may, for example, store operating systems, application programs, bootloaders, databases, and other programs, as well as various application programs and various data.

[0134] The processor 162 is configured to execute computer-readable instructions to implement the content generation method described in any of the foregoing embodiments.

[0135] The specific implementation of each step of the method can be found in the above embodiments, and the repeated parts will not be described again here.

[0136] Processor 162 can be configured to execute Figures 1 to 5 The processor 162 can be various processing devices, such as a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The central processing unit (CPU) can be based on x816 or ARM architecture, etc.

[0137] The processor 162 and the memory 161 can communicate with each other directly or indirectly. For example, the processor 162 and the memory 161 can communicate via a network. The network can include a wireless network, a wired network, and / or any combination of wireless and wired networks. The processor 162 and the memory 161 can also communicate with each other via a system bus, which is not limited in this disclosure.

[0138] It should be noted that Figure 16 The components of the electronic device 16 shown are merely exemplary and not limiting; the electronic device 16 may have other components as needed for the actual application. The processor 162 can control other components in the electronic device 16 to perform desired functions.

[0139] In the above embodiments, data instructions are stored in a memory and then processed by a processor, which improves creative motivation and efficiency.

[0140] Electronic device 16 can be implemented by software, firmware and / or hardware, and can be integrated into a device with relevant applications installed.

[0141] Figure 17 Block diagrams of electronic devices according to other embodiments of the present disclosure are shown.

[0142] Figure 17 The electronic device 17 shown can be a computer system with a dedicated hardware structure, which can perform corresponding functions when the relevant application is installed.

[0143] Electronic devices include, but are not limited to, mobile terminals such as smartphones, laptops, personal digital assistants (PDAs), tablet PCs, portable multimedia players (PMPs), in-vehicle terminals (such as in-vehicle navigation terminals), wearable devices, and fixed terminals such as digital televisions and desktop computers.

[0144] like Figure 17 As shown, the Central Processing Unit (CPU) 171 performs various processes based on a program stored in the Read-Only Memory (ROM) 172 or a program loaded from the storage section 178 into the Random Access Memory (RAM) 173. The RAM 173 stores data required as needed when the CPU 171 performs various processes, etc. The CPU is merely exemplary; it could also be other types of processors, such as the various processors described above. The ROM 172, RAM 173, and storage section 178 can be various forms of computer-readable storage media. It should be noted that although... Figure 17 The diagram shows ROM 172, RAM 173 and storage section 178, but one or more of them may be combined or located in the same or different memory or storage modules.

[0145] CPU 171, ROM 172 and RAM 173 are interconnected via bus 174. Input / output interface 175 is also connected to bus 174.

[0146] The following components are connected to the input / output interface 175: input section 176, such as a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output section 177, including displays such as cathode ray tube (CRT), liquid crystal display (LCD), speakers, vibrators, etc.; storage section 178, including hard disks, magnetic tapes, etc.; and communication section 179, including network interface cards such as LAN cards, modems, etc. Communication section 179 allows communication processing via a network such as the Internet. It is easy to understand that, although... Figure 17 The portion of the electronic device 17 shown communicates via bus 174, but it may also communicate via a network or other means, wherein the network may include a wireless network, a wired network, and / or any combination of wireless and wired networks.

[0147] As needed, drive 1710 is also connected to input / output interface 175. Removable media 1711, such as disks, optical disks, magneto-optical disks, semiconductor memories, etc., are installed on drive 1710 as needed, so that computer programs read from them can be installed into storage section 178 as needed.

[0148] When the above series of processes are implemented through software, the program constituting the software can be installed from a network such as the Internet or from a storage medium such as a removable medium 1711.

[0149] According to one aspect of this disclosure, a computer-readable storage medium is provided that stores computer instructions thereon, wherein the computer instructions, when executed by a processor, implement the above-described content generation method.

[0150] This computer-readable storage medium can enhance creative motivation and efficiency.

[0151] According to one aspect of this disclosure, a computer program product is provided, comprising: computer instructions that, when executed by a processor, implement the above-described content generation method.

[0152] This computer program product can enhance creative motivation and efficiency.

[0153] According to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of this disclosure include a computer program product that, when run on a computer, causes the computer to perform the methods described in any of the foregoing embodiments. The computer program product includes computer instructions carried on a computer-readable medium, containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer instructions can be downloaded and installed from a network via communication section 179, or installed from storage section 178, or installed from ROM 172. When the computer program is executed by CPU 171, the methods of embodiments of this disclosure are performed.

[0154] It should be noted that, in the context of this disclosure, a computer-readable medium can be a tangible medium that may contain or store programs for use by or in conjunction with an instruction execution system, apparatus, or device.

[0155] A computer-readable medium may be a computer-readable storage medium, a computer-readable signal medium, or any combination thereof, wherein the computer-readable storage medium is a non-transitory computer storage medium.

[0156] Computer-readable storage media include, but are not limited to, systems, apparatuses, or devices that are electrical, magnetic, optical, electromagnetic, infrared, or semiconductor, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. Computer instructions are stored on the computer-readable storage medium that, when executed by a processor, implement the methods described in any of the foregoing embodiments.

[0157] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, capable of sending, propagating, or transmitting programs for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0158] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0159] In some embodiments, a computer program is also provided, comprising: instructions that, when executed by a processor, cause the processor to perform the methods described in any of the foregoing embodiments. For example, the instructions may be embodied in computer program code.

[0160] In embodiments of this disclosure, computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof. These programming languages ​​include, but are not limited to, object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network (including a local area network (LAN) or a wide area network (WAN)), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0161] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0162] The functions described above can be performed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary hardware logic components that can be used include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0163] While specific embodiments of this disclosure have been described in detail by way of example, those skilled in the art should understand that the examples are for illustrative purposes only and not intended to limit the scope of this disclosure. Those skilled in the art should understand that modifications can be made to the above embodiments without departing from the scope and spirit of this disclosure. The scope of this disclosure is defined by the appended claims.

Claims

1. A content generation method, comprising: In response to user-triggered actions, the system captures the first-person view content from the video and displays the editing interface. In the editing interface, at least one candidate content related to the first visual content is recommended. The at least one candidate content is determined based on the understanding of the plot dimension, character dimension and scene dimension in the first visual content. The at least one candidate content comes from one or more of the video and interactive content related to the video. The at least one candidate content is recommended based on key plot points related to the first visual content. In response to the target content determined by the user among the at least one candidate content, a second visual content is generated based on the first visual content and the target content.

2. The content generation method of claim 1, wherein, The recommended candidate content related to the first visual content includes at least one of the following: Perform content understanding on the first visual content to determine the first key content of the first visual content in the plot dimension, the second key content in the character dimension, and the third key content in the scene dimension; Based on the first key content, the second key content, and the third key content, at least one candidate content is determined in the video and / or the interactive content.

3. The content generation method of claim 2, wherein, The step of performing content understanding on the first visual content and determining the first key content of the first visual content in the plot dimension, the second key content in the character dimension, and the third key content in the scene dimension includes: Perform content understanding on the first visual content to obtain the content focus; In response to the fact that the content focus is related to the development of events or the plot logic, the dimension related to the first visual content is determined as the plot dimension; Determine the first key content of the first visual content in the plot dimension.

4. The content generation method of claim 2, wherein, The step of performing content understanding on the first visual content and determining the first key content of the first visual content in the plot dimension, the second key content in the character dimension, and the third key content in the scene dimension includes: Perform content understanding on the first visual content to obtain the content focus; In response to the fact that the content focus is related to the character's image or personality psychology, the dimension related to the first visual content is determined as the character dimension; Determine the second key content of the first visual content in the character dimension.

5. The content generation method of claim 2, wherein, The step of performing content understanding on the first visual content and determining the first key content of the first visual content in the plot dimension, the second key content in the character dimension, and the third key content in the scene dimension includes: Perform content understanding on the first visual content to obtain the content focus; In response to the fact that the content focus is related to the environmental atmosphere or spatial setting, the dimension related to the first visual content is determined as the scene dimension; Determine the third key content of the first visual content in the scene dimension.

6. The content generation method of claim 2, wherein, The determination of at least one candidate content in the video based on the first key content, the second key content, and the third key content includes: Based on the first key content, the second key content, and the third key content, determine the key plot points related to the first visual content; Determine the target set of the key plot points in the video, the target set including one or more episodes; Based on the keyframes in the target set, at least one candidate content is recommended.

7. The content generation method of claim 2, wherein, Based on the first key content, the second key content, and the third key content, at least one candidate content is determined from the interactive content: Based on the first key content, the second key content, and the third key content, determine the key plot points related to the first visual content; In the interactive content, at least one candidate content related to the key plot is identified.

8. The content generation method according to claim 6 or 7, wherein, The determination of key plot points related to the first visual content based on the first key content, the second key content, and the third key content includes: Based on at least one of the following in the second key content: character identity, character actions, character dialogue, and character clothing, the current plot is determined, and the key plot related to the current plot is taken as the key plot related to the first visual content. Based on at least one of the environmental features, item features, prop features, and light and color features in the third key content, the current plot is determined, and the key plot related to the current plot is taken as the key plot related to the first visual content; The key plot points related to the first key content will be considered as the key plot points related to the first visual content.

9. The content generation method according to claim 1, wherein, Each of the at least one candidate content includes at least one of key people and key locations, and generating second visual content based on the first visual content and the target content in response to the target content determined by the user in the at least one candidate content includes: In response to the user selecting at least one of the key person and the key location, display reference visual content related to at least one of the key person and the key location; In response to the user determining the target content in the reference visual content, a second visual content is generated based on the first visual content and the target content, wherein the target content includes one or more of the reference visual content.

10. The content generation method according to claim 1, wherein, Each of the at least one candidate content includes time information, and generating the second visual content based on the first visual content and the target content includes: Based on the time information, the first visual content and the target content are synthesized to generate the second visual content.

11. The content generation method according to claim 1, wherein, The at least one candidate content includes at least one candidate text content, which is derived from one or more of dialogue in the video and interactive content related to the video. The step of generating second visual content based on the first visual content and the target content, in response to the target content determined by the user from the at least one candidate content, includes: In response to the user selecting target text content from the at least one candidate text content, the second visual content is generated based on the first visual content and the target text content.

12. The content generation method according to any one of claims 1 to 7, 9 to 11, further comprising: Based on at least one of the image style, image color, and image content of the first visual content, an editing tool is recommended to the user.

13. The content generation method according to any one of claims 1 to 7, 9 to 11, further comprising: In response to the user's triggering operation of the share control in the editing interface, the second visual content is shared.

14. The content generation method according to any one of claims 1 to 7, 9 to 11, wherein, The first visual content includes image information but excludes non-image information; and / or The first visual content is a frame of an image or a video clip from the video.

15. A content generation apparatus, comprising: The processing module is configured to respond to user-triggered operations, acquire first-view content from the video, and display the editing interface; The recommendation module is configured to recommend at least one candidate content related to the first visual content in the editing interface. The at least one candidate content is determined based on the understanding of the plot dimension, character dimension and scene dimension in the first visual content. The at least one candidate content comes from one or more of the video and interactive content related to the video. The at least one candidate content is recommended based on key plot points related to the first visual content. The content generation module is configured to generate second visual content based on the first visual content and the target content in response to a target content determined by the user among the at least one candidate content.

16. An electronic device comprising: processor; as well as A memory coupled to the processor is used to store instructions that, when executed by the processor, cause the processor to perform the content generation method as described in any one of claims 1 to 14.

17. A computer-readable storage medium having stored thereon computer instructions, wherein, When executed by a processor, the computer instructions implement the content generation method according to any one of claims 1 to 14.

18. A computer program product comprising: Includes computer instructions that, when executed by a processor, implement the content generation method according to any one of claims 1 to 14.

Citation Information

Patent Citations

  • Video processing method and device, electronic equipment and storage medium

    CN113645482A

  • Video processing method and device, electronic equipment and computer readable storage medium

    CN116471451A

  • Picture editing method, related equipment and system

    CN120196252A