Content generation method and device, electronic equipment, storage medium and program product

By recommending candidate content related to the plot, characters and scenes in the video and automatically generating second visual content, the problem of high difficulty in creating secondary works in the existing technology is solved, and efficient and low-cost creation effects are achieved.

CN120602745AActive Publication Date: 2025-09-05DOUYIN VISION CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510865574.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-25
Publication Date
2025-09-05
Estimated Expiration
2045-06-25

AI Technical Summary

Technical Problem

In the existing technology, when producing secondary creative works, a lot of effort needs to be spent on image editing, especially adding decorative elements that meet the needs of the plot, which increases the difficulty and cost of creation and reduces the efficiency of creation.

Method used

A content generation method is provided. By acquiring the visual content in the video, candidate content is recommended based on the plot dimension, character dimension and scene dimension, and the secondary visual content is automatically generated, reducing the user's dependence on third-party photo editing software.

Benefits of technology

It reduces the difficulty and cost of users' creation, improves their efficiency and motivation, and enables users to produce exquisite secondary creative content without the need for professional image editing skills.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120602745A_ABST
    Figure CN120602745A_ABST
Patent Text Reader

Abstract

The invention provides a content generation method and device, electronic equipment, a storage medium and a program product, and relates to the technical field of computers. The content generation method comprises the following steps: in response to a trigger operation of a user, obtaining first visual content in a video, and displaying an editing interface; at least one piece of candidate content related to the first visual content is recommended on the editing interface, and the at least one piece of candidate content is determined based on at least one of a plot dimension, a role dimension and a scene dimension in the first visual content, the at least one candidate content is from one or more of the video and interactive content related to the video; and in response to a target content determined by the user in the at least one candidate content, generating a second visual content according to the first visual content and the target content.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, and in particular to a content generation method, device, electronic device, storage medium, and program product. Background Art

[0002] With the development of the film and television industry and the rise in social media activity, more and more people are sharing and discussing derivative works of film and television works on social media. Related technologies require image editing using third-party photo editing software or painting tools such as Photoshop (image processing software), making production difficult and costly. Summary of the Invention

[0003] According to some embodiments of the present disclosure, a content generation method is provided, comprising: in response to a user's triggering operation, acquiring first visual content in a video and displaying an editing interface; on the editing interface, recommending at least one candidate content related to the first visual content, wherein the at least one candidate content is determined based on at least one of a plot dimension, a character dimension, and a scene dimension in the first visual content, and the at least one candidate content comes from one or more of the video and interactive content related to the video; in response to a target content determined by the user in the at least one candidate content, generating second visual content based on the first visual content and the target content.

[0004] According to other embodiments of the present disclosure, a content generation device is provided, including: a processing module, configured to obtain first visual content in a video and display an editing interface in response to a user's triggering operation; a recommendation module, configured to recommend at least one candidate content related to the first visual content on the editing interface, the at least one candidate content being determined based on at least one of a plot dimension, a character dimension, and a scene dimension in the first visual content, and the at least one candidate content being from one or more of the video and interactive content related to the video; a content generation module, configured to generate second visual content based on the first visual content and the target content in response to a target content determined by the user in the at least one candidate content.

[0005] According to some embodiments of the present disclosure, an electronic device is provided, comprising: a processor; and a memory coupled to the processor, for storing instructions, wherein when the instructions are executed by the processor, the processor executes the content generation method of any embodiment described in the present disclosure.

[0006] According to some embodiments of the present disclosure, a computer-readable storage medium is provided, on which computer instructions are stored. When the computer instructions are executed by a processor, the content generation method of any embodiment described in the present disclosure is implemented.

[0007] According to some embodiments of the present disclosure, a computer program product is provided, comprising: computer instructions, wherein when the computer instructions are executed by a processor, the content generating method of any embodiment described in the present disclosure is implemented.

[0008] Other features, aspects and advantages of the present disclosure will become apparent from the following detailed description of exemplary embodiments of the present disclosure with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] The following describes embodiments of the present disclosure with reference to the accompanying drawings. It should be understood that the drawings described below only relate to some embodiments of the present disclosure and do not constitute a limitation to the present disclosure. In the accompanying drawings:

[0010] Figure 1 A schematic diagram showing a flow chart of a content generation method according to some embodiments of the present disclosure;

[0011] Figure 2 A schematic diagram illustrating a process of recommending candidate content according to some embodiments of the present disclosure;

[0012] Figure 3 A schematic diagram showing a process of recommending candidate content according to some other embodiments of the present disclosure;

[0013] Figure 4 A schematic diagram illustrating a process of recommending candidate content according to yet other embodiments of the present disclosure;

[0014] Figure 5 A schematic diagram showing a flow chart of a content generation method according to some other embodiments of the present disclosure;

[0015] Figure 6 Showing video screenshots according to some embodiments of the present disclosure;

[0016] Figure 7 A schematic diagram illustrating an editing interface according to some embodiments of the present disclosure;

[0017] Figure 8 Schematic diagrams showing editing interfaces according to other embodiments of the present disclosure;

[0018] Figure 9 A schematic diagram showing an editing interface according to yet other embodiments of the present disclosure;

[0019] Figure 10 A schematic diagram showing an editing interface according to some further embodiments of the present disclosure;

[0020] Figure 11 A schematic diagram showing an editing interface according to some further embodiments of the present disclosure;

[0021] Figure 12 A schematic diagram showing an editing interface according to some further embodiments of the present disclosure;

[0022] Figure 13 A schematic diagram showing an editing interface according to some further embodiments of the present disclosure;

[0023] Figure 14 A schematic diagram showing an editing interface according to some further embodiments of the present disclosure;

[0024] Figure 15 A block diagram showing a content generation device according to some embodiments of the present disclosure;

[0025] Figure 16 A block diagram illustrating an electronic device according to some embodiments of the present disclosure is shown;

[0026] Figure 17 A block diagram of an electronic device according to some other embodiments of the present disclosure is shown.

[0027] It should be understood that, for ease of description, the dimensions of the various parts shown in the drawings are not necessarily drawn to scale. The same or similar reference numerals are used throughout the drawings to indicate the same or similar parts. Therefore, once an item is defined in one drawing, it may not be discussed further in subsequent drawings. DETAILED DESCRIPTION

[0028] The following will be combined with the accompanying drawings in the embodiments of the present disclosure to clearly and completely describe the technical solutions in the embodiments of the present disclosure. It should be understood that the present disclosure can be implemented in various forms and should not be interpreted as being limited to the embodiments described here.

[0029] It should be understood that the various steps described in the method embodiments of the present disclosure can be performed in different orders and / or performed in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect. Unless otherwise specifically stated, the relative arrangement, numerical expressions and numerical values ​​of the parts and steps set forth in these embodiments should be interpreted as being merely exemplary and do not limit the scope of the present disclosure.

[0030] The term “including” and its variations used in the present disclosure are open terms that include at least the following elements / features but do not exclude other elements / features, that is, “including but not limited to.” The term “based on” means “at least partially based on.”

[0031] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules, or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules, or units. Unless otherwise specified, concepts such as "first" and "second" are not intended to imply that the objects described in such a manner must be in a given order in time, space, ranking, or any other manner.

[0032] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".

[0033] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0034] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0035] The following detailed description of the embodiments of the present disclosure is provided in conjunction with the accompanying drawings, but the present disclosure is not limited to these specific embodiments. The following specific embodiments may be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments. In addition, in one or more embodiments, specific features, structures, or characteristics may be combined in any suitable manner that will be apparent to those skilled in the art from this disclosure.

[0036] In related technologies, when creating secondary creative works, users need to expend considerable effort to add decorative elements to images to suit the plot, and they also need to use third-party photo editing software to perform image editing, which increases the difficulty and cost of creation and reduces efficiency. This disclosure provides a content generation method that reduces the difficulty and cost of creation, allowing users without professional image editing skills to produce exquisite secondary creative content. The following describes the solution of this disclosure in conjunction with specific embodiments.

[0037] Figure 1 A flowchart illustrating a content generation method according to some embodiments of the present disclosure is shown.

[0038] like Figure 1As shown, the content generation method includes: step S1, in response to a user's triggering operation, obtaining the first visual content in the video and displaying an editing interface; step S2, in the editing interface, recommending at least one candidate content related to the first visual content, the at least one candidate content is determined based on at least one of the plot dimension, character dimension and scene dimension in the first visual content, and the at least one candidate content comes from the video, one or more interactive contents related to the video; step S3, in response to the target content determined by the user in the at least one candidate content, generating a second visual content based on the first visual content and the target content.

[0039] For example, when watching film and television content, the user captures a frame of image or a video and automatically enters the editing interface. In the editing interface, one or more candidate content is recommended and displayed. Each candidate content can be a frame of image or multiple frames in the video, or it can be a part of the interactive content. Interactive content includes comments, related posts, PGC (Professionally-generated Content), UGC (User-generated-Content), etc. Comments can be text, emoticons, pictures, animated pictures, videos, etc.

[0040] Among the recommended candidate contents, the user can select one or more target contents. The first visual content and the target content are synthesized to generate the second visual content, which can be a picture, a GIF, a video, etc.

[0041] In the above embodiment, the user does not need to spend a lot of energy looking for resources to edit images or videos. Instead, after the user triggers the operation, relevant candidate content is automatically recommended in the editing interface. After the user determines the target content among the candidate content, the creative content is generated. The user does not need to have professional image editing capabilities, and does not need to jump to a third-party platform for photo editing, which reduces the editing steps. Therefore, the difficulty and cost of creation are reduced, and the motivation and efficiency of creation are improved.

[0042] Next, we will combine Figures 2 to 14 , further introduces the content generation method disclosed in the present invention.

[0043] Figure 2 A schematic diagram illustrating a process of recommending candidate content according to some embodiments of the present disclosure is shown.

[0044] like Figure 2As shown, recommending at least one candidate content related to the first visual content includes: step S21, performing content understanding on the first visual content, and determining at least one of the first key content in the plot dimension, the second key content in the character dimension, and the third key content in the scene dimension for the first visual content; step S22, based on at least one of the first key content, the second key content, and the third key content, determining the at least one candidate content in the video and / or the interactive content.

[0045] For example, in the plot dimension, the first key content of the first visual content can be determined. This first key content may include, for example, the plot. In the character dimension, the second key content of the first visual content may be determined. This second key content may include, for example, the character's identity, character's actions, character's dialogue, character's clothing, etc. In the scene dimension, the third key content of the first visual content may be determined. This third key content may include, for example, environmental features, object features, prop features, lighting and color features, etc.

[0046] Based on the first key content, the second key content, and the third key content, multiple candidate content is determined. The candidate content can be from videos or interactive content. This recommendation process can be implemented using a large model.

[0047] In this embodiment, by understanding the content of the first visual content, the key content in the first visual content can be determined, and then candidate content can be obtained based on the key content, thereby realizing automatic recommendation of candidate content, facilitating the user's subsequent selection of target content, and realizing secondary creation of content.

[0048] In the above embodiment, the first visual content is a frame of image or a video clip in the video. The second visual content finally generated can also be a frame of image or a video clip, thereby enriching content creation.

[0049] Primary visual content includes image information but excludes non-image information. For example, after capturing images and video clips from film and television content, interactive features, navigation bars, episode information, and other content are automatically removed. This allows users to easily access relatively complete and unobstructed image or video resources, making subsequent editing and generating secondary visual content clear and complete, thereby improving the user's visual experience.

[0050] During the removal of non-image information, captured images and videos from the film and television content can be compared with the backend stored plot to obtain complete, unobstructed image or video resources. Alternatively, the client can hide interactive features, navigation bars, episode information, and other content, and then capture image or video clips to obtain relatively complete, unobstructed image or video resources.

[0051] In some embodiments, the performing content understanding on the first visual content and determining that the first visual content is at least one of the first key content in the plot dimension, the second key content in the character dimension, and the third key content in the scene dimension includes: performing content understanding on the first visual content to obtain a content focus; in response to the content focus being related to event development or plot logic, determining that the dimension related to the first visual content is the plot dimension; and determining the first key content of the first visual content in the plot dimension.

[0052] For example, understanding the primary visual content reveals a content focus, which reflects the dimension of the primary visual content that is prioritized. If this content focus can showcase story events, plot developments, or key turning points, emphasizing "what happened," meaning the content focus is related to the event development or plot logic, then the dimension associated with the primary visual content can be determined as the plot dimension.

[0053] For example, if the large-scale model analysis reveals that the captured image contains interactions between characters, such as dialogue, conflict, and cooperation; key actions that propel the storyline, such as decryption, combat, and escape; or important moments, such as memories, then the dimension associated with the primary visual content is the plot dimension. In this case, the primary visual content's key elements within the plot dimension need to be determined.

[0054] In the above embodiment, based on the content focus corresponding to the first visual content, it is determined that the dimension most relevant to the first visual content is the plot dimension, and then the key content of the plot dimension can be extracted to achieve recommendation of candidate content.

[0055] In some embodiments, the performing content understanding on the first visual content and determining that the first visual content is at least one of the first key content in the plot dimension, the second key content in the character dimension, and the third key content in the scene dimension includes: performing content understanding on the first visual content to obtain a content focus; in response to the content focus being related to the character image or personality psychology, determining that the dimension related to the first visual content is the character dimension; and determining the second key content of the first visual content in the character dimension.

[0056] For example, if the content focuses on shaping the character's image, personality or emotions, emphasizing "who the character is", that is, the focus content is related to the character's image or personality psychology, then the dimension related to the first visual content can be determined as the character dimension.

[0057] For example, if the large-scale model analysis reveals that the captured image contains a character's physical features, such as facial expressions, clothing, and iconic props; details that reflect the character's background or psychology; or information that reflects the character's growth stages or identity transitions, then the dimension associated with the primary visual content is the character dimension. In this case, the first key content of the primary visual content in the character dimension needs to be determined.

[0058] In the above embodiment, based on the content focus corresponding to the first visual content, the dimension most relevant to the first visual content is determined to be the role dimension, and then the key content of the role dimension can be extracted to achieve recommendation of candidate content.

[0059] In some embodiments, the performing content understanding on the first visual content and determining at least one of the first key content of the first visual content in the plot dimension, the second key content in the character dimension, and the third key content in the scene dimension includes: performing content understanding on the first visual content to obtain a content focus; in response to the content focus being related to the environmental atmosphere or spatial setting, determining that the dimension related to the first visual content is the scene dimension; and determining the third key content of the first visual content in the scene dimension.

[0060] For example, if the content focuses on constructing the environment, world view or atmosphere in which the story takes place, emphasizing "where it happens", the dimension related to the first visual content can be determined as the scene dimension.

[0061] For example, if a user captures an image through large-scale model analysis and finds that it contains the spatial structure of the environment, such as architectural layouts and natural landscapes; details of the historical context, such as the furnishings of ancient streets or the technological elements of future cities; or atmosphere, such as the oppressive feeling of rainy weather. This indicates that the dimension associated with the primary visual content is the scene dimension. In this case, it is necessary to determine the first key content of the primary visual content in the scene dimension.

[0062] In the above embodiment, based on the content focus corresponding to the first visual content, the dimension most relevant to the first visual content is determined to be the scene dimension, and then the key content of the scene dimension can be extracted to achieve recommendation of candidate content.

[0063] Those skilled in the art will appreciate that the content focus of an image or video may not be a single dimension, but rather a combination of multiple dimensions. For example, through content focus, the dimensions associated with the primary visual content can be determined to be the plot dimension and the character dimension, or the plot dimension and the scene dimension, or the character dimension and the scene dimension, or the plot dimension, the scene dimension, and the character dimension. Therefore, it is necessary to determine key content across multiple dimensions.

[0064] By using the proportion of screen elements, such as the proportion of character actions, environment, expressions, etc., we can determine the dimensions most relevant to the first visual content.

[0065] If the candidate content comes from a video, determine the candidate content as follows Figure 3 shown. Figure 3 A schematic diagram illustrating a process of recommending candidate content according to some other embodiments of the present disclosure.

[0066] like Figure 3 As shown, the determining of the at least one candidate content in the video based on at least one of the first key content, the second key content and the third key content includes: step S2211, determining the key plot related to the first visual content based on at least one of the first key content, the second key content and the third key content; step S2212, determining the target set of the key plot in the video, the target set including one or more episodes; step S2213, recommending the at least one candidate content based on the key frames in the target set.

[0067] For example, if the key content of a certain frame in a video is a campus sports meet, then the key plot related to that frame might be athlete selection, athlete awards, etc. We then determine a target set of key plots in the video. This target set can be a single episode or multiple episodes. The key frames in the target set that are related to the key plot are recommended as candidate content. This key frame can be a single frame or multiple frames.

[0068] In this embodiment, the position of the key plot related to the first visual content in the video can be quickly located, and then related candidate content can be recommended. In addition, since the key plot may exist in multiple episodes of the video, the cost of users taking screenshots across episodes is also reduced.

[0069] In some embodiments, determining the key plot related to the first visual content based on at least one of the first key content, the second key content, and the third key content includes: determining the current plot based on at least one of the character identity, character action, character dialogue, and character costume in the second key content, and using the key plot related to the current plot as the key plot related to the first visual content.

[0070] For example, identifying basic character information, such as identity and clothing, can lay the foundation for plot analysis; analyzing character actions or interactions can capture plot conflicts and developments; and character dialogue can link character motivations with plot logic. In this embodiment, based on character identity, action, dialogue, and clothing, the current plot of the primary visual content can be determined, and then key plot points related to the current plot can be identified. These key plot points may be similar to the current plot, reminiscent plot points, future development plot points, or plot points with twists and turns. In this way, candidate content can be recommended to users based on the character dimension.

[0071] In some embodiments, determining the key plot related to the first visual content based on at least one of the first key content, the second key content, and the third key content includes: determining the current plot based on at least one of the environmental features, object features, prop features, and light and color features in the third key content, and using the key plot related to the current plot as the key plot related to the first visual content.

[0072] For example, environmental features, object features, prop features, and lighting and color features can reflect the spatiotemporal coordinates, plot clues, plot logic, and capture the plot's driving force. Therefore, based on these features, we can determine the current plot and, in turn, identify key plot points related to the current plot. This allows us to recommend candidate content to users based on the context of the scene.

[0073] In some embodiments, determining the key plot related to the first visual content based on at least one of the first key content, the second key content and the third key content includes: using the key plot related to the first key content as the key plot related to the first visual content.

[0074] Since the first key content reflects the current plot, the key plots related to the current plot can be directly determined. In this way, candidate content can be recommended to the user from the plot dimension.

[0075] If the candidate content comes from interactive content, determine the candidate content as follows Figure 4 shown. Figure 4 A schematic diagram illustrating a process of recommending candidate content according to yet other embodiments of the present disclosure.

[0076] like Figure 4As shown, based on at least one of the first key content, the second key content and the third key content, the at least one candidate content is determined in the interactive content: Step S2221, based on at least one of the first key content, the second key content and the third key content, the key plot related to the first visual content is determined; Step S2222, in the interactive content, the at least one candidate content related to the key plot is determined.

[0077] The specific method for determining a key plot related to the first visual content based on at least one of the first key content, the second key content, and the third key content has been described in the above embodiments and will not be further elaborated here. Once the key plot has been determined, candidate content can be identified from commentary on the key plot, or from images and videos created by other users based on the key plot. This reduces the time users spend searching for relevant content, thereby improving creative efficiency.

[0078] After the large model analyzes the candidate content related to the first visual content, it can be displayed to the user. At this time, there are many candidate contents displayed, so the user can further filter and select them, thereby achieving key recommendations.

[0079] In some embodiments, each of the at least one candidate content includes at least one of a key person and a key place, and the generating of the second visual content according to the first visual content and the target content in response to the target content determined by the user in the at least one candidate content includes: in response to the user selecting at least one of the key person and the key place, displaying reference visual content related to at least one of the key person and the key place; in response to the user determining the target content in the reference visual content, generating the second visual content according to the first visual content and the target content, the target content including one or more of the reference visual contents.

[0080] For example, each candidate content item displays a key person or a key location. If the user selects a key person, all reference visual content items that are filtered out will include that key person. If the user selects a key location, all locations in the filtered reference visual content will be related to that key location. If the user selects both a key person and a key location, all locations in the filtered reference visual content will be related to that key location and include that key person, making filtering easier for the user.

[0081] In some embodiments, in response to a user selecting a candidate content from at least one candidate content, the candidate content is spliced ​​with the first visual content; in response to a user selecting at least one of a key person and a key place related to the candidate content, reference visual content related to at least one of the key person and the key place is displayed; in response to the user determining the target content in the reference visual content, the candidate content, the first visual content and the target content are spliced ​​to obtain the second visual content.

[0082] For example, after the system recommends multiple candidate contents, the user clicks on a candidate content, and the candidate content and the original screenshot are spliced ​​into a frame image. After the user clicks on information such as a person, it jumps to the secondary page and displays the recommended frames related to the person. The user can then select again, which facilitates user screening.

[0083] Each candidate content in the at least one candidate content includes time information. After the target content is determined, the first visual content and the target content are synthesized according to the time information to generate the second visual content.

[0084] For example, the episode number and time of the frame are displayed in the recommended key frame. In this way, multiple frames of images can be synthesized according to time, so that the synthesized second visual content has a time sequence and is more logical, which promotes narrative and story understanding.

[0085] In some embodiments, the at least one candidate content includes at least one candidate text content, and the at least one candidate text content comes from one or more of the lines in the video and the interactive content related to the video. The generating of the second visual content according to the first visual content and the target content in response to the target content determined by the user in the at least one candidate content includes: generating the second visual content according to the first visual content and the target text content in response to the user selecting the target text content in the at least one candidate text content.

[0086] The candidate text content is, for example, lines or internet buzzwords. The user can select the target text content, which is added to the first visual content to form the second visual content, so that the final generated content contains both image content and text content, making the content vivid and vivid while reducing the user's editing costs.

[0087] After generating the second visual content, the system can also recommend editing tools to the user to facilitate further editing of the content. For example, the editing tool is recommended to the user based on at least one of the image style, image color, and image content of the first visual content.

[0088] Editing tools include, for example, filters and painting styles. For example, if the image content in the first visual content includes scenery, filters that enhance color saturation and contrast can be recommended, or artistic painting styles can be recommended to give the image a unique artistic texture. For another example, if the image content in the first visual content includes images of people, filters that beautify skin tone and soften skin texture can be recommended, or painting styles that can transform people into cartoons, comics, and other personalized styles can be recommended. For another example, if the image content in the first visual content is retro in style, filters that add a retro atmosphere can be recommended, or retro painting styles can be recommended. For another example, if the image content in the first visual content is fresh in style, elegant filters can be recommended, or painting styles that make the image full of freshness and artistic atmosphere can be recommended. For another example, if the first visual content is a warm-toned image, warm-toned filters can be recommended to enhance the warm atmosphere of the image. If the first visual content is a cool-toned image, cool-toned filters can be recommended to give the image a cool aesthetic, and so on.

[0089] In the above embodiment, by adding editing tools to visual content, it is convenient for users to select appropriate editing tools to process images, thereby improving the emotional expression of the content and enhancing artistic creation.

[0090] After the system generates the second visual content, the system shares the second visual content in response to a user triggering an operation on a sharing control of the editing interface.

[0091] In related technologies, users need to edit images in third-party photo editing software and then share them on other social platforms. However, in this embodiment, after completing content editing in the editing interface, users can directly share it on the social platform through the sharing control, reducing cross-platform editing and sharing costs.

[0092] The content generation method of the present disclosure will be further introduced below with reference to a specific embodiment.

[0093] Figure 5 A flowchart illustrating a content generation method according to some other embodiments of the present disclosure is shown.

[0094] like Figure 5 As shown, in step 51, a screenshot of the video content is taken and the screen is automatically cleared.

[0095] For example, when watching a short drama, the user takes a screenshot, such as Figure 6 As shown, the image obtained after the screenshot automatically removes information such as the interactive function 61, the navigation bar 62, and the episode information 63, so that the user can obtain a relatively complete and unobstructed image resource.

[0096] In step S52, an editing interface is displayed.

[0097] For example, Figure 7 As shown, enter the picture editing process. Figure 7 A schematic diagram of an editing interface according to some embodiments of the present disclosure is shown. The editing interface can display related stills and can also add text. Recommended related stills are shown in 71.

[0098] Through the above operations, you can quickly and easily enter the drawing process, reducing the tedious process of users entering a third party to draw and then coming back to share.

[0099] In step S53, key frames related to the screenshot are recommended.

[0100] For example, based on a screenshot, the system intelligently identifies image elements and recommends relevant keyframes based on plot, characters, and scenes. The system not only selects keyframes near the current progress but also recommends related frames across episodes. For example, if the screenshot is of the plum blossom viewing scene from episode 3 of the work "**," the system can identify elements such as the red plum blossoms and characters in the image and identify subsequent key plot points such as XXXX, YYYY, and ZZZZ, and select and recommend the corresponding frames. Users can also select "Add stills" to stitch the images together.

[0101] Each keyframe displays the episode number and time of the series it appears in, as well as key characters or locations in the image. Clicking on a key character or location will take you to a secondary page displaying recommended frames related to that character or scene, making it easier for users to filter.

[0102] like Figure 8 As shown in FIG, after the user clicks on a recommended still photo 81, the two photos are spliced ​​and synthesized. After the user clicks on information such as characters, the still photos related to the content can be identified, such as Figure 9 As shown in , stills 91 related to the character are recommended. Figure 10 As shown, after the user clicks on the recommended still 101, image synthesis is further performed.

[0103] In this step, relevant stills are intelligently recommended, reducing the cost of users taking screenshots across episodes.

[0104] In step S54, text related to the screenshot is recommended.

[0105] For example, based on a screenshot, it can intelligently identify image elements and recommend relevant highlights, internet buzzwords, etc. based on the plot, characters, and scene information. It also supports adding corresponding text for secondary editing of images.

[0106] like Figure 11 As shown, the user clicks to add text, and the user is shown the highlights of the lines related to the picture, hot topics 111, etc. Figure 12 As shown, text 121 is added to the image.

[0107] In this step, relevant hot memes are intelligently recommended, reducing user editing costs.

[0108] In step S55, one-click sharing.

[0109] After the picture is edited, it can be shared directly to the community. Figure 12 The completion control in Figure 13 As shown, the user can publish the generated recommended image and can also carry some comments 131. Figure 14 As shown, after the recommended image is published in the community, other users can make further comments and interact with the image 141 .

[0110] In the above embodiment, image capture, image production, intelligent recommendation of related frames and texts, and one-click sharing links are closed into one product and can be completed without jumping to another end, thereby reducing cross-end editing and sharing costs. On the one hand, even if the user does not have professional image editing capabilities, they can still produce exquisite secondary creative content, which improves creative motivation and efficiency. On the other hand, it allows users to complete production and publishing more within the end, enriches the shared content, enhances the stickiness between users and the platform, and reduces user churn.

[0111] The above is the content generation method provided by the embodiment of the present disclosure. Figures 15 to 17 A content generating apparatus according to an embodiment of the present disclosure is described.

[0112] Figure 15 A block diagram of a content generating device according to some embodiments of the present disclosure is shown.

[0113] like Figure 15 As shown, the content generation device 15 includes a processing module 151, a recommendation module 152 and a content generation module 153. The processing module 151 is configured to obtain the first visual content in the video and display an editing interface in response to a user's triggering operation; the recommendation module 152 is configured to recommend at least one candidate content related to the first visual content on the editing interface, the at least one candidate content being determined based on at least one of the plot dimension, the character dimension and the scene dimension in the first visual content, and the at least one candidate content being from one or more of the video and the interactive content related to the video; the content generation module 153 is configured to generate second visual content based on the first visual content and the target content in response to the target content determined by the user in the at least one candidate content.

[0114] In this embodiment, after the user triggers the operation, relevant candidate content is automatically recommended in the editing interface. After the user determines the target content among the candidate content, the creative content is generated. The user does not need to have professional image editing capabilities and does not need to jump to a third-party platform for photo editing, which reduces the editing steps. Therefore, the difficulty and cost of creation are reduced, and the motivation and efficiency of creation are improved.

[0115] In some embodiments, the recommendation module 152 is configured to perform content understanding on the first visual content, determine at least one of the first key content in the plot dimension, the second key content in the character dimension, and the third key content in the scene dimension of the first visual content; and determine the at least one candidate content in the video and / or the interactive content based on at least one of the first key content, the second key content, and the third key content.

[0116] In some embodiments, the recommendation module 152 is configured to perform content understanding on the first visual content to obtain a content focus; in response to the content focus being related to event development or plot logic, determine that the dimension related to the first visual content is the plot dimension; and determine the first key content of the first visual content in the plot dimension.

[0117] In some embodiments, the recommendation module 152 is configured to perform content understanding on the first visual content to obtain a content focus; in response to the content focus being related to the character image or personality psychology, determine that the dimension related to the first visual content is the character dimension; and determine the second key content of the first visual content in the character dimension.

[0118] In some embodiments, the recommendation module 152 is configured to perform content understanding on the first visual content to obtain a content focus; in response to the content focus being related to the environmental atmosphere or spatial setting, determine that the dimension related to the first visual content is the scene dimension; and determine the third key content of the first visual content in the scene dimension.

[0119] In some embodiments, the recommendation module 152 is configured to determine a key plot related to the first visual content based on at least one of the first key content, the second key content, and the third key content; determine a target set of the key plot in the video, the target set including one or more episodes; and recommend the at least one candidate content based on the key frames in the target set.

[0120] In some embodiments, the recommendation module 152 is configured to determine a key plot related to the first visual content based on at least one of the first key content, the second key content, and the third key content; and to determine at least one candidate content related to the key plot in the interactive content.

[0121] In some embodiments, the recommendation module 152 is configured to determine the current plot based on at least one of the character identity, character action, character dialogue, and character costume in the second key content, and use the key plot related to the current plot as the key plot related to the first visual content.

[0122] In some embodiments, the recommendation module 152 is configured to determine the current plot based on at least one of the environmental features, object features, prop features, and light and color features in the third key content, and use the key plot related to the current plot as the key plot related to the first visual content.

[0123] In some embodiments, the recommendation module 152 is configured to use a key plot related to the first key content as the key plot related to the first visual content.

[0124] In some embodiments, each of the at least one candidate content includes at least one of a key person and a key place, and the content generation module 153 is configured to display reference visual content related to at least one of the key person and the key place in response to the user selecting at least one of the key person and the key place; and generate the second visual content based on the first visual content and the target content in response to the user determining the target content in the reference visual content, wherein the target content includes one or more of the reference visual contents.

[0125] In some embodiments, the content generation module 153 is configured to synthesize the first visual content and the target content according to the time information to generate the second visual content.

[0126] In some embodiments, the at least one candidate content includes at least one candidate text content, and the at least one candidate text content comes from one or more of the lines in the video and the interactive content related to the video. The content generation module 153 is configured to generate the second visual content based on the first visual content and the target text content in response to the user selecting the target text content from the at least one candidate text content.

[0127] In some embodiments, the content generation device further includes a tool recommendation module (not shown) configured to recommend the editing tool to the user based on at least one of the image style, image color, and image content of the first visual content.

[0128] In some embodiments, the content generating device further includes a sharing module configured to share the second visual content in response to a user triggering operation on a sharing control of the editing interface.

[0129] In some embodiments, the first visual content includes image information but does not include non-image information; and / or the first visual content is a frame of image or a video clip in the video.

[0130] It should be noted that the above-mentioned units are merely logical modules divided according to the specific functions they implement, and are not intended to limit specific implementation methods. For example, they can be implemented in software, hardware, or a combination of software and hardware. In actual implementation, the above-mentioned units can be implemented as independent physical entities, or can also be implemented by a single entity (for example, a processor (CPU or DSP, etc.), an integrated circuit, etc.). In addition, the above-mentioned units are shown with dotted lines in the accompanying drawings to indicate that these units may not actually exist, and the operations / functions they implement can be implemented by the processing circuit itself.

[0131] The system configuration device can be realized by electronic equipment, Figure 16 A block diagram of an electronic device according to some embodiments of the present disclosure is shown.

[0132] like Figure 16 As shown, the electronic device 16 includes: a processor 162; and a memory 161 coupled to the processor, for storing instructions, which, when executed by the processor, enable the processor to execute the above-mentioned content generation method.

[0133] Memory 161 is used to store one or more computer-readable instructions. Memory 161 may include any combination of various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory, including but not limited to random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), read-only memory (ROM), and flash memory. Memory 161 may, for example, store an operating system, application programs, a boot loader, a database, and other programs, as well as various application programs and data.

[0134] The processor 162 is configured to execute computer-readable instructions to implement the content generation method described in any of the aforementioned embodiments.

[0135] The specific implementation of each step of the method can be found in the above embodiments, and the repeated parts will not be repeated here.

[0136] The processor 162 may be configured to execute Figures 1 to 5 The processor 162 may be embodied as various processing devices, such as a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The central processing unit (CPU) may be an X816 or ARM architecture, etc.

[0137] The processor 162 and the memory 161 can communicate with each other directly or indirectly. For example, the processor 162 and the memory 161 can communicate via a network. The network can include a wireless network, a wired network, and / or any combination of wireless and wired networks. The processor 162 and the memory 161 can also communicate with each other via a system bus, which is not limited in this disclosure.

[0138] It should be noted that Figure 16 The components of the electronic device 16 shown are merely exemplary and non-limiting. The electronic device 16 may also have other components according to actual application requirements. The processor 162 may control the other components in the electronic device 16 to perform desired functions.

[0139] In the above embodiment, data instructions are stored in a memory and then processed by a processor, thereby improving creative power and efficiency.

[0140] The electronic device 16 may be implemented in software, firmware and / or hardware, and may be integrated into a device installed with relevant application programs.

[0141] Figure 17 A block diagram of an electronic device according to some other embodiments of the present disclosure is shown.

[0142] Figure 17 The electronic device 17 shown may be a computer system with a dedicated hardware structure, which can execute corresponding functions when a relevant application program is installed.

[0143] Electronic devices include but are not limited to mobile terminals such as smart phones, laptops, personal digital assistants (PDAs), tablet personal computers (Tablet PCs), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), wearable devices, etc., as well as fixed terminals such as digital televisions, desktop computers, etc.

[0144] like Figure 17 As shown, the central processing unit (CPU) 171 performs various processes according to the program stored in the read-only memory (ROM) 172 or the program loaded from the storage part 178 to the random access memory (RAM) 173. In the RAM 173, data required when the CPU 171 performs various processes is stored as needed. The central processing unit is only exemplary, and it can also be other types of processors, such as the various processors described above. The ROM 172, RAM 173 and the storage part 178 can be various forms of computer-readable storage media. It should be noted that although Figure 17 ROM 172, RAM 173 and storage portion 178 are shown separately in FIG, but one or more of them may be combined or located in the same or different memory or storage modules.

[0145] The CPU 171, the ROM 172, and the RAM 173 are connected to one another via a bus 174. An input / output interface 175 is also connected to the bus 174.

[0146] The following components are connected to the input / output interface 175: an input portion 176 such as a touch screen, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output portion 177 including a display such as a cathode ray tube (CRT), a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage portion 178 including a hard disk, a magnetic tape, etc.; and a communication portion 179 including a network interface card such as a LAN card, a modem, etc. The communication portion 179 allows communication processing to be performed via a network such as the Internet. It is easy to understand that although Figure 17 Some of the electronic devices 17 are shown to communicate via bus 174, but they can also communicate via a network or other means, where the network can include a wireless network, a wired network, and / or any combination of wireless networks and wired networks.

[0147] A drive 1710 is also connected to the input / output interface 175 as needed. A removable medium 1711 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is mounted on the drive 1710 as needed so that a computer program read therefrom is installed in the storage section 178 as needed.

[0148] When the series of processes described above is implemented by software, the program constituting the software can be installed from a network such as the Internet or a storage medium such as the removable medium 1711 .

[0149] According to one aspect of the present disclosure, a computer-readable storage medium is provided, on which computer instructions are stored. When the computer instructions are executed by a processor, the above-mentioned content generation method is implemented.

[0150] The computer-readable storage medium can enhance creative motivation and efficiency.

[0151] According to one aspect of the present disclosure, a computer program product is provided, comprising: computer instructions, wherein when the computer instructions are executed by a processor, the above-mentioned content generation method is implemented.

[0152] This computer program product can improve creative motivation and efficiency.

[0153] According to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, some embodiments of the present disclosure include a computer program product that, when the computer program product is run on a computer, causes the computer to implement the method described in any of the aforementioned embodiments. The computer program product includes computer instructions carried on a computer-readable medium, containing program code for executing the method shown in the flowchart. In such an embodiment, the computer instructions can be downloaded and installed from the network through the communication part 179, or installed from the storage part 178, or installed from the ROM 172. When the computer program is executed by the CPU 171, the method of the embodiment of the present disclosure is executed.

[0154] It should be noted that, in the context of the present disclosure, a computer-readable medium may be a tangible medium that may contain or store a program for use by an instruction execution system, apparatus, or device or for use in conjunction with an instruction execution system, apparatus, or device.

[0155] The computer-readable medium may be a computer-readable storage medium, or a computer-readable signal medium, or any combination thereof. The computer-readable storage medium is a non-transitory computer storage medium.

[0156] Computer-readable storage media include, but are not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices or components, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device or device. Computer instructions are stored on the computer-readable storage medium, and when the instructions are executed by the processor, the method described in any of the aforementioned embodiments is implemented.

[0157] A computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries a computer-readable program code. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.

[0158] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.

[0159] In some embodiments, a computer program is further provided, comprising: instructions, which, when executed by a processor, cause the processor to perform the method described in any of the aforementioned embodiments. For example, the instructions may be embodied as computer program codes.

[0160] In embodiments of the present disclosure, computer program code for performing the operations of the present disclosure may be written in one or more programming languages ​​or combinations thereof, including but not limited to object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0161] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0162] The functions described above may be performed at least in part by one or more hardware logic components. For example, and without limitation, exemplary hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.

[0163] Although some specific embodiments of the present disclosure have been described in detail by way of examples, those skilled in the art will appreciate that the above examples are for illustrative purposes only and are not intended to limit the scope of the present disclosure. Those skilled in the art will appreciate that modifications may be made to the above embodiments without departing from the scope and spirit of the present disclosure. The scope of the present disclosure is defined by the appended claims.

Claims

1. A content generation method, comprising: In response to a user's triggering operation, obtaining the first visual content in the video and displaying an editing interface; recommending, on the editing interface, at least one candidate content related to the first visual content, the at least one candidate content being determined based on at least one of a plot dimension, a character dimension, and a scene dimension in the first visual content, and the at least one candidate content being selected from one or more of the video and interactive content related to the video; In response to target content determined by the user from the at least one candidate content, second visual content is generated based on the first visual content and the target content.

2. The content generation method according to claim 1, wherein: The recommending at least one candidate content related to the first visual content includes: Performing content understanding on the first visual content to determine at least one of a first key content in the plot dimension, a second key content in the character dimension, and a third key content in the scene dimension; Based on at least one of the first key content, the second key content, and the third key content, the at least one candidate content is determined in the video and / or the interactive content.

3. The content generation method according to claim 2, wherein: The performing content understanding on the first visual content to determine at least one of a first key content in the plot dimension, a second key content in the character dimension, and a third key content in the scene dimension of the first visual content includes: Performing content understanding on the first visual content to obtain a content focus; In response to the content focus being related to event development or plot logic, determining the dimension related to the first visual content as the plot dimension; The first key content of the first visual content in the plot dimension is determined.

4. The content generation method according to claim 2, wherein: The performing content understanding on the first visual content to determine at least one of a first key content in the plot dimension, a second key content in the character dimension, and a third key content in the scene dimension of the first visual content includes: Performing content understanding on the first visual content to obtain a content focus; In response to the content focus being related to the character image or personality psychology, determining the dimension related to the first visual content as the character dimension; The second key content of the first visual content in the character dimension is determined.

5. The content generation method according to claim 2, wherein: The performing content understanding on the first visual content to determine at least one of a first key content in the plot dimension, a second key content in the character dimension, and a third key content in the scene dimension of the first visual content includes: Performing content understanding on the first visual content to obtain a content focus; In response to the content focus being related to an environmental atmosphere or a spatial setting, determining a dimension related to the first visual content as the scene dimension; Determine the third key content of the first visual content in the scene dimension.

6. The content generation method according to claim 2, wherein: The determining, in the video, the at least one candidate content based on at least one of the first key content, the second key content, and the third key content includes: determining a key plot related to the first visual content based on at least one of the first key content, the second key content, and the third key content; Determine a target set of the key plot in the video, where the target set includes one or more episodes; The at least one candidate content is recommended based on the key frames in the target set.

7. The content generation method according to claim 2, wherein: The step of determining, in the interactive content, at least one candidate content based on at least one of the first key content, the second key content, and the third key content: determining a key plot related to the first visual content based on at least one of the first key content, the second key content, and the third key content; In the interactive content, at least one candidate content related to the key plot is determined.

8. The content generation method according to claim 6 or 7, wherein: Determining a key plot related to the first visual content based on at least one of the first key content, the second key content, and the third key content includes at least one of the following: Determining a current plot based on at least one of the character identity, character action, character dialogue, and character clothing in the second key content, and using a key plot related to the current plot as the key plot related to the first visual content; Determining the current plot based on at least one of the environmental features, object features, prop features, and light and color features in the third key content, and using the key plot related to the current plot as the key plot related to the first visual content; The key plot related to the first key content is used as the key plot related to the first visual content.

9. The content generation method according to claim 1, wherein: Each candidate content in the at least one candidate content includes at least one of a key person and a key location, and in response to the target content determined by the user in the at least one candidate content, generating the second visual content according to the first visual content and the target content includes: In response to the user selecting at least one of the key person and the key place, displaying reference visual content related to at least one of the key person and the key place; In response to the user identifying the target content in the reference visual content, the second visual content is generated according to the first visual content and the target content, where the target content includes one or more of the reference visual contents.

10. The content generation method according to claim 1, wherein: Each candidate content in the at least one candidate content includes time information, and generating the second visual content according to the first visual content and the target content includes: The first visual content and the target content are synthesized according to the time information to generate the second visual content.

11. The content generation method according to claim 1, wherein: The at least one candidate content includes at least one candidate text content, the at least one candidate text content being selected from one or more of lines in the video and interactive content related to the video, and generating the second visual content according to the first visual content and the target content in response to target content determined by the user from the at least one candidate content includes: In response to the user selecting target text content from the at least one candidate text content, the second visual content is generated according to the first visual content and the target text content.

12. The content generation method according to any one of claims 1 to 7, 9 to 11, further comprising: The editing tool is recommended to the user based on at least one of image style, image color, and image content of the first visual content.

13. The content generation method according to any one of claims 1 to 7, 9 to 11, further comprising: In response to a user triggering an operation on a sharing control of the editing interface, the second visual content is shared.

14. The content generation method according to any one of claims 1 to 7, 9 to 11, wherein: The first visual content includes image information but does not include non-image information; and / or The first visual content is a frame of image or a video clip in the video.

15. A content generation device, comprising: a processing module configured to obtain first visual content in the video and display an editing interface in response to a user trigger operation; A recommendation module is configured to recommend, on the editing interface, at least one candidate content related to the first visual content, the at least one candidate content being determined based on at least one of a plot dimension, a character dimension, and a scene dimension in the first visual content, and the at least one candidate content being selected from one or more of the video and interactive content related to the video; The content generation module is configured to generate second visual content based on the first visual content and the target content in response to target content determined by the user from the at least one candidate content.

16. An electronic device comprising: processor; as well as A memory coupled to the processor, configured to store instructions, wherein when the instructions are executed by the processor, the processor executes the content generation method according to any one of claims 1 to 14.

17. A computer-readable storage medium having computer instructions stored thereon, wherein: When the computer instructions are executed by a processor, the content generation method according to any one of claims 1 to 14 is implemented.

18. A computer program product comprising: The method comprises computer instructions, which, when executed by a processor, implement the content generation method according to any one of claims 1 to 14.

Citation Information

Patent Citations

  • Video selection playing method, device and equipment and readable storage medium

    CN110225369A

  • Video processing method and device, electronic equipment and storage medium

    CN113645482A

  • Video processing method and device, electronic equipment and computer readable storage medium

    CN116471451A

  • Information processing method and device

    CN117519548A

  • Picture editing method, related equipment and system

    CN120196252A