Media content generation method and apparatus, and device, storage medium and program product

By automatically generating multimedia content through the client or server, the problem of cumbersome process for users to generate multimedia content is solved, and efficient one-click generation of multimedia content that matches the target text is achieved.

WO2026012008A1PCT designated stage Publication Date: 2026-01-15TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/098557
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-12
Filing Date
2025-05-30
Publication Date
2026-01-15

AI Technical Summary

Technical Problem

Users need to perform multiple steps when generating multimedia content, making the generation process cumbersome and inefficient.

Method used

This invention provides a method for generating media content that automatically generates multimedia content matching the target text, including episode scenes and dialogue, by responding to the text selected by the user through a client or server, thereby reducing the number of user operation steps.

Benefits of technology

It enables users to generate multimedia content that matches the target text with a single click without having to perform multiple operations, improving generation efficiency and reducing user workload.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025098557_15012026_PF_FP_ABST
    Figure CN2025098557_15012026_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of information processing. Disclosed are a media content generation method and apparatus, and a device, a storage medium and a program product. The method comprises: displaying text media content posted on a media account, wherein the text media content comprises several texts; in response to a selection operation performed on at least one of the several texts, determining a target text for generating target multimedia content; and displaying the target multimedia content generated on the basis of the target text. In the method, after the target text is determined, the target multimedia content can be automatically generated. A user can acquire the target multimedia content matching the target text, without the need to execute a plurality of operation steps, thereby facilitating an improvement in the efficiency of generating the target multimedia content.
Need to check novelty before this filing date? Find Prior Art

Description

Media content generation methods, apparatus, equipment, storage media and program products Technical Field

[0001] This application claims priority to Chinese Patent Application No. 202410942622X, filed on July 12, 2024, entitled “Media Content Generation Method, Apparatus, Device and Storage Medium”, the entire contents of which are incorporated herein by reference.

[0002] Technical Field

[0003] This application relates to the field of information processing technology, and in particular to a media content generation method, apparatus, device, storage medium, and program product. Background Technology

[0004] With the development of communication technology, people's social behaviors and needs are constantly changing. In the process of social interaction, users can actively share multimedia content such as animations, posters, audio and video.

[0005] In related technologies, users need to perform multiple steps to generate target multimedia content. For example, users need to manually select material information for creating the target multimedia content, and then edit the selected material information to generate the target multimedia content. Summary of the Invention

[0006] This application provides a method, apparatus, device, and storage medium for generating media content. The technical solution is as follows:

[0007] According to one aspect of this application, a media content generation method is provided, the method being executed by a client, the method comprising:

[0008] Displays text media content published by a media account, the text media content including several texts;

[0009] In response to a selection operation on at least one character among the plurality of texts, a target text for generating the target multimedia content is determined;

[0010] Display the target multimedia content generated based on the target text.

[0011] According to another aspect of this application, a media content generation method is provided, the method being executed by a server, the method comprising:

[0012] The client sends target text for generating target multimedia content, wherein the target text is selected from text media content published by the media account.

[0013] Generate the target multimedia content based on the target text;

[0014] The target multimedia content is sent to the client.

[0015] According to another aspect of this application, a media content generation apparatus is provided, the apparatus comprising:

[0016] The display module is used to display text media content published by a media account, wherein the text media content includes several texts;

[0017] A determination module is configured to determine the target text for generating the target multimedia content in response to a selection operation on at least one character among the plurality of texts.

[0018] The display module is also used to display the target multimedia content generated based on the target text.

[0019] According to another aspect of this application, a media content generation apparatus is provided, the apparatus comprising:

[0020] The receiving module is used to receive target text sent by the client for generating target multimedia content, wherein the target text is selected from the text media content published by the media account.

[0021] The generation module is used to generate the target multimedia content based on the target text;

[0022] The sending module is used to send the target multimedia content to the client.

[0023] According to another aspect of this application, a computer device is provided, the computer device comprising: a processor and a memory, the memory storing at least one program, the at least one program being loaded and executed by the processor to implement the media content generation method as described above.

[0024] According to another aspect of this application, a computer-readable storage medium is provided, which stores at least one program that is loaded and executed by a processor to implement the media content generation method as described above.

[0025] According to another aspect of this application, a computer program product is provided, comprising at least one program segment stored in a computer-readable storage medium; a processor of a computer device reads the at least one program segment from the computer-readable storage medium, and the processor executes the at least one program segment, causing the computer device to perform the media content generation method as described above.

[0026] The beneficial effects of the technical solution provided in this application include at least the following:

[0027] This invention provides a method for automatically generating target multimedia content. Once the target text is determined, the generated multimedia content can be directly displayed. Compared to methods where users manually generate target multimedia content, users do not need to perform multiple steps and can generate target multimedia content matching the target text with a single click, thus improving the efficiency of target multimedia content generation. Attached Figure Description

[0028] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0029] Figure 1 is a schematic diagram of a computer system provided in an exemplary embodiment of this application;

[0030] Figure 2 is a flowchart of a media content generation method provided in an exemplary embodiment of this application;

[0031] Figure 3 is a schematic diagram of a media content generation method provided in an exemplary embodiment of this application;

[0032] Figure 4 is a schematic diagram of a media content generation method provided in an exemplary embodiment of this application;

[0033] Figure 5 is a flowchart of a media content generation method provided in an exemplary embodiment of this application;

[0034] Figure 6 is a schematic diagram of a media content generation method provided in an exemplary embodiment of this application;

[0035] Figure 7 is a flowchart of a media content generation method provided in an exemplary embodiment of this application;

[0036] Figure 8 is a schematic diagram of a media content generation method provided in an exemplary embodiment of this application;

[0037] Figure 9 is a schematic diagram of a media content generation method provided in an exemplary embodiment of this application;

[0038] Figure 10 is a flowchart of a media content generation method provided in an exemplary embodiment of this application;

[0039] Figure 11 is a schematic diagram of a media content generation method provided in an exemplary embodiment of this application;

[0040] Figure 12 is a schematic diagram of a media content generation method provided in an exemplary embodiment of this application;

[0041] Figure 13 is a flowchart of a media content generation method provided in an exemplary embodiment of this application;

[0042] Figure 14 is a flowchart of a media content generation method provided in an exemplary embodiment of this application;

[0043] Figure 15 is a flowchart of a media content generation method provided in an exemplary embodiment of this application;

[0044] Figure 16 is a schematic diagram of a media content generation method provided in an exemplary embodiment of this application;

[0045] Figure 17 is a flowchart of a media content generation method provided in an exemplary embodiment of this application;

[0046] Figure 18 is a flowchart of a media content generation method provided in an exemplary embodiment of this application;

[0047] Figure 19 is a flowchart of a media content generation method provided in an exemplary embodiment of this application;

[0048] Figure 20 is a flowchart of a media content generation method provided in an exemplary embodiment of this application;

[0049] Figure 21 is a flowchart of a media content generation method provided in an exemplary embodiment of this application;

[0050] Figure 22 is a flowchart of a media content generation method provided in an exemplary embodiment of this application;

[0051] Figure 23 is a schematic diagram of a media content generation method provided in an exemplary embodiment of this application;

[0052] Figure 24 is a schematic diagram of a media content generation method provided in an exemplary embodiment of this application;

[0053] Figure 25 is a block diagram of a media content generation apparatus provided in an exemplary embodiment of this application;

[0054] Figure 26 is a block diagram of a media content generation apparatus provided in an exemplary embodiment of this application;

[0055] Figure 27 is a block diagram of a computer device provided in an exemplary embodiment of this application. Detailed Implementation

[0056] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0057] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.

[0058] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The singular forms “a,” “the,” and “the” used in this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.

[0059] It should be understood that although the terms first, second, etc., may be used in this application to describe various information, this information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this application, a first parameter may also be referred to as a second parameter, and similarly, a second parameter may also be referred to as a first parameter. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."

[0060] First, a brief introduction to the terms used in the embodiments of this application:

[0061] Multimedia content refers to information content that is processed through digital signals and presented in the form of text, images, audio, video, etc.

[0062] Text-based media content refers to media information carried by text. In this embodiment, text-based media content refers to text-based media information disseminated on the internet.

[0063] Screenshots: refers to video frame information in movies, TV series, short videos, and other videos.

[0064] Dialogue in TV series: refers to the dialogue information in videos such as movies, TV series, and short videos.

[0065] Figure 1 shows a structural block diagram of a computer system provided in an exemplary embodiment of this application. The computer system 100 includes a terminal device 120 and a server 140.

[0066] Terminal device 120 has a client installed and running that supports displaying text-based media content. The client supporting text-based media content display can be any of the following: video application, reading application, social application, live streaming application, virtual reality (VR) application, or augmented reality (AR) application. Terminal device 120 is the terminal used by the first user, who uses terminal device 120 to view or read text-based media content and to generate corresponding target multimedia content by triggering the text-based media content.

[0067] Terminal device 120 is connected to server 140 via wireless network or wired network.

[0068] Server 140 includes at least one of a single server, multiple servers, a cloud computing platform, and a virtualization center. For example, server 140 includes a processor 144 and a memory 142. Memory 142 includes a display module 1421, a determination module 1422, and a generation module 1423. The display module 1421 is used to display multimedia information, such as text media content; the determination module 1422 is used to determine the target text; and the generation module 1423 is used to generate target multimedia content including scene images and dialogue. Server 140 provides background services for clients supporting multimedia information. Optionally, server 140 undertakes the primary computing work, and terminal device 120 undertakes secondary computing work; or, server 140 undertakes secondary computing work, and terminal device 120 undertakes primary computing work; or, server 140 and terminal device 120 collaborate on computing using a distributed computing architecture.

[0069] Optionally, the device type of terminal device 120 includes at least one of the following: smartphone, tablet computer, e-book reader, laptop computer, and desktop computer. In this embodiment, a smartphone is used as an example to illustrate terminal device 120.

[0070] Those skilled in the art will understand that the number of terminals described above can be more or less. For example, there may be only one terminal, or there may be dozens or hundreds of terminals, or even more. This application does not limit the number of terminals or the type of device.

[0071] In related technologies, users need to perform multiple steps to generate target multimedia content. For example, users need to manually select material information for creating target multimedia content, and then edit the selected material information to generate the target multimedia content. This method of generating target multimedia content is relatively cumbersome for users. Based on this, this application proposes a media content generation method that can automatically generate target multimedia content, thereby reducing the user's operation process and improving the generation efficiency of target multimedia content. For example, as shown in FIG1, the user of terminal device 120 can directly select a portion of the text in the text content 12 of the text media content to achieve one-click triggering of the display of the target multimedia content 21 corresponding to the selected text.

[0072] Figure 2 is a schematic diagram of a media content generation method provided in an exemplary embodiment of this application. In this embodiment, the method is described using the terminal device 120 shown in Figure 1 as an example. The method includes:

[0073] Step 220: Text media content displayed on the terminal device.

[0074] In some embodiments, the text media content is published by a media account.

[0075] In some embodiments, a media account refers to a user account logged in on a client that supports editing and publishing text media content. Optionally, a media account refers to a user account corresponding to an application on a client that supports editing and publishing text media content. For example, a user account logged in on a social media application. Optionally, the social media application needs to support editing and publishing text media content.

[0076] In some embodiments, text-based media content includes a number of texts. Optionally, text-based media content also includes other forms of content such as images and videos. For example, text-based media content includes articles from public WeChat accounts, current affairs commentary articles, etc.

[0077] In some embodiments, the text media content is pre-created by the creator corresponding to the media account.

[0078] For example, as shown in Figure 3, text media content 11 is displayed. The text media content 11 includes at least text content 12. Optionally, the text media content 11 may also include image content 13.

[0079] Step 240: In response to a selection operation on at least one character among several texts, the terminal device determines the target text for generating the target multimedia content.

[0080] In some embodiments, the selection operation for at least one character among several texts includes at least one of the following: single click, double click, left / right swipe, up / down swipe, long press, hover, facial recognition, and voice recognition. It is worth noting that the selection operation for at least one character among several texts includes, but is not limited to, the aforementioned methods. Those skilled in the art should understand that any operation capable of achieving the above functions falls within the protection scope of the embodiments of this application.

[0081] In some embodiments, the target media content includes at least one of animation, poster, audio, and video.

[0082] In some embodiments, in response to a selection operation of at least one character among a plurality of texts, the terminal device sends a generation request to the server for generating target multimedia content. Optionally, the generation request for generating target multimedia content includes a request to determine target text based on the selected at least one character. Optionally, the generation request for generating target multimedia content further includes a request to generate target multimedia content based on the target text. Optionally, the target text includes key information from the selected at least one character.

[0083] In some embodiments, in response to a selection operation on at least one character among a plurality of texts, the terminal device switches the display state of the selected at least one character from a first display state to a second display state. The first display state is a state where at least one character is not selected, and the second display state is a state where at least one character is selected.

[0084] Step 260: The terminal device displays the target multimedia content generated based on the target text.

[0085] In some embodiments, the target multimedia content includes episode scenes and episode dialogue that match the target text.

[0086] In some embodiments, the content of the scene corresponding to the target text matches the text content expressed by the target text. Optionally, the meaning of the content of the scene corresponding to the target text is consistent with the meaning of the text content expressed by the target text. For example, assuming the target text is "xxx is so beautiful", and the content of scene 1 is "xxx has gorgeous makeup", then scene 1 is considered to match the target text.

[0087] In some embodiments, the text content corresponding to the TV drama dialogue that matches the target text matches the text content expressed by the target text. Optionally, the meaning of the text content corresponding to the TV drama dialogue that matches the target text is consistent with the meaning of the text content expressed by the target text. Optionally, the matching degree between the first text content corresponding to the TV drama dialogue that matches the target text and the second text content expressed by the target text is greater than or equal to a matching degree threshold. For example, assuming the target text is "xxx is really beautiful" and the TV drama dialogue is "xxx is really good-looking", since "beautiful" and "good-looking" express the same meaning, the TV drama dialogue is considered to match the target text.

[0088] In some embodiments, the target multimedia content is generated by the server based on a generation request sent by the terminal device. Based on the generation request, the server determines the episode scenes and dialogue that match the target text, and then generates the target multimedia content. Afterward, the server sends the generated target multimedia content to the terminal device, which receives and displays the target multimedia content.

[0089] In some embodiments, the target multimedia content displayed on the terminal device is the one that matches the target text most closely. That is, the server may generate multiple target multimedia content pieces based on the target text, but the matching degree between these multiple pieces of target multimedia content and the target text is different. For example, assuming that target multimedia content 1 matches the target text with a 90% matching degree and target multimedia content 2 matches the target text with an 80% matching degree, the server will only send target multimedia content 1, which has the highest matching degree, to the terminal device.

[0090] In some embodiments, the target multimedia content displayed on the terminal device is all target multimedia content that matches the target text. Optionally, the target multimedia content displayed on the terminal device is the target multimedia content whose matching degree with the target text is greater than a matching degree threshold. For example, assuming that the matching degree of target multimedia content 1 with the target text is 90%, the matching degree of target multimedia content 2 with the target text is 80%, and the matching degree of target multimedia content 3 with the target text is 60%, and the matching degree threshold is 75%, then the server will only send target multimedia content 1 and target multimedia content 2 to the terminal device.

[0091] In some embodiments, the source material corresponding to the target multimedia content is a static image. In some embodiments, the source material corresponding to the target multimedia content is a dynamic video frame.

[0092] In summary, the method provided in this embodiment offers an automatic way to generate target multimedia content. After determining the target text, the target multimedia content generated based on the target text can be directly displayed. Compared to methods where users manually generate target multimedia content, users do not need to perform multiple steps; they can generate target multimedia content matching the target text with a single click, reducing user workload and improving the efficiency of generating target multimedia content.

[0093] In some embodiments, the target multimedia content includes an upper half and a lower half arranged side by side. The upper half displays screen footage, and the lower half displays dialogue. For example, as shown in FIG4, the upper half 19 and the lower half 20 are arranged side by side to form the target multimedia content. The upper half 19 displays screen footage, and the lower half 20 displays dialogue.

[0094] In some embodiments, the target multimedia content includes a side-by-side upper and lower half-area, with the upper half-area displaying dialogue from the series and the lower half-area displaying visuals from the series.

[0095] In some embodiments, the target multimedia content includes a left half-area and a right half-area side by side, with the left half-area displaying dialogue from the series and the right half-area displaying images from the series.

[0096] In some embodiments, the target multimedia content includes a left half-area and a right half-area side by side, with the left half-area displaying the scene and the right half-area displaying the dialogue.

[0097] In some embodiments, the target multimedia content includes a content area. The content area displays screen images and dialogue overlaid on the screen images. Optionally, the dialogue may be displayed separately from the screen images. This application does not limit the display method of the screen images and dialogue in these embodiments.

[0098] In some embodiments, the target multimedia content further includes an information area. The information area is used to display information and / or target text related to the episode.

[0099] In some embodiments, the target multimedia content is automatically generated by a terminal device or a server. When the target multimedia content is generated by a server, or when the target multimedia content is generated by a terminal different from the terminal device, the following steps are included after step 240 and before step 260:

[0100] Step 251: The terminal device sends the target text used to generate the target multimedia content;

[0101] Step 252: The terminal device receives the target multimedia content generated based on the target text.

[0102] In some embodiments, the function of generating target multimedia content needs to be triggered by a generation control. Based on the embodiment shown in Figure 2 above, as shown in Figure 5, step 240 above can also be replaced by the following sub-steps:

[0103] Step 241: In response to the selection operation of at least one character among a plurality of texts, the terminal device displays the selected text and the selection function bar.

[0104] In some embodiments, in response to a selection operation on at least one character among a plurality of texts, the terminal device switches the selected text from a first display state to a second display state. The first display state is an unselected state, and the second display state is a selected state.

[0105] In some embodiments, the terminal device displays a selection function bar in response to the selection of at least one character among a plurality of texts. For example, as shown in FIG6, the terminal device displays the selection function bar 14 in response to the selection of at least one character.

[0106] In some embodiments, the selected function bar includes a generate control. The generate control is used to trigger the generation of target multimedia content. Optionally, in response to the triggering operation of the generate control, the terminal device sends a generation request to the server for generating target multimedia content. Optionally, the generation request for generating target multimedia content includes a request to determine target text based on at least one selected character. Optionally, the generation request for generating target multimedia content further includes a request to generate target multimedia content based on the target text.

[0107] In some embodiments, the selected function bar may further include other function controls besides the generate control. Optionally, the selected function bar may further include at least one of a copy control, a share control, and a favorite control. The copy control is used to copy at least one selected text. The share control is used to share at least one selected text. The favorite control is used to favorite at least one selected text.

[0108] Step 242: In response to the triggering operation of the generation control, the terminal device determines the selected text as the target text for generating the target multimedia content.

[0109] In some embodiments, the triggering operation for the generated control includes at least one of the following: single click, double click, left / right swipe, up / down swipe, long press, hover, facial recognition, and voice recognition. It is worth noting that the triggering operation for the generated control includes, but is not limited to, several of the above-mentioned operations. Those skilled in the art should understand that any operation capable of achieving the above functions falls within the protection scope of the embodiments of this application.

[0110] In some embodiments, the target text for generating the target multimedia content is determined directly by the terminal device based on the selected text. For example, the terminal device directly determines the selected text as the target text for generating the target multimedia content.

[0111] In some embodiments, the target text for generating the target multimedia content is indirectly determined by the terminal device based on the selected text. For example, the terminal device uses a sentence segmentation algorithm to analyze the selected text, segments the selected text into sentences, removes incomplete sentences from the selected text, and determines the remaining complete sentences as the target text for generating the target multimedia content.

[0112] In summary, the method provided in this embodiment allows the terminal device to display a selected function bar, enabling the user to determine the target text for generating target multimedia content based on triggering the generation control in the selected function bar. By setting the generation control, the user can choose whether to generate target multimedia content according to their needs.

[0113] In some embodiments, the plurality of texts includes a first text that is highlighted, the highlighting state indicating that the first text supports multimedia content generation functionality. Based on the embodiment shown in FIG5 above, as shown in FIG7, step 241 above can also be replaced by the following sub-step:

[0114] Step 2411: In response to the selection operation of the first text, the terminal device displays the first text as the selected text and displays the selected function bar.

[0115] In some embodiments, the selection operation for the first text includes at least one of the following: single click, double click, left / right swipe, up / down swipe, long press, hover, facial recognition, and voice recognition. It is worth noting that the selection operation for the first text includes, but is not limited to, several of the above-mentioned operations. Those skilled in the art should understand that any operation capable of achieving the above functions falls within the protection scope of the embodiments of this application.

[0116] In some embodiments, in response to a selection operation on the first text, the terminal device switches the first text from a first display state to a second display state. The first display state is an unselected state, and the second display state is a selected state.

[0117] In some embodiments, the terminal device displays a selection bar in response to a selection operation on the first text. The selection bar includes a generate control.

[0118] In some embodiments, creators of text-based media content need to enable the multimedia content generation function before publishing the text-based media content. With the multimedia content generation function enabled, other users can select at least one character from a plurality of texts within the text-based media content. If the creator of the text-based media content has not enabled the multimedia content generation function, it indicates that the creator does not wish their content to be used to generate relevant target multimedia content; in this case, it is impossible to generate target multimedia content by selecting at least one character from a plurality of texts within the text-based media content.

[0119] For example, taking the creation of a poster as an example of multimedia content, as shown in Figure 8, a function switch control 16 is displayed on the text media content editing interface 15. This function switch control 16 includes two options: "On" and "Off". In response to a trigger operation on the function switch control 16, the multimedia content generation function is turned on or off, i.e., the poster generation function is started or stopped. For example, in response to a trigger operation on the "On" option in the function switch control 16, the multimedia content generation function is turned on (i.e., the poster generation function is turned on). As another example, in response to a trigger operation on the "Off" option in the function switch control 16, the multimedia content generation function is turned off (i.e., the poster generation function is turned off).

[0120] In some embodiments, the editing interface 15 also displays a list of related materials 17. Optionally, at least one related material in the list of related materials 17 is automatically recommended by the backend server based on the text media content created by the creator. In some embodiments, the list of related materials 17 includes at least one related material, which is displayed in order of its matching degree with the text media content.

[0121] In some embodiments, the first text is determined by the backend server from the text media content based on the matching degree between the associated material and the text media content. For example, as shown in FIG9, at least one matching target text 18 is displayed on the text media content editing interface 15. The at least one target text 18 is text within the text media content. The at least one target text 18 includes the aforementioned first text.

[0122] In some embodiments, in response to an editing operation on any target text in at least one target text, either target text is retained or any target text is deleted.

[0123] In summary, the method provided in this embodiment enables users to quickly select text that can generate target multimedia content by directly triggering the first text that supports multimedia content generation, which helps to further accelerate the generation of target multimedia content.

[0124] Furthermore, when the first text is determined by the creator of the text-based media content, it enables users to adapt to the requirements of the text-based media creator when creating target multimedia content based on the text-based media content, thus preserving the creator's ideas to a certain extent and respecting the creator's choices.

[0125] In some embodiments, after the target multimedia content is generated, the user can choose to share the target multimedia content with the target object. As shown in Figure 10, the above method further includes the following steps:

[0126] Step 320: Display the sharing control;

[0127] In some embodiments, the share control is used to trigger the sharing of target multimedia content.

[0128] In some embodiments, the sharing control and the target multimedia content are displayed simultaneously. For example, as shown in FIG11, the sharing control 22 is displayed at the same time as the target multimedia content 21.

[0129] In some embodiments, the sharing control is displayed after the target multimedia content has already been displayed. That is, the target multimedia content and the sharing control are displayed sequentially.

[0130] Step 340: In response to the triggering operation of the sharing control, share the target multimedia content to the specified network location.

[0131] In some embodiments, the triggering operation for the sharing control includes at least one of the following: single click, double click, left / right swipe, up / down swipe, long press, hover, facial recognition, and voice recognition. It is worth noting that the triggering operation for the sharing control includes, but is not limited to, several of the above-mentioned operations. Those skilled in the art should understand that any operation capable of achieving the above functions falls within the protection scope of the embodiments of this application.

[0132] Sharing controls allow users to share target multimedia content, expanding its reach. Furthermore, sharing can facilitate the promotion of the corresponding TV series or shows.

[0133] In some embodiments, during the process of generating target multimedia content, the terminal device may also prompt the user with some information to indicate that the target multimedia content is being generated, or that the target multimedia content has been successfully generated, or that the target multimedia content has failed to be generated.

[0134] In some embodiments, a first prompt message is displayed in response to the selection of at least one character among a plurality of texts. In some embodiments, the first prompt message is displayed in response to a triggering operation on a generation control. The first prompt message is used to indicate that target multimedia content is being generated. For example, as shown in FIG12, the first prompt message 23 is displayed.

[0135] In some embodiments, if the generation of target multimedia content fails, a second prompt message is displayed to indicate that the generation of target multimedia content has failed. In some embodiments, the second prompt message instructs the user to reselect at least one character from a plurality of texts.

[0136] In some embodiments, if the target multimedia content is successfully generated, a third prompt message is displayed to indicate that the target multimedia content has been successfully generated.

[0137] By providing first, second, and third prompts, users can be promptly informed of the generation status of the target multimedia content. This helps users stay informed about the current progress and perform subsequent actions based on the corresponding prompts.

[0138] Next, Figure 13 is a schematic diagram of a media content generation method provided in an exemplary embodiment of this application. The generation process of the target multimedia content is described in detail. Optionally, the target multimedia content is generated by a server or a client. In this embodiment, the method is described using the server 140 shown in Figure 1 as an example. The method includes:

[0139] Step 420: Receive the target text sent by the client for generating the target multimedia content;

[0140] In some embodiments, at least one text selected by the user of the client is received. That is, the selected text sent by the client is received.

[0141] In some embodiments, the target text is selected from the text media content published by the media account. That is, the target text is the selected text.

[0142] In some embodiments, the target text is determined by the server based on the selected text. For example, the selected text is analyzed using a sentence segmentation algorithm, which segments the selected text into sentences, removes incomplete sentences from the selected text, and determines the remaining complete sentences as the target text for generating the target multimedia content.

[0143] In some embodiments, the target text is a first text highlighted in the text media content published by the media account, the highlighted state being used to indicate that the first text supports multimedia content generation functionality.

[0144] In some embodiments, the first text is determined based on the degree of matching between the associated episode material and the text media content.

[0145] Optionally, the associated TV series materials related to the text media content are determined based on the similarity between the semantic text of the text media content and the semantic text of the TV series dialogue subtitles in the associated TV series materials. The semantic similarity between the semantic text of the text media content and the semantic text of the TV series dialogue subtitles in the associated TV series materials is calculated, and TV series materials with a semantic similarity greater than or equal to a semantic similarity threshold are considered as associated TV series materials related to the text media content. Furthermore, the associated TV series materials can also be sorted according to the magnitude of semantic similarity.

[0146] Optionally, the text media content is segmented into sentences. The segmented sentences are then matched against subtitles from related TV series footage. The similarity between the segmented sentences and the subtitles is calculated. If the text similarity between the segmented sentences and the subtitles is greater than or equal to a text similarity threshold, the segmented sentence is designated as the first text. Furthermore, the first text can be sorted according to the magnitude of text similarity.

[0147] In some embodiments, the information related to the first text is sent synchronously to the relevant users when the media account publishes text media content. The display method of the first text on the client is different from the display method of other texts in the text media content. For example, the first text is highlighted, while other texts in the text media content are displayed normally.

[0148] In some embodiments, the information related to the first text includes at least one of the following: display information related to the first text, links related to the first text, associated TV series material related to the first text, and associated target multimedia content related to the first text. The display information related to the first text indicates how the first text is displayed on the client. The links related to the first text indicate that the first text supports multimedia content generation. For example, in response to a triggering operation on the first text, the user can be redirected to a multimedia content generation interface. The associated TV series material related to the first text indicates the associated TV series material corresponding to the first text, which is used to generate target multimedia content related to the first text. The associated target multimedia content related to the first text refers to pre-generated target multimedia content. In this case, when the user triggers the first text, the associated target multimedia content can be directly obtained.

[0149] It should be noted that the server or client that determines the first text described in the embodiments of this application may be the same as or different from the server or client that generates the multimedia content. This embodiment of the application does not impose any limitations on this.

[0150] In some embodiments, the target media content includes at least one of animation, poster, audio, and video.

[0151] Step 440: Generate target multimedia content based on the target text;

[0152] In some embodiments, the target text is matched with corresponding episode scenes and dialogue within a drama series database. This drama series database is pre-generated by the server. The server pre-segments different drama series to obtain a database containing episode scenes and dialogue. It should be understood that the episode scenes and dialogue in this embodiment are correlated. Typically, multiple episode scenes may correspond to the same dialogue.

[0153] In some embodiments, target multimedia content including episode images and episode dialogues is generated based on episode images and episode dialogues that match the target text.

[0154] Step 460: Send the target multimedia content to the client.

[0155] In summary, the method provided in this embodiment offers an automatic way to generate target multimedia content. Once the target text is determined, the target multimedia content generated based on the target text can be directly displayed. Compared to methods where users manually generate target multimedia content, users do not need to perform multiple steps and can generate target multimedia content matching the target text with a single click, thus improving the efficiency of target multimedia content generation.

[0156] In some embodiments, the target media content includes episode scenes and episode dialogue that match the target text.

[0157] In some embodiments, the target multimedia content may also include promotional content provided by the promoter. For example, the target media content may also include advertising information to be promoted provided by the advertiser.

[0158] In some embodiments, there is a preset correspondence between episode scenes and episode dialogue. Based on the embodiment shown in Figure 13 above, as shown in Figure 14, step 440 can also be replaced by the following sub-steps:

[0159] Step 441: Based on the target text, query the TV series images and lines that match the target text in the preset correspondence. The TV series images and lines are used to generate the target multimedia content.

[0160] In some embodiments, the preset correspondence is determined in advance by the server. For example, taking a first episode as an example, the preset correspondence between the episode scenes and dialogue included in the first episode is determined by the following method. As shown in Figure 15, before step 441 above, the method further includes:

[0161] Step 520: Obtain the first episode video and episode dialogue subtitles;

[0162] In some embodiments, the video of the first episode and the subtitles of the episode's dialogue corresponding to the first episode are obtained. Optionally, the video of the first episode and the subtitles of the episode's dialogue are obtained from a video database.

[0163] Step 540: Divide the first episode video into at least one video segment according to the start and end times of the subtitles in the episode dialogue;

[0164] In some embodiments, the first episode video is divided into at least one video segment according to the start and end times of the subtitles corresponding to the dialogue subtitles of the same episode. It should be understood that the first number of video segments is greater than or equal to the second number of episode dialogue subtitles.

[0165] For example, as shown in Figure 16, assume that the first episode video corresponds to two episode dialogue subtitles, including episode dialogue subtitle 1 and episode dialogue subtitle 2. The video segments corresponding to episode dialogue subtitle 1 and episode dialogue subtitle 2 are not consecutive. That is, as shown in Figure 16, there is a video segment without subtitles between episode dialogue subtitle 1 and episode dialogue subtitle 2.

[0166] In this embodiment, an example is given of a one-to-one correspondence between video segments and drama series subtitles. That is, each video segment corresponds to one drama series subtitle. For example, as shown in Figure 16, the first drama series video can be divided into two video segments, including video segment 1 and video segment 2. Among them, video segment 1 corresponds to drama series subtitle 1, and video segment 2 corresponds to drama series subtitle 2.

[0167] Step 560: Extract at least one episode frame from the video segments;

[0168] In some embodiments, the video is divided into segments, with each segment consisting of a video frame. Each video segment is divided into at least one video frame. Scenes from each video frame are then extracted. It is important to understand that there is a one-to-one correspondence between video frames and scenes from each video frame.

[0169] In some embodiments, at least one target video frame is extracted from the video slice, and the quality of each target video frame is greater than or equal to a quality threshold. For example, suppose video slice 1 includes video frame 1, video frame 2, and video frame 3. Where the quality of video frame 1 and video frame 2 is greater than the quality threshold, and the quality of video frame 3 is less than the quality threshold, then only video frame 1 and video frame 2 are extracted from video slice 1, while video frame 3 is ignored.

[0170] Step 580: Generate a preset correspondence based on at least one episode scene and episode dialogue subtitles.

[0171] In some embodiments, a preset correspondence between episode scenes and episode dialogue is generated based on at least one extracted episode scene. Examples are shown in Table 1 below:

[0172] Table 1

[0173] In some embodiments, there is a one-to-one correspondence between episode images and episode dialogue. For example, episode image 1 corresponds to episode dialogue 1.

[0174] In some embodiments, there is a many-to-one correspondence between episode images and episode dialogue. For example, episode image 2 and episode image 3 both correspond to episode dialogue 2.

[0175] Normally, there is no one-to-many correspondence between scenes and dialogue in a TV series. However, in some special cases, there may be a one-to-many correspondence between scenes and dialogue, which is not limited in this embodiment.

[0176] By pre-determining the correspondence between episode scenes and episode dialogue, the server can, after determining the target text, query the episode scenes and episode dialogue that match the target text within the pre-determined correspondence, which helps improve matching efficiency.

[0177] Optionally, based on the matching degree between the target text and the dialogue in the TV series, and the correspondence between TV series scenes and dialogue in the preset correspondence, the TV series scenes and dialogue that match the target text are determined. The TV series dialogue is used to connect the TV series dialogue and the target text.

[0178] In some embodiments, there is a preset correspondence between episode scenes and episode dialogue. Based on the embodiment shown in Figure 13 above, as shown in Figure 17, step 440 can also be replaced by the following sub-steps:

[0179] Step 442: Calculate the first feature vector corresponding to the target text, and calculate the vector similarity between the first feature vector and at least one second feature vector; determine the TV scene and TV line corresponding to the second feature vector with the highest vector similarity as the TV scene and TV line that match the target text, and use the TV scene and TV line to generate target multimedia content.

[0180] In some embodiments, the episode scenes and episode lines corresponding to the second feature vector with a vector similarity greater than or equal to the similarity threshold are all determined as episode scenes and episode lines that match the target text.

[0181] In some embodiments, the second feature vector is predetermined by the server. For example, taking the second episode as an example, the second feature vector is determined by the following method. As shown in Figure 18, prior to step 442 above, the method further includes:

[0182] Step 610: Obtain the video of the second episode;

[0183] In some embodiments, the second episode video is obtained from a video database.

[0184] Step 620: Divide the second episode video into at least one video segment according to the division step size;

[0185] In some embodiments, the step size is random. In some embodiments, the step size is preset. For example, the step size is 3 seconds.

[0186] In some embodiments, the second episode video is continuously divided according to a division step size. That is, at least one video segment in the embodiments of this application is continuous. It can be understood that combining at least one video segment in sequence can obtain a complete second episode video.

[0187] Step 630: Extract computer vision information for each video segment from at least one video segment;

[0188] In some embodiments, computer vision algorithms are used to extract computer vision information from each video segment. Optionally, the computer vision information includes subject information and / or text information in each video segment. Subject information includes human body information, object information, landscape information, etc.

[0189] Step 640: Based on the computer vision information of each video segment, generate content description information for each video segment;

[0190] In some embodiments, content description information for each video segment is generated based on the extracted computer vision information for each video segment.

[0191] In some embodiments, content description information for each video segment in at least one video segment is obtained based on a multimodal visual recognition model. In this case, it is unnecessary to perform step 630 above to obtain the computer vision information for each video segment. That is, steps 630 and 640 can be implemented directly based on the multimodal visual recognition model.

[0192] Step 650: Generate a second feature vector for each video segment based on the content description information of each video segment.

[0193] In some embodiments, a second feature vector for each video segment is generated based on the content description information of each video segment and a vector database. The vector database includes the correspondence between the content description information and the feature vectors.

[0194] In some embodiments, a second feature vector for each video segment is generated based on a multimodal embedding algorithm. Optionally, based on a multimodal embedding algorithm, at least two of the audio, image, and text information corresponding to each video segment are input to generate a second feature vector corresponding to each video segment. In this case, steps 630 and 640 above are unnecessary. That is, steps 630 to 650 above can be implemented directly based on the multimodal embedding algorithm.

[0195] By matching the first and second feature vectors, the server can determine the target text and then identify the matching episode scenes and dialogue based on the target text, which helps improve matching efficiency.

[0196] In some embodiments, based on the embodiment shown in FIG18 above, as shown in FIG19, the method further includes:

[0197] Step 720: Extract at least one episode frame from the video segments;

[0198] In some embodiments, the video is divided into segments, with each segment consisting of a video frame. Each video segment is divided into at least one video frame. Scenes from each video frame are then extracted. It is important to understand that there is a one-to-one correspondence between video frames and scenes from each video frame.

[0199] In some embodiments, at least one target video frame is extracted from the video slice, and the quality of each target video frame is greater than or equal to a quality threshold. For example, suppose video slice 1 includes video frame 1, video frame 2, and video frame 3. Where the quality of video frame 1 and video frame 2 is greater than the quality threshold, and the quality of video frame 3 is less than the quality threshold, then only video frame 1 and video frame 2 are extracted from video slice 1, while video frame 3 is ignored.

[0200] Step 740: Store the correspondence between at least one episode frame and the second feature vector of the video segment.

[0201] In some embodiments, given that the second feature vector corresponding to each video segment and at least one episode frame corresponding to each video segment are known, there is a correspondence between the at least one episode frame and the second feature vector. Examples are shown in Table 2 below:

[0202] Table 2

[0203] In some embodiments, there is a one-to-one correspondence between episode frames and the second feature vector. For example, episode frame 1 corresponds to second feature vector 1.

[0204] In some embodiments, there is a many-to-one correspondence between episode frames and the second feature vector. For example, episode frame 1 and episode frame 2 both correspond to the second feature vector 2.

[0205] In some embodiments, the relationship between episode frames and the second feature vector is one-to-many. For example, episode frame 1 corresponds to both second feature vector 1 and second feature vector 2.

[0206] In summary, the method provided in this embodiment, by storing the correspondence between at least one episode frame and the second feature vector of a video segment, enables the server to quickly find the episode frame corresponding to the target text after obtaining the target text, based on the correspondence between at least one episode frame and the second feature vector, as well as the first feature vector corresponding to the target text.

[0207] In some embodiments, when extracting at least one episode frame corresponding to a video segment, the quality is determined based on the quality of the video frames in each video segment. Optionally, as shown in FIG20, step 560 or step 720 above can be replaced by the following sub-steps:

[0208] Step 820: Score each video frame in the video segment according to at least one scoring factor to obtain the quality score corresponding to each video frame;

[0209] In some embodiments, each video segment includes multiple video frames, and each of the multiple video frames corresponds to a quality score.

[0210] In some embodiments, at least one scoring factor includes at least one of brightness, sharpness, and confidence in the identification of the subject in the video frame. Brightness refers to the light intensity in the video frame. Sharpness refers to the clarity of the image in the video frame. Confidence in the identification of the subject in the video frame refers to the probability that the subject in the video frame is identified.

[0211] Step 840: Based on the quality score, extract at least one episode frame from each video frame in the video segment.

[0212] In some embodiments, at least one episode frame is extracted from the target video frame in the video segment based on the quality score of each video frame. The quality score corresponding to the target video frame is greater than or equal to a quality threshold score.

[0213] In some embodiments, the video frames in the video segment are sorted from high to low or from low to high according to their quality scores, and at least one video frame with a quality score greater than or equal to a quality threshold score is selected. At least one episode scene is extracted from the at least one video frame with a quality score greater than or equal to the quality threshold score.

[0214] In summary, the method provided in this embodiment scores each video frame in each video segment, enabling the server to extract episode scenes based on preferred video frames. This helps ensure that the brightness, clarity, or confidence level of the main subject in the extracted video scenes are all better than the threshold.

[0215] In some embodiments, the target multimedia content is generated based on the corresponding episode scenes and dialogue that match the target text. Based on the embodiment shown in Figure 13 above, as shown in Figure 21, step 440 can also be replaced by the following sub-steps:

[0216] Step 443: If there are at least two candidate episode frames in the target text match, calculate the image quality score of at least two candidate episode frames;

[0217] In some embodiments, each candidate episode frame is scored according to at least one scoring factor to obtain a picture quality score for each candidate episode frame. Optionally, the at least one scoring factor includes at least one of brightness, sharpness, and confidence in the identification of the subject in the picture.

[0218] Step 444: Select the target episode from at least two candidate episodes based on the picture quality score;

[0219] In some embodiments, based on the image quality score of each candidate episode image, at least one target episode image is selected from at least two candidate episode images that match the target text. The image quality score corresponding to the target episode image is greater than or equal to an image quality score threshold.

[0220] Step 445: Generate target multimedia content based on the target episode images and the episode dialogue that matches the target text.

[0221] In some embodiments, the target episode image includes at least one, and the episode dialogue matching the target text may also include at least one. However, it is important to understand that each episode image corresponds to its own episode dialogue; that is, there is a one-to-one correspondence between episode images and episode dialogue. For example, suppose the target text matches the first episode image of a first series and the second episode image of a second series. Here, the first and second series are two different series. The first episode image corresponds to episode dialogue 1, and the second episode image corresponds to episode dialogue 2.

[0222] The one-to-one correspondence between scene and dialogue in the TV series mentioned above only applies when the scene corresponds to a specific line of dialogue. If no matching dialogue is found for the target text, the target text or a portion thereof will be used to generate matching dialogue.

[0223] In some embodiments, if there is a candidate episode frame that matches the target text, the candidate episode frame is determined as the target episode frame. That is, if there is one and only one matching candidate episode frame for the target text, that candidate episode frame is the target episode frame used to generate the target multimedia content.

[0224] In summary, the method provided in this embodiment calculates the picture quality score of each candidate TV series picture, enabling the server to generate target multimedia content based on the preferred candidate TV series picture, thereby helping to ensure that the quality of the TV series picture corresponding to the generated target multimedia content is relatively good.

[0225] In some embodiments, the target multimedia content is generated based on cropped episode footage. Based on the embodiment shown in FIG13 above, as shown in FIG22, prior to step 460, the method further includes:

[0226] Step 920: Identify the bounding box of the main subject in the scene;

[0227] In some embodiments, the main subjects in a scene are assigned a priority. Based on the priority of the main subjects in the scene, the main subjects in the scene are identified, and a bounding box is used to select the corresponding main subject.

[0228] In some embodiments, the main subject of the image includes at least one of a face, a human body, an animal, a plant, furniture, clutter, and scenery. Optionally, the priority of a face is higher than that of a human body, the priority of a human body is higher than that of an animal, the priority of an animal is higher than that of a plant, the priority of a plant is higher than that of furniture, the priority of furniture is higher than that of clutter, and the priority of clutter is higher than that of scenery.

[0229] For example, as shown in Figure 23, the main subject 24 in the scene of the drama is identified, and the corresponding main subject 24 is selected by the bounding box 25.

[0230] In some embodiments, when there are multiple subjects of the same priority in the scene, the bounding boxes of the subjects of the same priority in the scene are identified simultaneously. For example, assuming that there are multiple faces in the same scene, multiple faces are identified simultaneously, and multiple faces are selected by bounding boxes.

[0231] Step 940: Based on the bounding box of the main subject of the image, crop the screen of the series to obtain the cropped screen of the series.

[0232] In some embodiments, the episode frame is cropped based on the bounding box of the identified main subject to obtain the cropped episode frame. Typically, the cropped episode frame is smaller than or equal to the original episode frame.

[0233] In some embodiments, when at least two bounding boxes are identified based on the main subject in the episode frame, step 940 above can be replaced by the following sub-steps:

[0234] Step 941: If at least two bounding boxes are identified based on the main subject of the image, merge the at least two bounding boxes to obtain a merged bounding box;

[0235] In some embodiments, if at least two bounding boxes are identified based on the main subject in the scene, the identified at least two bounding boxes are merged to obtain a merged bounding box. Optionally, the area of ​​the merged bounding box is determined based on the merged at least two bounding boxes.

[0236] In some embodiments, the merged bounding box covers all or part of the bounding boxes in at least two bounding boxes.

[0237] In some embodiments, when bounding boxes of at least two different subjects are identified and the total number of bounding boxes is greater than a first threshold, at least one group of bounding boxes to be merged is selected based on the priorities corresponding to the at least two different subjects. The number of bounding boxes in a group does not exceed a second threshold, and the priority of the bounding boxes in the group is higher than a third threshold.

[0238] By limiting the number of bounding boxes in the bounding box group, the number of main subjects in the final episode frame cropped based on the bounding boxes is also limited to a certain threshold. This helps to ensure that the number of main subjects in the cropped episode frame is reasonable and avoids a cluttered overall picture.

[0239] In some embodiments, for each bounding box group in at least one group of bounding boxes to be merged, the maximum top-left coordinate, maximum top-right coordinate, minimum bottom-left coordinate, and minimum bottom-right coordinate in the bounding box group are determined. Based on the maximum top-left coordinate, maximum top-right coordinate, minimum bottom-left coordinate, and minimum bottom-right coordinate, a merged bounding box for each bounding box group is generated.

[0240] In some embodiments, each bounding box corresponds to top-left, top-right, bottom-left, and bottom-right coordinates. For example, each bounding box can be represented by (x... start y start ) and (x end y end The combination is determined by (x). start y start (x) represents the lower left coordinate, (x) start y end (x) represents the top-left coordinate, (x) end y start (x) represents the lower right coordinate, (x) end y end () represents the upper right coordinate.

[0241] In some embodiments, when a group of bounding boxes includes multiple bounding boxes, the merging of bounding boxes can be achieved by min(x1). start x2 start , ..., xn start ), min(y1) start y2 start , ...,yn start ), max(x1) end x2 end , ..., xn end ), max(y1) end y2 end , ...,yn end )Sure.

[0242] Step 942: Crop the TV series image based on the merged bounding box to obtain the cropped TV series image.

[0243] In some embodiments, the episode frame is cropped based on a defined fusion bounding box to obtain a cropped episode frame. Typically, the cropped episode frame is smaller than or equal to the original episode frame.

[0244] By cropping the episode footage, the server can generate target multimedia content based on the cropped episode footage, thus ensuring that the episode footage corresponding to the generated target multimedia content is based on the optimized episode footage obtained from the cropping.

[0245] In some embodiments, after obtaining the cropped episode frame, the method further includes: determining whether the layout of the target multimedia content is a portrait or landscape layout based on the aspect ratio of the cropped episode frame. For example, as shown in FIG24, if the aspect ratio of the cropped episode frame is less than or equal to a first ratio, the layout of the target multimedia content is determined to be a portrait layout. For example, as shown in FIG24(a), the first area 26 is the area for placing information and / or target text related to the episode, the second area 27 is the border area of ​​the episode frame, and the third area 28 is the area for placing the episode frame. If the aspect ratio of the cropped episode frame is greater than a second ratio, the layout of the target multimedia content is determined to be a landscape layout. For example, as shown in FIG24(b).

[0246] Therefore, by cropping the screen of a TV series, the main subject of the screen can be displayed reasonably, making the generated target multimedia content more aesthetically pleasing and increasing its attractiveness.

[0247] In some embodiments, after determining the layout of the target multimedia content based on the aspect ratio of the cropped episode frame, the episode frame may not be able to completely fill the entire content area. In this case, the theme color of the episode frame can be extracted and used to fill the area outside the episode frame within the content area, which helps to make the generated target multimedia content more natural.

[0248] In some embodiments, the display size of the dialogue in a TV series is dynamically adjusted based on the text content corresponding to the dialogue.

[0249] In some embodiments, the display size of each character in the dialogue is dynamically adjusted based on the number of characters in the text content corresponding to the dialogue and the size of the text area used to display the dialogue. Optionally, the size of the text area occupied by each character in the dialogue is equal to the quotient of the size of the text area used to display the dialogue and the number of characters in the text content corresponding to the dialogue. For example, assuming the size of the text area used to display the dialogue is 10 square centimeters and the number of characters in the text content corresponding to the dialogue is 10, then the size of the text area occupied by each character in the dialogue is 1 square centimeter. The display size of each character in the dialogue is directly proportional to the size of the text area occupied by each character.

[0250] In some embodiments, the size of the text area occupied by each character in the dialogue is less than or equal to the maximum size of the text area. For example, suppose the maximum size of the text area is 3 square centimeters. The area size of the text area used to display the dialogue is 10 square centimeters, and the number of characters in the dialogue is 2. Since 10 / 2 = 5 square centimeters, and 5 is greater than 3, the size of the text area occupied by each character in the dialogue is 3 square centimeters.

[0251] In some embodiments, the size of the text area occupied by each character in the dialogue is greater than or equal to the minimum size of the text area. For example, suppose the minimum size of the text area is 1 square centimeter. If the size of the text area used to display the dialogue is 10 square centimeters, and the number of characters in the dialogue is 20, then since 10 / 20 = 0.5 square centimeters, and 0.5 is less than 1, the size of the text area occupied by each character in the dialogue is 1 square centimeter. However, if the number of characters in the dialogue exceeds the number of text areas, the dialogue is dynamically displayed according to the order of the characters in the dialogue.

[0252] The goal is to make the dialogue in the TV series adaptably displayed in the text area of ​​the target multimedia content. This means avoiding both excessively large display sizes of the dialogue, which would make the overall image of the target multimedia content look unattractive, and excessively small display sizes, which would result in poor clarity of the dialogue.

[0253] In some embodiments, the target multimedia content also includes promotional content provided by the promoter. For example, the target media content also includes advertising information to be promoted provided by the advertiser. That is, the promotional content includes dissemination content used for publicity or promotion. In some embodiments, the promotional content is dynamic display content, or the promotional content is static display content.

[0254] In some embodiments, the promotional content is pre-set. For example, before the function of generating target multimedia content with one click is enabled, the promoter pre-determines the promotional content to be promoted. Alternatively, after the function of generating target multimedia content with one click is enabled, the promotional content that the promoter needs to promote can be added in real time during the process of users using this function to generate target multimedia content. This facilitates the dynamic adjustment of the promotional content in the target multimedia content.

[0255] In some embodiments, the promotional content is displayed overlaid with the aforementioned TV series footage. Optionally, when the promotional content is displayed overlaid with the TV series footage, the promotional content is displayed in a manner that does not obscure the main subject of the TV series footage.

[0256] In some embodiments, promotional content is displayed as a display element within the aforementioned episode scene. For example, the promotional content is displayed in combination with at least one main subject within the episode scene. Optionally, the promotional content is a decorative element of at least one main subject. For example, the promotional content is a scene object within the scene depicted in the episode scene.

[0257] In some embodiments, the promotional content is displayed in combination with the aforementioned TV drama dialogue. Optionally, when the promotional content is an advertisement, the promotional content and the TV drama dialogue are merged to obtain merged text. The merged text is displayed in the display area where the TV drama dialogue is displayed.

[0258] In some embodiments, when the target multimedia content is dynamically displayed, the promotional content is displayed at the start of the target multimedia content, or at the end of the target multimedia content.

[0259] In some embodiments, a media content generation method proposed in this application can be used to generate electronic posters. In this application embodiment, the method is described using the terminal device 120 shown in FIG1 as an example. The method includes:

[0260] Step 240-1: In response to the selection of at least one character among several texts, determine the target text for generating the target e-poster;

[0261] In some embodiments, the selection operation for at least one character among several texts includes at least one of the following: single click, double click, left / right swipe, up / down swipe, long press, hover, facial recognition, and voice recognition. It is worth noting that the selection operation for at least one character among several texts includes, but is not limited to, the aforementioned methods. Those skilled in the art should understand that any operation capable of achieving the above functions falls within the protection scope of the embodiments of this application.

[0262] In some embodiments, in response to a selection operation of at least one character among a plurality of texts, a generation request for generating a target e-poster is sent to the server. Optionally, the generation request for generating the target e-poster includes a request to determine target text based on the selected at least one character. Optionally, the generation request for generating the target e-poster further includes a request to generate the target e-poster based on the target text. Optionally, the target text is key information in the selected at least one character.

[0263] In some embodiments, in response to a selection operation on at least one character among a plurality of texts, the display state of the selected at least one character is switched from a first display state to a second display state. The first display state is a state where at least one character is not selected, and the second display state is a state where at least one character is selected.

[0264] Step 260-1: Display the target e-poster generated based on the target text.

[0265] In some embodiments, the target digital poster includes images and dialogue from the TV series that match the target text.

[0266] In some embodiments, the content of the scene corresponding to the target text matches the text content expressed by the target text. Optionally, the meaning of the content of the scene corresponding to the target text is consistent with the meaning of the text content expressed by the target text. For example, assuming the target text is "xxx is so beautiful", and the content of scene 1 is "xxx has gorgeous makeup", then scene 1 is considered to match the target text.

[0267] In some embodiments, the text content corresponding to the TV drama dialogue that matches the target text matches the text content expressed by the target text. Optionally, the meaning of the text content corresponding to the TV drama dialogue that matches the target text is consistent with the meaning of the text content expressed by the target text. Optionally, the matching degree between the first text content corresponding to the TV drama dialogue that matches the target text and the second text content expressed by the target text is greater than or equal to a matching degree threshold. For example, assuming the target text is "xxx is really beautiful" and the TV drama dialogue is "xxx is really good-looking", since "beautiful" and "good-looking" express the same meaning, the TV drama dialogue is considered to match the target text.

[0268] In some embodiments, the target e-poster is generated by the server based on a generation request sent by the terminal device. Based on the generation request, the server determines the episode scenes and dialogue that match the target text, and then generates the target e-poster. Afterward, the server sends the generated target e-poster to the terminal device, which receives and displays the target e-poster.

[0269] In some embodiments, the target e-poster displayed on the terminal device is the one with the highest matching degree to the target text. That is, the server may generate multiple target e-posters based on the target text, but the matching degree between the multiple target e-posters and the target text is different. For example, assuming that the matching degree between target e-poster 1 and the target text is 90% and the matching degree between target e-poster 2 and the target text is 80%, the server will only send the target e-poster 1 with the highest matching degree to the terminal device.

[0270] In some embodiments, the target e-posters displayed on the terminal device are all target e-posters that match the target text. Optionally, the target e-posters displayed on the terminal device are those whose matching degree with the target text is greater than a matching degree threshold. For example, assuming that target e-poster 1 has a matching degree of 90% with the target text, target e-poster 2 has a matching degree of 80% with the target text, and target e-poster 3 has a matching degree of 60% with the target text, and the matching degree threshold is 75%, then the server will only send target e-poster 1 and target e-poster 2 to the terminal device.

[0271] In some embodiments, the material corresponding to the target electronic poster is a static image. In some embodiments, the material corresponding to the target electronic poster is a dynamic video frame.

[0272] In some embodiments, the targeted e-poster may also include promotional content provided by the promoter. For example, the targeted media content may also include advertising information to be promoted provided by the advertiser.

[0273] In some embodiments, the target electronic poster includes a side-by-side upper and lower half, with the upper half displaying images from the TV series and the lower half displaying dialogue from the TV series.

[0274] In some embodiments, the target electronic poster includes a side-by-side upper and lower half, with the upper half displaying dialogue from the TV series and the lower half displaying visuals from the TV series.

[0275] In some embodiments, the target electronic poster includes a left half and a right half side by side, with the left half displaying dialogue from the TV series and the right half displaying visuals from the TV series.

[0276] In some embodiments, the target electronic poster includes a left half and a right half side by side, with the left half displaying images from the show and the right half displaying dialogue from the show.

[0277] In some embodiments, the target electronic poster includes a content area. The content area displays screen images of the TV series, as well as dialogue overlaid on the screen images. Optionally, the dialogue may be displayed separately from the screen images. This application does not limit the display method of the screen images and dialogue in this embodiment.

[0278] In some embodiments, the target e-poster also includes an information area. The information area is used to display information and / or target text related to the show.

[0279] In some embodiments, the target electronic poster is automatically generated by a client or a server. When the target electronic poster is generated by a server, or by a terminal different from the terminal device, the following steps are included after step 240-1 and before step 260-1:

[0280] Step 251-1: Send the target text for generating the target electronic poster;

[0281] Step 252-1: Receive the target electronic poster generated based on the target text.

[0282] In some embodiments, the function of generating the target electronic poster needs to be triggered by a generation control. Step 240-1 above can also be replaced by the following sub-steps:

[0283] Step 241-1: In response to the selection of at least one character among a plurality of texts, display the selected text and the selection function bar;

[0284] In some embodiments, in response to a selection operation on at least one character among a plurality of texts, the selected text is switched from a first display state to a second display state. The first display state is an unselected state, and the second display state is a selected state.

[0285] In some embodiments, a selection bar is displayed in response to the selection of at least one character among a plurality of texts.

[0286] In some embodiments, the selected function bar includes a generate control. The generate control is used to trigger the generation of a target e-poster. Optionally, in response to the triggering operation of the generate control, a generation request for generating the target e-poster is sent to the server. Optionally, the generation request for generating the target e-poster includes a request to determine target text based on at least one selected character. Optionally, the generation request for generating the target e-poster further includes a request to generate the target e-poster based on the target text.

[0287] In some embodiments, the selected function bar may further include other function controls besides the generate control. Optionally, the selected function bar may further include at least one of a copy control, a share control, and a favorite control. The copy control is used to copy at least one selected text. The share control is used to share at least one selected text. The favorite control is used to favorite at least one selected text.

[0288] Step 242-1: In response to the triggering operation of the generation control, the selected text is determined as the target text for generating the target electronic poster.

[0289] In some embodiments, the triggering operation for the generated control includes at least one of the following: single click, double click, left / right swipe, up / down swipe, long press, hover, facial recognition, and voice recognition. It is worth noting that the triggering operation for the generated control includes, but is not limited to, several of the above-mentioned operations. Those skilled in the art should understand that any operation capable of achieving the above functions falls within the protection scope of the embodiments of this application.

[0290] In some embodiments, the target text for generating the target e-poster is determined directly based on the selected text. For example, the selected text is directly determined as the target text for generating the target e-poster.

[0291] In some embodiments, the target text for generating the target e-poster is determined indirectly based on the selected text. For example, the selected text is analyzed using a sentence segmentation algorithm, which segments the selected text into sentences, removes incomplete sentences from the selected text, and determines the remaining complete sentences as the target text for generating the target e-poster.

[0292] In some embodiments, the text includes a first text that is highlighted, the highlighting state indicating that the first text supports the electronic poster generation function. Step 241-1 above can also be replaced by the following sub-steps:

[0293] Step 2411-1: In response to the selection operation of the first text, display the first text as the selected text and display the selection function bar.

[0294] In some embodiments, the selection operation for the first text includes at least one of the following: single click, double click, left / right swipe, up / down swipe, long press, hover, facial recognition, and voice recognition. It is worth noting that the selection operation for the first text includes, but is not limited to, several of the above-mentioned operations. Those skilled in the art should understand that any operation capable of achieving the above functions falls within the protection scope of the embodiments of this application.

[0295] In some embodiments, in response to a selection operation on the first text, the first text is switched from a first display state to a second display state. The first display state is an unselected state, and the second display state is a selected state.

[0296] In some embodiments, in response to a selection operation on the first text, a selection bar is displayed. The selection bar includes a generation control.

[0297] In some embodiments, creators of text-based media content need to enable the e-poster generation function before publishing the content. With the e-poster generation function enabled, other users can select at least one character from a set of text within the text-based media content. If the creator of the text-based media content has not enabled the e-poster generation function, it indicates that the creator does not wish their content to be used to generate a target e-poster; in this case, it is impossible to generate a target e-poster by selecting at least one character from a set of text within the text-based media content.

[0298] In some embodiments, after the target e-poster is generated, the user can choose to share the target e-poster with the target object. The method further includes the following steps:

[0299] Step 320-1: Display the sharing control;

[0300] In some embodiments, the share control is used to trigger the sharing of the target e-poster.

[0301] In some embodiments, the sharing control and the target e-poster are displayed simultaneously.

[0302] In some embodiments, the sharing control is displayed after the target e-poster has already been displayed. That is, the target e-poster and the sharing control are displayed in a sequential order.

[0303] Step 340-1: In response to the triggering operation of the sharing control, share the target electronic poster to the specified network location.

[0304] In some embodiments, the triggering operation for the sharing control includes at least one of the following: single click, double click, left / right swipe, up / down swipe, long press, hover, facial recognition, and voice recognition. It is worth noting that the triggering operation for the sharing control includes, but is not limited to, several of the above-mentioned operations. Those skilled in the art should understand that any operation capable of achieving the above functions falls within the protection scope of the embodiments of this application.

[0305] In some embodiments, during the process of generating the target electronic poster, the terminal device will also prompt the user with some prompts indicating that the target electronic poster is being generated, the target electronic poster has been successfully generated, or the target electronic poster has failed to be generated.

[0306] In some embodiments, a first prompt message is displayed in response to the selection of at least one character among a plurality of texts. In some embodiments, the first prompt message is displayed in response to a triggering operation on a generation control. The first prompt message indicates that a target electronic poster is being generated.

[0307] In some embodiments, if the generation of the target e-poster fails, a second prompt message is displayed to indicate that the generation of the target e-poster has failed. In some embodiments, the second prompt message instructs the user to reselect at least one character from a plurality of texts.

[0308] In some embodiments, if the target electronic poster is successfully generated, a third prompt message is displayed to indicate that the target electronic poster has been successfully generated.

[0309] Next, the process of generating the target electronic poster will be described in detail. Optionally, the target electronic poster is generated by a server or a client. In this embodiment, the method is illustrated by example, with the method executed by the server 140 shown in Figure 1. The method includes:

[0310] Step 420-1: Receive the target text sent by the client for generating the target electronic poster;

[0311] In some embodiments, at least one text selected by the user of the client is received. That is, the selected text sent by the client is received.

[0312] In some embodiments, the target text is selected from the text media content published by the media account. That is, the target text is the selected text.

[0313] In some embodiments, the target text is determined by the server based on the selected text. For example, the selected text is analyzed using a sentence segmentation algorithm, which segments the selected text into sentences, removes incomplete sentences from the selected text, and determines the remaining complete sentences as the target text for generating the target e-poster.

[0314] In some embodiments, the target text is a first text highlighted in the text media content published by the media account, and the highlighted state is used to indicate that the first text supports the electronic poster generation function.

[0315] In some embodiments, the first text is determined based on the degree of matching between the associated episode material and the text media content.

[0316] Optionally, the associated TV series materials related to the text media content are determined based on the similarity between the semantic text of the text media content and the semantic text of the TV series dialogue subtitles in the associated TV series materials. The semantic similarity between the semantic text of the text media content and the semantic text of the TV series dialogue subtitles in the associated TV series materials is calculated, and TV series materials with a semantic similarity greater than or equal to a semantic similarity threshold are considered as associated TV series materials related to the text media content. Furthermore, the associated TV series materials can also be sorted according to the magnitude of semantic similarity.

[0317] Optionally, the text media content is segmented into sentences. The segmented sentences are then matched against subtitles from related TV series footage. The similarity between the segmented sentences and the subtitles is calculated. If the text similarity between the segmented sentences and the subtitles is greater than or equal to a text similarity threshold, the segmented sentence is designated as the first text. Furthermore, the first text can be sorted according to the magnitude of text similarity.

[0318] In some embodiments, the information related to the first text is sent synchronously to the relevant users when the media account publishes text media content. The display method of the first text on the client is different from the display method of other texts in the text media content. For example, the first text is highlighted, while other texts in the text media content are displayed normally.

[0319] In some embodiments, the information related to the first text includes at least one of the following: display information related to the first text, links related to the first text, associated TV series material related to the first text, and associated target e-posters related to the first text. The display information related to the first text indicates how the first text is displayed on the client. The links related to the first text indicate that the first text supports e-poster generation. For example, in response to a triggering operation on the first text, the user can be redirected to a poster generation interface. The associated TV series material related to the first text indicates the associated TV series material corresponding to the first text, which is used to generate the target e-poster related to the first text. The associated target e-poster related to the first text refers to a pre-generated target e-poster. In this case, when the user triggers the first text, the associated target e-poster can be directly obtained.

[0320] It should be noted that the server or client for determining the first text described in this application embodiment may be the same as or different from the server or client for generating the poster. This application embodiment does not impose any limitation on this.

[0321] In some embodiments, the target media content includes at least one of animation, poster, audio, and video.

[0322] Step 440-1: Generate a target electronic poster based on the target text;

[0323] In some embodiments, the target text is matched with corresponding episode scenes and dialogue within a drama series database. This drama series database is pre-generated by the server. The server pre-segments different drama series to obtain a database containing episode scenes and dialogue. It should be understood that the episode scenes and dialogue in this embodiment are correlated. Typically, multiple episode scenes may correspond to the same dialogue.

[0324] In some embodiments, a target e-poster including episode images and episode lines is generated based on episode images and episode lines that match the target text.

[0325] Step 460-1: Send the target electronic poster to the client.

[0326] In some embodiments, the target media content includes episode scenes and episode dialogue that match the target text.

[0327] In some embodiments, the targeted e-poster may also include promotional content provided by the promoter. For example, the targeted media content may also include advertising information to be promoted provided by the advertiser.

[0328] In some embodiments, there is a preset correspondence between episode scenes and episode dialogue. Step 440-1 above can also be replaced by the following sub-steps:

[0329] Step 441-1: Based on the target text, query the TV series images and lines that match the target text in the preset correspondence. The TV series images and lines are used to generate the target electronic poster.

[0330] In some embodiments, the preset correspondence is determined in advance by the server. For example, taking a first TV series as an example, the preset correspondence between the episode scenes and dialogue included in the first TV series is determined by the following method. Before step 441-1 above, the method further includes:

[0331] Step 520-1: Obtain the first episode video and episode dialogue subtitles;

[0332] In some embodiments, the video of the first episode and the subtitles of the episode's dialogue corresponding to the first episode are obtained. Optionally, the video of the first episode and the subtitles of the episode's dialogue are obtained from a video database.

[0333] Step 540-1: Divide the first episode video into at least one video segment according to the start and end times of the subtitles in the episode dialogue;

[0334] In some embodiments, the first episode video is divided into at least one video segment according to the start and end times of the subtitles corresponding to the dialogue subtitles of the same episode. It should be understood that the first number of video segments is greater than or equal to the second number of episode dialogue subtitles.

[0335] In this embodiment, an example is given of a one-to-one correspondence between video segments and TV series dialogue subtitles. That is, each video segment corresponds to one TV series dialogue subtitle.

[0336] Step 560-1: Extract at least one episode frame from the video segments;

[0337] In some embodiments, the video is divided into segments, with each segment consisting of a video frame. Each video segment is divided into at least one video frame. Scenes from each video frame are then extracted. It is important to understand that there is a one-to-one correspondence between video frames and scenes from each video frame.

[0338] In some embodiments, at least one target video frame is extracted from the video slice, and the quality of each target video frame is greater than or equal to a quality threshold.

[0339] Step 580: Generate a preset correspondence based on at least one episode scene and episode dialogue subtitles.

[0340] In some embodiments, a preset correspondence between episode scenes and episode dialogue is generated based on at least one extracted episode scene.

[0341] Optionally, based on the matching degree between the target text and the dialogue in the TV series, and the correspondence between TV series scenes and dialogue in the preset correspondence, the TV series scenes and dialogue that match the target text are determined. The TV series dialogue is used to connect the TV series dialogue and the target text.

[0342] In some embodiments, there is a preset correspondence between episode scenes and episode dialogue. Step 440-1 above can also be replaced by the following sub-steps:

[0343] Step 442-1: Calculate the first feature vector corresponding to the target text, and calculate the vector similarity between the first feature vector and at least one second feature vector; determine the TV scene and TV line corresponding to the second feature vector with the highest vector similarity as the TV scene and TV line matching the target text, and use the TV scene and TV line to generate the target electronic poster.

[0344] In some embodiments, the episode scenes and episode lines corresponding to the second feature vector with a vector similarity greater than or equal to the similarity threshold are all determined as episode scenes and episode lines that match the target text.

[0345] In some embodiments, the second feature vector is predetermined by the server. For example, taking a second episode as an example, the second feature vector is determined by the following method. Prior to step 442-1 above, the method further includes:

[0346] Step 610-1: Obtain the video of the second episode;

[0347] In some embodiments, the second episode video is obtained from a video database.

[0348] Step 620-1: Divide the second episode video into at least one video segment according to the division step size;

[0349] In some embodiments, the step size is random. In some embodiments, the step size is preset. For example, the step size is 3 seconds.

[0350] In some embodiments, the second episode video is continuously divided according to a division step size. That is, at least one video segment in the embodiments of this application is continuous. It can be understood that combining at least one video segment in sequence can obtain a complete second episode video.

[0351] Step 630-1: Extract computer vision information for each video segment from at least one video segment;

[0352] In some embodiments, computer vision algorithms are used to extract computer vision information from each video segment. Optionally, the computer vision information includes subject information and / or text information in each video segment. Subject information includes human body information, object information, landscape information, etc.

[0353] Step 640-1: Based on the computer vision information of each video segment, generate content description information for each video segment;

[0354] In some embodiments, content description information for each video segment is generated based on the extracted computer vision information for each video segment.

[0355] In some embodiments, content description information for each video segment in at least one video segment is obtained based on a multimodal visual recognition model. In this case, it is unnecessary to perform step 630-1 above to obtain the computer vision information for each video segment. That is, steps 630-1 and 640-1 can be implemented directly based on the multimodal visual recognition model.

[0356] Step 650-1: Generate a second feature vector for each video segment based on the content description information of each video segment.

[0357] In some embodiments, a second feature vector for each video segment is generated based on the content description information of each video segment and a vector database. The vector database includes the correspondence between the content description information and the feature vectors.

[0358] In some embodiments, a second feature vector for each video segment is generated based on a multimodal embedding algorithm. Optionally, based on a multimodal embedding algorithm, at least two of the audio, image, and text information corresponding to each video segment are input to generate a second feature vector corresponding to each video segment. In this case, steps 630-1 and 640-1 above do not need to be performed. That is, steps 630-1 to 650-1 above can be implemented directly based on the multimodal embedding algorithm.

[0359] In some embodiments, the above method further includes:

[0360] Step 720-1: Extract at least one episode frame from the video segments;

[0361] In some embodiments, the video is divided into segments, with each segment consisting of a video frame. Each video segment is divided into at least one video frame. Scenes from each video frame are then extracted. It is important to understand that there is a one-to-one correspondence between video frames and scenes from each video frame.

[0362] In some embodiments, at least one target video frame is extracted from the video slice, and the quality of each target video frame is greater than or equal to a quality threshold.

[0363] Step 740-1: Store the correspondence between at least one episode frame and the second feature vector of the video segment.

[0364] In some embodiments, given that the second feature vector corresponding to each video segment and at least one episode frame corresponding to each video segment are known, there is a correspondence between at least one episode frame and the second feature vector.

[0365] In some embodiments, when extracting at least one episode frame corresponding to a video segment, the quality is determined based on the quality of the video frames in each video segment. Optionally, step 560-1 or step 720-1 above can be replaced by the following sub-steps:

[0366] Step 820-1: Score each video frame in the video segment according to at least one scoring factor to obtain the quality score corresponding to each video frame;

[0367] In some embodiments, each video segment includes multiple video frames, and each of the multiple video frames corresponds to a quality score.

[0368] In some embodiments, at least one scoring factor includes at least one of brightness, sharpness, and confidence in the identification of the subject in the video frame. Brightness refers to the light intensity in the video frame. Sharpness refers to the clarity of the image in the video frame. Confidence in the identification of the subject in the video frame refers to the probability that the subject in the video frame is identified.

[0369] Step 840-1: Based on the quality score, extract at least one episode frame from each video frame in the video segment.

[0370] In some embodiments, at least one episode frame is extracted from the target video frame in the video segment based on the quality score of each video frame. The quality score corresponding to the target video frame is greater than or equal to a quality threshold score.

[0371] In some embodiments, the video frames in the video segment are sorted from high to low or from low to high according to their quality scores, and at least one video frame with a quality score greater than or equal to a quality threshold score is selected. At least one episode scene is extracted from the at least one video frame with a quality score greater than or equal to the quality threshold score.

[0372] In some embodiments, the target e-poster is generated based on the corresponding scenes and dialogue from the TV series that match the target text. Step 440-1 above can also be replaced by the following sub-steps:

[0373] Step 443-1: If there are at least two candidate episode frames in the target text match, calculate the image quality score of at least two candidate episode frames;

[0374] In some embodiments, each candidate episode frame is scored according to at least one scoring factor to obtain a picture quality score for each candidate episode frame. Optionally, the at least one scoring factor includes at least one of brightness, sharpness, and confidence in the identification of the subject in the picture.

[0375] Step 444-1: Select the target episode from at least two candidate episodes based on the picture quality score;

[0376] In some embodiments, based on the image quality score of each candidate episode image, at least one target episode image is selected from at least two candidate episode images that match the target text. The image quality score corresponding to the target episode image is greater than or equal to an image quality score threshold.

[0377] Step 445-1: Generate a target e-poster based on the target TV series images and the TV series dialogue that matches the target text.

[0378] In some embodiments, the target episode image includes at least one, and the episode dialogue matching the target text may also include at least one. However, it is important to understand that each episode image corresponds to its own episode dialogue; that is, there is a one-to-one correspondence between episode images and episode dialogue. For example, suppose the target text matches the first episode image of a first series and the second episode image of a second series. Here, the first and second series are two different series. The first episode image corresponds to episode dialogue 1, and the second episode image corresponds to episode dialogue 2.

[0379] The one-to-one correspondence between scene and dialogue in the TV series mentioned above only applies when the scene corresponds to a specific line of dialogue. If no matching dialogue is found for the target text, the target text or a portion thereof will be used to generate matching dialogue.

[0380] In some embodiments, if there is a candidate episode image that matches the target text, the candidate episode image is determined as the target episode image. That is, if there is one and only one matching candidate episode image, that candidate episode image is the target episode image used to generate the target e-poster.

[0381] In some embodiments, the target digital poster is generated based on cropped episode footage. Prior to step 460-1 above, the method further includes:

[0382] Step 920-1: Identify the bounding box of the main subject in the scene;

[0383] In some embodiments, the main subjects in a scene are assigned a priority. Based on the priority of the main subjects in the scene, the main subjects in the scene are identified, and a bounding box is used to select the corresponding main subject.

[0384] In some embodiments, the main subject of the image includes at least one of a face, a human body, an animal, a plant, furniture, clutter, and scenery. Optionally, the priority of a face is higher than that of a human body, the priority of a human body is higher than that of an animal, the priority of an animal is higher than that of a plant, the priority of a plant is higher than that of furniture, the priority of furniture is higher than that of clutter, and the priority of clutter is higher than that of scenery.

[0385] In some embodiments, when there are multiple subjects of the same priority in the scene, the bounding boxes of the subjects of the same priority in the scene are identified simultaneously. For example, assuming that there are multiple faces in the same scene, multiple faces are identified simultaneously, and multiple faces are selected by bounding boxes.

[0386] Step 940-1: Based on the bounding box of the main subject of the image, crop the screen of the series to obtain the cropped screen of the series.

[0387] In some embodiments, the episode frame is cropped based on the bounding box of the identified main subject to obtain the cropped episode frame. Typically, the cropped episode frame is smaller than or equal to the original episode frame.

[0388] In some embodiments, when at least two bounding boxes are identified based on the main subject in the episode frame, step 940-1 above can be replaced by the following sub-steps:

[0389] Step 941-1: If at least two bounding boxes are identified based on the main subject of the image, merge the at least two bounding boxes to obtain a merged bounding box;

[0390] In some embodiments, if at least two bounding boxes are identified based on the main subject in the scene, the identified at least two bounding boxes are merged to obtain a merged bounding box. Optionally, the area of ​​the merged bounding box is determined based on the merged at least two bounding boxes.

[0391] In some embodiments, the merged bounding box covers all or part of the bounding boxes in at least two bounding boxes.

[0392] In some embodiments, when bounding boxes of at least two different subjects are identified and the total number of bounding boxes is greater than a first threshold, at least one group of bounding boxes to be merged is selected based on the priorities corresponding to the at least two different subjects. The number of bounding boxes in a group does not exceed a second threshold, and the priority of the bounding boxes in the group is higher than a third threshold.

[0393] By limiting the number of bounding boxes in the bounding box group, the number of main subjects in the final episode frame cropped based on the bounding boxes is also limited to a certain threshold. This helps to ensure that the number of main subjects in the cropped episode frame is reasonable and avoids a cluttered overall picture.

[0394] In some embodiments, for each bounding box group in at least one group of bounding boxes to be merged, the maximum top-left coordinate, maximum top-right coordinate, minimum bottom-left coordinate, and minimum bottom-right coordinate in the bounding box group are determined. Based on the maximum top-left coordinate, maximum top-right coordinate, minimum bottom-left coordinate, and minimum bottom-right coordinate, a merged bounding box for each bounding box group is generated.

[0395] In some embodiments, each bounding box corresponds to top-left, top-right, bottom-left, and bottom-right coordinates. For example, each bounding box can be represented by (x... start y start ) and (x end y end The combination is determined by (x). start ystart (x) represents the lower left coordinate, (x) start y end (x) represents the top-left coordinate, (x) end y start (x) represents the lower right coordinate, (x) end y end () represents the upper right coordinate.

[0396] In some embodiments, when a group of bounding boxes includes multiple bounding boxes, the merging of bounding boxes can be achieved by min(x1). start x2 start , ..., xn start ), min(y1) start y2 start , ...,yn start ), max(x1) end x2 end , ..., xn end ), max(y1) end y2 end , ...,yn end )Sure.

[0397] Step 942-1: Crop the TV series image based on the merged bounding box to obtain the cropped TV series image.

[0398] In some embodiments, the episode frame is cropped based on a defined fusion bounding box to obtain a cropped episode frame. Typically, the cropped episode frame is smaller than or equal to the original episode frame.

[0399] By cropping the TV series footage, the server can generate the target e-poster based on the cropped footage, thus ensuring that the TV series footage corresponding to the generated target e-poster is based on the optimized footage obtained from the cropping.

[0400] In some embodiments, after obtaining the cropped TV series image, the method further includes: determining whether the layout of the target e-poster is a portrait or landscape layout based on the aspect ratio of the cropped TV series image. Thus, cropping the TV series image allows for a more reasonable display of the main subject within the image, making the generated target e-poster more aesthetically pleasing and increasing its appeal.

[0401] In some embodiments, after determining the layout of the target e-poster based on the aspect ratio of the cropped episode frame, the episode frame may not be able to completely fill the entire content area. In this case, the theme color of the episode frame can be extracted and used to fill the area outside the episode frame within the content area, which helps to make the generated target e-poster look more natural.

[0402] In some embodiments, the display size of the dialogue in a TV series is dynamically adjusted based on the text content corresponding to the dialogue.

[0403] In some embodiments, the display size of each character in the dialogue is dynamically adjusted based on the number of characters in the text content corresponding to the dialogue and the size of the text area used to display the dialogue. Optionally, the size of the text area occupied by each character in the dialogue is equal to the quotient of the size of the text area used to display the dialogue and the number of characters in the text content corresponding to the dialogue. For example, assuming the size of the text area used to display the dialogue is 10 square centimeters and the number of characters in the text content corresponding to the dialogue is 10, then the size of the text area occupied by each character in the dialogue is 1 square centimeter. The display size of each character in the dialogue is directly proportional to the size of the text area occupied by each character.

[0404] In some embodiments, the size of the text area occupied by each character in the dialogue is less than or equal to the maximum size of the text area. For example, suppose the maximum size of the text area is 3 square centimeters. The area size of the text area used to display the dialogue is 10 square centimeters, and the number of characters in the dialogue is 2. Since 10 / 2 = 5 square centimeters, and 5 is greater than 3, the size of the text area occupied by each character in the dialogue is 3 square centimeters.

[0405] In some embodiments, the size of the text area occupied by each character in the dialogue is greater than or equal to the minimum size of the text area. For example, suppose the minimum size of the text area is 1 square centimeter. If the size of the text area used to display the dialogue is 10 square centimeters, and the number of characters in the dialogue is 20, then since 10 / 20 = 0.5 square centimeters, and 0.5 is less than 1, the size of the text area occupied by each character in the dialogue is 1 square centimeter. However, if the number of characters in the dialogue exceeds the number of text areas, the dialogue is dynamically displayed according to the order of the characters in the dialogue.

[0406] The goal is to make the dialogue from the TV series adaptably displayed in the text area of ​​the target electronic poster. This means avoiding both excessively large display sizes that would make the overall image of the target electronic poster unattractive and excessively small display sizes that would result in poor clarity of the dialogue.

[0407] Figure 25 shows a block diagram of a media content generation apparatus provided in an exemplary embodiment of this application. The apparatus includes:

[0408] Display module 2510 is used to display text media content published by media accounts.

[0409] In some embodiments, a media account refers to a user account logged in on a client that supports editing and publishing text media content. Optionally, a media account refers to a user account corresponding to an application on a client that supports editing and publishing text media content. For example, a user account logged in on a social media application. Optionally, the social media application needs to support editing and publishing text media content.

[0410] In some embodiments, text-based media content includes a number of texts. Optionally, text-based media content also includes other forms of content such as images and videos. For example, text-based media content includes articles from public WeChat accounts, current affairs commentary articles, etc.

[0411] In some embodiments, the text media content is pre-created by the creator corresponding to the media account.

[0412] In some embodiments, text media content may also include other forms of content such as images and videos.

[0413] The determination module 2520 is used to determine the target text for generating target multimedia content in response to a selection operation of at least one character among a plurality of texts.

[0414] In some embodiments, the selection operation for at least one character among several texts includes at least one of the following: single click, double click, left / right swipe, up / down swipe, long press, hover, facial recognition, and voice recognition. It is worth noting that the selection operation for at least one character among several texts includes, but is not limited to, the aforementioned methods. Those skilled in the art should understand that any operation capable of achieving the above functions falls within the protection scope of the embodiments of this application.

[0415] In some embodiments, in response to a selection operation of at least one character among a plurality of texts, a generation request for generating target multimedia content is sent to the server. Optionally, the generation request for generating target multimedia content includes a request to determine target text based on the selected at least one character. Optionally, the generation request for generating target multimedia content further includes a request to generate target multimedia content based on the target text. Optionally, the target text is key information in the selected at least one character.

[0416] In some embodiments, in response to a selection operation on at least one character among a plurality of texts, the display state of the selected at least one character is switched from a first display state to a second display state. The first display state is a state where at least one character is not selected, and the second display state is a state where at least one character is selected.

[0417] Display module 2510 is also used to display target multimedia content generated based on target text.

[0418] In some embodiments, the target multimedia content includes episode scenes and episode dialogue that match the target text.

[0419] In some embodiments, the content of the scene corresponding to the episode that matches the target text matches the text content expressed by the target text. Optionally, the meaning of the content of the scene corresponding to the episode that matches the target text is consistent with the meaning of the text content expressed by the target text.

[0420] In some embodiments, the text content corresponding to the TV drama dialogue that matches the target text matches the text content expressed by the target text. Optionally, the meaning of the text content corresponding to the TV drama dialogue that matches the target text is consistent with the meaning of the text content expressed by the target text. Optionally, the matching degree between the first text content corresponding to the TV drama dialogue that matches the target text and the second text content expressed by the target text is greater than or equal to a matching degree threshold.

[0421] In some embodiments, the target multimedia content is generated by the server based on a generation request sent by the terminal device. Based on the generation request, the server determines the episode scenes and dialogue that match the target text, and then generates the target multimedia content. Afterward, the server sends the generated target multimedia content to the terminal device, which receives and displays the target multimedia content.

[0422] In some embodiments, the target multimedia content displayed on the terminal device is the one that matches the target text most closely. That is, the server may generate multiple target multimedia content pieces based on the target text, but the matching degree between these multiple pieces of target multimedia content and the target text is different. For example, assuming that target multimedia content 1 matches the target text with a 90% matching degree and target multimedia content 2 matches the target text with an 80% matching degree, the server will only send target multimedia content 1, which has the highest matching degree, to the terminal device.

[0423] In some embodiments, the target multimedia content displayed on the terminal device is all target multimedia content that matches the target text. Optionally, the target multimedia content displayed on the terminal device is the target multimedia content whose matching degree with the target text is greater than a matching degree threshold. For example, assuming that the matching degree of target multimedia content 1 with the target text is 90%, the matching degree of target multimedia content 2 with the target text is 80%, and the matching degree of target multimedia content 3 with the target text is 60%, and the matching degree threshold is 75%, then the server will only send target multimedia content 1 and target multimedia content 2 to the terminal device.

[0424] In some embodiments, the source material corresponding to the target multimedia content is a static image. In some embodiments, the source material corresponding to the target multimedia content is a dynamic video frame.

[0425] In some embodiments, the function of generating target multimedia content needs to be triggered by a generation control.

[0426] In some embodiments, the target multimedia content may also include promotional content provided by the promoter. For example, the target media content may also include advertising information to be promoted provided by the advertiser.

[0427] The display module 2510 is also used to display the selected text and the selection function bar in response to the selection operation of at least one character among a plurality of texts.

[0428] In some embodiments, in response to a selection operation on at least one character among a plurality of texts, the selected text is switched from a first display state to a second display state. The first display state is an unselected state, and the second display state is a selected state.

[0429] In some embodiments, a selection bar is displayed in response to the selection of at least one character among a plurality of texts.

[0430] In some embodiments, the selected function bar includes a generate control. The generate control is used to trigger the generation of target multimedia content. Optionally, in response to the triggering operation of the generate control, a generation request for generating target multimedia content is sent to the server. Optionally, the generation request for generating target multimedia content includes a request to determine target text based on at least one selected character. Optionally, the generation request for generating target multimedia content further includes a request to generate target multimedia content based on the target text.

[0431] In some embodiments, the selected function bar may further include other function controls besides the generate control. Optionally, the selected function bar may further include at least one of a copy control, a share control, and a favorite control. The copy control is used to copy at least one selected text. The share control is used to share at least one selected text. The favorite control is used to favorite at least one selected text.

[0432] The determination module 2520 is used to determine the selected text as the target text for generating target multimedia content in response to a trigger operation on the generation control.

[0433] In some embodiments, the triggering operation for the generated control includes at least one of the following: single click, double click, left / right swipe, up / down swipe, long press, hover, facial recognition, and voice recognition. It is worth noting that the triggering operation for the generated control includes, but is not limited to, several of the above-mentioned operations. Those skilled in the art should understand that any operation capable of achieving the above functions falls within the protection scope of the embodiments of this application.

[0434] In some embodiments, the target text for generating the target multimedia content is determined directly based on the selected text. For example, the selected text is directly determined as the target text for generating the target multimedia content.

[0435] In some embodiments, the target text for generating the target multimedia content is determined indirectly based on the selected text. For example, the selected text is analyzed using a sentence segmentation algorithm, which segments the selected text into sentences, removes incomplete sentences from the selected text, and determines the remaining complete sentences as the target text for generating the target multimedia content.

[0436] In some embodiments, the text includes a first text that is highlighted, the highlighting state indicating that the first text supports multimedia content generation functionality.

[0437] The display module 2510 is also used to respond to the selection operation of the first text, display the first text as the selected text, and display the selected function bar.

[0438] In some embodiments, the selection operation for the first text includes at least one of the following: single click, double click, left / right swipe, up / down swipe, long press, hover, facial recognition, and voice recognition. It is worth noting that the selection operation for the first text includes, but is not limited to, several of the above-mentioned operations. Those skilled in the art should understand that any operation capable of achieving the above functions falls within the protection scope of the embodiments of this application.

[0439] In some embodiments, in response to a selection operation on the first text, the first text is switched from a first display state to a second display state. The first display state is an unselected state, and the second display state is a selected state.

[0440] In some embodiments, in response to a selection operation on the first text, a selection bar is displayed. The selection bar includes a generation control.

[0441] In some embodiments, the first text is determined by the creator of the text media content before it is published.

[0442] In some embodiments, creators of text-based media content need to enable the multimedia content generation function before publishing the content. With the multimedia content generation function enabled, other users can select at least one character from a plurality of texts within the text-based media content. If the creator of the text-based media content has not enabled the multimedia content generation function, it indicates that the creator does not wish their content to be used to generate relevant target multimedia content; in this case, it is impossible to generate target multimedia content by selecting at least one character from a plurality of texts within the text-based media content.

[0443] In some embodiments, the first text is determined by the backend server from the text media content based on the matching degree between the associated material and the text media content.

[0444] In some embodiments, in response to an editing operation on any target text in at least one target text, either target text is retained or any target text is deleted.

[0445] In some embodiments, after the target multimedia content is generated, the user may choose to share the target multimedia content with the target object.

[0446] Display module 2510 is also used to display sharing controls.

[0447] In some embodiments, the share control is used to trigger the sharing of target multimedia content.

[0448] In some embodiments, the sharing control and the target multimedia content are displayed simultaneously.

[0449] In some embodiments, the sharing control is displayed after the target multimedia content has already been displayed. That is, the target multimedia content and the sharing control are displayed sequentially.

[0450] The sending module 2530 is used to share target multimedia content to a specified network location in response to a trigger operation of the sharing control.

[0451] In some embodiments, the triggering operation for the sharing control includes at least one of the following: single click, double click, left / right swipe, up / down swipe, long press, hover, facial recognition, and voice recognition. It is worth noting that the triggering operation for the sharing control includes, but is not limited to, several of the above-mentioned operations. Those skilled in the art should understand that any operation capable of achieving the above functions falls within the protection scope of the embodiments of this application.

[0452] In some embodiments, the target multimedia content includes a side-by-side upper and lower half-area, with the upper half-area displaying screen footage and the lower half-area displaying dialogue from the series.

[0453] In some embodiments, the target multimedia content includes a side-by-side upper and lower half-area, with the upper half-area displaying dialogue from the series and the lower half-area displaying visuals from the series.

[0454] In some embodiments, the target multimedia content includes a left half-area and a right half-area side by side, with the left half-area displaying dialogue from the series and the right half-area displaying images from the series.

[0455] In some embodiments, the target multimedia content includes a left half-area and a right half-area side by side, with the left half-area displaying the scene and the right half-area displaying the dialogue.

[0456] In some embodiments, the target multimedia content includes a content area. The content area displays screen images and dialogue overlaid on the screen images. Optionally, the dialogue may be displayed separately from the screen images. This application does not limit the display method of the screen images and dialogue in these embodiments.

[0457] In some embodiments, the target multimedia content further includes an information area. The information area is used to display information and / or target text related to the episode.

[0458] In some embodiments, during the process of generating target multimedia content, the terminal device may also prompt the user with some information to indicate that the target multimedia content is being generated, or that the target multimedia content has been successfully generated, or that the target multimedia content has failed to be generated.

[0459] In some embodiments, a first prompt message is displayed in response to the selection of at least one character among a plurality of texts. In some embodiments, the first prompt message is displayed in response to a triggering operation on a generation control. The first prompt message indicates that target multimedia content is being generated.

[0460] In some embodiments, if the generation of target multimedia content fails, a second prompt message is displayed to indicate that the generation of target multimedia content has failed. In some embodiments, the second prompt message instructs the user to reselect at least one character from a plurality of texts.

[0461] In some embodiments, if the target multimedia content is successfully generated, a third prompt message is displayed to indicate that the target multimedia content has been successfully generated.

[0462] In some embodiments, the sending module 2530 is further configured to send a generation request for generating target multimedia content to the server.

[0463] In some embodiments, the above-described apparatus further includes:

[0464] The receiving module 2540 is used to receive target multimedia content sent by the server.

[0465] Figure 26 shows a block diagram of a media content generation apparatus provided in an exemplary embodiment of this application. The apparatus includes:

[0466] The receiving module 2610 is used to receive the target text sent by the client for generating target multimedia content.

[0467] In some embodiments, at least one text selected by the user of the client is received. That is, the selected text sent by the client is received.

[0468] In some embodiments, the target text is selected from the text media content published by the media account. That is, the target text is the selected text.

[0469] In some embodiments, the target text is determined by the server based on the selected text. For example, the selected text is analyzed using a sentence segmentation algorithm, which segments the selected text into sentences, removes incomplete sentences from the selected text, and determines the remaining complete sentences as the target text for generating the target multimedia content.

[0470] In some embodiments, the target text is a first text highlighted in the text media content published by the media account, the highlighted state being used to indicate that the first text supports multimedia content generation functionality.

[0471] In some embodiments, the first text is determined based on the degree of matching between the associated episode material and the text media content.

[0472] Optionally, the associated TV series materials related to the text media content are determined based on the similarity between the semantic text of the text media content and the semantic text of the TV series dialogue subtitles in the associated TV series materials. The semantic similarity between the semantic text of the text media content and the semantic text of the TV series dialogue subtitles in the associated TV series materials is calculated, and TV series materials with a semantic similarity greater than or equal to a semantic similarity threshold are considered as associated TV series materials related to the text media content. Furthermore, the associated TV series materials can also be sorted according to the magnitude of semantic similarity.

[0473] Optionally, the text media content is segmented into sentences. The segmented sentences are then matched against subtitles from related TV series footage. The similarity between the segmented sentences and the subtitles is calculated. If the text similarity between the segmented sentences and the subtitles is greater than or equal to a text similarity threshold, the segmented sentence is designated as the first text. Furthermore, the first text can be sorted according to the magnitude of text similarity.

[0474] In some embodiments, the target media content includes at least one of animation, poster, audio, and video.

[0475] The generation module 2620 is used to generate target multimedia content based on target text.

[0476] In some embodiments, the target text is matched with corresponding episode scenes and dialogue within a drama series database. This drama series database is pre-generated by the server. The server pre-segments different drama series to obtain a database containing episode scenes and dialogue. It should be understood that the episode scenes and dialogue in this embodiment are correlated. Typically, multiple episode scenes may correspond to the same dialogue.

[0477] In some embodiments, target multimedia content including episode images and episode dialogues is generated based on episode images and episode dialogues that match the target text.

[0478] The sending module 2630 is used to send target multimedia content to the client.

[0479] In some embodiments, the target media content includes episode scenes and episode dialogue that match the target text.

[0480] In some embodiments, the target multimedia content may also include promotional content provided by the promoter. For example, the target media content may also include advertising information to be promoted provided by the advertiser.

[0481] In some embodiments, there is a preset correspondence between episode images and episode dialogue.

[0482] The generation module 2620 is also used to query TV series images and TV series lines that match the target text in a preset correspondence based on the target text. The TV series images and TV series lines are used to generate target multimedia content.

[0483] In some embodiments, the preset mapping relationship is predetermined by the server.

[0484] The generation module 2620 is also used to obtain the first episode video and episode dialogue subtitles.

[0485] In some embodiments, the video of the first episode and the subtitles of the episode's dialogue corresponding to the first episode are obtained. Optionally, the video of the first episode and the subtitles of the episode's dialogue are obtained from a video database.

[0486] The generation module 2620 is also used to divide the first episode video into at least one video segment according to the start and end times of the subtitles in the episode dialogue.

[0487] In some embodiments, the first episode video is divided into at least one video segment according to the start and end times of the subtitles corresponding to the dialogue subtitles of the same episode. It should be understood that the first number of video segments is greater than or equal to the second number of episode dialogue subtitles.

[0488] The generation module 2620 is also used to extract at least one episode frame from the video segments.

[0489] In some embodiments, the video is divided into segments, with each segment consisting of a video frame. Each video segment is divided into at least one video frame. Scenes from each video frame are then extracted. It is important to understand that there is a one-to-one correspondence between video frames and scenes from each video frame.

[0490] In some embodiments, at least one target video frame is extracted from the video slice, and the quality of each target video frame is greater than or equal to a quality threshold.

[0491] The generation module 2620 is also used to generate a preset correspondence based on at least one episode screen and episode dialogue subtitles.

[0492] In some embodiments, a preset correspondence between episode scenes and episode dialogue is generated based on at least one extracted episode scene.

[0493] In some embodiments, there is a one-to-one correspondence between episode images and episode dialogue.

[0494] In some embodiments, there is a many-to-one correspondence between episode images and episode dialogue.

[0495] Normally, there is no one-to-many correspondence between scenes and dialogue in a TV series. However, in some special cases, there may be a one-to-many correspondence between scenes and dialogue, which is not limited in this embodiment.

[0496] In some embodiments, there is a preset correspondence between episode images and episode dialogue.

[0497] The generation module 2620 is also used to calculate the first feature vector corresponding to the target text, and to calculate the vector similarity between the first feature vector and at least one second feature vector; the TV scene and TV line corresponding to the second feature vector with the highest vector similarity are determined as the TV scene and TV line matching the target text, and the TV scene and TV line are used to generate target multimedia content.

[0498] In some embodiments, the episode scenes and episode lines corresponding to the second feature vector with a vector similarity greater than or equal to the similarity threshold are all determined as episode scenes and episode lines that match the target text.

[0499] In some embodiments, the second feature vector is predetermined by the server.

[0500] The generation module 2620 is also used to obtain the video of the second episode.

[0501] In some embodiments, the second episode video is obtained from a video database.

[0502] The generation module 2620 is also used to divide the second episode video into at least one video segment according to the division step size.

[0503] In some embodiments, the partitioning step size is random. In some embodiments, the partitioning step size is preset.

[0504] In some embodiments, the second episode video is continuously divided according to a division step size. That is, at least one video segment in the embodiments of this application is continuous. It can be understood that combining at least one video segment in sequence can obtain a complete second episode video.

[0505] The generation module 2620 is also used to extract computer vision information from each video segment in at least one video segment.

[0506] In some embodiments, computer vision algorithms are used to extract computer vision information from each video segment. Optionally, the computer vision information includes subject information and / or text information in each video segment. Subject information includes human body information, object information, landscape information, etc.

[0507] The generation module 2620 is also used to generate content description information for each video segment based on the computer vision information of each video segment.

[0508] In some embodiments, content description information for each video segment is generated based on the extracted computer vision information for each video segment.

[0509] In some embodiments, based on a multimodal visual recognition model, content description information of each video segment in at least one video segment is obtained.

[0510] The generation module 2620 is also used to generate a second feature vector for each video segment based on the content description information of each video segment.

[0511] In some embodiments, a second feature vector for each video segment is generated based on the content description information of each video segment and a vector database. The vector database includes the correspondence between the content description information and the feature vectors.

[0512] In some embodiments, a second feature vector for each video segment is generated based on a multimodal embedding algorithm. Optionally, based on a multimodal embedding algorithm, at least two of the audio information, image information, and text information corresponding to each video segment are input to generate a second feature vector corresponding to each video segment.

[0513] The generation module 2620 is also used to extract at least one episode frame from the video segments.

[0514] In some embodiments, the video is divided into segments, with each segment consisting of a video frame. Each video segment is divided into at least one video frame. Scenes from each video frame are then extracted. It is important to understand that there is a one-to-one correspondence between video frames and scenes from each video frame.

[0515] In some embodiments, at least one target video frame is extracted from the video slice, and the quality of each target video frame is greater than or equal to a quality threshold. For example, suppose video slice 1 includes video frame 1, video frame 2, and video frame 3. Where the quality of video frame 1 and video frame 2 is greater than the quality threshold, and the quality of video frame 3 is less than the quality threshold, then only video frame 1 and video frame 2 are extracted from video slice 1, while video frame 3 is ignored.

[0516] The generation module 2620 is also used to store the correspondence between at least one episode frame of the video segment and the second feature vector.

[0517] In some embodiments, given that the second feature vector corresponding to each video segment and at least one episode frame corresponding to each video segment are known, there is a correspondence between at least one episode frame and the second feature vector.

[0518] In some embodiments, there is a one-to-one correspondence between episode scenes and the second feature vector.

[0519] In some embodiments, the correspondence between episode scenes and the second feature vector is many-to-one.

[0520] In some embodiments, the relationship between episode scenes and the second feature vector is one-to-many.

[0521] In some embodiments, when extracting at least one episode frame corresponding to a video segment, the quality is determined based on the quality of the video frames in each video segment.

[0522] The generation module 2620 is also used to score each video frame in the video segment according to at least one scoring factor to obtain a quality score corresponding to each video frame.

[0523] In some embodiments, each video segment includes multiple video frames, and each of the multiple video frames corresponds to a quality score.

[0524] In some embodiments, at least one scoring factor includes at least one of brightness, sharpness, and confidence in the identification of the subject in the video frame. Brightness refers to the light intensity in the video frame. Sharpness refers to the clarity of the image in the video frame. Confidence in the identification of the subject in the video frame refers to the probability that the subject in the video frame is identified.

[0525] The generation module 2620 is also used to extract at least one episode frame from each video frame in the video segment based on quality scoring.

[0526] In some embodiments, at least one episode frame is extracted from the target video frame in the video segment based on the quality score of each video frame. The quality score corresponding to the target video frame is greater than or equal to a quality threshold score.

[0527] In some embodiments, the video frames in the video segment are sorted from high to low or from low to high according to their quality scores, and at least one video frame with a quality score greater than or equal to a quality threshold score is selected. At least one episode scene is extracted from the at least one video frame with a quality score greater than or equal to the quality threshold score.

[0528] In some embodiments, the target multimedia content is generated based on the corresponding episode scenes and episode lines that match the target text.

[0529] The generation module 2620 is also used to calculate the picture quality score of at least two candidate episode scenes when there are at least two candidate episode scenes in the target text matching.

[0530] In some embodiments, each candidate episode frame is scored according to at least one scoring factor to obtain a picture quality score for each candidate episode frame. Optionally, the at least one scoring factor includes at least one of brightness, sharpness, and confidence in the identification of the subject in the picture.

[0531] The generation module 2620 is also used to select the target episode from at least two candidate episodes based on the episode quality score.

[0532] In some embodiments, based on the image quality score of each candidate episode image, at least one target episode image is selected from at least two candidate episode images that match the target text. The image quality score corresponding to the target episode image is greater than or equal to an image quality score threshold.

[0533] The generation module 2620 is also used to generate target multimedia content based on the target TV series images and TV series dialogue that matches the target text.

[0534] In some embodiments, the target episode image includes at least one, and the episode dialogue matching the target text may also include at least one. However, it is important to understand that each episode image corresponds to its own episode dialogue; that is, there is a one-to-one correspondence between episode images and episode dialogue. For example, suppose the target text matches the first episode image of a first series and the second episode image of a second series. Here, the first and second series are two different series. The first episode image corresponds to episode dialogue 1, and the second episode image corresponds to episode dialogue 2.

[0535] The one-to-one correspondence between scene and dialogue in the TV series mentioned above only applies when the scene corresponds to a specific line of dialogue. If no matching dialogue is found for the target text, the target text or a portion thereof will be used to generate matching dialogue.

[0536] In some embodiments, if there is a candidate episode frame that matches the target text, the candidate episode frame is determined as the target episode frame. That is, if there is one and only one matching candidate episode frame for the target text, that candidate episode frame is the target episode frame used to generate the target multimedia content.

[0537] In some embodiments, the target multimedia content is generated based on cropped episode footage.

[0538] The generation module 2620 is also used to identify the bounding box of the main subject in the scene of the series.

[0539] In some embodiments, the main subjects in a scene are assigned a priority. Based on the priority of the main subjects in the scene, the main subjects in the scene are identified, and a bounding box is used to select the corresponding main subject.

[0540] In some embodiments, the main subject of the image includes at least one of a face, a human body, an animal, a plant, furniture, clutter, and scenery. Optionally, the priority of a face is higher than that of a human body, the priority of a human body is higher than that of an animal, the priority of an animal is higher than that of a plant, the priority of a plant is higher than that of furniture, the priority of furniture is higher than that of clutter, and the priority of clutter is higher than that of scenery.

[0541] In some embodiments, when there are multiple subjects of the same priority in the scene, the bounding boxes of the subjects of the same priority in the scene are identified simultaneously. For example, assuming that there are multiple faces in the same scene, multiple faces are identified simultaneously, and multiple faces are selected by bounding boxes.

[0542] The generation module 2620 is also used to crop the TV series screen based on the bounding box of the main subject of the screen to obtain the cropped TV series screen.

[0543] In some embodiments, the episode frame is cropped based on the bounding box of the identified main subject to obtain the cropped episode frame. Typically, the cropped episode frame is smaller than or equal to the original episode frame.

[0544] The generation module 2620 is also used to merge at least two bounding boxes to obtain a merged bounding box when at least two bounding boxes are identified based on the main body of the image.

[0545] In some embodiments, if at least two bounding boxes are identified based on the main subject in the scene, the identified at least two bounding boxes are merged to obtain a merged bounding box. Optionally, the area of ​​the merged bounding box is determined based on the merged at least two bounding boxes.

[0546] In some embodiments, the merged bounding box covers all or part of the bounding boxes in at least two bounding boxes.

[0547] In some embodiments, when bounding boxes of at least two different subjects are identified and the total number of bounding boxes is greater than a first threshold, at least one group of bounding boxes to be merged is selected based on the priorities corresponding to the at least two different subjects. The number of bounding boxes in a group does not exceed a second threshold, and the priority of the bounding boxes in the group is higher than a third threshold.

[0548] By limiting the number of bounding boxes in the bounding box group, the number of main subjects in the final episode frame cropped based on the bounding boxes is also limited to a certain threshold. This helps to ensure that the number of main subjects in the cropped episode frame is reasonable and avoids a cluttered overall picture.

[0549] In some embodiments, for each bounding box group in at least one group of bounding boxes to be merged, the maximum top-left coordinate, maximum top-right coordinate, minimum bottom-left coordinate, and minimum bottom-right coordinate in the bounding box group are determined. Based on the maximum top-left coordinate, maximum top-right coordinate, minimum bottom-left coordinate, and minimum bottom-right coordinate, a merged bounding box for each bounding box group is generated.

[0550] In some embodiments, each bounding box corresponds to top-left, top-right, bottom-left, and bottom-right coordinates. For example, each bounding box can be represented by (x... start y start ) and (x end y end The combination is determined by (x). start y start (x) represents the lower left coordinate, (x) start y end (x) represents the top-left coordinate, (x) end y start(x) represents the lower right coordinate, (x) end y end () represents the upper right coordinate.

[0551] In some embodiments, when a group of bounding boxes includes multiple bounding boxes, the merging of bounding boxes can be achieved by min(x1). start x2 start , ..., xn start ), min(y1) start y2 start , ...,yn start ), max(x1) end x2 end , ..., xn end ), max(y1) end y2 end , ...,yn end )Sure.

[0552] The generation module 2620 is also used to crop the TV series screen based on the fused bounding box to obtain the cropped TV series screen.

[0553] In some embodiments, the episode frame is cropped based on a defined fusion bounding box to obtain a cropped episode frame. Typically, the cropped episode frame is smaller than or equal to the original episode frame.

[0554] Figure 27 illustrates a schematic diagram of a computer device provided in an exemplary embodiment of this application. Indicatively, the computer device 1100 includes a Central Processing Unit (CPU) 1101, a system memory 1104 including Random Access Memory (RAM) 1102 and Read-Only Memory (ROM) 1103, and a system bus 1105 connecting the system memory 1104 and the CPU 1101. The computer device 1100 also includes a basic input / output system 1106 that facilitates information transfer between various devices within the computer, and a mass storage device 1107 for storing an operating system 1113, application programs 1114, and other program modules 1115.

[0555] The basic input / output system 1106 includes a display 1108 for displaying information and an input device 1109 for user input, such as a mouse or keyboard. Both the display 1108 and the input device 1109 are connected to the central processing unit 1101 via an input / output controller 1110 connected to the system bus 1105. The basic input / output system 1106 may also include the input / output controller 1110 for receiving and processing input from multiple other devices such as a keyboard, mouse, or electronic stylus. Similarly, the input / output controller 1110 also provides output to a display screen, printer, or other types of output devices.

[0556] Mass storage device 1107 is connected to central processing unit 1101 via a mass storage controller (not shown) connected to system bus 1105. Mass storage device 1107 and its associated computer-readable media provide non-volatile storage for computer device 1100. That is, mass storage device 1107 may include computer-readable media (not shown) such as hard disk or compact disc read-only memory (CD-ROM) drive.

[0557] Computer-readable media can include computer storage media and communication media. Computer storage media include volatile and non-volatile, removable and non-removable media implemented using any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media include RAM, ROM, erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other solid-state storage technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape cassettes, magnetic tape, disk storage, or other magnetic storage devices. Of course, those skilled in the art will understand that computer storage media are not limited to the above-mentioned types. The system memory 1104 and mass storage device 1107 described above can be collectively referred to as memory.

[0558] According to various embodiments of this application, the computer device 1100 can also be connected to a remote computer on a network, such as the Internet. That is, the computer device 1100 can be connected to the network 1112 via the network interface unit 1111 connected to the system bus 1105, or the network interface unit 1111 can be used to connect to other types of networks or remote computer systems (not shown).

[0559] An exemplary embodiment of this application also provides a computer-readable storage medium storing at least one program, which is loaded and executed by a processor to implement the media content generation method provided in the above-described method embodiments.

[0560] An exemplary embodiment of this application also provides a computer program product, which includes at least one program segment stored in a readable storage medium; the processor of the communication device reads signaling from the readable storage medium, and the processor executes the signaling to cause the communication device to perform the media content generation method provided in the above-described method embodiments.

[0561] It should be understood that "a plurality of" as used herein refers to two or more. Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this application are indicated by the following claims.

[0562] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.

[0563] The above are merely optional embodiments of this application and are not intended to limit this application. Any modifications, equivalent switching, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A method for generating media content, characterized in that, The method is executed by a terminal device, and the method includes: Display text media content, which includes several texts; In response to a selection operation of at least one character among the plurality of texts, a target text for generating target multimedia content is determined, the target text including key information of the selected at least one character; Display the target multimedia content generated based on the target text.

2. The method according to claim 1, characterized in that, The step of determining the target text for generating the target multimedia content in response to a selection operation on at least one character among the plurality of texts includes: In response to a selection operation on at least one character among the plurality of texts, the selected text and a selection function bar are displayed, the selection function bar including a generation control; In response to a triggering operation on the generated control, the selected text is determined as the target text.

3. The method according to claim 2, characterized in that, The plurality of texts includes a first text that is highlighted, the highlighted state being used to indicate that the first text supports multimedia content generation functionality; The step of responding to a selection operation on at least one character among the plurality of texts, displaying the selected text and a selection function bar, includes: In response to the selection operation of the first text, the first text is displayed as the selected text and the selection function bar is displayed.

4. The method according to any one of claims 1 to 3, characterized in that, The target media content includes episode scenes and dialogue that match the target text.

5. The method according to claim 4, characterized in that, The target multimedia content also includes promotional content provided by the promoter.

6. A method for generating media content, characterized in that, The method is executed by the server, and the method includes: The client sends target text for generating target multimedia content, wherein the target text is selected from text media content published by the media account. Generate the target multimedia content based on the target text; The target multimedia content is sent to the client.

7. The method according to claim 6, characterized in that, The target media content includes episode scenes and dialogue that match the target text; The process of generating the target multimedia content based on the target text includes: Based on the target text in a preset correspondence, query for matching episode scenes and dialogue from the TV series; the episode scenes and dialogue are then used to generate the target multimedia content; or... Calculate the first feature vector corresponding to the target text, and calculate the vector similarity between the first feature vector and at least one second feature vector; The episode scenes and dialogues corresponding to the second feature vector with the highest vector similarity are determined as the episode scenes and dialogues that match the target text, and the episode scenes and dialogues are used to generate the target multimedia content.

8. The method according to claim 7, characterized in that, Before querying the TV series scenes and lines that match the target text in a preset correspondence based on the target text, the process further includes: Get the first episode's video and subtitles; According to the start and end times of the subtitles in the TV series, the first TV series video is divided into at least one video segment, and the video segment corresponds one-to-one with the subtitles in the TV series. Extract at least one episode frame from the video segment; The preset correspondence is generated based on at least one episode scene and the episode dialogue subtitles.

9. The method according to claim 7, characterized in that, Before calculating the vector similarity between the first feature vector and at least one second feature vector, the method further includes: Get the second episode video; The second series video is divided into at least one video segment according to the division step size; Extract computer vision information from each video segment in the at least one video segment; Based on the computer vision information of each video segment, generate content description information for each video segment; Based on the content description information of each video segment, a second feature vector is generated for each video segment.

10. The method according to claim 9, characterized in that, The method further includes: Extract at least one episode frame from the video segment; Store the correspondence between at least one episode frame of the video segment and the second feature vector.

11. The method according to claim 8 or 10, characterized in that, Extracting at least one episode frame from the video segment includes: Each video frame in the video segment is scored according to at least one scoring factor to obtain a quality score for each video frame. Based on the quality score, at least one episode scene is extracted from each video frame in the video segment; The at least one scoring factor includes at least one of brightness, sharpness, and confidence in the recognition of the subject in the image.

12. The method according to any one of claims 6 to 11, characterized in that, The process of generating the target multimedia content based on the target text includes: If the target text match yields at least two candidate episode scenes, calculate the image quality score for the at least two candidate episode scenes. Based on the picture quality score, the target episode picture is selected from the at least two candidate episode pictures; The target multimedia content is generated based on the target TV series footage and the TV series dialogue that matches the target text.

13. The method according to any one of claims 7 to 12, characterized in that, Before generating the target multimedia content based on the target text, the method further includes: Identify the bounding box of the main subject in the scene of the TV series; Based on the bounding box of the main subject of the image, the screen of the TV series is cropped to obtain the cropped screen of the TV series.

14. The method according to claim 13, characterized in that, The process of cropping the TV series image based on the bounding box of the main subject of the image to obtain the cropped TV series image includes: When at least two bounding boxes are identified based on the main body of the image, the at least two bounding boxes are merged to obtain a merged bounding box, which covers all or part of the bounding boxes in the at least two bounding boxes. The episode footage is cropped based on the fused bounding box to obtain the cropped episode footage.

15. The method according to claim 14, characterized in that, The main subject of the image is assigned a priority. The step of fusing at least two bounding boxes to obtain a fused bounding box when at least two bounding boxes are identified based on the main subject of the image includes: If at least two different subject bounding boxes are identified and the total number of bounding boxes is greater than a first threshold, at least one bounding box group to be merged is selected based on the priority corresponding to the at least two different subject bounding boxes; the number of bounding boxes in the bounding box group does not exceed a second threshold and the priority corresponding to the bounding boxes in the bounding box group is higher than a third threshold. For each of the at least one bounding box groups to be merged, determine the maximum top-left coordinate, maximum top-right coordinate, minimum bottom-left coordinate, and minimum bottom-right coordinate in the bounding box group; Based on the maximum top-left coordinate, maximum top-right coordinate, minimum bottom-left coordinate, and minimum bottom-right coordinate, a merged bounding box is generated for each bounding box group.

16. A media content generation device, characterized in that, The device includes: The display module is used to display text media content published by a media account, wherein the text media content includes several texts; A determination module is configured to determine the target text for generating the target multimedia content in response to a selection operation on at least one character among the plurality of texts. The display module is also used to display the target multimedia content generated based on the target text.

17. A media content generation device, characterized in that, The device includes: The receiving module is used to receive target text sent by the client for generating target multimedia content, wherein the target text is selected from the text media content published by the media account. The generation module is used to generate the target multimedia content based on the target text; The sending module is used to send the target multimedia content to the client.

18. A computer device, characterized in that, The computer device includes a processor and a memory, wherein the memory stores at least one program, which is loaded and executed by the processor to implement the media content generation method as described in any one of claims 1 to 15.

19. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one program, which is loaded and executed by a processor to implement the media content generation method as described in any one of claims 1 to 15.

20. A computer program product, characterized in that, The computer program product includes at least one program segment stored in a computer-readable storage medium; a processor of a computer device reads the at least one program segment from the computer-readable storage medium, and the processor executes the at least one program segment, causing the computer device to perform the media content generation method as described in any one of claims 1 to 15.

Citation Information

Patent Citations

  • Multimedia data processing method and device, electronic equipment and storage medium

    CN109756751A

  • Content processing method and device

    CN110704647A

  • Video display and processing method, device and system, equipment and medium

    CN112579826A