Video generation method and device, computer equipment and storage medium
By displaying text and image works in the video generation interface and generating associated video works, the problem of time-consuming and laborious video production by users is solved, achieving efficient one-click conversion while maintaining the consistency between video works and original text and image works.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-14
- Publication Date
- 2026-04-14
AI Technical Summary
In existing technologies, user-generated videos are time-consuming, labor-intensive, and inefficient.
This paper provides a video generation method that generates video works associated with published text and image works by displaying the text and image works, including the original images and images matching the text, and realizes one-click conversion using a video generation interface.
Users are not required to manually create videos, saving manpower and time, improving video generation efficiency, and ensuring the consistency between the converted video works and the original text and image works.
Smart Images

Figure CN121865049A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a video generation method, apparatus, computer device, and storage medium. Background Technology
[0002] With the development of internet technology, users can create their own videos on various platforms and publish or share them.
[0003] In related technologies, users typically select suitable video footage and edit it themselves to create a video. However, manual video production is inefficient due to the significant manpower and time required. Summary of the Invention
[0004] This application provides a video generation method, apparatus, computer device, and storage medium, which can improve the efficiency of video generation. The technical solution is as follows:
[0005] On the one hand, a video generation method is provided, the method comprising:
[0006] Based on the video generation interface, published graphic and text works are displayed. The video generation interface includes video generation options, and the graphic and text works include text and a first image.
[0007] In response to the triggering operation of the video generation option, a video work associated with the text and image work is generated based on the text and image work in the video generation interface;
[0008] The video work is displayed on the video generation interface. The video work includes the first image and a second image that matches the text.
[0009] On the other hand, a video generation apparatus is provided, the apparatus comprising:
[0010] The display module is used to display published graphic works based on a video generation interface, wherein the video generation interface includes video generation options and the graphic works include text and a first image;
[0011] The generation module is used to generate a video work associated with the text and image works based on the text and image works in the video generation interface in response to the triggering operation of the video generation option.
[0012] The display module is also used to display the video work on the video generation interface, the video work including the first image and a second image matching the text.
[0013] Optionally, the display module is used for:
[0014] The video generation interface displays a link input area and link parsing options;
[0015] Retrieve the image and text link entered in the link input area, wherein the image and text link is a link to any published image and text work;
[0016] In response to the triggering operation of the link parsing option, the image and text works associated with the image and text link are obtained and displayed on the video generation interface.
[0017] Optionally, the display module is used for:
[0018] The display interface for the published graphic works is shown, and the display interface for the graphic works includes an entry point for converting graphic works to video.
[0019] In response to the triggering operation of the text-to-video entry point, the video generation interface is displayed, and the text-to-image work is displayed on the video generation interface.
[0020] Optionally, the text in the graphic work includes multiple text segments, and the graphic work includes at least one first image interspersed among the multiple text segments; the device further includes:
[0021] The modification module is used to respond to the modification operation of the graphic work in the video generation interface, modify the graphic work, and display the modified graphic work in the video generation interface.
[0022] The modification operations on the graphic works include at least one of the following: adjusting the position of any text segment, editing any text segment, deleting any text segment, adjusting the position of any first image, deleting any first image, adding text segments, or adding images.
[0023] Optionally, each of the first images is associated with one of the text segments; the display module is configured to display a plurality of content areas on the video generation interface, each content area including one of the text segments, and if the text segment is associated with a first image, the content area also includes at least one first image associated with the text segment;
[0024] The modification module is used to determine at least one selected target content area among the plurality of content areas;
[0025] The modification module is also configured to delete the text fragment and the first image in the at least one target content area in response to the deletion operation.
[0026] Optionally, the device further includes:
[0027] A replacement module is configured to, in response to a replacement request for a second image in the video work, obtain a third image that matches the text and is different from the second image, and replace the second image in the video work with the third image.
[0028] Optionally, the device further includes:
[0029] The publishing module is used to respond to the publishing operation of the video work, publish the video work carrying the image and text link, and add the video link to the published image and text work;
[0030] The text and image links are links to the display interfaces of the published text and image works, and the video links are links to the display interfaces of the published video works.
[0031] Optionally, the display module is used for:
[0032] The video generation interface of the first client displays the graphic and textual works published on the second client, and the first client and the second client are different.
[0033] Optionally, the display module is used for:
[0034] The graphic and textual work is obtained, wherein the text in the graphic and textual work includes multiple text fragments, and the graphic and textual work includes multiple first images;
[0035] The graphic and textual works are identified to obtain the type of each text fragment and the type of each first image in the graphic and textual works;
[0036] The text fragments belonging to the target type and the first image belonging to the target type are removed from the graphic work to obtain the preprocessed graphic work;
[0037] The pre-processed graphic artwork is displayed on the video generation interface.
[0038] Optionally, the generation module is used for:
[0039] In response to a triggering operation of the video generation option, a second image matching the text is acquired;
[0040] The first image and the second image are merged to obtain a video frame;
[0041] Based on the video footage, a video work associated with the text and image work is generated.
[0042] Optionally, the generation module is used for:
[0043] Using the text-to-text model, descriptive text is generated based on the text in the graphic work. The semantics of the descriptive text are the same as the semantics of the text in the graphic work. The descriptive text is used to describe the image content.
[0044] The second image is generated based on the descriptive text using the text-based image model, and the image content of the second image is consistent with the image content described by the descriptive text.
[0045] Optionally, the generation module is used for:
[0046] Extract keywords from the text;
[0047] Search for images that match the keywords, and identify the searched images as the second images that match the text.
[0048] Optionally, the generation module is used for:
[0049] The video generation interface displays an image upload entry and a prompt message, which prompts users to upload an image that matches the text.
[0050] The image uploaded in the image upload portal is retrieved, and the uploaded image is identified as the second image that matches the text.
[0051] Optionally, the generation module is further configured to:
[0052] The text is divided into segments according to a preset number of characters to obtain multiple text fragments, and the number of characters in each text fragment does not exceed the preset number of characters;
[0053] For any text fragment, obtain a second image that matches the text fragment;
[0054] The acquired second images are fused with the first image to obtain the video frame.
[0055] Optionally, the generation module is used for:
[0056] According to the arrangement order of the text fragments corresponding to the multiple second images in the graphic work, the multiple second images are merged to obtain an initial video frame. The display order of the multiple second images in the initial video frame is consistent with the arrangement order of the text fragments corresponding to the multiple second images in the graphic work.
[0057] The first image is added to the initial video frame to obtain the video frame.
[0058] Optionally, the first image is interspersed among the plurality of text fragments; the generation module is configured to:
[0059] Among the plurality of text fragments, a target text fragment that is adjacent to and located in front of the first image is identified;
[0060] In the initial video frame, the target video segment containing the second image that matches the target text fragment is determined;
[0061] The first image is added to the target video segment of the initial video frame to obtain the video frame.
[0062] Optionally, the generation module is configured to perform any of the following:
[0063] The first image is overlaid on the second image to obtain the video frame, wherein the layer containing the first image is located above the layer containing the second image;
[0064] The second image is overlaid on the first image to obtain the video frame, wherein the layer containing the second image is located above the layer containing the first image.
[0065] The first image and the second image are stitched together to obtain the video frame.
[0066] Optionally, the generation module is further configured to:
[0067] Based on the text in the graphic works, generate a video dubbing that matches the text;
[0068] The video work is generated based on the video footage and the video voiceover.
[0069] Optionally, the generation module is further configured to:
[0070] The text is divided into multiple sentences, and the audio segment corresponding to each sentence is determined in the video dubbing. The sentences are then identified as subtitles for the audio segments corresponding to the sentences.
[0071] The video work is generated based on the video footage, the video dubbing, and the subtitles for each audio segment in the video dubbing.
[0072] Optionally, the graphic work includes multiple first images; the generation module is used to:
[0073] In response to a triggering operation of the video generation option, a target image is determined among a plurality of first images, the target image being a first image containing a graphic code;
[0074] Based on the text in the graphic work and the target image, the video work is generated, and the video work includes the target image and a second image that matches the text.
[0075] On the other hand, a computer device is provided, the computer device including a processor and a memory, the memory storing at least one computer program, the at least one computer program being loaded and executed by the processor to perform the operations performed by the video generation method as described above.
[0076] On the other hand, a computer-readable storage medium is provided that stores at least one computer program, which is loaded and executed by a processor to perform the operations performed by the video generation method as described above.
[0077] On the other hand, a computer program product is provided, including a computer program loaded and executed by a processor to perform the operations performed by the video generation method as described above.
[0078] The solution provided in this application, for published text and image works, can generate video works associated with those works based on the text and images within them. This achieves one-click conversion of published text and image works into video works, eliminating the need for users to manually create video works, saving manpower and time, and improving the efficiency of video generation. Furthermore, the video works converted from text and image works include two types of images: one image is the original image from the text and image work, and the other image is an image that matches the text in the text and image work. Therefore, it ensures that the original images in the text and image work are not lost, and it also reflects the content of the text in the text and image work, improving the fit between the one-click converted video works and the text and image works, and improving the quality of the video works. Attached Figure Description
[0079] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0080] Figure 1 This is a schematic diagram of an implementation environment provided in an embodiment of this application;
[0081] Figure 2 This is a flowchart of a video generation method provided in an embodiment of this application;
[0082] Figure 3 This is a flowchart of another video generation method provided in the embodiments of this application;
[0083] Figure 4This is a schematic diagram of a video generation interface provided in an embodiment of this application;
[0084] Figure 5 This is a schematic diagram of another video generation interface provided in an embodiment of this application;
[0085] Figure 6 This is a schematic diagram of another video generation interface provided in an embodiment of this application;
[0086] Figure 7 This is a schematic diagram of another video generation interface provided in an embodiment of this application;
[0087] Figure 8 This is a schematic diagram of another video generation interface provided in an embodiment of this application;
[0088] Figure 9 This is a schematic diagram of another video generation interface provided in an embodiment of this application;
[0089] Figure 10 This is a schematic diagram of another video generation interface provided in an embodiment of this application;
[0090] Figure 11 This is a schematic diagram of another video generation interface provided in an embodiment of this application;
[0091] Figure 12 This is a schematic diagram of another video generation interface provided in an embodiment of this application;
[0092] Figure 13 This is a flowchart of another video generation method provided in the embodiments of this application;
[0093] Figure 14 This is a flowchart of another video generation method provided in the embodiments of this application;
[0094] Figure 15 This is a flowchart of another video generation method provided in the embodiments of this application;
[0095] Figure 16 This is a flowchart of another video generation method provided in the embodiments of this application;
[0096] Figure 17 This is a flowchart of another video generation method provided in the embodiments of this application;
[0097] Figure 18 This is a schematic diagram of the structure of a video generation device provided in an embodiment of this application;
[0098] Figure 19 This is a schematic diagram of another video generation device provided in an embodiment of this application;
[0099] Figure 20 This is a schematic diagram of the structure of a terminal provided in an embodiment of this application;
[0100] Figure 21 This is a schematic diagram of the structure of a server provided in an embodiment of this application. Detailed Implementation
[0101] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.
[0102] It is understood that the terms "first," "second," etc., used in this application may be used herein to describe various concepts, but unless otherwise stated, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of this application, a first image may be referred to as a second image, and similarly, a second image may be referred to as a first image.
[0103] "At least one" refers to one or more images. For example, at least one image can be one image, two images, three images, or any integer number of images greater than or equal to one. "Multiple" refers to two or more images. For example, multiple images can be two images, three images, or any integer number of images greater than or equal to two. "Each" refers to each of the at least one images. For example, each image refers to each of the multiple images. If the multiple images consist of three images, then each image refers to each of the three images.
[0104] It should be noted that the information (including but not limited to user equipment information, user personal information, etc.), data (including but not limited to data used for analysis, data stored, data displayed, etc.) and signals (including but not limited to signals transmitted between user terminals and other devices) involved in this application have all been fully authorized by the user or relevant parties, and the collection, use and processing of the relevant data shall comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0105] For example, the graphic works, texts, images, and video works involved in this application are all fully authorized by the users or relevant parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions.
[0106] The video generation method provided in this application can be used in computer devices. Optionally, the computer device is a terminal or a server. Optionally, the server is an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. Optionally, the terminal is a smartphone, tablet computer, laptop computer, desktop computer, smart speaker, smartwatch, smart voice interaction device, smart home appliance, vehicle terminal, aircraft, etc., but is not limited to these. This application can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, smart transportation, and assisted driving.
[0107] Figure 1 This is a schematic diagram of an implementation environment provided in an embodiment of this application. See also... Figure 1 The implementation environment includes: terminal 101 and server 102. Terminal 101 and server 102 are connected via a wireless or wired network.
[0108] Terminal 101 has at least one client installed and running. The client can be a social application client, online payment client, online shopping client, game client, medical service client, video client, etc. When terminal 101 runs a client, the client's user interface is displayed on the screen of terminal 101. Terminal 101 is the terminal used by user 121.
[0109] Optionally, terminal 101 may refer to one of a number of terminals, including: smartphones, tablets, laptops, desktop computers, smart voice interaction devices, smart home appliances, in-vehicle terminals, aircraft, VR (Virtual Reality) devices, AR (Augmented Reality) devices, etc., but not limited to these.
[0110] Those skilled in the art will understand that the number of terminals described above can be more or less. For example, there may be only one terminal, or there may be six, eight, or more terminals. This application does not limit the number of terminals or the type of device.
[0111] Figure 1Only one terminal is shown in the diagram, but in different embodiments, multiple other terminals 103 can access the server 102. Optionally, one or more terminals 103 may also be terminals corresponding to developers, on which a client development and editing platform is installed. Developers can edit and update the client on the terminal 103 and transmit the updated client installation package to the server 102 via wired or wireless network. Terminal 101 can download the client installation package from the server 102 to update the client.
[0112] Terminal 101 and other terminals 103 are connected to server 102 via wired or wireless networks.
[0113] Server 102 includes at least one of a single server, multiple servers, a cloud computing platform, and a virtualization center. Server 102 is used to provide backend services to clients. Optionally, server 102 undertakes the primary computing work, and terminal 101 undertakes the secondary computing work; or, server 102 undertakes the secondary computing work, and terminal 101 undertakes the primary computing work; or, server 102 and terminal 101 collaborate on computing using a distributed computing architecture.
[0114] In this embodiment, in response to a video generation request for a published graphic work, terminal 101 sends a video generation task for the graphic work to server 102. After receiving the video generation task, server 102 generates a video work associated with the graphic work based on the graphic work, and returns the video work to terminal 101. Terminal 101 receives and displays the video work.
[0115] It should be noted that the above implementation environment is only an example. The method provided in this application embodiment can also be executed by the terminal 101 alone, or by other computer devices. This application embodiment does not limit this.
[0116] Figure 2 This is a flowchart illustrating a video generation method provided in an embodiment of this application. This embodiment is executed by a computer device, such as a terminal, or jointly by a terminal and a server. See also... Figure 2 The method includes:
[0117] 201. The terminal displays published graphic works based on the video generation interface. The video generation interface includes video generation options, and the graphic works include text and a first image.
[0118] The terminal displays a video generation interface, which is used to convert text and image works into video works. The video generation interface includes video generation options, which are used to trigger the conversion of text and image works into video works.
[0119] In this embodiment of the application, the video generation interface displays published graphic works, which include text and a first image. These graphic works can be published on any platform, the text length can be arbitrary, and the number of first images in the graphic works can be one or more.
[0120] 202. In response to the triggering operation of the video generation option, the terminal generates a video work associated with the text and image works in the video generation interface.
[0121] If a user wants to convert a text and image work into a video work, they can trigger the video generation option in the video generation interface. In response to this triggering operation, the terminal will convert the text and image work into a video work associated with the text and image work.
[0122] 203. The terminal displays the video work on the video generation interface. The video work includes a first image and a second image that matches the text.
[0123] The first image is an original image in the graphic work, while the second image is not an image in the graphic work. Matching the second image with the text means that the content of the second image matches the content of the text.
[0124] The method provided in this application, for published text and image works, can generate video works associated with the published text and image works based on the text and images in the published text and image works. This achieves one-click conversion of published text and image works into video works, eliminating the need for users to manually create video works, saving manpower and time, and improving the efficiency of video generation. Furthermore, the video works converted from text and image works include two types of images: one image is the original image from the text and image work, and the other image is an image that matches the text in the text and image work. Therefore, it ensures that the original images in the text and image work are not lost, and it can reflect the content of the text and image work, improving the fit between the one-click converted video works and the text and image works, and improving the quality of the video works.
[0125] Figure 3 This is a flowchart of another video generation method provided in this application embodiment. This application embodiment is executed by a computer device. See also... Figure 3 The method includes:
[0126] 301. The computer device displays published graphic works based on a video generation interface. The video generation interface includes video generation options, and the graphic works include text and a first image.
[0127] In this embodiment of the application, the published graphic and textual works can be any type of graphic and textual works published on any platform. Graphic and textual works refer to works that include both images and text. For example, the graphic and textual works can be articles from public accounts, news information, graphic novels, etc.
[0128] In one possible implementation, a computer device displays a graphic and textual work published on a second client on the video generation interface of a first client, and the first client and the second client are different.
[0129] The video generation interface is the interface of the first client, and the computer device displays the video generation interface on the first client. Published text and image works are those published on the second client; that is, text and image works published on other clients can be uploaded to the video generation interface of the first client. For example, the first client is a video editing client, and the second client is a social media client, etc.
[0130] It should be noted that the embodiments of this application are only illustrated by the example that the first client and the second client are different clients. In another embodiment, the first client and the second client may also be the same client.
[0131] In this embodiment of the application, the published text and image works in the second client can be converted into video works in the video generation interface of the first client, realizing cross-platform conversion of text and image works into video works, and improving the convenience and versatility of the video generation method.
[0132] In one possible implementation, the computer device displays a link input area and a link resolution option on the video generation interface; retrieves the image / text link entered in the link input area, where the image / text link is a link to any published image / text work; and in response to a triggering operation on the link resolution option, retrieves the image / text work associated with the image / text link and displays the image / text work on the video generation interface.
[0133] The user enters a link in the link input area, and the computer device displays the entered image and text link in the link input area based on the input. After completing the input of the image and text link, the user triggers an action to select the link parsing option, and the computer device responds to this action by retrieving the image and text artwork associated with the link.
[0134] Among them, the image and text link is the address that points to the image and text work or the interface where the image and text work is located. By parsing the image and text link, the image and text work indicated by the image and text link can be obtained.
[0135] the following Figures 4-6 Taking text and image works as WeChat official account articles as an example, this article demonstrates the process of displaying WeChat official account articles on the video generation interface.
[0136] Figure 4 This is a schematic diagram of a video generation interface provided in an embodiment of this application. Figure 4 The video generation interface shown is used to generate video works based on WeChat official account articles. For example... Figure 4 As shown, the video generation interface displays a link input area 401 and a link parsing option 402. The link input area 401 is used to input the article link of a WeChat official account article, and the link parsing option 402 is used to parse the article link to obtain the WeChat official account article.
[0137] Figure 5 This is a schematic diagram of another video generation interface provided in an embodiment of this application, such as... Figure 5 As shown, when a user triggers the link input area 401, a virtual keyboard appears on the video generation interface, allowing the user to enter an article link in the link input area 401. Alternatively, the user can paste the article link into the link input area 401.
[0138] Figure 6 This is a schematic diagram of another video generation interface provided in this application embodiment. When the user enters an article link in the link input area 401, the link parsing option 402 is triggered, so that the computer device parses the article link to obtain the WeChat official account article. Figure 6 As shown, after obtaining the WeChat official account article, the article is displayed on the video generation interface. The article includes text and images. The text consists of multiple text segments, with each paragraph being a text segment. The article also contains multiple images, which are interspersed among the text segments.
[0139] In this embodiment, by providing and parsing image and text links, published image and text works can be obtained, simplifying the method of obtaining image and text works and improving the efficiency of obtaining them. Furthermore, by parsing image and text works, works published on any platform can be obtained, without being limited to a specific platform, reducing the limitations of obtaining image and text works, and consequently reducing the limitations of converting image and text works into video works. Here, "platform" refers to web pages on a website, interfaces in an application client, etc.
[0140] In one possible implementation, the computer device displays a display interface for published text and graphic works, which includes a text-to-video conversion entry point; in response to a triggering operation on the text-to-video conversion entry point, a video generation interface is displayed, on which the text and graphic works are displayed.
[0141] The computer device displays published text and image works on the display interface. This display interface also includes a text-to-video conversion entry. When a user triggers the text-to-video conversion entry, the user is redirected from the text and image work display interface to a video generation interface that includes the text and image work.
[0142] Optionally, the computer device displays a presentation interface for the graphic work on the second client, which is a graphic work published on the second client. In response to a trigger operation on the graphic-to-video conversion entry, the computer device jumps from the presentation interface of the second client to the video generation interface of the first client. The first client and the second client are different clients.
[0143] In this embodiment of the application, by requesting the conversion of the graphic work into a video work on the display interface of the graphic work, the user can be directly redirected to the video generation interface, which has already uploaded the graphic work itself. The user does not need to perform the operation of uploading the graphic work to the video generation interface again, thus simplifying the user operation and improving the efficiency of human-computer interaction.
[0144] In one possible implementation, the video generation interface displays an input area, and the computer device displays the input graphic and text works in the input area of the video generation interface based on the input operation in the input area, that is, the user can input graphic and text works by himself.
[0145] the following Figure 7 Taking illustrated novels as an example, this paper demonstrates the process of displaying illustrated works on the video generation interface. Figure 7 This is a schematic diagram of another video generation interface provided in an embodiment of this application, such as... Figure 7 As shown, the video generation interface includes an input area 701, a settings area 702, and a video generation option 703. The input area 701 is used to input the text and image novel, the settings area 702 is used to set the illustrations, timbre, and speed of the dubbing for the generated video work, and the video generation option 703 is used to trigger the conversion of the text and image novel in the input area 701 into a video work.
[0146] like Figure 7 As shown, settings area 702 includes a switch option for the AI automatic image matching function. By triggering this switch, the AI automatic image matching function can be enabled or disabled. When AI automatic image matching is enabled, the style of the automatically matched image can be selected, such as choosing an ancient Chinese style. In addition, settings area 702 also allows setting the tone and speech rate. After completing the settings, the video generation option 703 is triggered to request the generation of a video.
[0147] In one possible implementation, a computer device acquires a graphic work, wherein the text in the graphic work includes multiple text fragments and the graphic work includes multiple first images; the graphic work is identified to obtain the type of each text fragment and the type of each first image in the graphic work; text fragments belonging to the target type and first images belonging to the target type are removed from the graphic work to obtain a preprocessed graphic work; and the preprocessed graphic work is displayed on a video generation interface.
[0148] In this graphic work, each text fragment and each first image corresponds to its own type. By identifying the graphic work, the type of each text fragment and the type of each first image can be obtained. Based on the type of the text fragment and the type of the first image, the graphic work can be cleaned by deleting text fragments and first images belonging to specific types.
[0149] For example, the types of text snippets and the first image include introductions, prefaces, body text, advertising snippets, and donation requests. The target type can be advertising snippets and donation requests, thus automatically removing advertising content and donation requests from the text and image work.
[0150] In this embodiment of the application, text fragments or images belonging to the target type are removed from the published graphic works. Therefore, by flexibly setting the target type, the target type content in the graphic works can be intelligently removed and then converted into video, eliminating the need for users to manually search for and delete specific types of content, simplifying user operations and improving operational efficiency.
[0151] 302. The computer device responds to the modification operation of the graphic work in the video generation interface, modifies the graphic work, and displays the modified graphic work in the video generation interface.
[0152] After the computer device displays the published graphic and textual work on the video generation interface, the user can also modify the published graphic and textual work. In this embodiment of the application, the text in the graphic and textual work includes multiple text segments, and the graphic and textual work includes at least one first image interspersed among the multiple text segments.
[0153] Therefore, the modification operations for graphic works include at least one of the following: adjusting the position of any text fragment, editing any text fragment, deleting any text fragment, adjusting the position of any first image, deleting any first image, adding a text fragment, or adding an image.
[0154] The operation of adjusting the position of a text segment refers to moving the position of the text segment up or down, and the operation of adjusting the position of the first image refers to moving the position of the first image up or down, etc.
[0155] In one possible implementation, each first image is associated with one of the text segments. The terminal displays multiple content areas on the video generation interface, each content area including a text segment. If a text segment is associated with a first image, the content area also includes the first image associated with the text segment. Then, step 302 includes: identifying at least one selected target content area among the multiple content areas; and deleting the text segment and the first image from the at least one target content area in response to a deletion operation.
[0156] A text segment may be associated with one or more first images, or it may not be associated with any of them. For example, first images may be interspersed among multiple text segments. For any text segment, if there is at least one first image between the text segment and the next text segment, then the text segment is associated with that at least one first image. That is, the text segment is associated with the first image following it. Each text segment is displayed in its own content area. If a text segment is not associated with a first image, then the content area containing that text segment only displays that text segment. If a text segment is associated with at least one first image, then the content area containing that text segment also displays the at least one first image associated with that text segment.
[0157] For any text segment, if the user wants to delete it, they select the content area containing that text segment. After selecting the content area containing at least one text segment to be deleted, the user executes the delete operation, which deletes these text segments and the first image associated with them with a single click. Optionally, the video generation interface displays a delete option; this delete operation refers to triggering the delete option.
[0158] In this embodiment, text fragments and at least one first image associated with the text fragments are displayed in their respective content areas. By performing a selection operation on the content area, all text fragments and first images in the selected content area can be deleted with one click, eliminating the need for the user to delete text fragments and first images one by one. This simplifies user operations and helps improve the efficiency of human-computer interaction.
[0159] Figure 8 This is a schematic diagram of another video generation interface provided in an embodiment of this application, such as... Figure 8 As shown, the video generation interface can display an editing area for any text segment in the graphic work, where users can edit the text segment.
[0160] Figure 9 This is a schematic diagram of another video generation interface provided in an embodiment of this application, such as... Figure 9As shown, the video generation interface displays multiple content areas. Each content area displays a text segment, or a text segment and at least one first image associated with that text segment. Users can select at least one content area in the video generation interface to select the content to be deleted, and then execute the trigger operation of the delete option 901 to batch delete all text segments and first images of the selected at least one content area.
[0161] Figure 10 This is a schematic diagram of another video generation interface provided in an embodiment of this application, such as... Figure 10 As shown, users can move the first image in a graphic work to the previous or next paragraph, etc.
[0162] In this embodiment, after the published text and image works are displayed on the video generation interface, they can be modified. Subsequently, associated video works can be generated based on the modified text and image works. By providing users with the space to modify published text and image works during the conversion to video works, users can make modifications based on existing works without needing to create a new text and image work, thus improving operational flexibility and convenience, and enhancing the user experience.
[0163] It should be noted that the embodiments of this application only illustrate the modification of published graphic works. In another embodiment, the published graphic works may not be modified, that is, step 302 may not be performed.
[0164] 303. In response to a triggering operation of the video generation option, the computer device generates a video work associated with the text and image works in the video generation interface.
[0165] If a user wants to convert a text and image work into a video work, they can trigger the video generation option in the video generation interface. In response to this triggering operation, the terminal will convert the text and image work into a video work associated with the text and image work.
[0166] The video work includes a first image and a second image that matches the text in the graphic work. The first image is an original image in the graphic work, while the second image is not an image in the graphic work. Matching the second image with the text means that the content of the second image is compatible with the content of the text.
[0167] In one possible implementation, the graphic work includes a plurality of first images. In response to a triggering operation of a video generation option, a computer device determines a target image among the plurality of first images. The target image is a first image containing a graphic code. A video work is generated based on the text in the graphic work and the target image. The video work includes the target image and a second image matching the text.
[0168] Alternatively, the graphic code can be a QR code, etc.
[0169] Optionally, the content indicated by the graphic code is related to the graphic work, such as supplementing the graphic work.
[0170] Optionally, the graphic code can be used for advertising, such as as an ad link, where recognizing the code redirects to the ad page. For example, the ad page might showcase the same item featured in the text or image work. Optionally, the graphic code can be used for website redirection, such as as a website link, where recognizing the code redirects to video websites, shopping websites, or social networking sites. Optionally, the graphic code can be used to display a personal profile, such as as a homepage link, where recognizing the code redirects to the account's homepage, which could be the account that published the text or image work, or the account featured in the text or image work. In addition, the graphic code can be applied to various scenarios such as information retrieval and account login.
[0171] In this embodiment of the application, the first image containing the graphic code carries more information. When requesting the generation of a video work, the first image containing the graphic code is intelligently identified. The generated video work only includes the first image containing the graphic code and does not include the first image that does not contain the graphic code. This makes it easier to draw the user's attention to the graphic code in the video work so as to obtain more information indicated by the graphic code.
[0172] 304. The computer device displays the video work on the video generation interface.
[0173] After a video work is generated that is associated with the text and image work, the video work is displayed on the video generation interface.
[0174] Figure 11 This is a schematic diagram of a video generation interface provided in an embodiment of this application, such as... Figure 11 As shown, the video generation interface displays a generated video work 1101, which includes a second image 1102, as well as first images 1103 and 1104. First images 1103 and 1104 are original images from the text-based work, while the second image 1102 is an additional image obtained to match the text in the text-based work. Optionally, as... Figure 11As shown, the size of the first image 1103 and the size of the first image 1104 are smaller than the size of the second image 1102, and the first image 1103 and the second image 1104 are superimposed on the second image 1102.
[0175] 305. In response to a request to replace a second image in a video work, a computer device obtains a third image that matches the text and is different from the second image, and replaces the second image in the video work with the third image.
[0176] Since the second image in the video work is matched with the text, rather than being an original image in the text and image work, if the user is not satisfied with the second image, after the video work is generated, the second image in the video work can be replaced with another third image that matches the text in the text and image work. This third image is different from the second image.
[0177] In one possible implementation, the third image can be an image generated using AI (Artificial Intelligence), an image selected by the user from a local photo album, an image selected by the user from a media library, or an image obtained by searching for keywords extracted from the text of a video work.
[0178] Figure 12 This is a schematic diagram of another video generation interface provided in an embodiment of this application, such as... Figure 12 As shown, after the video is generated, the computer displays it in a video editing interface, which includes volume settings 1201, aspect ratio settings 1202, and a progress bar 1203. The video generation interface also displays an album selection entry 1204, an AI generation entry 1205, and a media library selection area 1206. Users can drag the progress bar 1203 to control the playback progress of the currently displayed video frame. If a user wants to replace the second image in the currently displayed video frame, they can select the second image and then replace the third image using the album selection entry 1204, the AI generation entry 1205, or the media library selection area 1206.
[0179] In this embodiment of the application, after converting the text and image work into an associated video work, the second image in the video work can also be replaced, that is, the image matching the text can be replaced. This provides users with the operational space to make secondary modifications to the generated video work without having to regenerate a new video work, which improves the flexibility and convenience of the operation and helps to improve the user's operating experience.
[0180] It should be noted that the embodiments of this application only take replacing the second image in the video work with the third image as an example for illustration. In another embodiment, the second image in the video work may not be replaced, that is, step 305 may not be performed.
[0181] 306. In response to the publishing operation of a video work, the computer device publishes the video work with a text and image link, and adds the video link to the published text and image work.
[0182] Among them, the image and text links are links to the display interfaces of published image and text works, and the video links are links to the display interfaces of published video works.
[0183] In other words, by adding a text / image link when publishing a video, users can easily access the related text / image content by triggering the link when viewing the published video. Similarly, by adding a video link to an existing text / image content, users can also easily access the related video content by triggering the video link when viewing the text / image content.
[0184] In this embodiment of the application, after converting a text and image work into a video work, if the video work is published, a link to the display interface of the video work is added to the text and image work, and a link to the display interface of the text and image work is added to the video work. Therefore, users can jump to the display interface of the video work with one click from the display interface of the text and image work, or vice versa, which improves the convenience of viewing related text and image works or video works and helps to improve the efficiency of human-computer interaction.
[0185] It should be noted that this application embodiment only illustrates the example of publishing a video work. In another embodiment, after generating the video work, it is not necessary to publish the video work, that is, step 306 may not be performed. Even without publishing the video work, a storage operation on the video work can still be performed. In response to the storage operation, the computer device adds the image / text link to the video work and stores the video work carrying the image / text link.
[0186] Furthermore, when publishing a stored video work later, the published video work will still include the image and text link. Optionally, in this case, the computer device can determine the image and text work associated with the image and text link based on the image and text link, and add the video link corresponding to the video work to that image and text work.
[0187] It should be noted that the solutions in this application embodiment can convert not only text-based works into video works, but also plain text works into video works. For example, a plain text novel can be converted into a video work.
[0188] The method provided in this application, for published text and image works, can generate video works associated with the published text and image works based on the text and images in the published text and image works. This achieves one-click conversion of published text and image works into video works, eliminating the need for users to manually create video works, saving manpower and time, and improving the efficiency of video generation. Furthermore, the video works converted from text and image works include two types of images: one image is the original image from the text and image work, and the other image is an image that matches the text in the text and image work. Therefore, it ensures that the original images in the text and image work are not lost, and it can reflect the content of the text and image work, improving the fit between the one-click converted video works and the text and image works, and improving the quality of the video works.
[0189] Figure 13 This is a flowchart of another video generation method provided in this application embodiment. This application embodiment is executed by a computer device. See also... Figure 13 The method includes:
[0190] 1301. A computer device displays published graphic works based on a video generation interface, the video generation interface including video generation options, and the graphic works including text and a first image.
[0191] The process of displaying the graphic and textual works in step 1301 is the same as the process of displaying the graphic and textual works in step 301 above, and will not be described again here.
[0192] 1302. In response to a triggering operation of a video generation option, a computer device acquires a second image that matches the text.
[0193] In one possible implementation, the method of obtaining the second image that matches the text includes any one of the following three methods.
[0194] The first method involves using a text-to-text model to generate descriptive text based on the text in the graphic work. The semantics of the descriptive text are the same as those of the text in the graphic work, and the descriptive text is used to describe the image content. Then, using a text-to-image model, a second image is generated based on the descriptive text, and the image content of the second image is consistent with the image content described by the descriptive text.
[0195] Specifically, the text-to-text model is used to generate another text based on one text, and the text-to-image model is used to generate an image based on text. Optionally, both the text-to-text model and the text-to-image model can be large language models.
[0196] It should be noted that although the semantics of the descriptive text generated by the text-to-text model are the same as those of the text in the graphic work, the descriptive text can more appropriately and accurately describe the image content compared to the text in the graphic work. For example, the text in the graphic work focuses on narration, while the descriptive text focuses on description. Consequently, the images generated using the descriptive text are more detailed, richer, and more accurate, while retaining the semantics expressed by the text in the graphic work.
[0197] Optionally, the computer device acquires the text-to-text prompt text, inputs the text in the graphic artwork and the text-to-text prompt text into the text-to-text model, and the text-to-text model outputs the descriptive text. This text-to-text prompt text is used to prompt the generation of text with the same semantics as the text in the graphic artwork and to describe the content of the image.
[0198] For example, the prompt text is: "You are a visual describer. You can summarize the composition and scene of a scene based on a piece of text. Please generate a similar visual description text based on this sentence. The description needs to include quantifiers, details, and background. The text should be concise, and the output should be within 50 characters." The sentence is: xxxxx. The generated description text has a photographic style and uses a panoramic lens. The image parameters are: aspect ratio 9:16, image quality HD.
[0199] In this process, after adding the text from the graphic work at the "xxxxx" position above, the entire text can be input into the text-to-text model so that the model can output descriptive text.
[0200] In this embodiment of the application, text-to-text model and text-to-image model can be used to generate images that match the text in the text and image works. That is, artificial intelligence is used to obtain images that match the text in the text and image works, which can improve the diversity and convenience of obtaining images that match the text.
[0201] The second method is to extract keywords from the text; search for images that match the keywords, and then identify the searched images as the second images that match the text.
[0202] Computer equipment extracts keywords from the text of a graphic work. Keywords in the text refer to words or phrases that can summarize, express, or indicate the main content or meaning of the text. Furthermore, the image that matches the keyword can summarize, express, or indicate the main content or meaning of the text, so the image that matches the keyword can be used as a second image that matches the text.
[0203] Optionally, the computer device searches an image database for images that match the keywords. This image database stores multiple images.
[0204] Optionally, the image database stores multiple images and the semantic features of each image. The computer device extracts the semantic features of the keywords, searches the image database for the image semantic features with the highest similarity to the keyword semantic features, and identifies the image corresponding to the image semantic features as the image that matches the text.
[0205] In this embodiment, keywords are automatically extracted from the text and images are searched for. This reduces the manpower and time required for manual image searching and ensures that the obtained images are highly relevant to the text in the graphic works, thereby improving the efficiency and accuracy of obtaining images that match the text.
[0206] The third method: Display the image upload entry and prompt information on the video generation interface. The prompt information is used to prompt the upload of an image that matches the text; retrieve the image uploaded in the image upload entry, and determine the uploaded image as the second image that matches the text.
[0207] The video generation interface displays an image upload entry, through which users can upload a second image that matches the text.
[0208] Optionally, in response to a trigger operation on the image upload entry, the computer device displays a local photo album, which includes multiple images. The user can search for and select an image that matches the text from among the multiple images in the local photo album. In response to the selection of any image in the local photo album, the computer device uploads the selected image to the video generation interface, using the uploaded image as a second image that matches the text.
[0209] It should be noted that since the image was uploaded by the user, the matching degree between the image and the text of the graphic work may be high or low. However, for computer devices, it is not necessary to pay attention to the matching degree between the uploaded image and the text. Instead, the uploaded image is directly used as the image that matches the text, and the image is processed according to the logic that the image belongs to the image that matches the text.
[0210] In this embodiment, users can upload images that match the text in the graphic works. By allowing users to select and provide images that match the text, user needs can be met to the greatest extent, enhancing user control over the video generation process and improving the personalization and accuracy of the generated video works.
[0211] 1303. The computer equipment fuses the first image with the second image to obtain a video image.
[0212] After obtaining the first image and the second image, the first image and the second image are merged to obtain the video frame, which is the frame of the video work.
[0213] In one possible implementation, the dimensions of the first image and the second image can be the same or different. Optionally, the first image and the second image are located on different layers in the video frame, and the images in different layers are independent of each other.
[0214] In one possible implementation, if a graphic work includes multiple first images, the computer device merges the multiple first images with second images to obtain a video frame, wherein the order of the multiple first images in the video frame is the same as the order of the multiple first images in the graphic work.
[0215] In one possible implementation, the first image and the second image are fused, including any one of the following three methods.
[0216] The first method involves overlaying the first image onto the second image to obtain the video frame, with the layer containing the first image located above the layer containing the second image.
[0217] In this method, the size of the second image is larger than the size of the first image. Therefore, displaying the first image on top of the second image will create a picture-in-picture effect, that is, the first image is embedded in the second image.
[0218] Optionally, the computer device adjusts the transparency of the first image to a preset transparency, and overlays the first image with the preset transparency onto the second image. For example, the preset transparency can be semi-transparent or opaque.
[0219] The second method involves overlaying a second image onto the first image to obtain a video frame, with the layer containing the second image located above the layer containing the first image.
[0220] In this method, the size of the first image is larger than the size of the second image. The second image is superimposed on the first image, which also creates a picture-in-picture effect, that is, the second image is embedded on the first image.
[0221] The third method: stitch the first image and the second image together to obtain the video image.
[0222] In this case, the size of the first image and the size of the second image can be the same or different, and there is no limitation on the size of the first image and the size of the second image.
[0223] For example, the first image can be stitched on top of the second image, or the first image can be stitched on the left side of the second image, etc., but this application does not limit this.
[0224] In this embodiment, the first and second images in the video work can be merged by overlaying or splicing, which improves the diversity and flexibility of the presentation of content in the generated video work.
[0225] In one possible implementation, in step 1302 above, the computer device segments the text according to a preset number of characters to obtain multiple text fragments, each text fragment having no more than the preset number of characters; for any text fragment, a second image matching the text fragment is obtained.
[0226] In step 1303, the computer device fuses the acquired second images with the first image to obtain a video frame, which includes the second images and the first image.
[0227] In a video work, a second image matching the text in the text-based work is displayed. However, if only one second image matching the complete text in the text-based work is obtained, or multiple second images matching existing paragraphs of the text are obtained, the number of second images will be insufficient, leading to a monotonous video work. Therefore, the computer device divides the text into multiple text segments according to a preset number of characters, and obtains a second image matching each text segment. This allows the second image matching each text segment to be displayed sequentially in the second image section of the video work, making the content of the video work richer and more varied.
[0228] In this embodiment, the text is divided into multiple text segments according to the number of characters, and each text segment corresponds to its own matching second image, so that the final generated video work includes images matching multiple text segments, thereby improving the content richness of the generated video work.
[0229] Optionally, the computer device merges multiple second images according to the arrangement order of the text fragments corresponding to the multiple second images in the graphic work to obtain an initial video frame, wherein the display order of the multiple second images in the initial video frame is consistent with the arrangement order of the text fragments corresponding to the multiple second images in the graphic work; a first image is added to the initial video frame to obtain a video frame.
[0230] Optionally, the display duration of each second image in the initial video frame is positively correlated with the length of the text segment corresponding to that second image. That is, the longer the text segment corresponding to the second image, the longer the display duration of the second image in the initial video frame; conversely, the shorter the text segment corresponding to the second image, the shorter the display duration of the second image in the initial video frame.
[0231] For example, the preset word count is 50 characters, meaning that the number of characters in the divided text segment is no more than 50 characters. For a text segment of length 50 characters, the second image matched by the text segment can be displayed in the initial video frame for 10 seconds.
[0232] Optionally, the first image is interspersed among multiple text segments. Adding the first image to the initial video frame to obtain a video frame includes: identifying a target text segment adjacent to and preceding the first image among the multiple text segments; identifying a target video segment containing a second image that matches the target text segment in the initial video frame; and adding the first image to the target video segment in the initial video frame to obtain the video frame.
[0233] The multiple text segments in the text can be natural paragraphs that have already been divided in the text of the graphic work, or the text in the graphic work can be re-segmented to obtain multiple text segments.
[0234] The first image is interspersed among multiple text segments. It serves as an illustration for a specific text segment within a graphic work. Typically, the first image is an illustration for a target text segment adjacent to and preceding the first image. Therefore, the first image and its corresponding target text segment are strongly correlated. Similarly, the second image matching this target text segment is also strongly correlated. Thus, the first image and its matching second image can be considered strongly correlated. Furthermore, to ensure optimal video display, the first image is added to the target video segment containing the matching second image, allowing the strongly correlated first and second images to be displayed synchronously in the video.
[0235] Optionally, the number of first images in a graphic work can be one or more, with each first image interspersed among multiple text segments. When there are multiple first images in a graphic work, for any given first image, the above method is used to add the first image to its corresponding target video segment, thereby obtaining the video frame.
[0236] Optionally, the display duration of each first image in the video frame is a first preset duration, and the display duration of each second image in the video frame is a second preset duration. For example, the second preset duration is longer than the first preset duration, such as the first preset duration being 3 seconds and the second preset duration being 10 seconds.
[0237] In this embodiment of the application, the image in a graphic work is usually strongly correlated with the content of the text segment preceding the image. Therefore, when generating a video work, the image is merged into the video segment where the image matching the text segment is located, so that the image following the text segment can be displayed correspondingly to the image matching the text segment, thereby improving the smoothness and rigor of the intelligently generated video content.
[0238] 1304. Computer equipment generates video works that are associated with text and graphic works based on video footage.
[0239] In one possible implementation, computer devices add dynamic effects, static effects, filters, stickers, etc., to video footage to create a video work that is associated with the text and image work.
[0240] The method provided in this application, for published text and image works, can generate video works associated with the published text and image works based on the text and images in the published text and image works. This achieves one-click conversion of published text and image works into video works, eliminating the need for users to manually create video works, saving manpower and time, and improving the efficiency of video generation. Furthermore, the video works converted from text and image works include two types of images: one image is the original image from the text and image work, and the other image is an image that matches the text in the text and image work. Therefore, it ensures that the original images in the text and image work are not lost, and it can reflect the content of the text and image work, improving the fit between the one-click converted video works and the text and image works, and improving the quality of the video works.
[0241] Figure 14 This is a flowchart of another video generation method provided in this application embodiment. This application embodiment is executed by a computer device. See also... Figure 14 The method includes:
[0242] 1401. A computer device displays published graphic works based on a video generation interface, the video generation interface including video generation options, and the graphic works including text and a first image.
[0243] The process of displaying the graphic and textual works in step 1401 is the same as the process of displaying the graphic and textual works in step 301 above, and will not be repeated here.
[0244] 1402. In response to a triggering operation of a video generation option, the computer device acquires a second image that matches the text.
[0245] The process of obtaining the second image that matches the text in step 1402 is the same as the process of obtaining the second image that matches the text in step 1302 above, and will not be repeated here.
[0246] 1403. The computer equipment fuses the first image with the second image to obtain a video image.
[0247] The process of acquiring video footage in step 1403 is the same as that in step 1303 above, and will not be repeated here.
[0248] 1404. Computer devices generate video dubbing that matches the text in a graphic work.
[0249] In this embodiment of the application, a computer device acquires the text in a graphic work, converts the text into audio, and uses the audio as the video dubbing in a video work.
[0250] In one possible implementation, the computer device stores multiple audio conversion models, each used to convert any text into audio. Different audio conversion models have different parameters such as intonation, tone, timbre, and language to generate audio with different characteristics. For example, the audio converted by the audio conversion model could be the voice of a lively cartoon character, a deep male voice, or a gentle female voice. Optionally, the audio conversion model in this embodiment is a TTS (Text To Speech) model or other models.
[0251] The computer device acquires a target model identifier, determines the audio conversion model indicated by the target model identifier from multiple audio conversion models, inputs the text from the graphic work into the audio conversion model, and the audio conversion model converts the text to output the corresponding video dubbing. Optionally, the target model identifier can be a randomly determined model identifier, a preset model identifier, or a model identifier determined based on the user's selection operation.
[0252] In one possible implementation, a computer device segments the text in a graphic work to obtain multiple text fragments. For any given text fragment, the computer device generates a video voiceover that matches that text fragment, thereby obtaining video voiceovers corresponding to multiple text fragments.
[0253] Optionally, to ensure the generated video dubbing is more natural, the length of a single video dubbing needs to be limited, that is, the length of the text segment used to generate the video dubbing needs to be limited. Therefore, the computer device can segment the text in the graphic work according to a preset number of characters, resulting in multiple text segments, each with no more than the preset number of characters. For example, the preset number of characters is 130, meaning that when generating the video dubbing, the number of characters in each text segment will not exceed 130.
[0254] It should be noted that the preset number of characters used to divide the text segments when generating the video dubbing in step 1404 can be the same as or different from the preset number of characters used to divide the text segments when obtaining the second image in step 1302.
[0255] For example, in order to improve generation efficiency, the preset number of characters used to divide the text segments when generating video dubbing and the preset number of characters used to divide the text segments when acquiring the second image can be set to the same value. In this way, only one text segmentation is required, which simplifies the operation process.
[0256] For example, when acquiring a second image, to improve its diversity, the number of text segments needs to be increased appropriately. Therefore, the preset word count used to segment the text when acquiring the second image can be set to a relatively small value. Conversely, when generating video dubbing, to improve its smoothness, the number of text segments needs to be reduced appropriately; that is, the preset word count used to segment the text when generating the video dubbing can be set to a relatively large value. For instance, the preset word count used to segment the text when acquiring the second image can be set to 50 characters, and the preset word count used to segment the text when generating the video dubbing can be set to 140 characters.
[0257] 1405. Computer equipment generates video works based on video footage and video audio.
[0258] After acquiring video footage and video audio, computer equipment adds video audio to the video footage to obtain a video work, which includes the video footage and video audio.
[0259] In one possible implementation, the computer device further divides the text into multiple sentences, identifies the audio segment corresponding to each sentence in the video dubbing, and determines the sentences as subtitles for the corresponding audio segments. Then, in step 1405, the computer device generates a video work based on the video frame, the video dubbing, and the subtitles for each audio segment in the video dubbing.
[0260] Since there is a one-to-one correspondence between each word in the text and each word in the video dubbing, the audio segment corresponding to each sentence in the text can be found in the video dubbing. Multiple words in the sentence correspond to multiple words in the audio segment, so the sentence can serve as a subtitle for that audio segment. Furthermore, when generating the video, the subtitles for the audio segment are added to the video segment corresponding to the playback time of that audio segment, thus adding subtitles to the video.
[0261] Optionally, in the process of converting text into video dubbing, since it is necessary to convert word by word or sentence by sentence according to the text order, the audio segment corresponding to each sentence in the text can be determined directly and synchronously.
[0262] Optionally, if the audio segment corresponding to each sentence in the text is determined after obtaining the video dubbing, the following method can be used: The computer device divides the text into multiple first sentences and identifies the video dubbing to obtain multiple second sentences. For any second sentence, the computer device performs word segmentation to obtain multiple second words in the second sentence. In the text of the graphic work, the first word that matches the second word in the second sentence is searched, thus obtaining the matching relationship between the first word in the graphic work and the second word in the video dubbing. For any first sentence, the second word that matches the first first word and the second word that matches the last first word in the first sentence is determined, and the audio segment between the second word that matches the first first word and the second word that matches the last first word is determined as the audio segment corresponding to that first sentence.
[0263] For example, each second word in the video dubbing carries a timestamp indicating the playback time of that second word in the video dubbing. The computer device uses the timestamp carried by the second word that matches the first first word as the start time of the first statement, and the timestamp carried by the second word that matches the last first word as the end time of the first statement. The audio segment in the video dubbing between the start time and the end time is determined as the audio segment corresponding to the first statement.
[0264] For example, the process of finding the first word that matches the second word in the second sentence within the text of a graphic work is as follows: Following the order of the first word in the text and the order of the second word in the video dubbing, first obtain the first second word, determine the similarity between the first first word and the first second word. If the similarity is greater than a preset threshold, then the first first word matches the first second word. If the similarity is not greater than the preset threshold, then the first first word does not match the first second word. Continue determining the similarity between the first second word and the next first word, until a first word matching the first second word is found, or until the number of first words searched reaches a preset number. Stop searching for a first word matching the first second word and start searching for a first word matching the next second word, and so on, until all second words have been traversed.
[0265] If no first word matching the first second word is found when the number of first words searched reaches the preset number, then it is confirmed that no first word matching the first second word has been found.
[0266] In this context, when searching for a first word that matches another second word, the search begins after the last confirmed match of the first word. For example, if a first word that matches the previous second word is successfully found, the search for the current second word begins after the first word that matches the previous second word.
[0267] For example, when determining the second word matching the first word and the second word matching the last word in the first statement, taking the first word as an example, if no second word matching the first word is found, then the timestamp carried by the second word matching the last word in the previous first statement is determined, and this timestamp is used as the start time of the first statement. Similarly, taking the last word as an example, if no second word matching the last word is found, then the timestamp carried by the second word matching the first word in the next first statement is determined, and this timestamp is used as the end time of the first statement.
[0268] For example, when determining the second word matching the first word and the second word matching the last word in a first statement, taking the first word as an example, if no second word matching the first word is found, a first reference statement is searched. This first reference statement is the statement preceding the first statement with a determined end time point and closest to the first statement. Based on the number of words between the last word in the first reference statement and the first word in the first statement, an offset duration is determined. The time point obtained by adding this offset duration to the start time point of the first reference statement is taken as the start time point of the first statement. Optionally, the offset duration is equal to the product of the number and a preset average duration, where the preset average duration is the average playback duration per word. Taking the last word as an example, if no second word matching the last word is found, a second reference statement is searched. This second reference statement is the statement following the first statement with a determined start time point and closest to the first statement. The offset duration is determined based on the number of words between the first word in the second reference statement and the last word in the first statement. The starting time of the first statement is then obtained by adding this offset duration to the starting time of the second reference statement. Optionally, the offset duration is equal to the product of this number and a preset average duration.
[0269] The reason for not directly using the second statement identified from the video dubbing as subtitles in the above method is as follows: Since the second statement is identified, there will be some error. The second statement is not completely consistent with the statement in the graphic work, and the accuracy is not high enough. For example, the identified second statement is "I call you Xiaoming", while the actual text in the graphic work is "I call you Li Xiaoming", or the identified second statement is "Hello Xiaoming", while the actual text in the graphic work is "Hello, Xiaoming". Therefore, in this embodiment, after identifying the second statement, the matching first statement is determined in the text of the graphic work using the second statement, and then the first statement is used as the subtitle to ensure the accuracy of the subtitles.
[0270] Optionally, the process of segmenting the second sentence to obtain multiple words is as follows: The second sentence is segmented to obtain multiple initial words. High-frequency words are identified from these initial words; high-frequency words are those that appear a preset number of times. The remaining words (excluding the high-frequency words) from the initial sentences are then divided into characters to obtain multiple characters. These multiple characters and the high-frequency words are then identified as the segmented multiple words.
[0271] The method provided in this application embodiment can intelligently generate video dubbing based on the text in the graphic works when converting them into video works. This saves time and resources compared to manually recording dubbing, and can enhance the information transmission effect and emotional expression of the generated video works, thereby improving the attractiveness and interactivity of the video works.
[0272] Furthermore, when converting text and image works into video works, subtitles can be intelligently generated based on the text in the text and image works, saving time and resources for manually creating subtitles. This can enhance the information transmission effect and emotional expression of the generated video works, which is conducive to improving the attractiveness of the video works and improving the user's viewing experience.
[0273] Figure 15 This is a flowchart of another video generation method provided in the embodiments of this application, such as... Figure 15 As shown, the user enters an image and text link in the video generation interface. The link is parsed to generate an image and text video. The user can choose whether to modify the video. The computer device acquires the first image and text from the video, converts the text into video narration and subtitles, and acquires a second image based on the text. This second image can be obtained through AI-generated images, image searches, or user uploads. Based on the first image, the second image, the narration, and the subtitles, the computer device generates an initial video. The user can then modify the initial video to obtain the final video.
[0274] Figure 16This is a flowchart of another video generation method provided in the embodiments of this application, such as... Figure 16 As shown, the computer device downloads background music, sound effects, and special effects; acquires the text and first image from the graphic work; segments the text to obtain text fragments; acquires multiple second images based on the text fragments; converts the multiple text fragments into video narration; divides the text into multiple sentences; determines the audio segment corresponding to each sentence in the video narration; and uses the sentences corresponding to the audio segments as subtitles for the audio segments. It acquires second images that match each text fragment in the graphic work, creates a video frame including multiple second images, adjusts the aspect ratio of the video frame, and adds video narration, subtitles, and background music to the adjusted video frame. According to the position of the first image in the graphic work, it adds the first image to the video frame to obtain the final video frame; and adds sound effects, special effects, a title, and a cover to the video frame to obtain the video work.
[0275] The video generation method provided in this application can be applied to any scenario of converting text and images into video. For example, it can be applied to the scenario of converting published WeChat official account articles into video works. See below for details. Figure 17 Examples of implementations. Figure 17 This is a flowchart of another video generation method provided in the embodiments of this application, such as... Figure 17 As shown, the method includes the following steps:
[0276] 1701. Display the video generation interface on the first client. The video generation interface includes a link input area and link parsing options.
[0277] 1702. Retrieve the article link entered in the link input area. The article link is the link to any WeChat Official Account article that has been published on the second client.
[0278] In this system, a first account is logged in on the first client, and a second account is logged in on the second client. These two accounts are linked; for example, they may share the same real-name authentication, or they may have been registered using the same mobile phone number or email address. Essentially, the user of the first account and the user of the second account are the same person. Articles published on the second client are from the second account. This means that a user can paste the link to an article published on the second client into the video generation interface of the first client to generate a video associated with that article. Furthermore, after generating the article, the first account can subsequently publish the video on the first client.
[0279] 1703. In response to the triggering operation of the link parsing option, retrieve the WeChat Official Account article associated with the article link and display the WeChat Official Account article on the video generation interface.
[0280] 1704. In response to the modification operation of the official account article in the video generation interface, the official account article is modified and the modified official account article is displayed in the video generation interface.
[0281] 1705. In response to the triggering operation of the video generation option, generate a video work associated with the WeChat official account article based on the WeChat official account article in the video generation interface.
[0282] The video works include a first image and a second image that matches the text.
[0283] Step 1705 includes: acquiring a second image that matches the text in the WeChat Official Account article; fusing the first image and the second image in the WeChat Official Account article to obtain a video frame; generating a video voiceover that matches the text based on the text in the WeChat Official Account article; dividing the text in the WeChat Official Account article into multiple sentences; determining the audio segment corresponding to each sentence in the video voiceover; and identifying the sentences as subtitles for the corresponding audio segments. Based on the video frame, the video voiceover, and the subtitles for each audio segment in the video voiceover, a video work associated with the WeChat Official Account article is generated. Optionally, the first image is overlaid on the second image.
[0284] Optionally, after generating the video work, a link to the corresponding WeChat official account article can be added to the video work, thereby linking the WeChat official account article and video work created by the same author, which is conducive to improving the conversion rate between different works by the same author.
[0285] This application proposes a method for converting WeChat Official Account articles into videos with one click. This method assists WeChat Official Account authors and other text and image creators in converting published articles into video works with a single click. The method parses the article link entered by the user to obtain the WeChat Official Account article, converts the text in the article into video narration and subtitles, uses multiple images matching the text in the article as the video background, and overlays the images from the article onto the video background to create a picture-in-picture effect. This results in a video work associated with the WeChat Official Account article, enabling one-click video production, lowering the barrier to entry for creators, and improving the efficiency of video generation.
[0286] It should be noted that the above explanation only uses the conversion of WeChat official account articles into video works as an example. In addition, the video generation method provided in this application embodiment can also be applied to other scenarios, such as converting news information into video works with one click, or converting text-based novels into video works with one click. Alternatively, it can also convert novels containing only text into video works with one click.
[0287] Figure 18 This is a schematic diagram of the structure of a video generation device provided in an embodiment of this application. See also... Figure 18The device includes:
[0288] Display module 1801 is used to display published graphic works based on a video generation interface, the video generation interface including video generation options, and the graphic works including text and a first image.
[0289] The generation module 1802 is used to generate a video work associated with the text and image works based on the text and image works in the video generation interface in response to the trigger operation of the video generation option.
[0290] The display module 1801 is also used to display a video work on the video generation interface, the video work including a first image and a second image matching the text.
[0291] The video generation apparatus provided in this application can generate a video work associated with a published text and image work based on the text and images in the published text and image work. This achieves one-click conversion of published text and image works into video works, eliminating the need for users to manually create video works, saving manpower and time, and improving the efficiency of video generation. Furthermore, the video work converted from a text and image work includes two types of images: one image is the original image from the text and image work, and the other image is an image that matches the text in the text and image work. Therefore, it ensures that the original images in the text and image work are not lost, and it can reflect the content of the text and image work, improving the fit between the one-click converted video work and the text and image work, and improving the quality of the video work.
[0292] Optionally, see Figure 19 Display module 1801, used for:
[0293] Display the link input area and link parsing options in the video generation interface;
[0294] Retrieves the image and text link entered in the link input area. The image and text link can be a link to any published image and text work.
[0295] In response to the triggering of the link parsing option, retrieve the image and text works associated with the image and text links, and display the image and text works on the video generation interface.
[0296] Optionally, see Figure 19 Display module 1801, used for:
[0297] The interface displays the published text and image works, including an entry point for converting text and image works to video.
[0298] In response to the trigger operation of the text-to-video conversion entry, the video generation interface is displayed, showing the text-to-image works.
[0299] Optionally, see Figure 19The text in the graphic work comprises multiple text fragments, and the graphic work includes at least one first image interspersed among the multiple text fragments; the installation also includes:
[0300] The modification module 1803 is used to respond to the modification operation of the graphic work in the video generation interface, modify the graphic work, and display the modified graphic work in the video generation interface.
[0301] The modification operations for graphic works include at least one of the following: adjusting the position of any text fragment, editing any text fragment, deleting any text fragment, adjusting the position of any first image, deleting any first image, adding a text fragment, or adding an image.
[0302] Optionally, each first image is associated with one of the text segments; the display module 1801 is used to display multiple content areas in the video generation interface, each content area including a text segment, and when the text segment is associated with a first image, the content area also includes at least one first image associated with the text segment;
[0303] Modify module 1803 to determine at least one selected target content area among multiple content areas;
[0304] Modification module 1803 is also configured to, in response to a deletion operation, delete a text fragment and a first image from at least one target content area.
[0305] Optionally, see Figure 19 The device also includes:
[0306] Replacement module 1804 is used to respond to a replacement request for a second image in a video work, obtain a third image that matches the text and is different from the second image, and replace the second image in the video work with the third image.
[0307] Optionally, see Figure 19 The device also includes:
[0308] The publishing module 1805 is used to respond to the publishing operation of video works, publish video works with image and text links, and add video links to published image and text works.
[0309] Among them, the image and text links are links to the display interfaces of published image and text works, and the video links are links to the display interfaces of published video works.
[0310] Optionally, see Figure 19 Display module 1801, used for:
[0311] The video generation interface on the first client displays the graphic and text works published on the second client, which are different from the first client.
[0312] Optionally, see Figure 19 Display module 1801, used for:
[0313] Retrieve text and image works, where the text in a text and image work includes multiple text fragments and multiple first images;
[0314] The text and image works are identified to obtain the type of each text fragment and the type of each first image in the text and image works;
[0315] The text fragments belonging to the target type and the first image belonging to the target type are removed from the graphic and text work to obtain the preprocessed graphic and text work;
[0316] The pre-processed graphic works are displayed on the video generation interface.
[0317] Optionally, see Figure 19 Module 1802 is used for:
[0318] In response to a trigger action on the video generation option, acquire a second image that matches the text;
[0319] The first image and the second image are merged to obtain the video footage;
[0320] Based on video footage, generate video works that are associated with text and image works.
[0321] Optionally, see Figure 19 Module 1802 is used for:
[0322] The text-to-text model generates descriptive text based on the text in the graphic work. The semantics of the descriptive text are the same as the semantics of the text in the graphic work. The descriptive text is used to describe the content of the image.
[0323] Using the text-based image model, a second image is generated based on the descriptive text, and the content of the second image is consistent with the image content described in the descriptive text.
[0324] Optionally, see Figure 19 Module 1802 is used for:
[0325] Extracting keywords from text;
[0326] Search for images that match the keywords, and identify the searched images as the second images that match the text.
[0327] Optionally, see Figure 19 Module 1802 is used for:
[0328] The video generation interface displays an image upload entry and a prompt message, which prompts users to upload an image that matches the text.
[0329] Retrieve the image uploaded from the image upload portal and identify the uploaded image as the second image that matches the text.
[0330] Optionally, see Figure 19 Module 1802 is also used for:
[0331] The text is divided into segments according to a preset word count, resulting in multiple text fragments, each with no more than the preset word count.
[0332] For any text fragment, obtain the second image that matches the text fragment;
[0333] The acquired second images are fused with the first image to obtain the video footage.
[0334] Optionally, the generation module 1802 is used for:
[0335] According to the arrangement order of the text fragments corresponding to the multiple second images in the graphic work, the multiple second images are merged to obtain the initial video frame. The display order of the multiple second images in the initial video frame is consistent with the arrangement order of the text fragments corresponding to the multiple second images in the graphic work.
[0336] Add the first image to the initial video frame to obtain the video frame.
[0337] Optionally, the first image is interspersed among multiple text fragments; the generation module 1802 is used for:
[0338] Among multiple text fragments, identify the target text fragment that is adjacent to and in front of the first image;
[0339] In the initial video frame, identify the target video segment containing the second image that matches the target text fragment;
[0340] Add the first image to the target video segment of the initial video frame to obtain the video frame.
[0341] Optionally, see Figure 19 Module 1802 is used to perform any of the following:
[0342] The first image is overlaid on the second image to obtain the video frame, with the layer containing the first image located above the layer containing the second image.
[0343] The second image is overlaid on the first image to obtain the video image, with the layer containing the second image located above the layer containing the first image.
[0344] The first image and the second image are stitched together to obtain the video footage.
[0345] Optionally, see Figure 19 Module 1802 is also used for:
[0346] Generate video dubbing that matches the text in the graphic works;
[0347] Video works are generated based on video footage and audio.
[0348] Optionally, see Figure 19 Module 1802 is also used for:
[0349] The text is divided into multiple sentences, and the audio segment corresponding to each sentence is determined in the video dubbing. The sentences are then identified as subtitles for the corresponding audio segments.
[0350] A video production is generated based on the video footage, the video dubbing, and the subtitles for each audio segment in the video dubbing.
[0351] Optionally, see Figure 19 The graphic work includes multiple first images; the generation module 1802 is used for:
[0352] In response to a triggering operation on the video generation option, a target image is determined among a plurality of first images, wherein the target image refers to the first image containing the graphic code;
[0353] Based on the text and target image in the graphic work, a video work is generated, which includes the target image and a second image that matches the text.
[0354] It should be noted that the video generation apparatus provided in the above embodiments is only an example of the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the computer device can be divided into different functional modules to complete all or part of the functions described above. In addition, the video generation apparatus and the video generation method embodiments provided in the above embodiments belong to the same concept, and their specific implementation process can be found in the method embodiments, which will not be repeated here.
[0355] This application also provides a computer device, which includes a processor and a memory. The memory stores at least one computer program, which is loaded and executed by the processor to perform the operations performed in the video generation method of the above embodiments.
[0356] Optionally, the computer device is provided as a terminal. Figure 20A schematic diagram of the structure of a terminal 2000 provided in an exemplary embodiment of this application is shown.
[0357] Terminal 2000 includes: processor 2001 and memory 2002.
[0358] Processor 2001 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. Processor 2001 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field Programmable Gate Array), and PLA (Programmable Logic Array). Processor 2001 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 2001 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, processor 2001 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.
[0359] The memory 2002 may include one or more computer-readable storage media, which may be non-transitory. The memory 2002 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory 2002 are used to store at least one computer program, which is used by the processor 2001 to implement the video generation method provided in the method embodiments of this application.
[0360] In some embodiments, the terminal 2000 may also optionally include: a peripheral device interface 2003 and at least one peripheral device. The processor 2001, memory 2002, and peripheral device interface 2003 can be connected via a bus or signal line. Each peripheral device can be connected to the peripheral device interface 2003 via a bus, signal line, or circuit board. Optionally, the peripheral device includes at least one of: a radio frequency circuit 2004, a display screen 2005, a camera assembly 2006, an audio circuit 2007, and a power supply 2008.
[0361] Peripheral device interface 2003 can be used to connect at least one I / O (Input / Output) related peripheral device to processor 2001 and memory 2002. In some embodiments, processor 2001, memory 2002 and peripheral device interface 2003 are integrated on the same chip or circuit board; in some other embodiments, any one or two of processor 2001, memory 2002 and peripheral device interface 2003 can be implemented on separate chips or circuit boards, which is not limited in this embodiment.
[0362] The radio frequency (RF) circuit 2004 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The RF circuit 2004 communicates with communication networks and other communication devices via electromagnetic signals. The RF circuit 2004 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals back into electrical signals. Optionally, the RF circuit 2004 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, etc. The RF circuit 2004 can communicate with other devices through at least one wireless communication protocol. This wireless communication protocol includes, but is not limited to: metropolitan area networks (MANs), various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area networks (WLANs), and / or WiFi (Wireless Fidelity) networks. In some embodiments, the RF circuit 2004 may also include circuitry related to NFC (Near Field Communication), which is not limited in this application.
[0363] Display screen 2005 is used to display a UI (User Interface). This UI may include graphics, text, icons, video, and any combination thereof. When display screen 2005 is a touch display screen, it also has the ability to collect touch signals on or above its surface. These touch signals can be input as control signals to processor 2001 for processing. In this case, display screen 2005 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one display screen 2005, disposed on the front panel of terminal 2000; in other embodiments, there may be at least two display screens, disposed on different surfaces of terminal 2000 or in a folded design; in still other embodiments, display screen 2005 may be a flexible display screen, disposed on a curved or folded surface of terminal 2000. Furthermore, display screen 2005 may also be configured as a non-rectangular, irregular shape, i.e., a non-rectangular screen. The display screen 2005 can be made of materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).
[0364] The camera assembly 2006 is used to acquire images or videos. Optionally, the camera assembly 2006 includes a front-facing camera and a rear-facing camera. The front-facing camera is disposed on the front panel of the terminal 2000, and the rear-facing camera is disposed on the back of the terminal 2000. In some embodiments, there are at least two rear-facing cameras, which are any one of a main camera, a depth-sensing camera, a wide-angle camera, and a telephoto camera, to achieve background blurring by fusion of the main camera and the depth-sensing camera, panoramic shooting by fusion of the main camera and the wide-angle camera, VR (Virtual Reality) shooting, or other fusion shooting functions. In some embodiments, the camera assembly 2006 may also include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm light flash and a cool light flash, which can be used for light compensation at different color temperatures.
[0365] The audio circuit 2007 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, converting them into electrical signals that are input to the processor 2001 for processing, or to the radio frequency circuit 2004 for voice communication. For stereo sound acquisition or noise reduction purposes, multiple microphones may be used, each positioned at a different location on the terminal 2000. The microphone may also be an array microphone or an omnidirectional microphone. The speaker is used to convert electrical signals from the processor 2001 or the radio frequency circuit 2004 into sound waves. The speaker may be a conventional diaphragm speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can convert electrical signals not only into audible sound waves but also into inaudible sound waves for purposes such as distance measurement. In some embodiments, the audio circuit 2007 may also include a headphone jack.
[0366] The power supply 2008 is used to power the various components in the terminal 2000. The power supply 2008 can be AC power, DC power, a disposable battery, or a rechargeable battery. When the power supply 2008 includes a rechargeable battery, the rechargeable battery can support wired or wireless charging. The rechargeable battery can also be used to support fast charging technology.
[0367] In some embodiments, the terminal 2000 further includes one or more sensors 2009. The one or more sensors 2009 include, but are not limited to: an acceleration sensor 2010, a gyroscope sensor 2011, a pressure sensor 2012, an optical sensor 2013, and a proximity sensor 2014.
[0368] Accelerometer 2010 can detect the magnitude of acceleration along the three coordinate axes of a coordinate system established by terminal 2000. For example, accelerometer 2010 can be used to detect the components of gravitational acceleration along the three coordinate axes. Processor 2001 can control display screen 2005 to display the user interface in either a landscape or portrait view based on the gravitational acceleration signal acquired by accelerometer 2010. Accelerometer 2010 can also be used for games or for acquiring user motion data.
[0369] The gyroscope sensor 2011 can detect the orientation and rotation angle of the terminal 2000. The gyroscope sensor 2011, in conjunction with the accelerometer sensor 2010, can collect 3D motion data from the user on the terminal 2000. Based on the data collected by the gyroscope sensor 2011, the processor 2001 can perform the following functions: motion sensing (e.g., changing the UI based on the user's tilt), image stabilization during shooting, game control, and inertial navigation.
[0370] The pressure sensor 2012 can be installed on the side bezel of the terminal 2000 and / or on the lower layer of the display screen 2005. When the pressure sensor 2012 is installed on the side bezel of the terminal 2000, it can detect the user's grip signal on the terminal 2000, and the processor 2001 can perform left / right hand recognition or quick operation based on the grip signal collected by the pressure sensor 2012. When the pressure sensor 2012 is installed on the lower layer of the display screen 2005, the processor 2001 can control the operable controls on the UI interface based on the user's pressure operation on the display screen 2005. The operable controls include at least one of button controls, scroll bar controls, icon controls, and menu controls.
[0371] An optical sensor 2013 is used to collect ambient light intensity. In one embodiment, a processor 2001 can control the display brightness of a display screen 2005 based on the ambient light intensity collected by the optical sensor 2013. Optionally, when the ambient light intensity is high, the display brightness of the display screen 2005 is increased; when the ambient light intensity is low, the display brightness of the display screen 2005 is decreased. In another embodiment, the processor 2001 can also dynamically adjust the shooting parameters of a camera assembly 2006 based on the ambient light intensity collected by the optical sensor 2013.
[0372] The proximity sensor 2014, also known as a distance sensor, is installed on the front panel of the terminal 2000. The proximity sensor 2014 is used to detect the distance between the user and the front of the terminal 2000. In one embodiment, when the proximity sensor 2014 detects that the distance between the user and the front of the terminal 2000 is gradually decreasing, the processor 2001 controls the display screen 2005 to switch from a screen-on state to a screen-off state; when the proximity sensor 2014 detects that the distance between the user and the front of the terminal 2000 is gradually increasing, the processor 2001 controls the display screen 2005 to switch from a screen-off state to a screen-on state.
[0373] Those skilled in the art will understand that Figure 20 The structure shown does not constitute a limitation on the terminal 2000, and may include more or fewer components than shown, or combine certain components, or use different component arrangements.
[0374] Optionally, the computer device is provided as a server. Figure 21This is a schematic diagram of a server structure provided in an embodiment of this application. The server 2100 can vary significantly due to different configurations or performance. It may include one or more Central Processing Units (CPUs) 2101 and one or more memories 2102. The memories 2102 store at least one computer program, which is loaded and executed by the processor 2101 to implement the methods provided in the above-described method embodiments. Of course, the server may also have wired or wireless network interfaces, a keyboard, and input / output interfaces for input and output. The server may also include other components for implementing device functions, which will not be elaborated upon here.
[0375] This application also provides a computer-readable storage medium storing at least one computer program, which is loaded and executed by a processor to implement the operations performed by the video generation method of the above embodiments.
[0376] This application also provides a computer program product, including a computer program loaded and executed by a processor to perform the operations performed by the video generation method of the above embodiments.
[0377] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.
[0378] The above description is only an optional embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present application should be included within the protection scope of the present application.
Claims
1. A video generation method, characterized in that, The method includes: Based on the video generation interface, published graphic and text works are displayed. The video generation interface includes video generation options, and the graphic and text works include text and a first image. In response to the triggering operation of the video generation option, a video work associated with the text and image work is generated based on the text and image work in the video generation interface; The video work is displayed on the video generation interface. The video work includes the first image and a second image that matches the text.
2. The method according to claim 1, characterized in that, The video-based interface displays published text and image works, including: The video generation interface displays a link input area and link parsing options; Retrieve the image and text link entered in the link input area, where the image and text link is a link to any published image and text work; In response to the triggering operation of the link parsing option, the image and text works associated with the image and text link are obtained and displayed on the video generation interface.
3. The method according to claim 1, characterized in that, The video-based interface displays published text and image works, including: The display interface for the published graphic works is shown, and the display interface for the graphic works includes an entry point for converting graphic works to video. In response to the triggering operation of the text-to-video entry point, the video generation interface is displayed, and the text-to-image work is displayed on the video generation interface.
4. The method according to claim 1, characterized in that, The text in the graphic work includes multiple text segments, and the graphic work includes at least one first image interspersed among the multiple text segments; after the video work is displayed on the video generation interface, the method further includes: In response to the modification operation of the graphic work in the video generation interface, the graphic work is modified and the modified graphic work is displayed in the video generation interface; The modification operations on the graphic works include at least one of the following: adjusting the position of any text segment, editing any text segment, deleting any text segment, adjusting the position of any first image, deleting any first image, adding text segments, or adding images.
5. The method according to claim 4, characterized in that, Each of the first images is associated with one of the text fragments; the video-based interface displays published graphic and text works, including: The video generation interface displays multiple content areas, each content area including a text segment. If the text segment is associated with the first image, the content area also includes at least one first image associated with the text segment. The step of responding to a modification operation on the graphic work in the video generation interface, modifying the graphic work, and displaying the modified graphic work in the video generation interface includes: Among the plurality of content areas, at least one target content area is selected; In response to the deletion operation, the text fragment and the first image in the at least one target content area are deleted.
6. The method according to claim 1, characterized in that, After the video work is displayed on the video generation interface, the method further includes: In response to a request to replace a second image in the video work, a third image that matches the text and is different from the second image is obtained, and the second image in the video work is replaced with the third image.
7. The method according to claim 1, characterized in that, After the video work is displayed on the video generation interface, the method further includes: In response to the publishing operation of the video work, the video work carrying the image and text link is published, and the video link is added to the published image and text work; The text and image links are links to the display interfaces of the published text and image works, and the video links are links to the display interfaces of the published video works.
8. The method according to claim 1, characterized in that, The video-based interface displays published text and image works, including: The video generation interface of the first client displays the graphic and textual works published on the second client, and the first client and the second client are different.
9. The method according to any one of claims 1-8, characterized in that, The video-based interface displays published text and image works, including: The graphic and textual work is obtained, wherein the text in the graphic and textual work includes multiple text fragments, and the graphic and textual work includes multiple first images; The graphic and textual works are identified to obtain the type of each text fragment and the type of each first image in the graphic and textual works; The text fragments belonging to the target type and the first image belonging to the target type are removed from the graphic work to obtain the preprocessed graphic work; The pre-processed graphic artwork is displayed on the video generation interface.
10. The method according to any one of claims 1-8, characterized in that, In response to a triggering operation of the video generation option, generating a video work associated with the text and image work based on the text and image work in the video generation interface includes: In response to a triggering operation of the video generation option, a second image matching the text is acquired; The first image and the second image are merged to obtain a video frame; Based on the video footage, a video work associated with the text and image work is generated.
11. The method according to claim 10, characterized in that, The step of obtaining a second image that matches the text includes: Using the text-to-text model, descriptive text is generated based on the text in the graphic work. The semantics of the descriptive text are the same as the semantics of the text in the graphic work. The descriptive text is used to describe the image content. The second image is generated based on the descriptive text using the text-based image model, and the image content of the second image is consistent with the image content described by the descriptive text.
12. The method according to claim 10, characterized in that, The step of obtaining a second image that matches the text includes: Extract keywords from the text; Search for images that match the keywords, and identify the searched images as the second images that match the text.
13. The method according to claim 10, characterized in that, The step of obtaining a second image that matches the text includes: The video generation interface displays an image upload entry and a prompt message, which prompts users to upload an image that matches the text. The image uploaded in the image upload portal is retrieved, and the uploaded image is identified as the second image that matches the text.
14. The method according to claim 10, characterized in that, The step of obtaining a second image that matches the text includes: The text is divided into segments according to a preset number of characters to obtain multiple text fragments, and the number of characters in each text fragment does not exceed the preset number of characters; For any text fragment, obtain a second image that matches the text fragment; The step of fusing the first image and the second image to obtain a video frame includes: The acquired second images are fused with the first image to obtain the video frame.
15. The method according to claim 14, characterized in that, The step of fusing the acquired multiple second images with the first image to obtain the video frame includes: According to the arrangement order of the text fragments corresponding to the multiple second images in the graphic work, the multiple second images are merged to obtain an initial video frame. The display order of the multiple second images in the initial video frame is consistent with the arrangement order of the text fragments corresponding to the multiple second images in the graphic work. The first image is added to the initial video frame to obtain the video frame.
16. The method according to claim 15, characterized in that, The first image is interspersed among the multiple text fragments; the step of adding the first image to the initial video frame to obtain the video frame includes: Among the plurality of text fragments, a target text fragment that is adjacent to and located in front of the first image is identified; In the initial video frame, the target video segment containing the second image that matches the target text fragment is determined; The first image is added to the target video segment of the initial video frame to obtain the video frame.
17. The method according to claim 10, characterized in that, The step of fusing the first image and the second image to obtain a video frame includes any one of the following: The first image is overlaid on the second image to obtain the video frame, wherein the layer containing the first image is located above the layer containing the second image; The second image is overlaid on the first image to obtain the video frame, wherein the layer containing the second image is located above the layer containing the first image. The first image and the second image are stitched together to obtain the video frame.
18. The method according to claim 10, characterized in that, The method further includes: Based on the text in the graphic works, generate a video dubbing that matches the text; The process of generating a video work associated with the text and image work based on the video footage includes: The video work is generated based on the video footage and the video voiceover.
19. The method according to claim 18, characterized in that, The method further includes: The text is divided into multiple sentences, and the audio segment corresponding to each sentence is determined in the video dubbing. The sentences are then identified as subtitles for the audio segments corresponding to the sentences. The process of generating the video work based on the video footage and the video voiceover includes: The video work is generated based on the video footage, the video dubbing, and the subtitles for each audio segment in the video dubbing.
20. The method according to any one of claims 1-8, characterized in that, The graphic and textual work includes multiple first images; the step of generating a video work associated with the graphic and textual work based on the graphic and textual work in the video generation interface, in response to a trigger operation on the video generation option, includes: In response to a triggering operation of the video generation option, a target image is determined among a plurality of first images, the target image being a first image containing a graphic code; Based on the text in the graphic work and the target image, the video work is generated, and the video work includes the target image and a second image that matches the text.
21. A video generation apparatus, characterized in that, The device includes: The display module is used to display published graphic works based on a video generation interface, wherein the video generation interface includes video generation options and the graphic works include text and a first image; The generation module is used to generate a video work associated with the text and image works based on the text and image works in the video generation interface in response to the triggering operation of the video generation option. The display module is also used to display the video work on the video generation interface, the video work including the first image and a second image matching the text.
22. A computer device, characterized in that, The computer device includes a processor and a memory, the memory storing at least one computer program, which is loaded and executed by the processor to perform the operations of the video generation method as described in any one of claims 1 to 20.
23. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one computer program, which is loaded and executed by a processor to perform the operations of the video generation method as described in any one of claims 1 to 20.
24. A computer program product, comprising a computer program, characterized in that, The computer program is loaded and executed by a processor to perform the operations performed by the video generation method as described in any one of claims 1 to 20.