Video generation method and apparatus

CN117009581BActive Publication Date: 2025-10-10SHANGHAI BILIBILI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311148982.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-06
Publication Date
2025-10-10
Estimated Expiration
2043-09-06

AI Technical Summary

Technical Problem

现有技术中制作故事视频的流程繁琐,用户需要手动调整音频、视频与文字的匹配关系,创作门槛高,导致用户体验感差。

Method used

By determining the image tags of the initial image and/or video, combining them with story style parameters, and using AI models to generate target story text and text-to-speech, personalized story videos are automatically generated.

Benefits of technology

降低了用户制作故事视频的门槛,提升了生成速度和视频质量,提供了更好的用户体验。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117009581B_ABST
    Figure CN117009581B_ABST
Patent Text Reader

Abstract

The video generation method and device provided in the application are applied to a video generation client, wherein the video generation method comprises the following steps: determining an initial picture and / or a video, and determining a picture label of the initial picture and / or the video; determining a story style parameter, and generating a target story text according to the story style parameter and the picture label; determining a text voice of the target story text, and generating a target story video according to the target story text, the text voice, the initial picture and / or the video. The above-mentioned video generation method realizes automatic generation of a story video only through a picture and / or a video, reduces the threshold for a user to make a story video, and guarantees the quality of the generated target story video.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of video generation technology, and in particular to a video generation method, a video generation device, a computing device, and a computer-readable storage medium. Background Art

[0002] In the prior art, when people make story videos, they often do so by manually inputting story text, video materials, and music materials into a video editing platform, and they also need to manually adjust the playback speed or matching relationship between the three to obtain a story video with better effects.

[0003] However, the process for users to make story videos is cumbersome, and since they need to manually adjust the matching relationship between audio, video and text, as well as manually edit the story text, the creation threshold is high, resulting in a poor user experience in making story videos. Therefore, it is necessary to provide a technical solution that can help users generate personalized story videos more conveniently without reducing the quality of the story videos. Summary of the Invention

[0004] In view of this, an embodiment of the present application provides a video generation method. The present application also relates to a video generation apparatus, a computing device, and a computer-readable storage medium to solve the above technical problems.

[0005] According to a first aspect of an embodiment of the present application, a video generation method is provided, which is applied to a video generation client, including:

[0006] Determining an initial image and / or video, and determining image tags for the initial image and / or video;

[0007] Determining story style parameters, and generating target story text according to the story style parameters and the image labels;

[0008] The text and speech of the target story text are determined, and a target story video is generated according to the target story text, the text and speech, the initial picture and / or the video.

[0009] According to a second aspect of an embodiment of the present application, there is provided an apparatus, including:

[0010] a label obtaining module, configured to determine an initial image and / or video, and determine an image label of the initial image and / or video;

[0011] a text generation module configured to determine story style parameters and generate a target story text according to the story style parameters and the image tags;

[0012] a video generation module configured to determine a text voice of the target story text, and generate a target story video according to the target story text, the text voice, the initial picture and / or the video.

[0013] According to a third aspect of the embodiments of the present application, a computing device is provided, which includes a memory, a processor, and computer instructions stored in the memory and executable on the processor, and the processor implements the steps of the video generation method when executing the computer instructions.

[0014] According to a fourth aspect of the embodiments of the present application, a computer readable storage medium is provided, which stores computer instructions, and the computer instructions implement the steps of the video generation method when executed by a processor.

[0015] The video generation method provided by the present application is applied to a video generation client, and the initial picture and / or the video are determined, and the picture label of the initial picture and / or the video is determined; the story style parameter is determined, and the target story text is generated according to the story style parameter and the picture label; the text voice of the target story text is determined, and the target story video is generated according to the target story text, the text voice, the initial picture and / or the video.

[0016] Specifically, the video generation method can automatically generate the story text and the text voice corresponding to the story text only by selecting the picture and / or the video, that is, the personalized story video can be automatically generated according to the story text, the text voice corresponding to the story text, the picture and / or the video, which reduces the threshold for users to make the story video, improves the processing speed of the generation process of the story video, and greatly guarantees the video quality of the generated target story video. BRIEF DESCRIPTION OF DRAWINGS

[0017] Figure 1 is a specific application scenario diagram of a video generation method provided by an embodiment of the present application;

[0018] Figure 2 is a flowchart of a video generation method provided by an embodiment of the present application;

[0019] Figure 3 is a specific implementation mode diagram of a video generation method provided by an embodiment of the present application;

[0020] Figure 4 is a structural diagram of a video generation device applied to a video generation client provided by an embodiment of the present application;

[0021] Figure 5 is a structural block diagram of a computing device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0022] The following description sets forth many specific details to facilitate a thorough understanding of the present application. However, the present application can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the scope of the present application. Therefore, the present application is not limited to the specific implementations disclosed below.

[0023] The terms used in one or more embodiments of the present application are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of the present application. The singular forms "a", "the" and "the" used in one or more embodiments of the present application and the appended claims are also intended to include plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of the present application refers to and includes any or all possible combinations of one or more associated listed items.

[0024] It should be understood that although the terms first, second, etc. may be used to describe various information in one or more embodiments of the present application, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of one or more embodiments of the present application, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".

[0025] In one or more embodiments of this specification, an AI model refers to a mathematical model that uses methods from fields such as mathematics, statistics, computer science, and machine learning to analyze, process, predict, and optimize data with certain regularity and predictability. Compared with traditional mathematical models, AI models are more powerful, efficient, and flexible. AI models include machine learning models and deep neural network models.

[0026] Large models refer to deep learning models with large-scale model parameters, typically containing hundreds of millions, tens of billions, hundreds of billions, trillions, or even more than ten trillion model parameters. Large models, also known as foundation models, are pre-trained on large-scale unlabeled corpora to produce pre-trained models with parameters exceeding 100 million. Such models are adaptable to a wide range of downstream tasks and exhibit good generalization capabilities, such as large language models (LLMs) and multi-modal pre-training models.

[0027] When large models are used in practice, only a small number of samples are needed to fine-tune the pre-trained model and it can be applied to different tasks. Large models can be widely used in natural language processing (NLP), computer vision and other fields. Specifically, they can be applied to computer vision tasks such as visual question answering (VQA), image caption (IC), and image generation, as well as natural language processing tasks such as text-based sentiment classification, text summary generation, and machine translation. The main application scenarios of large models include digital assistants, intelligent robots, search, online education, office software, e-commerce, and intelligent design.

[0028] First, the terms involved in one or more embodiments of the present application are explained.

[0029] AI model: Artificial Intelligence; AI model refers to a mathematical model that uses methods from mathematics, statistics, computer science, and machine learning to analyze, process, predict, and optimize data with certain regularity and predictability; simply put, an AI model is a mathematical model that converts "data" into "intelligence."

[0030] TimeLine: Timeline refers to the project data obtained by expanding the video in chronological order, including video track, audio track, text track, sticker track, etc.

[0031] In the present application, a video generation method is provided. The present application also relates to a video generation device, a computing device, and a computer-readable storage medium, which are described in detail one by one in the following embodiments.

[0032] See also Figure 1 , Figure 1 A specific application scenario diagram of a video generation method provided according to an embodiment of the present application is shown.

[0033] Figure 1 The system includes a video generation client 102, a video generation server 104, and a model deployment server 106, wherein the video generation client 102 includes but is not limited to mobile phones, tablet computers, laptop computers, desktop computers and other clients, and the video generation server 104 and the model deployment server 106 include but are not limited to physical servers, cloud servers, etc.; the following is a video generation system formed by the video generation client 102, the video generation server 104, and the model deployment server 106, and a detailed description of the video generation method provided in an embodiment of the present application.

[0034] In actual applications, the video generation client 102 provides users with a story video generation interface, which includes but is not limited to interface controls such as a guide page, tutorials, albums, AI homepages, and additional pages. The guide page control may include a function introduction and instructions for the story video generation interface, the tutorial control may include a text tutorial or video tutorial to guide users in generating story videos, the album control may include some pictures or videos so that subsequent users can select pictures and / or videos to generate story videos, the AI ​​homepage control may include a text introduction to the AI ​​model that assists in generating story videos, and the additional page control may include an introduction to some other functions of the story video generation interface.

[0035] In addition, the video generation method provides functional services such as frame extraction, frame upload, label acquisition, story expansion, intelligent sentence segmentation, text-to-speech, script matching, and project data construction. Among them, general AI services include AI labeling (i.e., label acquisition), large language models (i.e., story expansion), intelligent sentence segmentation, TTS services (i.e., text-to-speech), etc., and business services include guide pages, tutorials, theme packages, etc.; in addition, the video generation method also includes the rendering engine Meishe Nvs for the final story video rendering and generation.

[0036] When generating a story video, the user triggers the album control in the story video generation interface of the video generation client 102 to select a target video. The video generation client 102 extracts frames from the target video to obtain target video frames, and then uploads the frames to obtain image tags of the target video frames according to the image processing model. At the same time, the story style parameters selected by the user in the story generation interface are received (for example, the atmosphere of the story selected in the theme package is cheerful, a family story, etc.), and the target story text is generated by the text processing model according to the story style parameters and the image tags (that is, an expanded story is generated through a large language model). After the target story text is intelligently segmented into sentences, the text-to-speech model is used to obtain the text and speech of the target story text. Finally, the target story video is generated according to the target story text, its corresponding text and speech, and the video. That is, after the target story text and the video are script-matched, the matched target story text and video, and the text and speech corresponding to the target story text are laid out on the timeline to generate the target story video.

[0037] Among them, the image processing model, text processing model and text-to-speech model are all AI models, and these AI models are all deployed on the model deployment server 106. In actual applications, they are called through the video generation server 104 for functional use.

[0038] The video generation method provided in the embodiment of the present application can automatically generate a story text only through the selected video, with the assistance of an AI model that can realize image processing and text processing, and use an AI model that can convert text to speech to generate the text speech corresponding to the story text. In this way, a personalized story video can be automatically generated based on the story text, the text speech corresponding to the story text, and the video, which lowers the threshold for users to make story videos and inspires inspiration. It also uses the AI ​​model to empower creation and assist users in generating more creative story videos.

[0039] See also Figure 2 , Figure 2 A flowchart of a video generation method provided by an embodiment of the present application is shown, wherein the video generation method is applied to a video generation client and specifically includes the following steps.

[0040] Step 202: Determine an initial picture and / or video, and determine picture tags of the initial picture and / or video.

[0041] Among them, the initial picture and / or video can be understood as the material used to generate the story video. The initial picture can be understood as any number of pictures containing any picture content and any format. For example, the initial picture is a picture of a puppy; the video can be understood as any number of pictures containing any video content and any format. For example, the video can be understood as a video of a puppy sleeping on the grass in the sun.

[0042] In specific implementation, if there are many video frames in a video, obtaining the image labels of all video frames will require a lot of calculations. Therefore, in order to save calculations and improve the efficiency of generating the target story video, the image labels can be determined only for the target images determined based on the initial image and / or video. The specific implementation method is as follows:

[0043] The determining of the image tag of the initial image and / or the video includes:

[0044] Determining a target image based on the determined initial image and / or the video, and obtaining an image label of the target image based on an image processing model;

[0045] Determine the image tag of the target image as the image tag of the initial image and / or the video,

[0046] Among them, the image processing model is an AI model.

[0047] Specifically, the target image can be understood as a target image determined according to the determined initial image and / or the video.

[0048] The image processing model can be understood as an AI model used to extract image labels. After the target image is input into the image processing model, the image processing model can output the image label of the target image. For example, if a target image contains a puppy sleeping in a cave, the image label of the target image can be puppy.

[0049] The video generation method provided in the embodiments of this specification can automatically generate image tags by only selecting pictures and / or videos, using an AI model that can process pictures and text, thereby improving the speed and accuracy of obtaining image tags. In practical applications, in order to provide users with a better experience, the initial pictures and / or videos can be selected or uploaded by the user on the video generation client. The specific implementation method is as follows:

[0050] Determining the target image based on the determined initial image and / or the video includes:

[0051] Receive the initial picture and / or the video selected by the user in the story video generation interface, and determine the target picture based on the initial picture and / or the video.

[0052] Among them, the story video generation interface can be understood as an interface that the video generation client displays to users and can realize user interaction. Users can interact with the video generation client by performing operations (such as clicking, sliding, etc.) on the story video generation interface. For example, the user can select the initial picture and / or video by clicking the album control of the story video generation interface, so that the target picture can be determined later based on the initial picture and / or video selected by the user.

[0053] In actual applications, users can also upload their favorite initial pictures and / or videos by clicking the upload control in the story video generation interface, so that the target picture can be determined later based on the initial pictures and / or videos uploaded by them.

[0054] The video generation method provided in the embodiments of this specification is a video generation client that can provide users with a story video generation interface, so that users can select initial pictures and / or videos based on the story video generation interface, so that the subsequently generated target story video is closer to the user's preferences, thereby improving the user experience.

[0055] In actual applications, the content to be determined is different depending on whether it is an initial image, a video, or both, and the corresponding target image is also different. The specific implementation method is as follows:

[0056] The determining of the target image according to the initial image and / or the video includes:

[0057] Determining the initial image as a target image;

[0058] extract video frames of the video according to the preset frame extraction rule and the length of the video, and determine the extracted video frames as the target pictures; or

[0059] extract video frames of the video according to the preset frame extraction rule and the length of the video, and determine the initial pictures and the extracted video frames as the target pictures.

[0060] The preset frame extraction rule can be understood as a preset rule for extracting frames of a video. For example, the preset frame extraction rule can be extracting one frame every 5 seconds of the video, extracting two frames every 10 seconds of the video, etc. The length of the video can be understood as the playing time of a single video, i.e., the time from the start of playing to the end of playing of a single video. For example, the length of the video 1 is 40 seconds.

[0061] Then, in the case where the user only selects pictures, the video generation client determines the initial pictures as the target pictures.

[0062] Or, in the case where the user only selects videos, the video generation client extracts video frames of the video according to the preset frame extraction rule and the length of the video, and determines the extracted video frames as the target pictures. For example, the preset frame extraction rule is extracting one frame every 5 seconds, and the length of the video 1 is 40 seconds. Then, one frame is extracted every 5 seconds of the video 1, and a total of 8 video frames are obtained, which are taken as the target pictures.

[0063] Or, in the case where the user selects both pictures and videos, the video generation client extracts video frames of the video according to the above implementation manner of extracting video frames of the video according to the preset frame extraction rule and the length of the video, and then determines the video frames and the initial pictures as the target pictures. In the above example, if the initial pictures include the picture 1 and the picture 2, and 8 video frames are obtained by extracting the video, then the 8 video frames and the picture 1 and the picture 2 are taken as the target pictures.

[0064] The video generation method provided by the embodiments of the present specification can obtain target pictures by extracting frames of a video when the video generation client determines that the user selects the video, so that only the target pictures can be used to obtain picture tags to generate a story video, which greatly improves the efficiency of generating the story video.

[0065] In actual application, in the case where the video generation client has sufficient resources, the picture processing model can be deployed on the video generation client, so that the video generation client can directly obtain picture tags of the target pictures according to the picture processing model deployed thereon.

[0066] In practice, to save computing resources and storage space on the video generation client, the image processing model can be deployed on the video generation server or on the model deployment server. The image label of the target image can then be obtained by calling the image processing model deployed on the video generation server or requesting to call the image processing model deployed on the model deployment server. The specific implementation is as follows:

[0067] The obtaining of the image label of the target image according to the image processing model includes:

[0068] Calling an image processing model deployed on the video generation server to perform image processing on the target image to obtain an image label of the target image; or

[0069] Generate a tag acquisition request, wherein the tag acquisition request carries the target image;

[0070] Sending the label acquisition request to the video generation server, so that the video generation server responds to the label acquisition request, calls the image processing model deployed on the model deployment server, performs image processing on the target image, and obtains the image label of the target image;

[0071] Receive the image tag of the target image returned by the video generation server.

[0072] The tag acquisition request may be understood as a request sent by the video generation client to the video generation server for acquiring the image tag of the target image, and the request at least includes the target image.

[0073] The video generation server can be understood as the server corresponding to the video generation client. The video generation client can call the image processing model deployed on the video generation server, or it can interact with the video generation server to instruct the video generation server to call the image processing model on the model deployment server.

[0074] The model deployment server can be understood as a server on which an image processing model is deployed outside the video generation client and the video generation server. The video generation server can call the image processing model on the model deployment server to process the target image in the video generation server.

[0075] Then, when the image processing model is deployed on the video generation server, the video generation client can directly call the image processing model deployed on the video generation server, perform image processing on the target image, and obtain the image label of the target image.

[0076] When the image processing model is deployed on the model deployment server, the video generation client generates a label acquisition request and sends the label acquisition request to the video generation server. After receiving the label acquisition request, the video generation server parses the label acquisition request and obtains the target image contained in the label acquisition request, calls the image processing model deployed on the model deployment server, inputs the target image into the image processing model, and the image processing model processes the input target image and outputs an image label for the target image. Then, the video generation server sends the image label to the video generation client.

[0077] For example, the video generation client determines a target image (for example, B1, B2, B3, B4, and B5, where B1 is a picture of a puppy sleeping in a cave, B2 is a picture of sunlight shining on a grassy mountain, B3 is a picture of a little rabbit, B4 is a picture of a little rabbit eating a carrot, and B5 is a picture of a big bad wolf on the grass). The video generation client obtains an image label (B1: puppy, cave; B2: sunlight, grass; B3: little rabbit; B4: little rabbit, carrot; B5: big bad wolf) of the target image output by the image processing model through any one of the three image processing methods mentioned above (using the image processing model deployed on the local end to process the target image, or calling the image processing model deployed on the video generation server to perform image processing on the target image, or generating a label acquisition request and sending it to the video generation server, so that the video generation server calls the image processing model deployed on the model deployment server, performs image processing on the target image, and returns the image label).

[0078] In the video generation method provided by the embodiments of this specification, when the image processing model is deployed on the video generation server, the video generation client obtains the image label of the target image by calling the image processing model deployed on the video generation server, thereby reducing the storage pressure and operating pressure of the video generation client and ensuring the speed of obtaining the image label; and when the image processing model is deployed on the model deployment server, the video generation client receives the image label obtained by the video generation client and calling the image processing model deployed on the model deployment server to process the target image, so as to realize the use of a more accurate image processing model stored on the model deployment server (that is, the image processing model deployed in the model deployment server can be updated more timely) to process the target image, thereby improving the accuracy of the image label.

[0079] Step 204: Determine story style parameters, and generate target story text according to the story style parameters and the image tags.

[0080] Specifically, in order to improve the efficiency and accuracy of generating the target story text based on the story style parameters and image labels, the target story text can be generated based on the text processing model. The specific implementation method is as follows:

[0081] Generating a target story text according to the story style parameters and the image tags includes:

[0082] A target story text is generated using a text processing model according to the story style parameters and the image labels.

[0083] The story style parameter is used to characterize the style of the generated target story text. The style of the target story text includes but is not limited to sad, happy, scary, etc.

[0084] In actual applications, users can choose their own story style to enhance user experience. The specific implementation is as follows:

[0085] Determining the story style parameters includes:

[0086] Determine the story style parameters selected by the user in the story style list of the story video generation interface.

[0087] Among them, the story style list can be understood as a list in the story video generation interface that enables users to select story styles; the video generation client can obtain the story style parameters selected by the user by obtaining the user's operations in the story style list.

[0088] Continuing with the above example, the story style list in the story video generation interface includes option 1 "sad", option 2 "exciting", option 3 "scary", and option 4 "happy". The user selects option 4 "happy", and thus the video generation client determines the story style parameter selected by the user as "happy".

[0089] The video generation method provided in the embodiment of this specification increases the user's operating space and further enhances the user's experience by determining the story style parameters selected by the user in the story style list of the story video generation interface on the video generation client, and using them as the style basis for the subsequent story text.

[0090] The text processing model can be understood as an AI model that generates stories based on keywords (such as the above-mentioned image tags, input words, etc.), story style and other information. Information such as image tags and story style parameters are input into the text processing model, and the text processing model can output the target story text for the image tags and story style.

[0091] Thus, based on the story style parameters and image tags, the target story text is generated using a text processing model. This means that after determining the story style parameters, the video generation client inputs these and other information, such as the image tags, into the text processing model. The text processing model then calculates these inputs and outputs the target story text. In practical applications, to improve the readability of the story text and enhance the viewing experience of the subsequently generated target story video, the video generation client may also consider the length of the story text as a factor in story text generation.

[0092] Before determining the story style parameters and generating the target story text using a text processing model based on the story style parameters and the image tags, the method further includes:

[0093] Determining a preset playback duration of the initial image and / or a video playback duration of the video, and the number of texts played per second;

[0094] Determining the text length of the target story text according to the preset playback duration, the video playback duration, and the number of texts played per second;

[0095] Accordingly, generating a target story text using a text processing model according to the story style parameters and the image tags includes:

[0096] A target story text is generated using a text processing model according to the story style parameters, the image tags, and the text length.

[0097] The preset playback duration of the initial image can be understood as the preset playback duration of the initial image in the subsequently generated target story video, such as 3 seconds.

[0098] The number of texts played per second can be understood as the number of words contained in the audio played per second in the subsequently generated target story video, such as 5 words played per second.

[0099] The text length can be understood as the maximum number of words contained in the text that can be played in a story video. For example, the text length of a story video can be 50 words, that is, the video can play a maximum of a 50-word story text; therefore, the text length of the target story text can be understood as the maximum number of words of the target story text that can be contained in the target story video.

[0100] For ease of understanding, in some embodiments, t can be used to represent the total playback time of the initial picture and / or video, the number of initial pictures can be represented by n, and the video playback time of the video can be represented by x (then, the video playback time of the first video is x1, the video playback time of the second video is x2, and so on, the video playback time of the nth video is xn), then the total playback time of the initial picture and video is t=n*3+x1+x2+...xn, and the number of texts played per second is, for example, preset to 5 words (the specific number can be adjusted according to actual conditions), then the maximum number of texts in the target story text is t*5.

[0101] Specifically, the video generation client calculates the playback time of all initial images plus the video playback time of all videos as the total playback time of all initial images and / or videos, and calculates the product of the total playback time and the number of texts played per second. The product is the text length of the target story text.

[0102] The video generation method provided in the embodiment of this specification is a method in which the video generation client generates a target story text by using a text processing model based on story style parameters, image tags, and text length to obtain a more readable target story text, thereby improving the viewing effect of the subsequently generated target story video.

[0103] In practical applications, in order to further improve the readability of the target story text, the video generation client can also consider the number of sentences when generating the target story text. The specific implementation method is as follows:

[0104] After determining the text length of the target story text, the method further includes:

[0105] determining a maximum number of texts per sentence of the target story text;

[0106] Determining the number of sentences in the target story text according to the maximum number of texts and the text length;

[0107] Accordingly, generating a target story text using a text processing model according to the story style parameters, the image tags, and the text length includes:

[0108] A target story text is generated using a text processing model according to the story style parameters, the image labels, the text length, and the number of sentences.

[0109] The maximum number of texts per story text can be understood as the maximum number of texts that each single sentence in the target story text can contain to ensure the readability of the single sentence. For example, the maximum number of texts per story text is 20, which means that each single sentence in the target story text can contain a maximum of 20 words.

[0110] The number of sentences of the target story text can be understood as the number of single sentences contained in the target story. For example, if the number of sentences of the target story text is 10, it means that there are 10 single sentences in the target story text.

[0111] Continuing with the above example, for example, the maximum text quantity of each story text in the target story text is 20, then the number of sentences of the target story text is t*5 / 20.

[0112] Therefore, according to the story style parameter, the picture label, the text length, and the number of sentences, the target story text is generated by using the text processing model. It can be understood that the video generation client inputs the story style parameter, the picture label, the text length, and the number of sentences into the text processing model, and the text processing model performs operations thereon to output the target story text. The video generation client obtains the story text output by the text processing model.

[0113] The video generation method provided by the embodiments of the present specification can further improve the readability of the target story text and improve the viewing effect of the target story video by generating the target story text by the video generation client according to the story style parameter, the picture label, the text length, and the number of sentences using the text processing model.

[0114] In actual application, when the video generation client has sufficient resources, the text processing model can be deployed on the video generation client. Therefore, the video generation client can directly obtain the target story text according to the text processing model deployed thereon.

[0115] In specific implementation, in order to save the computing resources and storage space of the video generation client, the text processing model can be deployed on the video generation server or the model deployment server. Therefore, the target story text obtained according to the text processing model can be obtained by calling the text processing model deployed on the video generation server or requesting to call the text processing model deployed on the model deployment server. The specific implementation manner is as follows:

[0116] The target story text is generated by using the text processing model according to the story style parameter and the picture label, comprising:

[0117] calling the text processing model deployed on the video generation server to process the story style parameter and the picture label to generate the target story text; or

[0118] generating a text acquisition request, wherein the text acquisition request carries the story style parameter and the picture label;

[0119] Sending the text acquisition request to the video generation server, so that the video generation server responds to the text acquisition request, calls the text processing model deployed on the model deployment server, processes the story style parameters and the image tags, and generates the target story text;

[0120] Receive the target story text returned by the video generation server.

[0121] The generated text acquisition request can be understood as a request sent by the video generation client to the video generation server for requesting the generation of the target story text, and the request at least includes story style parameters and image tags.

[0122] Specifically, the calling process of the text processing model deployed on the video generation server is the same as the calling process of the image processing model deployed on the video generation server mentioned above, and the request calling process of the text processing model deployed on the model deployment server is the same as the request calling process of the image processing model deployed on the model deployment server mentioned above. The only difference lies in the input, output and specific operations of the model. Therefore, the calling process or request calling process of the model deployed on the video generation server or the model deployment server can refer to the specific implementation method of the text processing model deployed on the video generation client mentioned above, and will not be repeated here.

[0123] The video generation method provided in the embodiments of this specification is that the video generation client obtains the target story text by calling the text processing model deployed on the video generation server, thereby reducing the storage pressure and operating pressure of the video generation client and ensuring the generation speed of the target story text; and the video generation client receives the target story text by calling the text processing model deployed on the model deployment server to process the story style parameters and picture labels, so that the story style parameters and picture labels can be processed using a more accurate text processing model stored on the model deployment server, thereby improving the accuracy of the target story text.

[0124] In practical applications, in order to reduce the occurrence of sensitive words in the target story text, the video generation model can also perform sensitive word detection on the story style parameters, the image labels, and the target story text. The specific implementation method is as follows:

[0125] Before generating the target story text using the text processing model according to the story style parameters and the image tags, the method further includes:

[0126] Performing sensitive word detection on the story style parameters and the image labels according to the risk control model;

[0127] Accordingly, after generating the target story text using the text processing model according to the story style parameters and the image tags, the method further includes:

[0128] Sensitive word detection is performed on the target story text according to the risk control model.

[0129] Among them, the risk control model can be understood as an AI model that detects whether there are sensitive words in the text. The video generation client inputs the story style parameters and the image labels or target story text into the risk control model, and the risk control model can determine whether there are sensitive words in the input content.

[0130] In the video generation method provided in the embodiments of this specification, the video generation client uses a risk control model to detect sensitive words in story style parameters, image tags, and target story text to reduce the occurrence of sensitive words in the target story text.

[0131] In actual applications, when the resources of the video generation client are large, the risk control model can be deployed on the video generation client. Thus, the video generation client can quickly detect sensitive words in the target story text based on the risk control processing model deployed.

[0132] In addition, the risk control model can also be deployed on the video generation server and the model deployment server. The specific calling process can be found in the specific implementation method of the above-mentioned AI model, which will not be repeated here.

[0133] In practical applications, in order to increase the readability of the target story text, the video generation client can also perform sentence segmentation on the target story text after generating the target story text. The specific implementation method is as follows:

[0134] After generating the target story text using the text processing model, the method further includes:

[0135] Sentence segmentation is performed on the target story text according to a sentence segmentation model to obtain a sentence-divided target story text.

[0136] Among them, sentence segmentation can be understood as the process of dividing a story text into one or more single sentences; the sentence segmentation model can be understood as a model that can perform sentence segmentation on the input text content.

[0137] Therefore, the target story text is segmented according to the sentence segmentation model to obtain the target story text after sentence segmentation. It can be understood that the video generation client inputs the target story text into the sentence segmentation model and obtains the target story text after sentence segmentation output by the sentence segmentation model for the target story text.

[0138] For example, the target story text is "The sun is shining, a little rabbit was eating carrots on the grass, and suddenly a big bad wolf came, and the little rabbit was scared and went into a cave. There was a puppy sleeping in the cave. The puppy helped the little rabbit drive away the big bad wolf, and the little rabbit and the puppy became good friends." The video generation client inputs the target story text into the sentence segmentation model, and obtains the sentence segmentation model output for the target story text, which is a sentence-divided target story text ["The sun is shining, a little rabbit was eating carrots on the grass", "Suddenly a big bad wolf came, and the little rabbit was scared and went into a cave", "There was a puppy sleeping in the cave", "The puppy helped the little rabbit drive away the big bad wolf", "The little rabbit and the puppy became good friends"].

[0139] The video generation method provided in the embodiments of this specification increases the readability of the target story text by segmenting the target story text, thereby further improving the viewing effect of the target story video.

[0140] In practical applications, the sentence segmentation model can be deployed on the video generation client, video generation server or model deployment server. Thus, the target story text can be segmented according to the sentence segmentation model, and the target story text after sentence segmentation can be obtained by the sentence segmentation model deployed on the video generation client, video generation server or model deployment server.

[0141] In practical applications, when the resources of the video generation client are large, the sentence segmentation model can be deployed on the video generation client, so that the video generation client can directly perform sentence segmentation on the target story text based on the sentence segmentation model deployed by it.

[0142] In the case where the sentence segmentation model is deployed on the video generation server or the model deployment server, the target story text is segmented according to the sentence segmentation model to obtain the target story text after sentence segmentation. The specific implementation method is as follows:

[0143] The sentence segmentation of the target story text according to the sentence segmentation model to obtain the sentence-divided target story text includes:

[0144] Calling a sentence segmentation model deployed on the video generation server to perform sentence segmentation on the target story text to obtain a sentence-divided target story text; or

[0145] generating a sentence segmentation request, wherein the sentence segmentation request carries the target story text;

[0146] Sending the sentence segmentation request to the video generation server, so that the video generation server responds to the sentence segmentation request, calls the sentence segmentation model deployed on the model deployment server, performs sentence segmentation on the target story text, and obtains the target story text after sentence segmentation;

[0147] Receive the target story text after the sentence segmentation returned by the video generation server.

[0148] Among them, the sentence segmentation request can be understood as a request sent by the video generation client to the video generation server for requesting sentence segmentation of the target story text. The request at least contains the target story text, so that the video generation server responds to the sentence segmentation request and receives the target story text carried in the sentence segmentation request.

[0149] To improve accuracy, in some embodiments, the image tags may be displayed to the user so that the user can modify and adjust the image tags to regenerate the story text. The specific implementation is as follows: After generating the target story text, the method further includes:

[0150] When it is determined that the picture tag is updated, the updated picture tag is determined, and an updated target story text is generated according to the story style parameter and the updated picture tag.

[0151] Among them, the update of the image label can come from the modification of the image label by the user, and the updated target story text is generated according to the modification of the image label by the user and the story style parameter; for example, a picture containing a bird, a puppy and grass, the image label is "bird, grass", the story style parameter is "happy", and the generated story is "bird sleeping on the grass". The user modifies the image label to "puppy, grass", and the video generation client generates the story "puppy sleeping on the grass" according to the updated image label and the story style parameter of "happy".

[0152] The video generation method provided in the embodiments of this specification utilizes the user's modification of image tags to provide a better personalized user experience and improve the accuracy of the generated target story text.

[0153] In actual applications, users can also modify the story style parameters to regenerate the target story text. The specific implementation is as follows:

[0154] After generating the target story text, the method further includes:

[0155] When it is determined that the story style parameters are updated, the updated story style parameters are determined, and an updated target story text is generated according to the updated story style parameters and the picture tags.

[0156] Among them, the update of the story style parameters can also come from the user's modification of the story style parameters. According to the user's modification of the story style parameters and the picture label, the updated target story text is generated; for example, a picture containing a puppy and grass, the picture label is "puppy, grass", the story style parameter is "happy", and the generated target story text is "puppy sleeping on the grass". The user modifies the story style parameter to "sad", and the video generation client generates the target story text "puppy fell on the grass" according to the updated story style parameters and picture label.

[0157] The video generation method provided in the embodiments of this specification regenerates the story text according to the user's modification of the story style parameters, thereby further enhancing the user's personalized experience.

[0158] The video generation method provided in the embodiments of this specification can achieve the beneficial effect of improving the speed of sentence segmentation of the target story text when the sentence segmentation model is deployed on the video generation client. When the sentence segmentation model is deployed on the video generation server, it can achieve the beneficial effect of reducing the storage pressure and operation pressure of the video generation client and ensuring the speed of sentence segmentation of the target story text. In addition, when the sentence segmentation model is deployed on the model deployment server, it can achieve the beneficial effect of improving the accuracy of sentence segmentation of the target story text.

[0159] Step 206: Determine the text and speech of the target story text, and generate a target story video based on the target story text, the text and speech, the initial picture and / or the video.

[0160] Specifically, the text-to-speech model can be used to obtain the text-to-speech corresponding to the target story text. The specific implementation method is as follows:

[0161] Determining the text-to-speech of the target story text includes:

[0162] According to the text-to-speech model, the text-to-speech of the target story text is obtained.

[0163] Among them, the image processing model, the text processing model and the text-to-speech model are all AI models.

[0164] Among them, the text-to-speech model can be understood as an AI model that generates the audio corresponding to the input text. Therefore, according to the text-to-speech model, obtaining the text speech of the target story text can be understood as the video generation client inputting the target story text into the text-to-speech model and obtaining the audio corresponding to the target story text output by the text-to-speech model.

[0165] In practical applications, the text-to-speech model can be deployed on the video generation client, the video generation server, or the model deployment server. Thus, obtaining the text-to-speech of the target story text can be achieved by using the text-to-speech model deployed on the video generation client, the video generation server, or the model deployment server. The specific implementation method is as follows:

[0166] When the text-to-speech model is deployed on the video generation client, the video generation client can directly call the deployed text-to-speech model to perform text-to-speech processing on the target story text to obtain the text-to-speech of the target story text.

[0167] In the case where the text-to-speech model is deployed on the video generation server or the model deployment server, the specific implementation method of obtaining the text-to-speech of the target story text according to the text-to-speech model is as follows:

[0168] The step of obtaining the text-to-speech of the target story text according to the text-to-speech model includes:

[0169] Calling a text-to-speech model deployed on the video generation server to perform speech processing on the target story text according to speech generation rules to obtain the text-to-speech of the target story text; or

[0170] Generate a voice acquisition request, wherein the voice acquisition request carries the target story text and preset voice generation rules;

[0171] Sending the speech acquisition request to the video generation server, so that the video generation server responds to the speech acquisition request, calls the text-to-speech model deployed on the model deployment server, performs speech processing on the target story text according to the speech generation rule, and obtains the text speech of the target story text;

[0172] Receive the text speech of the target story text returned by the video generation server.

[0173] Among them, the text-to-speech request can be understood as a request sent by the video generation client to the video generation server for requesting text-to-speech processing of the target story text. The request at least contains the target story text, so that the video generation server responds to the text-to-speech request and receives the target story text carried in the text-to-speech request.

[0174] The video generation method provided in the embodiments of this specification can achieve the beneficial effect of improving the speed of converting text to speech of the target story text when the text-to-speech model is deployed on the video generation client. When the text-to-speech model is deployed on the video generation server, it can achieve the beneficial effect of reducing the storage pressure and operating pressure of the video generation client and ensuring the speed of converting text to speech of the target story text. In addition, when the text-to-speech model is deployed on the model deployment server, it can achieve the beneficial effect of improving the accuracy of converting text to speech of the target story text.

[0175] Generating a target story video based on the target story text, text voice, initial picture and / or video can be understood as the video generation client synthesizing the target story text, text voice, initial picture and / or video to generate the target story video.

[0176] In practical applications, the target story video can be generated based on the target story text, text-to-speech, initial image and / or video by matching them according to preset matching rules to obtain a target story video with better viewing effect. The specific implementation method is as follows:

[0177] Generating a target story video according to the target story text, the text voice, the initial picture and / or the video includes:

[0178] Matching the target story text with the initial image and / or the video according to a preset matching rule;

[0179] A target story video is generated based on the matched target story text, the initial picture and / or the video, and the text voice.

[0180] Among them, the preset matching rules can be understood as rules for matching the target story text with the initial picture and / or video. The preset matching rules include but are not limited to picture matching rules, video matching rules or mixed matching rules.

[0181] Thus, according to the preset matching rules, the target story text and the initial image and / or video are matched; and the target story video is generated based on the matched target story text, initial image and / or video, and text and speech. It can be understood that the video generation client matches the target story text and the initial image and / or video according to the preset matching rules and generates the matched target story text, initial image and / or video. Thereafter, the matched target story text, initial image and / or video are synthesized with the text and speech to generate the target story video.

[0182] The video generation method provided in the embodiments of the present specification is that the video generation client first matches the target story text and the initial picture and / or video according to a preset matching rule, and then synthesizes the matched target story text and the initial picture and / or video with text voice to generate a target story video, thereby improving the quality of the target story video.

[0183] In actual application, the video generation client can adopt different matching rules in the case that the initial picture and / or video only contains the initial picture, only contains the video, or contains both the initial picture and the video, so as to further improve the personalized matching effect of the matched initial picture and / or video and the target story text, thereby improving the quality of the subsequently generated target story video. The specific implementation manners are as follows:

[0184] The matching of the target story text and the initial picture and / or video according to the preset matching rule comprises:

[0185] matching the target story text and the initial picture according to a picture matching rule;

[0186] matching the target story text and the video according to a video matching rule;

[0187] matching the target story text and the initial picture and the video according to a mixed matching rule.

[0188] The picture matching rule can be understood as a rule of matching a picture with a text, the video matching rule can be understood as a rule of matching a video with a text, and the mixed matching rule can be understood as a rule of matching a picture and a video with a text.

[0189] Specifically, in the case that the user selects only the initial picture, the video generation client matches the target text and the initial picture according to the picture matching rule; in the case that the user selects only the video, the video generation client matches the target text and the video according to the video matching rule; and in the case that the user selects both the initial picture and the video, the video generation client matches the target text and the initial picture and the video according to the mixed matching rule.

[0190] The video generation method provided in the embodiments of the present specification matches the initial picture and / or video in different cases, i.e., only containing the initial picture, only containing the video, or containing both the initial picture and the video, by adopting different matching rules, so that the user can generate a story video only according to a picture, or generate a story video only according to a video, or generate a story video according to both a picture and a video, thereby increasing the personalized experience of the user, and the initial picture and / or video in different cases are matched by adopting different matching rules, thereby improving the quality of the subsequently generated target story video.

[0191] Then, if the initial image and / or video only contains the initial image, the video generation client matches the target story text with the initial image according to the image matching rules. Specifically, the matching can be performed based on the image tags in the initial image to achieve a better matching effect between the matched text and the initial image. The specific implementation method is as follows:

[0192] Matching the target story text with the initial picture according to the picture matching rule includes:

[0193] Determining a picture tag of the initial picture;

[0194] When it is determined that the target story text contains a keyword corresponding to the image tag of the initial image, matching the target story text with the initial image according to a first image matching rule;

[0195] When it is determined that the keyword corresponding to the image tag of the initial image does not exist in the target story text, the target story text is matched with the initial image according to a second image matching rule.

[0196] Among them, the first image matching rule can be understood as a rule that matches the target story text with the initial image (that is, the above-mentioned initial image, the target story text has keywords corresponding to the image tag of the initial image) if the target story text has keywords corresponding to the image tag of the initial image.

[0197] The second image matching rule can be understood as follows: if the target story text does not contain keywords corresponding to the image tags of the initial image, the target story text is matched with the initial image in the order in which the user uploads the initial image.

[0198] In specific implementation, each sentence in the target story text can be matched with the initial picture. For ease of understanding, the matching sentence of the current target story text is marked as the sentence to be matched, and the sentence to be matched is used as the sentence to be matched in the current target story text in the subsequent steps.

[0199] Furthermore, in the case where the to-be-matched single sentence contains a keyword corresponding to the image tag of the initial image, the keywords in the to-be-matched single sentence are deduplicated, and the number of keywords contained in the to-be-matched single sentence after deduplication is determined;

[0200] If the deduplicated sentence contains only one keyword, determine whether the initial image corresponding to the keyword has been matched.

[0201] If so, the next initial image will be matched with the sentence to be matched according to the order in which the user uploaded the initial images.

[0202] If not, then match the sentence to be matched with the initial picture;

[0203] If the single sentence to be matched after deduplication contains multiple keywords, the keyword ranked first is used as the keyword to be confirmed, and it is determined whether the initial image corresponding to the keyword to be confirmed has been matched;

[0204] If yes, then according to the order in which the keywords appear in the sentence to be matched, determine whether the keyword is the last keyword in the sentence to be matched.

[0205] If so, the sentence to be matched is matched with the next unmatched initial image found in the order in which the user uploaded the initial images;

[0206] If not, the next keyword is used as the keyword to be confirmed, and the above step of determining whether the initial image corresponding to the keyword to be confirmed has been matched is performed;

[0207] If not, the sentence to be matched is matched with the initial picture corresponding to the keyword to be confirmed.

[0208] The above steps are the specific implementation of the above first image matching rule, so as to match the target story text with the initial image when the target story text contains keywords corresponding to the image tag of the initial image.

[0209] Furthermore, if the keyword corresponding to the image tag of the initial image does not exist in the single sentence to be matched, it is determined whether there is an initial image that has not been matched;

[0210] If so, determine whether there is an initial image for which all keywords in the target story text do not correspond to their image labels;

[0211] If so, match the sentence to be matched with the initial picture of the keyword corresponding to the non-existent expected picture tag according to the order of the initial pictures uploaded by the user;

[0212] If not, the sentence to be matched is matched with the next unmatched initial picture in the order of the initial pictures uploaded by the user.

[0213] If not, the sentence to be matched is matched with the initial picture corresponding to the previous sentence of the sentence to be matched and the next initial picture corresponding to the order in which the user uploads the initial pictures.

[0214] The above steps are the specific implementation of the above second image matching rule, so as to match the target story text with the initial image when there is no keyword corresponding to the image tag of the initial image in the target story text.

[0215] The video generation method provided in the embodiments of this specification, when the initial image and / or video only contains the initial image, matches the image tags in the initial image with the keywords of the target story text so that the matched text matches the initial image better, thereby improving the quality of the subsequently generated target story video.

[0216] If the initial image and / or video only contains video, the video generation client matches the target story text with the video according to the video matching rules. Specifically, different matching rules can be used based on the relationship between the playback time of the target story text and the playback time of the video to achieve a better matching effect between the target story text and the video. The specific implementation method is as follows:

[0217] Matching the target story text with the video according to the video matching rule includes:

[0218] When it is determined that the playing time of the target story text is greater than the playing time of the video, matching the target story text with the video according to a first video matching rule;

[0219] When it is determined that the playing duration of the target story text is equal to the playing duration of the video, matching the target story text with the video according to a second video matching rule;

[0220] When it is determined that the playing time of the target story text is less than the playing time of the video, the target story text is matched with the video according to a third video matching rule.

[0221] Among them, the first video matching rule can be understood as the rule for matching the target story text with the video when the playback time of the target story text is longer than the playback time of all videos; the second video matching rule can be understood as the rule for matching the target story text with the video when the playback time of the target story text is equal to the playback time of all videos; the third video matching rule can be understood as the rule for matching the target story text with the video when the playback time of the target story text is less than the playback time of all videos.

[0222] During specific implementation, in the case where the playback duration of the target story text is longer than the playback duration of all videos, the target story text and the video are matched according to the first video matching rule. Specifically, the video generation client arranges the videos in the order in which the users upload the videos. The arrangement includes the case where the time from all videos to the target story text is less than 1 second after they are played, and the case where the time from all videos to the target story text is more than 1 second after they are played.

[0223] If the target story text is less than 1 second after all videos have finished playing, the video generation client will display a black screen after playing the video.

[0224] If the time between all videos and the target story text being played is greater than 1 second after all videos are played, the video generation client will play the target story text and then play the video again from the beginning.

[0225] In the case where the playback duration of the target story text is equal to the playback duration of all videos, the target story text and the video are matched according to the second video matching rule. Specifically, the video generation client plays the video in the order in which the user uploaded the video materials.

[0226] In the case where the playback time of the target story text is less than the playback time of all videos, the target story text and the video are matched according to the third video matching rule. Specifically, the video generation client plays the video in the order in which the user uploads the video materials, and the redundant target story text is matched with the background music.

[0227] The video generation method provided in the embodiments of this specification, when the initial image and / or video only contains video, uses different matching rules to match the target story text based on the size relationship between the playback time of the video, so that the matching effect between the target story text and the video is better, thereby improving the quality of the subsequently generated target story video.

[0228] When the initial image and / or video contains the initial image and video, the video generation client matches the target story text with the initial image and video according to a mixed matching rule. Specifically, different matching rules may be used based on the relationship between the image tags in the initial image and the keywords in the target story text to achieve a better matching effect between the target story text and the initial image and video. The specific implementation method is as follows:

[0229] Matching the target story text, the initial image, and the video according to a mixed matching rule includes:

[0230] Determining a picture tag of the initial picture;

[0231] In a case where it is determined that the keyword corresponding to the picture label of the initial picture exists in the target story text, the target story text is matched with the initial picture and the video according to a first mixed matching rule.

[0232] In a case where it is determined that the keyword corresponding to the picture label of the initial picture does not exist in the target story text, the target story text is matched with the initial picture and the video according to a second mixed matching rule.

[0233] The first mixed matching rule can be understood as a rule for matching the target story text with the initial picture and the video in a case where the keyword corresponding to the picture label of the initial picture exists in the target story text, and the second mixed matching rule can be understood as a rule for matching the target story text with the initial picture and the video in a case where the keyword corresponding to the picture label of the initial picture does not exist in the target story text.

[0234] In a specific implementation, in a case where the keyword corresponding to the picture label of the initial picture exists in the target story text, the video generation client first arranges the video in order, and inserts the picture at a position where the target story text first appears the keyword corresponding to the initial picture (or first fills the initial picture according to the initial picture matching logic, and sequentially fills the remaining part of the video).

[0235] In the above step, if the position of the inserted picture is less than 1 second from the end of the inserted video, the video at the end is automatically ignored.

[0236] In a case where the keyword corresponding to the picture label of the initial picture does not exist in the target story text, the video generation client plays the picture in advance, and each picture occupies one single sentence.

[0237] Further, in a case where the total duration of all single sentences of the target story text is greater than the total duration of all initial pictures and videos, the video generation client arranges the videos in the order of the user-uploaded videos, and matches the excess target story text with background music.

[0238] In a case where the total duration of all single sentences of the target story text is less than the total duration of all initial pictures and videos, the video is played from the beginning, and if the video is less than 1 second from the end of the single sentence, the video is played to the end and a black screen is played.

[0239] The video generation method provided in the specification of the embodiment can adopt different matching rules according to the relationship between the picture label in the initial picture and the keyword of the target story text, so that the matching effect of the matched target story text, the initial picture and the video is better, thereby improving the quality of the subsequently generated target story video.

[0240] In actual applications, in order to make the generated target story video more effective, the video generation client can also render the story video according to the video effect selected by the user. The specific implementation method is as follows:

[0241] Generating a target story video according to the matched target story text, the initial picture and / or the video, and the text voice includes:

[0242] Determine the video effect parameters selected by the user in the story style list in the story video generation interface;

[0243] The matched target story text, the initial image and / or the video, and the text voice are laid out on a timeline, and rendered using the video effect parameters to generate a target story video.

[0244] Among them, the story style list can be understood as a list in the story video generation interface that enables users to select video effects; the video generation client can obtain the video effect parameters selected by the user by obtaining the user's operations in the story style list; the video effect parameters are used to characterize the video effects of the generated target story video, including but not limited to video transition effect parameters, filter parameters, subtitle color, subtitle font, etc.

[0245] Specifically, after the video generation client determines the video effect parameters of the target story video selected by the user in the story style list in the story video generation interface, it lays the matched target story text, initial picture and / or video, and the text voice on the timeline, and uses the video effect parameters for rendering to generate the target story video.

[0246] The video generation method provided in the embodiments of this specification is that the video generation client determines the video effect parameters selected by the user in the story style list in the story video generation interface as the basis for the video effect of the target story video, and uses the video effect parameters for rendering, thereby increasing the user's operating space and further improving the user experience.

[0247] In practical applications, in order to make the generated target story video have a better viewing effect, the target story text can be combined with the target story video generated in the above steps according to the sentence segmentation. The specific implementation method is as follows:

[0248] The step of obtaining the text-to-speech of the target story text according to the text-to-speech model includes:

[0249] According to the text-to-speech model, the text-to-speech of the target story text after the sentence segmentation is obtained,

[0250] Accordingly, generating a target story video according to the target story text, the text voice, the initial picture and / or the video includes:

[0251] A target story video is generated based on the sentence-segmented target story text, the text voice, the initial picture and / or the video.

[0252] The target story text after sentence segmentation is obtained by the video generation client through the above-mentioned sentence segmentation model.

[0253] Then, the video generation client inputs the sentence-divided target story text into the text-to-speech model, obtains the text-to-speech of the sentence-divided target story text output by the text-to-speech model, and generates the target story video based on the text-to-speech of the sentence-divided target story text, the text-to-speech, the initial picture and / or video.

[0254] The video generation method provided in the embodiment of this specification generates a target story video based on the target story text after sentence segmentation, thereby ensuring the fluency of the text speech corresponding to the target story text in the generated target story video, and further improving the playback effect of the generated target story video.

[0255] The video generation method provided in the embodiments of this specification can automatically generate a story text and generate the text voice corresponding to the story text only through selected pictures and / or videos. It can automatically generate a personalized story video based on the story text, the text voice, pictures and / or videos corresponding to the story text, which lowers the threshold for users to make story videos, improves the processing speed of the story video generation process, and greatly guarantees the video quality of the generated target story video.

[0256] The embodiment of this specification also provides a video generation method applied to a video generation server, and the specific implementation method is as follows:

[0257] Receiving a target image sent by a video generation client, and obtaining an image tag of the target image according to an image processing model, wherein the target image is obtained by the video generation client according to the determined initial image and / or video;

[0258] Generate target story text using a text processing model based on the image tags and the story style parameters sent by the received video generation client;

[0259] According to the text-to-speech model, the text speech of the target story text is obtained, and the text speech is returned to the video generation client, so that the video generation client generates a target story video based on the target story text, the text speech, the initial picture and / or the video.

[0260] Among them, the specific implementation of the video generation method applied to the video generation server can refer to the corresponding description of the specific implementation of the video generation method applied to the video generation client mentioned above, and will not be repeated here.

[0261] The video generation method provided in the embodiments of this specification is that the video generation server uses an AI model that can realize image processing and text processing to assist in generating story text, and uses an AI model that can convert text into speech to assist in generating text speech corresponding to the story text. This allows the video generation client to automatically generate personalized story videos based on the story text, the text speech, pictures and / or videos corresponding to the story text, thereby lowering the threshold for users to make story videos. Moreover, since the AI ​​model is used as an auxiliary implementation in the above technical solution, it not only improves the processing speed of the video generation client in generating story videos, but also greatly guarantees the video quality of the target story video generated by the video generation client.

[0262] See also Figure 3 , Figure 3 A schematic diagram of a specific implementation of a video generation method provided according to an embodiment of this specification is shown.

[0263] Specifically, the video generation method is applied to a video generation system, which includes a video generation client, a video generation server, and a model deployment server. The video generation client includes an AI story generation component, an image selection component, an application component, a frame extraction component, and a custom service SDK. Specifically, the video generation method includes the following steps:

[0264] Step 302: The user generates an AI story generation component of the client based on the video and enters the AI ​​story generation interface.

[0265] Among them, the AI ​​story generation interface can be understood as the story video generation interface of the above embodiment.

[0266] Step 304: The AI ​​story generation interface of the video generation client includes a novice guide control, and the user can learn about the functions of the AI ​​story generation interface by clicking on the novice guide control.

[0267] Step 306: When the user triggers the album control of the AI ​​story generation interface, the video generation client enters the album through the picture selection component to select materials, such as pictures and / or videos.

[0268] Step 308: When it is determined that the material selected by the user includes a video, the frame extraction component extracts frames based on the video selected by the picture selection component.

[0269] Among them, the specific implementation method of frame extraction by the frame extraction component is shown in the above embodiment and will not be repeated here.

[0270] Step 310: The video frames extracted by the frame extraction component are displayed to the user through the picture selection component.

[0271] If multiple videos are selected, the above steps will be repeated to extract frames for the second video.

[0272] Step 312: After all video frames are extracted, the user-selected pictures and the extracted video frames are sent to the AI ​​story generation component through the picture selection component.

[0273] Step 314: The AI ​​story generation component uploads the target image package formed by the user-selected image and the extracted video frames to the application component.

[0274] Step 316: The application component feeds back identification information of the uploaded material to the AI ​​story generation component.

[0275] Step 318: The AI ​​story generation component sends a labeling request for the target image package to the custom service SDK.

[0276] Step 320: The custom service SDK generates a marking task in response to the marking request, and sends the marking task to the video generation server, wherein the marking task carries the target image package.

[0277] Step 322: The video generation server generates a corresponding task identifier for the marking task, and returns the task identifier, polling interval, timeout period, etc. corresponding to the marking task to the custom service SDK.

[0278] Step 324: The custom service SDK sends the polling result for the marking task to the video generation server according to the task identifier, polling interval, timeout period, etc. corresponding to the marking task.

[0279] Step 326: The video generation server returns the task processing result for the labeling task to the custom service SDK, that is, the image label of each target image in the target image package.

[0280] In actual applications, since users may select multiple pictures and / or videos, in order to ensure that the picture labels are not confused, the video generation server will perform the above operation each time it receives a labeling task. For the specific implementation of the picture label of the target picture, the video generation server calls the picture processing model deployed by the model deployment server. For details, please refer to the detailed introduction of the above embodiment.

[0281] Step 328: The custom service SDK returns the image label of each target image in the target image package to the AI ​​story generation component.

[0282] Step 330: The AI ​​story generation component performs risk control filtering on the image tags through the video generation server.

[0283] Among them, the video generation server performs risk control filtering on image tags, which can also be achieved by calling the risk control model in the model deployment server.

[0284] Step 332: The video generation server returns the filtered image tags to the AI ​​story generation component.

[0285] Step 334: The AI ​​story generation component enters the picture label editing page so that the user can edit the picture label on this page.

[0286] Step 336: The AI ​​story generation component requests a style list from the video generation server.

[0287] Step 338: The video generation server sends the style list to the AI ​​story generation component.

[0288] In actual applications, there is no limit on the time for obtaining the style list. For example, the style list can be obtained when the user selects the material.

[0289] Step 340: The AI ​​story generation component displays the selected pictures and / or videos to the user, as well as the picture tags of the target pictures in the target picture package.

[0290] Step 342: The user can edit and modify the image tags based on the image tag editing page, and can also delete, add, etc. the selected images and / or videos.

[0291] Step 344: The AI ​​story generation component receives a trigger instruction generated by the user clicking a button.

[0292] Step 346: The AI ​​story generation component sends a story expansion request to the custom service SDK.

[0293] Step 348: The custom service SDK returns the expanded text generated by the video generation server to the AI ​​story generation component.

[0294] Step 350: The AI ​​story generation component sends an intelligent sentence segmentation request for the expanded text to the custom service SDK.

[0295] Step 352: The custom service SDK returns the expanded text after the sentence segmentation determined by the video generation server to the AI ​​story generation component.

[0296] Step 354: The AI ​​story generation component sends a text-to-speech request for the expanded text after the sentence to the custom service SDK.

[0297] Step 356: The custom service SDK returns the text-to-speech result of the expanded text after sentence segmentation determined by the video generation server to the AI ​​story generation component.

[0298] Step 358: The AI ​​story generation component sends a request to the video generation server to download the text speech of the expanded text after the sentence.

[0299] Step 360: The video generation server returns the text speech of the expanded text after the sentence to the AI ​​story generation component.

[0300] Step 362: The AI ​​story generation component matches the expanded text after the sentence, the corresponding text voice, and the selected pictures and / or videos to generate the target story video.

[0301] Step 364: The AI ​​story generation component displays the target story video to the user.

[0302] The video generation method provided in the embodiment of the present application can automatically generate a story text only through the selected pictures and / or videos, assisted by an AI model that can realize picture processing and text processing, and use an AI model that can convert text to speech to generate the text speech corresponding to the story text. In addition, a personalized story video can be automatically generated based on the story text, the text speech, pictures and / or videos corresponding to the story text, thereby lowering the threshold for users to make story videos. Moreover, since the AI ​​model is used as an auxiliary implementation in the above technical solution, it can not only improve the processing speed of the story video generation process, but also greatly guarantee the video quality of the generated target story video.

[0303] Corresponding to the above method embodiment, the present application also provides a video generation device embodiment, Figure 4 FIG. 1 shows a schematic diagram of the structure of a video generating device provided by an embodiment of the present application. Figure 4 As shown, the device is applied to a video generation client, comprising:

[0304] The tag obtaining module 402 is configured to determine an initial image and / or video, and determine an image tag of the initial image and / or video;

[0305] The text generation module 404 is configured to determine story style parameters and generate a target story text according to the story style parameters and the image tags;

[0306] The video generation module 406 is configured to determine the text and speech of the target story text, and generate a target story video according to the target story text, the text and speech, the initial picture and / or the video.

[0307] Optionally, the label obtaining module 402 is further configured to:

[0308] Calling an image processing model deployed on the video generation server to perform image processing on the target image to obtain an image label of the target image; or

[0309] Generate a tag acquisition request, wherein the tag acquisition request carries the target image;

[0310] Sending the label acquisition request to the video generation server, so that the video generation server responds to the label acquisition request, calls the image processing model deployed on the model deployment server, performs image processing on the target image, and obtains the image label of the target image;

[0311] Receive the image tag of the target image returned by the video generation server.

[0312] Optionally, the device further includes:

[0313] The text length determination module is configured as follows:

[0314] Determining a preset playback duration of the initial image and / or a video playback duration of the video, and the number of texts played per second;

[0315] Determining the text length of the target story text according to the preset playback duration, the video playback duration, and the number of texts played per second;

[0316] Accordingly, the text generation module 404 is further configured to:

[0317] A target story text is generated using a text processing model according to the story style parameters, the image tags, and the text length.

[0318] Optionally, the device further includes:

[0319] The statement quantity determination module is configured as follows:

[0320] determining a maximum number of texts per sentence of the target story text;

[0321] Determining the number of sentences in the target story text according to the maximum number of texts and the text length;

[0322] Correspondingly, the text generation module 404 is further configured to:

[0323] generate the target story text according to the story style parameter, the picture label, the text length, and the number of sentences, by using a text processing model.

[0324] Optionally, the text generation module 404 is further configured to:

[0325] invoke a text processing model deployed on a video generation server to process the story style parameter and the picture label, and generate the target story text; or

[0326] generate a text acquisition request, wherein the text acquisition request carries the story style parameter and the picture label;

[0327] send the text acquisition request to the video generation server, so that the video generation server responds to the text acquisition request, invokes a text processing model deployed on a model deployment server to process the story style parameter and the picture label, and generates the target story text;

[0328] receive the target story text returned by the video generation server.

[0329] Optionally, the apparatus further comprises:

[0330] a segmentation module configured to:

[0331] segment the target story text according to a sentence segmentation model, to obtain a segmented target story text.

[0332] Optionally, the segmentation module is further configured to:

[0333] invoke a sentence segmentation model deployed on a video generation server to segment the target story text, to obtain a segmented target story text; or

[0334] generate a sentence segmentation request, wherein the sentence segmentation request carries the target story text;

[0335] send the sentence segmentation request to the video generation server, so that the video generation server responds to the sentence segmentation request, invokes a sentence segmentation model deployed on a model deployment server to segment the target story text, and obtains a segmented target story text;

[0336] receive the segmented target story text returned by the video generation server.

[0337] Optionally, the video generation module 406 is further configured to:

[0338] obtain a text voice of the segmented target story text according to the text-to-voice model;

[0339] Correspondingly, the video generation module 406 is further configured to:

[0340] generate a target story video according to the segmented target story text, the text voice, the initial picture and / or the video.

[0341] Optionally, the video generation module 406 is further configured to:

[0342] invoke a text-to-voice model deployed on a video generation server to perform voice processing on the target story text according to a voice generation rule to obtain a text voice of the target story text; or

[0343] generate a voice acquisition request, wherein the voice acquisition request carries the target story text and a preset voice generation rule;

[0344] send the voice acquisition request to the video generation server, so that the video generation server invokes a text-to-voice model deployed on a model deployment server to perform voice processing on the target story text according to the voice generation rule to obtain a text voice of the target story text;

[0345] receive the text voice of the target story text returned by the video generation server.

[0346] Optionally, the device further comprises:

[0347] a first sensitive word detection module configured to:

[0348] perform sensitive word detection on the story style parameters and the picture label according to a risk control model;

[0349] Correspondingly, the device further comprises:

[0350] a second sensitive word detection module configured to:

[0351] perform sensitive word detection on the target story text according to the risk control model.

[0352] Optionally, the video generation module 406 is further configured to:

[0353] match the target story text with the initial picture and / or the video according to a preset matching rule;

[0354] generate a target story video according to the matched target story text, the initial picture and / or the video, and the text voice.

[0355] Optionally, the video generation module 406 is further configured to:

[0356] match the target story text with the initial picture according to a picture matching rule;

[0357] match the target story text with the video according to a video matching rule;

[0358] match the target story text with the initial picture and the video according to a hybrid matching rule.

[0359] Optionally, the video generation module 406 is further configured to:

[0360] determine a picture label of the initial picture;

[0361] in a case where it is determined that the target story text contains a keyword corresponding to the picture label of the initial picture, match the target story text with the initial picture according to a first picture matching rule;

[0362] in a case where it is determined that the target story text does not contain a keyword corresponding to the picture label of the initial picture, match the target story text with the initial picture according to a second picture matching rule.

[0363] Optionally, the video generation module 406 is further configured to:

[0364] in a case where it is determined that the playing time of the target story text is greater than the playing time of the video, match the target story text with the video according to a first video matching rule;

[0365] in a case where it is determined that the playing time of the target story text is equal to the playing time of the video, match the target story text with the video according to a second video matching rule;

[0366] in a case where it is determined that the playing time of the target story text is less than the playing time of the video, match the target story text with the video according to a third video matching rule.

[0367] Optionally, the video generation module 406 is further configured to:

[0368] determine a picture label of the initial picture;

[0369] When it is determined that the target story text contains a keyword corresponding to the image tag of the initial image, matching the target story text with the initial image and the video according to a first mixed matching rule;

[0370] When it is determined that the target story text does not contain the keyword corresponding to the image tag of the initial image, the target story text is matched with the initial image and the video according to a second hybrid matching rule.

[0371] Optionally, the video generation module 406 is further configured to:

[0372] Determine the video effect parameters selected by the user in the story style list in the story video generation interface;

[0373] The matched target story text, the initial image and / or the video, and the text voice are laid out on a timeline, and rendered using the video effect parameters to generate a target story video.

[0374] Optionally, the label obtaining module 402 is further configured to:

[0375] Receive the initial picture and / or the video selected by the user in the story video generation interface, and determine the target picture based on the initial picture and / or the video.

[0376] Optionally, the label obtaining module 402 is further configured to:

[0377] Determining the initial image as a target image;

[0378] Extracting video frames from the video according to a preset frame extraction rule and the duration of the video, and determining the extracted video frames as target images; or

[0379] Video frames are extracted from the video according to a preset frame extraction rule and the duration of the video, and the initial picture and the extracted video frames are determined as target pictures.

[0380] Optionally, the text generation module 404 is further configured to:

[0381] Determine the story style parameters selected by the user in the story style list of the story video generation interface.

[0382] Optionally, the device further includes:

[0383] The user modifies the label module and is configured as follows:

[0384] When it is determined that the picture tag is updated, the updated picture tag is determined, and an updated target story text is generated according to the story style parameter and the updated picture tag.

[0385] Optionally, the device further includes:

[0386] User-modified style modules are configured as follows:

[0387] When it is determined that the story style parameters are updated, the updated story style parameters are determined, and an updated target story text is generated according to the updated story style parameters and the picture tags.

[0388] The video generation device provided in the embodiment of the present application can automatically generate a story text only through the selected pictures and / or videos, assisted by an AI model that can realize picture processing and text processing, and generate the text speech corresponding to the story text using an AI model that can convert text to speech. It can automatically generate a personalized story video based on the story text, the text speech, pictures and / or videos corresponding to the story text, thereby lowering the threshold for users to make story videos. Moreover, since the AI ​​model is used as an auxiliary implementation in the above technical solution, it can not only improve the processing speed of the story video generation process, but also greatly guarantee the video quality of the generated target story video.

[0389] The above is a schematic diagram of a video generation device according to this embodiment. It should be noted that the technical solution of the video generation device and the technical solution of the above-mentioned video generation method are based on the same concept. For details not described in detail in the technical solution of the video generation device, please refer to the description of the technical solution of the above-mentioned video generation method.

[0390] Figure 5 The block diagram of a computing device 500 according to one embodiment of the present disclosure is shown. Components of the computing device 500 include, but are not limited to, a memory 510 and a processor 520. The processor 520 is connected to the memory 510 via a bus 530, and a database 550 is used to store data.

[0391] The computing device 500 also includes an access device 540 that enables the computing device 500 to communicate via one or more networks 560. Examples of such networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. The access device 540 may include one or more of any type of network interface (e.g., a network interface card (NIC)) whether wired or wireless, such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a universal serial bus (USB) interface, a cellular network interface, a Bluetooth interface, a near field communication (NFC) interface, and the like.

[0392] In one embodiment of the present specification, the above components of the computing device 500 and Figure 5 Other components not shown in the figure may also be connected to each other, for example, via a bus. Figure 5 The computing device structure block diagram shown is for illustrative purposes only and is not intended to limit the scope of this specification. Those skilled in the art may add or replace other components as needed.

[0393] Computing device 500 may be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, personal digital assistant, laptop computer, notebook computer, netbook computer, etc.), a mobile phone (e.g., a smartphone), a wearable computing device (e.g., a smartwatch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or personal computer (PC). Computing device 500 may also be a mobile or stationary server.

[0394] The processor 520 implements the steps of the video generation method when executing the computer instructions.

[0395] The above is a schematic diagram of a computing device according to this embodiment. It should be noted that the technical solution of the computing device and the technical solution of the above-mentioned video generation method are based on the same concept. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solution of the above-mentioned video generation method.

[0396] An embodiment of the present application further provides a computer-readable storage medium storing computer instructions, which, when executed by a processor, implement the steps of the video generation method as described above.

[0397] The above is a schematic diagram of a computer-readable storage medium of this embodiment. It should be noted that the technical solution of the storage medium and the technical solution of the video generation method described above are based on the same concept. For details not described in detail in the technical solution of the storage medium, please refer to the description of the technical solution of the video generation method described above.

[0398] The foregoing description describes specific embodiments of the present application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0399] The computer instructions include computer program code, which may be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal, and software distribution medium. It should be noted that the content contained in the computer-readable medium may be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media does not include electric carrier signals and telecommunication signals.

[0400] It should be noted that for the aforementioned method embodiments, for ease of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this application.

[0401] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0402] The preferred embodiments of the present application disclosed above are intended only to help illustrate the present application. The optional embodiments do not describe all details in detail, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made based on the content of this application. This application selects and describes these embodiments in detail in order to better explain the principles and practical applications of this application, so that those skilled in the art can better understand and utilize this application. This application is limited only by the claims and their full scope and equivalents.

Claims

1. A video generation method, characterized in that: Applicable to video generation clients, including: Determining an initial image and / or video, and determining image tags for the initial image and / or video; Determining story style parameters, and generating target story text according to the story style parameters and the image labels; Determining the text and speech of the target story text, and generating a target story video based on the target story text, the text and speech, the initial picture and / or the video; The step of generating a target story text according to the story style parameters and the image tags includes: Generate target story text using a text processing model according to the story style parameters and the image labels; Accordingly, determining the text-to-speech of the target story text includes: According to the text-to-speech model, the text-to-speech of the target story text is obtained, Among them, the text processing model and the text-to-speech model are both AI models.

2. The video generation method according to claim 1, wherein: The determining of the image tag of the initial image and / or the video includes: Determining a target image based on the determined initial image and / or the video, and obtaining an image label of the target image based on an image processing model; The image label of the target image is determined to be the image label of the initial image and / or the video, wherein the image processing model is an AI model.

3. The video generation method according to claim 2, characterized in that The obtaining of the image label of the target image according to the image processing model includes: Calling an image processing model deployed on the video generation server to perform image processing on the target image to obtain an image label of the target image; or Generate a tag acquisition request, wherein the tag acquisition request carries the target image; Sending the label acquisition request to the video generation server, so that the video generation server responds to the label acquisition request, calls the image processing model deployed on the model deployment server, performs image processing on the target image, and obtains the image label of the target image; Receive the image tag of the target image returned by the video generation server.

4. The video generation method according to claim 1, wherein: Before determining the story style parameters and generating the target story text according to the story style parameters and the picture tags, the method further includes: Determining a preset playback duration of the initial image and / or a video playback duration of the video, and the number of texts played per second; Determining the text length of the target story text according to the preset playback duration, the video playback duration, and the number of texts played per second; Accordingly, generating a target story text using a text processing model according to the story style parameters and the image tags includes: A target story text is generated using a text processing model according to the story style parameters, the image tags, and the text length.

5. The video generation method according to claim 4, characterized in that: After determining the text length of the target story text, the method further includes: determining a maximum number of texts per sentence of the target story text; Determining the number of sentences in the target story text according to the maximum number of texts and the text length; Accordingly, generating a target story text using a text processing model according to the story style parameters, the image tags, and the text length includes: A target story text is generated using a text processing model according to the story style parameters, the image labels, the text length, and the number of sentences.

6. The video generation method according to any one of claims 1, 4 or 5, characterized in that: Generating a target story text using a text processing model according to the story style parameters and the image tags includes: Call the text processing model deployed on the video generation server to analyze the story style parameters and the image tags Processing to generate target story text; or Generate a text acquisition request, wherein the text acquisition request carries the story style parameter and the image tag; Sending the text acquisition request to the video generation server, so that the video generation server responds to the text acquisition request, calls the text processing model deployed on the model deployment server, processes the story style parameters and the image tags, and generates the target story text; Receive the target story text returned by the video generation server.

7. The video generation method according to claim 1, characterized in that: After generating the target story text using the text processing model, the method further includes: Sentence segmentation is performed on the target story text according to the sentence segmentation model to obtain the target story text after sentence segmentation.

8. The video generation method according to claim 7, characterized in that: The sentence segmentation of the target story text according to the sentence segmentation model to obtain the sentence-divided target story text includes: Calling a sentence segmentation model deployed on the video generation server to perform sentence segmentation on the target story text to obtain a sentence-divided target story text; or generating a sentence segmentation request, wherein the sentence segmentation request carries the target story text; Sending the sentence segmentation request to the video generation server, so that the video generation server responds to the sentence segmentation request, calls the sentence segmentation model deployed on the model deployment server, performs sentence segmentation on the target story text, and obtains the target story text after sentence segmentation; Receive the target story text after the sentence segmentation returned by the video generation server.

9. The video generation method according to claim 7, characterized in that: The step of obtaining the text-to-speech of the target story text according to the text-to-speech model includes: According to the text-to-speech model, the text-to-speech of the target story text after the sentence segmentation is obtained, Accordingly, generating a target story video according to the target story text, the text voice, the initial picture and / or the video includes: A target story video is generated based on the sentence-segmented target story text, the text voice, the initial picture and / or the video.

10. The video generation method according to claim 1 or 9, characterized in that: The step of obtaining the text-to-speech of the target story text according to the text-to-speech model includes: Calling a text-to-speech model deployed on the video generation server to perform speech processing on the target story text according to speech generation rules to obtain the text-to-speech of the target story text; or Generate a voice acquisition request, wherein the voice acquisition request carries the target story text and preset voice generation rules; Sending the speech acquisition request to the video generation server, so that the video generation server responds to the speech acquisition request, calls the text-to-speech model deployed on the model deployment server, performs speech processing on the target story text according to the speech generation rule, and obtains the text speech of the target story text; Receive the text speech of the target story text returned by the video generation server.

11. The video generation method according to claim 1, wherein: Before generating the target story text using the text processing model according to the story style parameters and the image tags, the method further includes: Performing sensitive word detection on the story style parameters and the image labels according to the risk control model; Accordingly, after generating the target story text using the text processing model according to the story style parameters and the image tags, the method further includes: Sensitive word detection is performed on the target story text according to the risk control model.

12. The video generation method according to claim 1, wherein: Generating a target story video according to the target story text, the text voice, the initial picture and / or the video includes: Matching the target story text with the initial image and / or the video according to a preset matching rule; A target story video is generated based on the matched target story text, the initial picture and / or the video, and the text voice.

13. The video generation method according to claim 12, characterized in that: Matching the target story text with the initial image and / or the video according to a preset matching rule includes: Matching the target story text with the initial picture according to picture matching rules; Matching the target story text with the video according to a video matching rule; According to a mixed matching rule, the target story text is matched with the initial picture and the video.

14. The video generation method according to claim 13, characterized in that: Matching the target story text with the initial picture according to the picture matching rule includes: Determining an image tag of the initial image; When it is determined that the target story text contains a keyword corresponding to the image tag of the initial image, matching the target story text with the initial image according to a first image matching rule; When it is determined that the keyword corresponding to the image tag of the initial image does not exist in the target story text, the target story text is matched with the initial image according to a second image matching rule.

15. The video generation method according to claim 13, wherein matching the target story text with the video according to a video matching rule comprises: When it is determined that the playing time of the target story text is greater than the playing time of the video, matching the target story text with the video according to a first video matching rule; When it is determined that the playing time of the target story text is equal to the playing time of the video, matching the target story text with the video according to a second video matching rule; When it is determined that the playing time of the target story text is less than the playing time of the video, the target story text is matched with the video according to a third video matching rule.

16. The video generation method according to claim 13, wherein matching the target story text with the initial image and the video according to a mixed matching rule comprises: Determining an image tag of the initial image; When it is determined that the target story text contains a keyword corresponding to the image tag of the initial image, matching the target story text with the initial image and the video according to a first mixed matching rule; When it is determined that the target story text does not contain the keyword corresponding to the image tag of the initial image, the target story text is matched with the initial image and the video according to a second hybrid matching rule.

17. The video generation method according to claim 12, wherein generating a target story video based on the matched target story text, the initial image and / or the video, and the text speech comprises: Determine the video effect parameters selected by the user in the story style list in the story video generation interface; The matched target story text, the initial image and / or the video, and the text voice are laid out on a timeline, and rendered using the video effect parameters to generate a target story video.

18. The video generation method according to claim 2, wherein determining the target image based on the determined initial image and / or the video comprises: Receive the initial picture and / or the video selected by the user in the story video generation interface, and determine the target picture based on the initial picture and / or the video.

19. The video generation method according to claim 18, characterized in that: The determining of the target image according to the initial image and / or the video includes: determining the initial picture as the target picture; or Extracting video frames from the video according to a preset frame extraction rule and the duration of the video, and determining the extracted video frames as target images; or Video frames are extracted from the video according to a preset frame extraction rule and the duration of the video, and the initial picture and the extracted video frames are determined as target pictures.

20. The video generation method according to claim 1, wherein: Determining the story style parameters includes: Determine the story style parameters selected by the user in the story style list of the story video generation interface.

21. The video generation method according to claim 1, wherein: After generating the target story text, the method further includes: When it is determined that the picture tag is updated, the updated picture tag is determined, and an updated target story text is generated according to the story style parameter and the updated picture tag.

22. The video generation method according to claim 1, characterized in that: After generating the target story text, the method further includes: When it is determined that the story style parameters are updated, the updated story style parameters are determined, and an updated target story text is generated according to the updated story style parameters and the picture tags.

23. A video generation method, characterized in that: Applied to the video generation server, including: Receiving a target image sent by a video generation client, and obtaining an image tag of the target image according to an image processing model, wherein the target image is obtained by the video generation client according to the determined initial image and / or video; Generate target story text using a text processing model based on the image tags and the story style parameters sent by the received video generation client; According to the text-to-speech model, the text-to-speech of the target story text is obtained, and the text-to-speech is returned to the video generation client, so that the video generation client can generate the target story text, the text-to-speech, ... The initial image and / or the video are used to generate a target story video; The video generation client generates the target story text according to the story style parameters and the image tags, including: Generate a target story text using the text processing model according to the story style parameters and the image labels; Accordingly, determining the text-to-speech of the target story text includes: According to the text-to-speech model, the text-to-speech of the target story text is obtained, Among them, the text processing model and the text-to-speech model are both AI models.

24. A video generating device, characterized in that: Applicable to video generation clients, including: a label obtaining module, configured to determine an initial image and / or video, and determine an image label of the initial image and / or video; a text generation module configured to determine story style parameters and generate a target story text according to the story style parameters and the image tags; a video generation module configured to determine the text and speech of the target story text, and generate a target story video based on the target story text, the text and speech, the initial picture and / or the video; The step of generating a target story text according to the story style parameters and the image tags includes: Generate target story text using a text processing model according to the story style parameters and the image labels; Accordingly, determining the text-to-speech of the target story text includes: According to the text-to-speech model, the text-to-speech of the target story text is obtained, Among them, the text processing model and the text-to-speech model are both AI models.

25. A computing device comprising a memory, a processor, and computer instructions stored in the memory and executable on the processor, wherein: When the processor executes the computer instructions, the steps of the video generation method according to any one of claims 1 to 23 are implemented.

26. A computer-readable storage medium storing computer instructions, characterized in that: When the computer instructions are executed by a processor, the steps of the video generation method described in any one of claims 1 to 23 are implemented.

27. A computer program product comprising computer instructions, characterized in that: When the computer instructions are executed by a processor, the steps of the method described in any one of claims 1 to 23 are implemented.

Citation Information

Patent Citations

  • Picture copywriting generation method and device and storage medium

    CN116303972A

  • Video processing method and electronic equipment

    CN116668768A