System for assisting AI video production

By assisting the AI ​​video production system, terminals and servers are used to generate and splice finished video products, which solves the problem of being unable to generate complete VLOG videos in existing technologies and achieves the effect of quickly generating high-quality video products.

CN120640100APending Publication Date: 2025-09-12NANJING AIZHAO FEIDA IMAGING TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510987476.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-17
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing AI video generation platforms are unable to generate complete VLOG video products that can be directly published. Users need to post-produce them themselves, which cannot meet user needs.

Method used

An auxiliary AI video production system is provided. The product parameters are configured through the first terminal, the second terminal imports photos and generates synthesis instructions, and the AI ​​video generation and synthesis server is used to generate and splice the finished video, including pre-shot opening, ending and transition videos, and supports generating multiple videos from one photo.

Benefits of technology

It realizes the automatic generation of complete VLOG videos that include the characteristics of scenic spots and the images of tourists, eliminating the actual shooting steps, quickly delivering high-quality video products, and improving user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120640100A_ABST
    Figure CN120640100A_ABST
Patent Text Reader

Abstract

The invention discloses a system for assisting AI video production. The system comprises a first terminal, a second terminal, an AI video generation server and an AI video synthesis server, the first terminal is used for a user to carry out product parameter configuration according to characteristics and requirements of different scenes to generate a plurality of products; the second terminal is used for importing or uploading a plurality of character photos and correspondingly selecting a required product, then generating a video generation synthesis instruction, and sending the video generation synthesis instruction to the AI video synthesis server; the AI video synthesis server receives a video generation synthesis instruction sent by the second terminal, analyzes out character photos, video parameters and prompt words and sends the character photos, the video parameters and the prompt words to an AI video generation server, and a plurality of AI videos are generated; and the AI video synthesis server splices the plurality of AI videos according to the video template and the video parameters so as to synthesize a video finished product set according to the product. According to the invention, the actual shot video can be omitted, and the video content can be creatively enabled, so that the purpose of quick delivery is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video production, and in particular to a system for assisting AI video production. Background Art

[0002] Currently, major AI video generation service platforms offer the following video generation capabilities: 1. Text-based video: Using a text prompt, the AI ​​engine and a large model generate a video. 2. Image-based video: Using one (or more) photos and a prompt, the AI ​​engine and a large model generate a video. The videos generated in these ways are generally limited to 5-10 seconds in length. Some platforms offer AI sound generation, such as the sound of gurgling water, chirping insects, and birdsong, based on the AI-generated video content. AI can also generate music based on the prompt. However, AI videos and AI sounds generated in these ways are not yet considered ready-to-deliver products. Generally, users desire a complete video product, a complete vlog video that can be published immediately on private platforms without the need for post-production. Summary of the Invention

[0003] The purpose of this invention is to address the shortcomings of the existing technology and provide a system that assists in the automatic post-production of AI videos, so that a complete vlog video containing scenic spot feature clips and multiple video clips containing the tourists' own images, music and other content can be automatically generated with just one photo.

[0004] To achieve the above-mentioned object, the present invention provides a system for assisting AI video production, comprising a first terminal, a second terminal, an AI video generation server, and an AI video synthesis server; The first terminal allows professional users on the B side to configure product parameters according to the characteristics and needs of different scenarios to generate multiple products and generate product names and product descriptions. The product parameter configuration includes setting video parameters, editing multiple different prompt words, and selecting the desired video template name; The second terminal is used to obtain the products configured by the first terminal and is used by general C-end users to import or upload several photos of people in the current scene and select the required products based on the product name and product description, and then generate a video generation and synthesis instruction, and send the video generation and synthesis instruction to the AI ​​video synthesis server; The AI ​​video synthesis server receives the video generation synthesis instruction sent by the second terminal, and parses the video generation synthesis instruction to obtain the video parameters, prompt words and video template names corresponding to the character photo and the selected product, and sends the video parameters and prompt words corresponding to the character photo and the selected product to the AI ​​video generation server, and extracts the video template according to the video template name for standby; The AI ​​video generation server receives the character photos, video parameters corresponding to the selected product, and prompt words sent by the AI ​​video synthesis server, and generates a plurality of AI videos according to the character photos, video parameters corresponding to the selected product, and prompt words; The AI ​​video synthesis server splices several AI videos according to the pre-shot video content fragments of the extracted video template based on the video template and video parameters to synthesize a finished video, and delivers the synthesized finished video as a product to the user for preview and saving.

[0005] Furthermore, the aforementioned AI videos can extract photo frames by means of frame extraction and deliver them to C-end users as photos with replaced backgrounds.

[0006] Furthermore, the person's photo is obtained by scanning a QR code on a ticket containing a photo link through a second terminal or by obtaining it from a cloud server through face recognition in a mini program.

[0007] Furthermore, the finished video includes a pre-shot opening and a pre-shot ending, the several AI videos are arranged between the pre-shot opening and the pre-shot ending, and a pre-shot transition video is provided between two adjacent AI videos.

[0008] Furthermore, the second terminal is also used for general users to pay for the fees required for video synthesis, and after successful payment, it generates video synthesis instructions based on the character photos uploaded by the general users on the C end and the selected required products.

[0009] Furthermore, the second terminal is also used to generate a status query instruction and send the status query instruction to the AI ​​video synthesis server. After receiving the status query instruction, the AI ​​video synthesis server returns the current status of the AI ​​video generation server and the AI ​​video synthesis server to the second terminal.

[0010] Furthermore, the second terminal is also used to edit the extended prompt words so that the AI ​​video generation server generates several AI videos according to the prompt words and the extended prompt words.

[0011] Furthermore, the prompt words include positive prompt words and negative prompt words.

[0012] Furthermore, the video template is generated via a video template editor and stored in an AI video synthesis server with different template names, and is provided to the first terminal user through different template names for selection when setting.

[0013] Furthermore, the product parameters are generated by the product editor of the first terminal, recorded in the first terminal with different product names and corresponding product descriptions, and provided to the first terminal user for selection through several product options.

[0014] Furthermore, if the size, resolution, and frame rate of the video template are different from those of the AI ​​video generated by the AI ​​video generation server, the AI ​​video synthesis server automatically adjusts the parameters of the AI-generated video to meet the specifications of the video template.

[0015] Beneficial effects: The present invention allows professional users to configure product parameters on the first terminal, and general users use the second terminal to scan codes and compare faces to find photos that have been pre-taken and saved in the system by professional users on the B side, or upload photos by themselves. Other parameters are confirmed through product selection, and then several AI video clips are generated using AI technology. The video module configured by the AI ​​video synthesis server is called by the template set by the first terminal and then selected by the second terminal to synthesize the finished video. This saves the need for actual video shooting and can better empower the creativity of video content, achieving the purpose of rapid delivery. Through template management, different templates can be provided to different professional users on the B side, thereby serving more general users to have a better experience when traveling. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 2 is a schematic diagram of a system for assisting AI video production according to an embodiment of the present invention. DETAILED DESCRIPTION

[0017] The present invention will be further illustrated below with reference to the accompanying drawings and specific embodiments. These embodiments are implemented based on the technical solutions of the present invention. It should be understood that these embodiments are only used to illustrate the present invention and are not used to limit the scope of the present invention.

[0018] like Figure 1 As shown, an embodiment of the present invention provides a system for assisting AI video production, including a first terminal 1, a second terminal 2, an AI video generation server 3 and an AI video synthesis server 4.

[0019] First Terminal 1 is a computer or mobile phone used by professional business users, such as photography vendors, within the scenic area. Each business user is assigned a unique user ID. For example, 001715014 represents a photography vendor user in Heishantou. Users can use the "Product Editor" on First Terminal 1 to configure product parameters based on the characteristics and needs of different scenarios, generating multiple products and assigning corresponding product names and descriptions. Different product names and corresponding product descriptions are recorded on First Terminal 1 and presented to users on Second Terminal 2 through a selection of product options. For example, the product name is "30-second composite video of two segments of a single image" and the description is "Requires a vertical photo and two sets of prompt words." Parameter configuration includes setting video parameters, editing several different prompt words, selecting the desired video template name (e.g., "001715015_Summer vertical"), and selling price. Video parameters include image processing method, size, resolution, frame rate, horizontal and vertical orientation, etc., as follows: Image processing method: Parameters are configured according to the scene through the first terminal 1. There are multiple combinations of the relationship between the image and the prompt words required when calling the AI ​​video generation server 3 through the AI ​​video synthesis server 4 and the final generated video clip. For example: single-image single-segment video (each photo corresponds to a set of positive and negative prompt words, and AI can generate a video), single-image multiple-segment video (this same photo corresponds to multiple sets of positive and negative prompt words multiple times, and each time AI can generate a different video), first and last frame video (a video is generated for every two photos), etc. The first terminal 1 will pre-set the template format of each product parameter to determine how many AI videos are required and how many pictures are required for different photo processing methods to pre-set multiple products, and match the number of video clips required for the video template of the product finally selected when used by the user of the second terminal 2. After the data is returned to each video from the AI ​​video generation server 4 through interaction between servers, the final video can be generated after the AI ​​video synthesis server 3 is perfectly embedded in the template.

[0020] How video templates are processed: Video templates are generated via the "Video Template Editor" and recorded and stored in the AI ​​video synthesis server 3 with different template names. Different template names are provided to B-side professional users for selection when setting up. The template editing module is placed on the AI ​​video synthesis server 3. Each scene will pre-edit the template and take a "template name". The template name is generally set to be distinguished by the user number and Chinese characters. For example: 001715015_Summer vertical version of the product name and description is "Vertical version single picture two-segment 30-second composite video", which means that one photo and two sets of prompt words are required to generate two AI videos. Video template editing and structure are all existing technologies, including pre-shot opening, pre-shot ending, transitions and background music, etc., which will not be repeated here. Actual case description: According to the settings of the "Product Editor" of the first terminal 1, in addition to the preset opening, ending, music and other elements, a single-image, multi-video segment is required. The case is a single image with two videos. The general user on the C-end will see a picture found from the system at the time of this login on the interface of the second terminal 2. In addition, two sets of prompt words preset in the first terminal are displayed. Without changing any settings, just press OK to generate two AI video segments from the same photo, and finally synthesize the following combined video: pre-shot opening + AI video 1 + pre-shot transition video + AI video 2 + pre-shot ending. To increase user flexibility, the second terminal 2 will prompt on the interface that the photo can be replaced by other photos imported from the local device, and the two sets of prompt words can also be spliced ​​or replaced with extended prompt words.

[0021] Size: The size of the generated video, generally 9:16, 16:9, 4:3, 3:4, 2:3, 3:2, 1:1, etc., and is matched with the resolution dpi pixel parameter, such as 1080, 720, etc. The video template size and resolution are preferably exactly the same as the AI-generated video size and resolution to obtain the final high-quality video. If they are different, the AI ​​video synthesis server 3 will automatically adjust the length, width, and resolution of the AI-generated video to match the specifications of the video template.

[0022] The above prompt words include positive prompt words and negative prompt words. Positive prompt words describe what the AI ​​video you want to generate looks like, and negative prompt words describe what the AI ​​video you do not want to generate looks like.

[0023] Second Terminal 2 can be a mobile phone for a general C-end user, such as a tourist taking photos at a scenic spot. A general C-end user (tourist) can use system-provided methods to find photos stored in the system and import or upload several photos of people in the current scene into Second Terminal 2. Based on the imported or uploaded photos, and according to the product name and description, they select a product pre-configured on First Terminal 1, such as a "30-second composite video of two segments of a single image." The parameters contained in the product are combined into a video generation and synthesis instruction, which is transmitted to AI video synthesis server 3. If payment is required, payment is made at this time. Finally, a video generation and synthesis instruction is generated and sent to AI video synthesis server 3.

[0024] The AI ​​video synthesis server 3 receives the video generation synthesis instruction sent by the second terminal 2, and parses the received video generation synthesis instruction to obtain the character photos uploaded by the tourist, the video parameters, prompt words and video template names corresponding to the selected products, and sends the character photos, video parameters and prompt words corresponding to the selected products to the AI ​​video generation server 4, and extracts the video template (for example: "001715015_Summer Vertical Edition") at the same time, and waits for the AI ​​video generation server 3 to return the result.

[0025] The AI ​​video generation server 4 receives the video parameters and prompt words corresponding to the product selected by the person photo sent by the AI ​​video synthesis server 3, and generates several AI videos based on the received person photo, the video parameters and prompt words corresponding to the selected product. The AI ​​video generation server 4 can be a third-party AI server. In this case, the person photo can be used to generate an AI video through an API call. The specific generation methods of AI videos include single-image single-segment video (multiple different single images can generate multiple single-image single-segment videos), single-image multiple-segment video (the same image corresponds to multiple sets of different prompt words, generating multiple corresponding AI videos) and first and last frame video (every two photos are used as the first and last frames to generate a video). Finally, according to the settings of the template, each AI-generated video corresponds to a different position in the video template.

[0026] After receiving the AI ​​video generated by the AI ​​generation server 4, the AI ​​synthesis server 3 can also extract several photo frames by means of frame extraction and deliver them to the user as AI photo products to increase the user's product options.

[0027] The AI ​​video synthesis server 3 splices several AI videos according to the pre-shot video content segments of the extracted video template based on the video template and video parameters to synthesize a finished video, and delivers the synthesized video finished product as a product to the user for preview and storage. In the finished video, the AI ​​video is set between the pre-shot opening and pre-shot ending, and a pre-shot transition video is set between two adjacent AI videos to enrich the creativity and content. Taking the AI ​​video generation server 4 as an example, the specific structure of the finished video is as follows: pre-shot opening + AI video 1 + pre-shot transition video 1 + AI video 2 + pre-shot transition video 2 + AI video 3 + pre-shot ending.

[0028] The above-mentioned photos of tourists are preferably taken by photographers of photography businesses in the scenic area at designated locations within the scenic area and delivered through a cloud server. This delivery method is a prior art. Specifically, tourists can obtain their photos stored in the cloud server by scanning the QR code on the receipt containing the photo link through the second terminal 2 or by facial recognition in the mini program. After obtaining their own photos, tourists can choose to save the photos to the second terminal 2 or choose to further create an AI video. After clicking on the "Create AI Video" button, an AI video is created using the currently acquired photos. When obtaining a tourist's photo by scanning the QR code on the receipt containing the photo link or by facial recognition in the mini program, the first terminal 1 can determine which photography business took the photo by the photo's file name (including the photography business's user number), and then automatically bind to the photography business's first terminal 1.

[0029] The second terminal 2 is also used for tourists to pay the fees required for video synthesis, and after successful payment, it generates video synthesis instructions based on the person photos uploaded by the tourists and the selected required products, thereby realizing paid services.

[0030] The second terminal 2 is also used to generate a status query instruction and send the status query instruction to the AI ​​video synthesis server 3. After receiving the status query instruction, the AI ​​video synthesis server 3 returns the current status to the second terminal 2. The above status includes success, failure, processing and waiting time.

[0031] Visitors can also use the second terminal 2 to edit extended prompts, allowing the AI ​​video generation server 4 to generate multiple AI videos based on the prompts and extended prompts. Typically, the extended prompts are merged after the prompts set by the first terminal 1 to influence the generated AI video. To allow for more flexible video generation for general users, the extended prompts can also completely replace the prompts set by the first terminal 1.

[0032] The application scenarios of this application include but are not limited to: AI generates a smooth drifting process and an exciting drifting process from a drifting photo, and the two videos are then synthesized into a finished video through a video template.

[0033] A skiing photo generates a smooth skiing process + a skiing process that stirs up snowflakes. The two videos are then synthesized into a finished video through a video template.

[0034] A photo of horse riding generates a surround shooting effect and a galloping horse effect, and the two videos are then synthesized into a finished video through a video template.

[0035] The above description is merely a preferred embodiment of the present invention. It should be noted that any other aspects not specifically described are considered prior art or common knowledge to those skilled in the art. Improvements and modifications may be made without departing from the principles of the present invention, and such improvements and modifications are also within the scope of protection of the present invention.

Claims

1. A system for assisting AI video production, characterized in that: It includes a first terminal, a second terminal, an AI video generation server, and an AI video synthesis server; The first terminal allows professional users on the B side to configure product parameters according to the characteristics and needs of different scenarios to generate multiple products and generate product names and product descriptions. The product parameter configuration includes setting video parameters, editing multiple different prompt words, and selecting the desired video template name; The second terminal is used to obtain the products configured by the first terminal and is used by general C-end users to import or upload several photos of people in the current scene and select the required products based on the product name and product description, and then generate a video generation and synthesis instruction, and send the video generation and synthesis instruction to the AI ​​video synthesis server; The AI ​​video synthesis server receives the video generation synthesis instruction sent by the second terminal, and parses the video generation synthesis instruction to obtain the video parameters, prompt words and video template names corresponding to the character photo and the selected product, and sends the video parameters and prompt words corresponding to the character photo and the selected product to the AI ​​video generation server, and extracts the video template according to the video template name for standby; The AI ​​video generation server receives the character photos, video parameters corresponding to the selected product, and prompt words sent by the AI ​​video synthesis server, and generates a plurality of AI videos according to the character photos, video parameters corresponding to the selected product, and prompt words; The AI ​​video synthesis server splices several AI videos according to the pre-shot video content fragments of the extracted video template based on the video template and video parameters to synthesize a finished video, and delivers the synthesized finished video as a product to the user for preview and saving.

2. The system for assisting AI video production according to claim 1, characterized in that: The person's photo is obtained by scanning the QR code on the ticket containing the photo link through the second terminal or by obtaining it from the cloud server through face recognition in the mini program.

3. The system for assisting AI video production according to claim 1, characterized in that: The finished video includes a pre-shot opening and a pre-shot ending, the several AI videos are arranged between the pre-shot opening and the pre-shot ending, and a pre-shot transition video is provided between two adjacent AI videos.

4. The system for assisting AI video production according to claim 1, characterized in that: The second terminal is also used for general users to pay the fees required for video synthesis, and after successful payment, it generates video synthesis instructions based on the character photos uploaded by the general users on the C end and the selected required products.

5. The system for assisting AI video production according to claim 1, characterized in that: The second terminal is also used to generate a status query instruction and send the status query instruction to the AI ​​video synthesis server. After receiving the status query instruction, the AI ​​video synthesis server returns the current status of the AI ​​video generation server and the AI ​​video synthesis server to the second terminal.

6. The system for assisting AI video production according to claim 1, characterized in that: The second terminal is further used to edit the extended prompt words so that the AI ​​video generation server generates a plurality of AI videos according to the prompt words and the extended prompt words.

7. The system for assisting AI video production according to claim 1, characterized in that: The prompt words include positive prompt words and negative prompt words.

8. The system for assisting AI video production according to claim 1, characterized in that: The video template is generated via a video template editor and stored in an AI video synthesis server with different template names, and is provided to the first terminal user through different template names for selection when setting.

9. The system for assisting AI video production according to claim 1, characterized in that: The product parameters are generated by the product editor of the first terminal, recorded in the first terminal with different product names and corresponding product descriptions, and provided to the second terminal user for selection through several product options.

10. The system for assisting AI video production according to claim 1, characterized in that: If the size, resolution, and frame rate of the video template are different from those of the AI ​​video generated by the AI ​​video generation server, the AI ​​video synthesis server automatically adjusts the parameters of the AI-generated video to meet the specifications of the video template.