Video creation content generation method and device, equipment, storage medium and product

By calling the content generation model to generate video creation information, the problems of low efficiency and uncertainty in the generation of video creation information are solved, and efficient and reliable video creation information generation is achieved.

CN121547658APending Publication Date: 2026-02-17BEIJING YOUZHUJU NETWORK TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511713904.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-20
Publication Date
2026-02-17

AI Technical Summary

Technical Problem

In existing technologies, the generation efficiency of video creation information is low, the consistency of results is poor, and the credibility is uncertain, mainly due to excessive reliance on manual processing.

Method used

By responding to input operations, the content generation model is invoked to generate video creation information, including video creation dimensions, elements, and credibility indicators, thereby reducing human intervention and subjectivity.

Benefits of technology

It improves the efficiency and consistency of video creation information generation, provides credibility indicators, and enhances the reliability of the generated content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121547658A_ABST
    Figure CN121547658A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to a video creation content generation method and device, equipment, a storage medium and a product. The method comprises the following steps: in response to a first input operation acting on an input box, determining first input content corresponding to the first input operation; calling a content generation model based on the first input content, and determining and displaying video creation information corresponding to the first input content; the video creation information comprises at least one video creation dimension, at least one video creation element corresponding to the video creation dimension, a credibility index corresponding to the video creation element, and at least one execution content associated with the video creation element and used for guiding video creation. Therefore, the generation efficiency, consistency and credibility of the video creation information are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of data processing technology, and in particular to a method, apparatus, device, storage medium, and product for generating video content. Background Technology

[0002] With the rapid development of video technology, many businesses promote their products through videos, which can be called promotional videos (such as advertising videos). The effectiveness of promotional videos is closely related to the quality of the video's creative concept, which in turn depends on the inspiration behind the video (also known as video creation information).

[0003] Currently, the main method for obtaining video creation information is through keyword searches to acquire a small number of promotional videos. Then, either manually or semi-automatically using simple video analysis tools, the selling points / highlights and video presentation techniques involved in these promotional videos are summarized as video creation information. However, this method relies excessively on manual labor, resulting in low efficiency in generating video creation information, poor consistency of results, and significant uncertainty in reliability. Summary of the Invention

[0004] To address the aforementioned technical problems, this disclosure provides a method, apparatus, device, storage medium, and product for generating video content.

[0005] In a first aspect, embodiments of this disclosure provide a method for generating video content, the method comprising: In response to a first input operation applied to the input box, determine the first input content corresponding to the first input operation; Based on the first input content, the content generation model is invoked to determine and display the video creation information corresponding to the first input content; wherein, the video creation information includes at least one video creation dimension, at least one video creation element corresponding to the video creation dimension, a credibility index corresponding to the video creation element, and at least one execution content associated with the video creation element for guiding video creation.

[0006] Secondly, this disclosure also provides a video content creation generation apparatus, the apparatus comprising: The first input content determination module is used to respond to a first input operation applied to the input box and determine the first input content corresponding to the first input operation. The video creation information determination module is used to, based on the first input content, invoke a content generation model to determine and display the video creation information corresponding to the first input content; wherein, the video creation information includes at least one video creation dimension, at least one video creation element corresponding to the video creation dimension, a credibility index corresponding to the video creation element, and at least one execution content associated with the video creation element for guiding video creation.

[0007] Thirdly, embodiments of this disclosure also provide an electronic device, the electronic device comprising: processor; Memory, used to store executable instructions; The processor is used to read executable instructions from memory and execute the executable instructions to implement the video creation content generation method described in any embodiment of this disclosure.

[0008] Fourthly, embodiments of this disclosure also provide a computer-readable storage medium storing a computer program that, when executed by a processor, causes the processor to implement the video creation content generation method described in any embodiment of this disclosure.

[0009] Fifthly, embodiments of this disclosure also provide a computer program product for executing the video creation content generation method described in any embodiment of this disclosure.

[0010] The video creation content generation method, apparatus, device, storage medium, and product of this disclosure are capable of responding to a first input operation applied to an input box, determining the first input content corresponding to the first input operation; based on the first input content, invoking a content generation model, determining and displaying video creation information corresponding to the first input content; the video creation information includes at least one video creation dimension, at least one video creation element corresponding to the video creation dimension, a credibility index corresponding to the video creation element, and at least one execution content associated with the video creation element for guiding video creation; it realizes the automatic generation of required video creation information using a generative model of a certain scale, greatly reducing human intervention and subjectivity, thereby improving the generation efficiency and consistency of video creation information; furthermore, the generated video creation information includes a credibility index corresponding to each video creation element, providing a quantitative index value for the reliability of the video creation information, thereby improving the credibility of the video creation information.

[0011] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in the embodiments of this disclosure are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse. Attached Figure Description

[0012] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and elements are not necessarily drawn to scale.

[0013] Figure 1 A flowchart illustrating a video content creation method provided in this embodiment of the disclosure; Figure 2 A schematic diagram showing a dialog interface for video content creation provided in an embodiment of this disclosure; Figure 3 for Figure 1 The diagram shows a detailed process flow diagram of S120 in the video creation content generation method. Figure 4 This is a schematic diagram illustrating the process of generating video creation information according to an embodiment of the present disclosure; Figure 5 This is a schematic diagram of the structure of a video content creation device provided in an embodiment of the present disclosure; Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation

[0014] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0015] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.

[0016] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.

[0017] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.

[0018] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0019] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0020] The video content creation method provided in this disclosure is applicable to video creation-related scenarios. This method can be executed by a video content creation device, which can be implemented in software and / or hardware. The device can be integrated into an electronic device with display functionality, human-computer interaction functionality, and a certain data processing capability. This electronic device may include, but is not limited to, smartphones, personal digital assistants (PDAs), tablet computers (Tablet PCs), laptops, desktop computers, or servers.

[0021] Figure 1 A flowchart illustrating a video content creation method provided in an embodiment of this disclosure is shown. Figure 1 As shown, the video content creation method may include the following steps: S110, Respond to the first input operation applied to the input box, and determine the first input content corresponding to the first input operation.

[0022] The first input operation is an interactive operation performed on the input box to trigger the generation request of video creation information.

[0023] Specifically, to simplify the user interaction process and improve the efficiency of video creation information generation, electronic devices can rely on a generative model of a certain scale (i.e., a content generation model) to provide users with a dialog interface that generates video-related information, and display input boxes in this dialog interface. When the user performs a first input operation through the input box, the electronic device can receive the initial input content corresponding to the first input operation, and obtain the first input content based on the initial input content. This first input content is used to trigger the content generation model to generate video creation information.

[0024] For example, if the first input operation is text input into the input box, the electronic device can determine the initial input content as the first input content. If the first input operation is voice input into the input box, the electronic device can process the initial input content into speech-to-text and determine the conversion result as the first input content. If the first input operation is uploading content from the input box, and the uploaded file contains text content, the electronic device can parse the initial input content into the file and determine the parsing result as the first input content. If the first input operation is uploading content from the input box, and the uploaded image, the electronic device can perform image recognition on the initial input content and determine the recognition result as the first input content.

[0025] S120. Based on the first input content, call the content generation model to determine and display the video creation information corresponding to the first input content; the video creation information includes at least one video creation dimension, at least one video creation element corresponding to the video creation dimension, a credibility index corresponding to the video creation element, and at least one execution content associated with the video creation element for guiding video creation.

[0026] Among them, video creation information refers to inspirational data for video creation, which guides the generation of videos with good creative quality. Video creation dimensions are the core classification direction of video creation information. They categorize the core aspects of video creation, such as content, form, and objectives, providing a clear framework for video creation and covering key aspects such as content presentation, image building, and audience adaptation. These dimensions can be set according to product characteristics or business requirements. Video creation elements are the specific constituent units under the video creation dimensions, the core details supporting the implementation of the dimensions. They have a clear direction, such as "food lovers" or "office workers" under the audience adaptation dimension (e.g., the target audience dimension). Execution content is the specific operation and presentation form that transforms the creative elements into the actual video. It is the practical content that can be directly implemented, such as specific execution details like delivery techniques, shooting methods, editing rhythm, and visual presentation. Credibility metrics are pre-defined indicators that characterize the overall credibility of video creation information. For example, they can be at least one of the following: overall click-through rate (CTR), overall conversion rate (CVR), overall consumption / cost ratio, average number of views, and completion rate in the first n seconds.

[0027] Specifically, the electronic device constructs model prompts based on the first input content, generates a model from the input content, processes the model, and outputs structured video creation information adapted to the first input content. The electronic device can then display this video creation information as feedback to the first input content in the dialog box.

[0028] It should be noted that if the electronic device has model execution capabilities, the content generation model can be deployed locally. This allows the electronic device to call the content generation model locally to perform the aforementioned video creation information generation process, directly generating the video creation information locally. If the electronic device does not have model execution capabilities, the content generation model can be deployed on a server and an API can be provided. In this case, the electronic device can use the content as the first input parameter to remotely call the content generation model to perform the aforementioned video creation information generation process, and the electronic device can receive the video creation information fed back from the server.

[0029] For example, see Figure 2The electronic device can display a dialog interface with video creation content generation capabilities to the user, and show an input box 210 at the bottom of the dialog interface. The user can input their needs through the input box 210, and the electronic device can display the first input content 220 obtained in the dialog interface. Then, the electronic device internally uses the first input content to call the content generation model to perform a series of processes, and finally outputs video creation information 230, such as an access URL or structured chart pointing to video creation information. Figure 2 (Specific details are not shown in the text).

[0030] For example, the video creation dimensions in this disclosure may include, for instance, a value presentation dimension (such as a selling point dimension), a presenter image dimension (such as a voice-over persona dimension), an opening attention guidance dimension (such as a captivating first n seconds dimension), an audience fit dimension (such as a target audience dimension), a video structure dimension, or a scene dimension, among other core video element dimensions. The aforementioned value presentation dimension focuses on presenting the core value of the video content, conveying the key advantages and unique value of things through concrete expression, allowing the audience to clearly perceive its core appeal, which is a key basis for driving user decisions and achieving video conversion. The aforementioned presenter image dimension revolves around shaping the image of the video content presenter, covering core elements such as language style, expression logic, and temperament, forming a stable, perceptible personalized characteristic for the audience. The aforementioned opening attention guidance dimension targets the content design direction of the key opening period of the video (such as 3 seconds or 5 seconds), optimizing elements such as scene, information, and rhythm to quickly capture the audience's attention and guide them to continue watching. The aforementioned audience adaptation dimensions are the positioning direction for video content adaptation / promotion targets. They focus on the cognitive habits, attention points, and receiving preferences of specific groups, providing accurate adaptation basis for video content creation.

[0031] Taking the keyword "chicken leg" as an example of the first input content, the following examples illustrate the video creation information. The value presentation dimension of this example's video creation information can include video creation elements such as taste and flavor, convenience, health / nutrition, and price / discounts. Each video creation element can include the delivery of the message and shooting techniques, as well as its corresponding credibility metrics (such as overall consumption ratio, overall CTR, overall CVR, and average views). See Table 1, which shows an example of the taste and flavor video creation element and its corresponding credibility metrics and delivery of the message. The opening attention-grabbing dimension (such as the first 3 seconds to grab attention) of this example's video creation information can include video creation elements such as effect demonstration, curiosity arousal, pain point attraction, scene immersion, and benefit highlighting. Each video creation element can include the delivery of the message and shooting techniques, as well as its corresponding credibility metrics (such as 3-second completion rate, overall consumption ratio, overall CTR, overall CVR, and average views). See Table 2, which shows an example of the effect demonstration video creation element and its corresponding credibility metrics and delivery of the effect demonstration. The example video creation information includes audience-adaptation dimensions (such as the target audience dimension), which can include video creation elements such as home cooks, late-night snackers, office workers, food enthusiasts, and fitness and weight loss enthusiasts. Each video creation element can include audience descriptions and pain points, as well as corresponding credibility metrics (such as 3-second completion rate, overall consumption ratio, overall CTR, overall CVR, and average views). See Table 3 for an example showing the video creation elements for home cooks and their corresponding credibility metrics and content. The example video creation information also includes the presenter image dimension (such as the voice-over persona dimension), which can include video creation elements for different personas such as factory owners / source manufacturers, domain experts, product recommendation influencers, and ordinary bloggers. Each video creation element can include persona descriptions and speaking techniques, as well as corresponding credibility metrics (such as 3-second completion rate, overall consumption ratio, overall CTR, overall CVR, and average views). See Table 4 for an example showing the video creation elements for factory owners / source manufacturers and their corresponding credibility metrics and content.

[0032] Table 1. Video creation information for value presentation dimensions (e.g., selling points).

[0033] Table 2: Video Creation Information Based on Opening Attention Guidance Dimensions (e.g., Captivating the First 3 Seconds)

[0034] Table 3. Video creation information based on audience fit dimensions (e.g., target audience).

[0035] Table 4. Video Creation Information Based on the Speaker Image Dimension (e.g., Voiceover Persona Dimension)

[0036] It should be noted that the video creation information mentioned above may also include video identifiers (such as video name, video number, etc.) corresponding to the video creation elements. These video identifiers can be the identifiers of the videos that participated in the analysis to obtain the execution content corresponding to the video creation elements, or they can be the identifiers of the videos that participated in the analysis to obtain the credibility index corresponding to the video creation elements.

[0037] In some embodiments, S120 includes: based on the first input content, invoking a content generation model to determine and display video retrieval results; based on the video retrieval results and the first input content, continuing to invoke the content generation model to determine and display video creation information corresponding to the first input content.

[0038] Specifically, the process of generating video creation information mainly includes two stages: one is retrieving relevant videos, and the other is analyzing the retrieved videos and summarizing video creation information. Both stages require a certain amount of time, so users will be in a continuous waiting state during the processing of the content generation model. To shorten the user's continuous waiting time and improve the user's interactive experience, electronic devices can decouple the video retrieval stage and the video creation information summarization stage, and provide feedback to the user on the execution results of each stage after completion. Based on this, the electronic device can first call the content generation model based on the first input content to trigger its video retrieval. After the video retrieval is completed, the content generation model can first output the video retrieval results, such as the number of videos retrieved.

[0039] See also Figure 2 The electronic device can output the execution content and results in real time in stages. For example, when the model starts running, it can output "Starting to search for 'chicken leg' related videos..." to inform the user of the current execution. Then, after the video search is completed, it can continue to output "xxxx videos found, time taken xx seconds, starting analysis...".

[0040] Next, the electronic device uses the retrieved videos and the initial input content to generate model prompts, which are then fed into the content generation model. This triggers the model to analyze and summarize the retrieved videos, outputting video creation information. See also... Figure 2 After the electronic device finishes video analysis, it can continue to output "Video search term 'chicken leg', analysis complete, time taken xx seconds + URL or chart of generated video creation information" in the dialog interface.

[0041] The video creation content generation method provided in this disclosure can respond to a first input operation applied to an input box, determine the first input content corresponding to the first input operation, and based on the first input content, call a content generation model to determine and display the video creation information corresponding to the first input content. The video creation information includes at least one video creation dimension, at least one video creation element corresponding to the video creation dimension, a credibility index corresponding to the video creation element, and at least one execution content associated with the video creation element for guiding video creation. This method realizes the automatic generation of required video creation information using a generative model of a certain scale, greatly reducing human intervention and subjectivity, thereby improving the generation efficiency and consistency of video creation information. Furthermore, the generated video creation information contains a credibility index corresponding to each video creation element, providing a quantitative index value for the reliability of the video creation information, thereby improving the credibility of the video creation information.

[0042] In some embodiments, after S120, the method further includes: responding to a second input operation applied to the input box, determining the second input content corresponding to the second input operation; and based on the second input content and video creation information, invoking a video generation model to generate a target video.

[0043] The second input operation is an interactive action performed on the input box to trigger the video generation request. The video generation model is a generative model of a certain scale, based on video modalities.

[0044] Specifically, see [link to relevant documentation] Figure 2 Users can perform a second input operation through input box 210 to input their request to the electronic device to generate a video using video creation information. The electronic device can continue to display the second input content 250 in the dialog interface, then generate corresponding model prompts based on the second input content, and input them into the video generation model. After processing by the model, at least one target video is obtained. Afterwards, the electronic device can continue to display the generated target video 260 in the dialog interface. This allows for the acquisition of target videos with good creative ideas that meet user needs, improving the quality of the target videos and laying the foundation for obtaining good post-project data / metrics.

[0045] Figure 3 yes Figure 1 The diagram illustrates a detailed process flow of step S120 in the video content creation method. (See also...) Figure 2 S120, "Based on the first input content, call the content generation model to determine and display the video creation information corresponding to the first input content," specifically includes the following steps: S310. Based on the first input content, call the content generation model to determine the video search terms corresponding to the first input content, and determine each video creation dimension.

[0046] Among them, video search terms are keywords used for video retrieval.

[0047] Specifically, for the internal logic of video creation information generation, see [link to relevant documentation]. Figure 4 The electronic device can use the content generation model 401 to understand the first input content and extract keywords to obtain the video search terms 402 contained in the first input content.

[0048] In addition, the electronic device can determine the various video creation dimensions that need to be analyzed during this video analysis process 403. For example, if the default is to perform full-dimensional video analysis, the electronic device can read the configuration information to obtain the pre-configured various video creation dimensions. Furthermore, if the user has specific analysis needs, they can specify them through the first input content, and the electronic device can determine the user's specified / preferred video creation dimensions by understanding the first input content.

[0049] In some embodiments, determining each video creation dimension includes: if the video search term contains at least one target keyword, then based on the mapping relationship between multiple preset keywords and multiple preset creation dimensions, determining the preset creation dimension corresponding to the target keyword as the video creation dimension.

[0050] The target keyword is one of the preset keywords. Preset keywords are keywords related to video analysis that are set in advance. Preset creation dimensions are pre-defined dimensions for video creation.

[0051] Specifically, for certain industries or scenarios, the required videos may have a specific focus on certain video creation dimensions. For example, for industries with high content homogeneity or that rely on conveying core information in a short time, the required videos may place more emphasis on the opening attention-grabbing dimension (such as the eye-catching dimension in the first 3 seconds). Therefore, in this embodiment of the disclosure, multiple preset keywords can be constructed in advance according to the industry or scenario, and a mapping relationship (which can be referred to as the first mapping relationship) can be established between these preset keywords and their respective preset creation dimensions.

[0052] After obtaining video search terms, the electronic device can match each video search term with each preset keyword in the first mapping relationship mentioned above. If all matches fail, it indicates that the user does not have specific video analysis needs, and the subsequent steps can continue. If at least one match is successful, the successfully matched preset keyword can be determined as the target keyword. Then, the preset creation dimension corresponding to the target keyword is determined from the first mapping relationship, which serves as the video creation dimension required for this video analysis. This allows for more targeted generation and processing of video creation information, reducing the need for video analysis and summarization of video creation dimensions that users may not pay much attention to, thereby further improving the efficiency of video creation information generation.

[0053] In some embodiments, based on the first input content, a content generation model is invoked to determine the video search terms corresponding to the first input content, including: based on the first input content, the content generation model is invoked to perform search target identification, keyword extraction, and non-core keyword filtering on the first input content to determine the video search terms corresponding to the first input content.

[0054] Specifically, regarding the internal logic of video creation information generation, the electronic device can invoke a content generation model based on the first input content. This model can then perform the following processing: identify the retrieval target / purpose from the first input content, extract keywords that match the retrieval target from the first input content, and then filter out non-core keywords (such as brand names, product specifications, etc.) from the extracted keywords to finally obtain video search terms. This can improve the retrieval success rate, thereby increasing the quantity and quality of retrieved videos.

[0055] S320. Perform video retrieval based on video search terms to determine the target video set that matches the video creation dimensions.

[0056] Specifically, the electronic device can use video search terms to search for videos in the video database 404, obtaining multiple videos directly from the search results without filtering. This set of videos can be called the initial video set. Then, the electronic device can filter the initial video set to obtain multiple videos adapted to the video creation dimensions, forming the target video set.

[0057] It should be noted that the video database is adapted to the type of video creation information to be generated. For example, if the video creation information is for a promotional video, then the video database could be a database that stores many already generated promotional videos.

[0058] In some embodiments, S320 includes the following steps A to C: Step A: Based on video search terms, the content generation model calls the video search tool to select an initial set of videos from the video database that are compatible with the first input content.

[0059] Specifically, in addition to the aforementioned ability to determine video search terms, the content generation model in this disclosure embodiment can also plan the processing steps for generating video creation information, such as first calling a video retrieval tool and then calling a video understanding model. Therefore, see... Figure 4 After obtaining the video search term 402, the electronic device can invoke the video retrieval tool 405 through the content generation model 401. The video retrieval tool 405 then searches the video database 404 based on the video search term to obtain an initial set of videos.

[0060] Step B: Based on multiple video evaluation metrics, filter the initial video set to generate an intermediate video set.

[0061] Video evaluation metrics are used to assess the effectiveness of individual videos. These can be metrics such as click-through rate (CTR), conversion rate, spend, number of plays, or completion rate (in n seconds). Each video evaluation metric must at least cover the dimensions of the credibility metric. For example, if the credibility metric includes the overall CTR, then the video evaluation metrics must at least include the click-through rate of individual videos.

[0062] Specifically, see [link to relevant documentation] Figure 4 Electronic devices or video retrieval tools can first filter the initial video set based on the availability of video evaluation metrics, removing videos that lack any of these metrics. Then, the electronic devices or video retrieval tools further filter the resulting videos based on the median threshold corresponding to each video evaluation metric. For example, if every video evaluation metric for a video is greater than or equal to the median threshold of its respective metric, the video is retained; otherwise, if any video evaluation metric for a video is less than its corresponding median threshold, the video is filtered out. In this way, several relatively high-quality videos can be selected, forming an intermediate video set 406.

[0063] It should be noted that the median threshold mentioned above can be set empirically in advance, or it can be determined by statistically analyzing the video evaluation metrics of each video in the initial video set.

[0064] Step C: Based on the target video evaluation metrics corresponding to the video creation dimension, filter the intermediate video set to generate a target video set that fits the video creation dimension.

[0065] The target video evaluation metric is one of the various video evaluation metrics.

[0066] Specifically, different video creation dimensions may emphasize different video evaluation metrics. Therefore, in addition to pre-establishing a first mapping relationship between preset keywords and preset creation dimensions, this embodiment can also pre-establish a second mapping relationship between preset creation dimensions and their emphasized video evaluation metrics. Thus, after the electronic device determines the video creation dimension in the aforementioned steps, it can determine the corresponding video evaluation metric, called the target video evaluation metric, through the preset creation dimension that matches the video creation dimension in the second mapping relationship. Then, the electronic device can sort the videos in the intermediate video set according to the target video evaluation metric in descending order, and extract the top m videos from the sorting results to form the target video set 407. This provides a more targeted target video set with specific video creation dimensions, thereby further improving the accuracy of video creation information for the corresponding dimensions.

[0067] The value of m can be set according to the processing power of the video understanding model and / or the processing power of the device running the video understanding model. For example, a larger value can be set if the processing power is strong, and a smaller value can be set if the processing power is weak. For example, m can be set to a value of around 100 to balance processing power and model accuracy.

[0068] If multiple video creation dimensions exist, then the same number of target videos can be obtained. For example... Figure 4 There are four video creation dimensions, so four target video sets can be obtained from the intermediate video set, namely target video set 1, target video set 2, target video set 3 and target video set 4.

[0069] For example, if the value presentation dimension (such as the selling point dimension), the audience fit dimension (such as the target audience dimension), and the expresser image dimension (such as the voice-over persona dimension) all focus on the consumption ratio of a single video, then the aforementioned target video set 1, target video set 3, and target video set 4 can be the same video set. However, the opening attention guidance dimension (such as the first 3 seconds to grab attention dimension) focuses on the completion rate of a single video in the first 3 seconds, then the aforementioned target video set 2 is a different video set.

[0070] S330. Based on the target video set and the model prompt words adapted to the video creation dimension, the content generation model calls the video understanding model to understand and analyze each video in the target video set and generate initial creation information. The initial creation information includes the video creation dimension, the video creation elements corresponding to the video creation dimension, and the execution content associated with the video creation elements.

[0071] Among them, the video understanding model is a generative model with the ability to understand, analyze, and summarize videos.

[0072] Specifically, see [link to relevant documentation] Figure 4 For each video creation dimension, the electronic device constructs a corresponding model prompt using relevant information from each scene / first scene of each video in the target video set. Then, the electronic device uses the target video set and the model prompt as input data, and calls the video understanding model through the content generation model to understand, analyze, and summarize the video footage, spoken content, scenes, and shooting techniques from multiple perspectives, generating initial creation information for the corresponding dimension.

[0073] S340. Based on the initial creation information and the intermediate video set, determine the credibility index corresponding to the video creation elements, and generate video creation information from the initial creation information and the credibility index.

[0074] The intermediate video set includes the target video set and other videos. The other videos are those that match the video search terms and are not adapted to the video creation dimensions.

[0075] Specifically, see [link to relevant documentation] Figure 4 For each video creation dimension, after obtaining the initial creation information, the electronic device can parse and structure it to obtain initial creation information with a uniform format. Then, using each video creation element in the initial creation information as a statistical unit, the electronic device performs statistical calculations on the video evaluation indicators corresponding to each video in the intermediate video set corresponding to the video creation dimension, according to each credibility index, to obtain a comprehensive credibility index for each video creation element under the video creation dimension. Afterwards, the electronic device can anonymize the initial creation information and credibility index, and the results constitute the video creation information for the corresponding video creation dimension. The collection of video creation information for all video creation dimensions serves as the total video creation information corresponding to the first input content.

[0076] The video creation content generation method provided in the above embodiments of this disclosure can, based on a first input content, call a content generation model to determine the video search terms corresponding to the first input content, and perform video retrieval based on the video search terms to determine a target video set suitable for the video creation dimension; it expands the video retrieval scale from hundreds to tens of thousands, enhancing the coverage and representativeness of the videos, thereby providing a good data foundation for the generation of video creation information; then, based on the target video set and model prompts suitable for the video creation dimension, the content generation model calls a video understanding model to understand and analyze each video in the target video set, generating initial creation information, and based on the initial creation information and intermediate video set, determining the credibility index corresponding to the video creation elements, and generating video creation information from the initial creation information and credibility index; it realizes the automatic analysis and inspiration summary of each video using the video understanding model, and combined with the calculation of credibility index, greatly reducing human intervention, thereby further improving the generation efficiency, consistency and credibility of video creation information.

[0077] In some embodiments, if the initial video set is cached in an object storage service, the method further includes, before step B above, retrieving the initial video set from the object storage service based on video search terms.

[0078] Specifically, if multiple users have the same or similar needs for generating video creation information, each user's first input content will trigger the content generation model to perform the same video retrieval and obtain the same initial video set. However, the video retrieval process is relatively time-consuming and consumes certain computing resources. Therefore, in this embodiment, the video search terms and initial video set obtained from the first video retrieval of the first input content within a certain period (e.g., one day to one week) can be cached in the object storage service. In this way, after obtaining the first input content, the electronic device can first use its corresponding video search terms to query the object storage service to determine whether the required initial video set is already stored there. If it is, the video retrieval process can be saved, and the video retrieval result can be obtained directly, thereby shortening the video retrieval time and further reducing the waiting time for video creation information, thus further improving the efficiency of video creation information generation.

[0079] In some embodiments, if there are multiple video creation dimensions, after step A above, the method further includes: caching the initial video set to local storage space, and concurrently performing filtering and screening processes on the initial video set according to each video creation dimension to generate a target video set.

[0080] Specifically, as described in the foregoing embodiments, if there are multiple video creation dimensions in the generation process of video creation information, then multiple target video sets need to be obtained. Each target video set needs to obtain an initial video set and perform video filtering and selection processes. Therefore, in order to reduce read operations on object storage services, this embodiment can cache the initial video set to local storage space after obtaining it for the first time. Then, according to the number of video creation dimensions, the aforementioned video filtering and selection processes can be performed concurrently on the initial video set to obtain the corresponding target video sets, further shortening the generation time of each target video set, thereby further shortening the waiting time of video creation information and further improving the generation efficiency of video creation information.

[0081] In other embodiments, if there are multiple video creation dimensions, after step A above, the method further includes: caching the intermediate video set to local storage space, and concurrently performing filtering processing on the intermediate video set according to each video creation dimension to generate the target video set.

[0082] Specifically, in cases where multiple video creation dimensions exist during the generation of video creation information, given that the intermediate video sets are generated in the same way, this embodiment can execute a process of generating intermediate video sets from the initial video set once, and then cache the intermediate video sets in local storage. Afterwards, the aforementioned filtering process can be performed concurrently on the intermediate video sets according to the number of video creation dimensions to obtain the corresponding target video sets, further shortening the generation time of each target video set, thereby further shortening the waiting time for video creation information and further improving the generation efficiency of video creation information.

[0083] In some embodiments, after S120, the method further includes: storing video creation information in a key-value pair storage system, using video search terms and preset keywords corresponding to video search terms as independent storage key information and video creation information as storage value information.

[0084] Specifically, to avoid redundant computational overhead caused by repeatedly generating the same video creation information, electronic devices can store the video creation information in a key-value pair storage system with characteristics such as large capacity, high throughput, low latency, high availability, and easy scalability after obtaining it. To improve the query hit rate of video creation information, in this embodiment, the obtained video search terms and their corresponding preset keywords can be used as independent storage keys, and the video creation information can be used as the storage value information, storing the video creation information in the key-value pair storage system. This allows upstream and downstream services to efficiently query video creation information. To further improve the query efficiency of video creation information, electronic devices can store key-value pair information at the granularity of the video creation dimension.

[0085] The following are embodiments of the video content creation device provided in this disclosure. This device and the video content creation method described above belong to the same inventive concept. For details not described in detail in the embodiments of the video content creation device, please refer to the embodiments of the video content creation method described above.

[0086] Figure 5 A schematic diagram of the structure of a video content creation device provided in an embodiment of this disclosure is shown. Figure 5 As shown, the video content creation device 500 may include: The first input content determination module 510 is used to respond to the first input operation applied to the input box and determine the first input content corresponding to the first input operation. The video creation information determination module 520 is used to determine and display the video creation information corresponding to the first input content by calling the content generation model based on the first input content. The video creation information includes at least one video creation dimension, at least one video creation element corresponding to the video creation dimension, a credibility index corresponding to the video creation element, and at least one execution content associated with the video creation element to guide video creation.

[0087] The video creation content generation device provided in this embodiment can respond to a first input operation applied to an input box, determine the first input content corresponding to the first input operation, and, based on the first input content, call a content generation model to determine and display the video creation information corresponding to the first input content. The video creation information includes at least one video creation dimension, at least one video creation element corresponding to the video creation dimension, a credibility index corresponding to the video creation element, and at least one execution content associated with the video creation element for guiding video creation. This realizes the automatic generation of required video creation information using a generative model of a certain scale, greatly reducing human intervention and subjectivity, thereby improving the generation efficiency and consistency of video creation information. Furthermore, the generated video creation information contains a credibility index corresponding to each video creation element, providing a quantitative index value for the reliability of the video creation information, thereby improving the credibility of the video creation information.

[0088] In some embodiments, the video creation information determination module is specifically used for: Based on the first input content, the content generation model is invoked to determine and display the video retrieval results; Based on the video retrieval results and the first input content, the content generation model is invoked to determine and display the video creation information corresponding to the first input content.

[0089] In some embodiments, the video creation information determination module includes: The video search term determination submodule is used to call the content generation model based on the first input content, determine the video search terms corresponding to the first input content, and determine each video creation dimension; The target video set determination submodule is used to perform video retrieval based on video search terms and determine the target video set that is suitable for video creation dimensions. The initial creation information generation submodule is used to generate initial creation information based on the target video set and model prompts adapted to the video creation dimensions. It calls the video understanding model through the content generation model to understand and analyze each video in the target video set and generate initial creation information. The initial creation information includes the video creation dimension, the video creation elements corresponding to the video creation dimension, and the execution content associated with the video creation elements. The video creation information generation submodule is used to determine the credibility index corresponding to the video creation elements based on the initial creation information and the intermediate video set, and to generate video creation information from the initial creation information and the credibility index. The intermediate video set includes the target video set and other videos, which are videos that match the video search terms but are not adapted to the video creation dimensions.

[0090] In some embodiments, the video search term determination submodule is specifically used for: Based on the first input content, the content generation model is invoked to identify the retrieval target, extract keywords, and filter non-core keywords to determine the video search terms corresponding to the first input content.

[0091] In some embodiments, the video search term determination submodule is further specifically used for: If the video search term contains at least one target keyword, then based on the mapping relationship between multiple preset keywords and multiple preset creation dimensions, the preset creation dimension corresponding to the target keyword is determined as the video creation dimension; wherein, the target keyword is one of the preset keywords.

[0092] In some embodiments, the target video set determination submodule is specifically used for: Based on video search terms, a content generation model calls a video search tool to select an initial set of videos from the video database that are compatible with the first input content. Based on multiple video evaluation metrics, the initial video set is filtered to generate an intermediate video set; Based on the target video evaluation metrics corresponding to the video creation dimensions, the intermediate video set is filtered and processed to generate a target video set that is adapted to the video creation dimensions.

[0093] In some embodiments, the target video set determination submodule is further specifically used for: If the initial video set is cached in the object storage service, then before filtering the initial video set based on multiple video evaluation metrics to generate the intermediate video set, the initial video set is retrieved from the object storage service based on video search terms.

[0094] In some embodiments, the target video set determination submodule is further specifically used for: If there are multiple video creation dimensions, then based on video search terms, the content generation model calls the video search tool to select an initial set of videos from the video database that are compatible with the first input content. After that, the initial set of videos is cached in the local storage space, and filtering and selection processes are performed concurrently on the initial set of videos according to each video creation dimension to generate the target set of videos. Alternatively, if there are multiple video creation dimensions, after using the content generation model to call the video retrieval tool based on the video search terms to select an initial set of videos from the video database that are compatible with the first input content, the intermediate video set is cached in the local storage space, and the intermediate video set is concurrently filtered according to each video creation dimension to generate the target video set.

[0095] In some embodiments, the video content creation device 500 further includes a video creation information storage module, used for: After calling the content generation model based on the first input content, determining and displaying the video creation information corresponding to the first input content, the video creation information is stored in the key-value pair storage system, with the video search term and the preset keywords corresponding to the video search term as independent storage key information and the video creation information as storage value information.

[0096] In some embodiments, the video content creation device 500 further includes a video generation module, used for: After calling the content generation model based on the first input content, determining and displaying the video creation information corresponding to the first input content, responding to the second input operation applied to the input box, and determining the second input content corresponding to the second input operation; Based on the second input content and video creation information, the video generation model is invoked to generate the target video.

[0097] The video content creation device provided in this disclosure can execute the video content creation method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects of executing the method.

[0098] It is worth noting that in the embodiments of the video creation content generation device described above, the various modules and sub-modules are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of each functional module / sub-module are only for easy differentiation and are not used to limit the scope of protection of this disclosure.

[0099] This disclosure also provides an electronic device that may include a processor and a memory, the memory being used to store executable instructions. The processor can be used to read the executable instructions from the memory and execute them to implement the video creation content generation method described above.

[0100] Figure 6 A schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure is shown.

[0101] like Figure 6As shown, the electronic device 600 may include a processing unit 601 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 602 or a program loaded from storage device 608 into random access memory (RAM) 603. The RAM 603 also stores various programs and data required for the operation of the electronic device 600. The processing unit 601, ROM 602, and RAM 603 are interconnected via a bus 604. An input / output interface (I / O interface) 605 is also connected to the bus 604.

[0102] Typically, the following devices can be connected to I / O interface 605: input devices 606 including, for example, touch screens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 607 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 608 including, for example, magnetic tapes, hard disks, etc.; and communication devices 609. Communication device 609 allows electronic device 600 to communicate wirelessly or wiredly with other devices to exchange data.

[0103] It should be noted that, Figure 6 The illustrated electronic device 600 is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein. That is, although... Figure 6 An electronic device 600 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0104] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 609, or installed from a storage device 608, or installed from a ROM 602. When the computer program is executed by the processing device 601, it performs the functions defined in the video creation content generation method of any embodiment of this disclosure.

[0105] This disclosure also provides a computer-readable storage medium storing a computer program that, when executed by a processor, causes the processor to implement the video creation content generation method in any embodiment of this disclosure.

[0106] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. Computer-readable storage media can be, for example, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include data signals propagated in baseband or as part of a carrier wave, carrying computer-readable program code. The transmitted data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, radio frequency (RF), etc., or any suitable combination thereof.

[0107] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol, such as Hypertext Transfer Protocol (HTTP), and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include Local Area Networks (LANs), Wide Area Networks (WANs), the Internet (e.g., the Internet), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.

[0108] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0109] The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the video creation content generation method described in any embodiment of this disclosure.

[0110] In embodiments of this disclosure, computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof. These programming languages ​​include, but are not limited to, object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0111] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of devices, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0112] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Array (FPGA), Application Specific Integrated Circuit (ASIC), Application Specific Standard Parts (ASSP), System on Chip (SOC), Complex Programmable Logic Device (CPLD), and so on.

[0113] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.

[0114] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.

[0115] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.

Claims

1. A video authoring content generation method, characterized by, The method comprises: in response to a first input operation on an input box, determining first input content corresponding to the first input operation; based on the first input content, calling a content generation model to determine and display video creation information corresponding to the first input content; wherein the video creation information comprises at least one video creation dimension, at least one video creation element corresponding to the video creation dimension, a credibility index corresponding to the video creation element, and at least one execution content associated with the video creation element for guiding video creation.

2. The method of claim 1, wherein, The method further comprises: based on the first input content, calling the content generation model to determine and display a video retrieval result; based on the video retrieval result and the first input content, continuing to call the content generation model to determine and display the video creation information corresponding to the first input content.

3. The method of claim 1, wherein, The method further comprises: based on the first input content, calling the content generation model to determine a video retrieval keyword corresponding to the first input content and determine each video creation dimension; based on the video retrieval keyword, performing video retrieval to determine a target video set adapted to the video creation dimension; based on the target video set and a model prompt word adapted to the video creation dimension, calling a video understanding model through the content generation model to understand and analyze each video in the target video set to generate initial creation information; wherein the initial creation information comprises the video creation dimension, the video creation element corresponding to the video creation dimension, and the execution content associated with the video creation element; based on the initial creation information and an intermediate video set, determining the credibility index corresponding to the video creation element, and generating the video creation information from the initial creation information and the credibility index; wherein the intermediate video set comprises the target video set and other videos that meet the video retrieval keyword and are not adapted to the video creation dimension.

4. The method of claim 3, wherein, The method further comprises: based on the first input content, calling the content generation model to perform retrieval target identification, keyword extraction, and non-core keyword filtering on the first input content to determine the video retrieval keyword corresponding to the first input content.

5. The method according to claim 3 or 4, characterized in that, The method further comprises: if the video retrieval keyword comprises at least one target keyword, determining a preset creation dimension corresponding to the target keyword as the video creation dimension based on a mapping relationship between a plurality of preset keywords and a plurality of preset creation dimensions; wherein the target keyword is one of the preset keywords.

6. The method of claim 3, wherein, The method further comprises: based on the video retrieval keyword, performing video retrieval to determine a target video set adapted to the video creation dimension, comprising: based on the video search term, calling a video search tool through the content generation model to screen an initial video set from a video database, the initial video set being adapted to the first input content; based on a plurality of video evaluation indexes, performing filtering processing on the initial video set to generate the intermediate video set; based on a target video evaluation index corresponding to the video creation dimension, performing screening processing on the intermediate video set to generate a target video set adapted to the video creation dimension.

7. The method of claim 6, wherein, If the initial video set is cached in an object storage service, before the filtering processing on the initial video set based on a plurality of video evaluation indexes to generate the intermediate video set, the method further comprises: based on the video search term, querying the initial video set from the object storage service.

8. The method of claim 6, wherein, If there are a plurality of video creation dimensions, after the initial video set is screened from the video database based on the video search term through the content generation model calling the video search tool, the method further comprises: caching the initial video set to a local storage space, and performing the filtering processing and the screening processing on the initial video set concurrently according to each video creation dimension to generate the target video set; or, caching the intermediate video set to the local storage space, and performing the screening processing on the intermediate video set concurrently according to each video creation dimension to generate the target video set.

9. The method of claim 3, wherein, After the video creation information corresponding to the first input content is determined and displayed based on the first input content calling the content generation model, the method further comprises: respectively taking the video search term and a preset keyword corresponding to the video search term as independent storage key information, taking the video creation information as storage value information, and storing the video creation information in a key-value pair storage system.

10. The method of claim 1 or 2, wherein, After the video creation information corresponding to the first input content is determined and displayed based on the first input content calling the content generation model, the method further comprises: in response to a second input operation acting on the input box, determining second input content corresponding to the second input operation; based on the second input content and the video creation information, calling a video generation model to generate a target video.

11. A video authoring content generation apparatus characterized by comprising: comprising: a first input content determination module configured to determine first input content corresponding to a first input operation acting on an input box; a video creation information determination module configured to determine and display video creation information corresponding to the first input content based on the first input content calling a content generation model; wherein the video creation information comprises at least one video creation dimension, at least one video creation element corresponding to the video creation dimension, a credibility index corresponding to the video creation element, and at least one execution content associated with the video creation element for guiding video creation.

12. An electronic device, comprising: comprising: a processor; a memory configured to store executable instructions; The processor is configured to read the executable instructions from the memory and execute the executable instructions to implement the video creation content generation method according to any one of claims 1-10.

13. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, and when the computer program is executed by the processor, the processor implements the video creation content generation method according to any one of claims 1-10.

14. A computer program product, characterised in that, The computer program product is configured to implement the video creation content generation method according to any one of claims 1-10.