Movie shooting planning method and system based on deep learning

By using a deep learning-based film and television shooting planning method, and generating target video main story scripts and function demonstration scripts using user viewing behavior data, the problem of lack of data support in traditional film and television advertising planning is solved, thereby improving the accuracy and market adaptability of advertising content.

CN120769130BActive Publication Date: 2026-02-27GUANGZHOU SHUANGYANG ADVERTISING CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511063445.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-31
Publication Date
2026-02-27
Estimated Expiration
2045-07-31

AI Technical Summary

Technical Problem

Traditional product video advertising planning lacks systematic quantitative data support, resulting in a lack of objective basis for emphasizing product selling points in advertisements. This may lead to a mismatch with the needs of the target audience, and the market adaptability of the planning scheme is difficult to predict.

Method used

By employing a deep learning-based film and television shooting planning method, user viewing behavior data is acquired, and video main story generation model and function demonstration generation model are used to generate target video main story scripts and function demonstration scripts, ensuring that advertising content is close to the audience's interests and reducing subjective bias.

Benefits of technology

It improves user engagement and appeal of ads, reduces lengthy segments, increases ad conversion rates, ensures ad content is more relevant to audience interests, and enhances market adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120769130B_ABST
    Figure CN120769130B_ABST
Patent Text Reader

Abstract

The application relates to the field of film and television shooting planning, and discloses a film and television shooting planning method and system based on deep learning, which comprises the following steps: obtaining a target commodity category corresponding to a target film and television shooting task, obtaining a plurality of promotion videos corresponding to the target commodity category, and determining video main line segments, function demonstration segments and user viewing behavior data corresponding to the plurality of promotion videos respectively; determining a target video main line script corresponding to the target film and television shooting task according to the video main line segments in the plurality of promotion videos and the user viewing behavior data corresponding to the video main line segments; and determining a target function demonstration script according to the target video main line script, the function demonstration segments in the plurality of promotion videos and the user viewing behavior data corresponding to the function demonstration segments. The application introduces a systematic quantitative data support system in the planning of commodity film and television video advertisements, thereby improving the effect of film and television shooting planning.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of film and television shooting planning, more specifically, to a film and television shooting planning method and system based on deep learning. BACKGROUND

[0002] In the planning method of traditional commercial film and television video advertisements, the generation of ideas and the decision of plans highly depend on the personal experience of creators or the subjective evaluation of teams, and lack a systematic quantitative data support system. For example, when planning an advertisement for a skin care product, the selection of the product's core selling point is often determined by the creative team based on "market trend judgment" or "customer subjective preference", rather than quantitative analysis through consumer research data (such as user comment keyword frequency, search hot word distribution).

[0003] The subjective decision-making mode of existing commercial film and television video advertisements leads to a lack of objective basis for the emphasis on the selling points of the product in the advertisement, which may be misaligned with the real needs of the target audience. In addition, in the design of advertising scenes and the selection of actor images, traditional methods often rely on the artistic intuition of directors or planners, such as selecting a forest scene as the background of the advertisement based only on "visual aesthetics", without verifying the matching degree of the scene and the product tone through user behavior data, which makes it difficult to predict the market adaptability of the planning scheme. Therefore, the existing shooting mode of commercial film and television video advertisements not only leads to the difficulty in improving the input-output ratio of product advertisements, but also makes it difficult for small and medium-sized brands to form a scientific planning methodology due to the lack of data accumulation. SUMMARY

[0004] The purpose of the present application is to provide a film and television shooting planning method and system based on deep learning, which solves the technical problem of the lack of a systematic quantitative data support system in the planning method of traditional commercial film and television video advertisements, and achieves the technical effect of introducing a systematic quantitative data support system in the planning of commercial film and television video advertisements and improving the effect of film and television shooting planning.

[0005] The embodiment of the application provides a movie shooting planning method based on deep learning, which comprises the following steps: obtaining a target commodity category corresponding to a target movie shooting task, and obtaining a plurality of promotional videos corresponding to the target commodity category, and determining a video main line segment and a function demonstration segment corresponding to each promotional video; obtaining user viewing behavior data corresponding to the video main line segment and the function demonstration segment in each promotional video; wherein the user viewing behavior data comprises video playing operation data and comment data of the user; determining a target video main line script corresponding to the target movie shooting task by a video main line generation model according to the video main line segment in the plurality of promotional videos and the user viewing behavior data corresponding to the video main line segment; determining a target function demonstration script by a function demonstration generation model according to the target video main line script, the function demonstration segment in the plurality of promotional videos and the user viewing behavior data corresponding to the function demonstration segment; wherein a time period corresponding to the target function demonstration script is inserted into a time period corresponding to the target video main line script.

[0006] In a possible implementation, the method further comprises: clustering the video main line segments of the plurality of promotional videos according to the user viewing behavior data corresponding to the video main line segments respectively, to obtain a plurality of video main line segment clustering groups; obtaining a plurality of function demonstration segment clustering groups by clustering the function demonstration segments of the plurality of promotional videos according to the user viewing behavior data corresponding to the function demonstration segments; determining the average of the similarity of the target commodity category and the commodity categories corresponding to the plurality of video main line segment clustering groups respectively as a video main line similarity, and determining the average of the similarity of the target commodity category and the commodity categories corresponding to the plurality of function demonstration segment clustering groups respectively as a function demonstration segment similarity; determining the maximum video main line similarity in the plurality of video main line similarities, and determining the maximum function demonstration similarity in the plurality of function demonstration segment similarities; determining the target video main line script corresponding to the target movie shooting task by the video main line generation model according to the video main line segments in the plurality of promotional videos in the video main line segment clustering group corresponding to the maximum video main line similarity and the user viewing behavior data corresponding to the video main line segments; and determining the target function demonstration script by the function demonstration generation model according to the target video main line script, the function demonstration segments in the plurality of promotional videos in the function demonstration segment clustering group corresponding to the maximum function demonstration similarity and the user viewing behavior data corresponding to the function demonstration segments.

[0007] In another possible implementation, the method further includes: obtaining target evaluation indicator values corresponding to the target film and television shooting task in multiple evaluation dimensions respectively; obtaining evaluation indicator values corresponding to each video mainline segment clustering group in multiple evaluation dimensions respectively, and obtaining evaluation indicator values corresponding to each function demonstration segment clustering group in multiple evaluation dimensions respectively; wherein the multiple evaluation dimensions include sales data, play quantity data, and user sentiment tendency data; determining a target video mainline segment clustering group from the multiple video mainline segment clustering groups and determining a target function demonstration segment clustering group from the multiple function demonstration segment clustering groups, with the sum of the evaluation indicator values corresponding to the video mainline segment clustering groups in multiple evaluation dimensions respectively and the evaluation indicator values corresponding to the function demonstration segment clustering groups in multiple evaluation dimensions respectively being closest to the target evaluation indicator values; determining a target video mainline script corresponding to the target film and television shooting task through a video mainline generation model according to the video mainline segments in the multiple promotion videos in the target video mainline segment clustering group and the user viewing behavior data corresponding to the video mainline segments; and determining a target function demonstration script through a function demonstration generation model according to the target video mainline script, the function demonstration segments in the multiple promotion videos in the target function demonstration segment clustering group, and the user viewing behavior data corresponding to the function demonstration segments.

[0008] In another possible implementation, the determining of the target video mainline segment clustering group from the multiple video mainline segment clustering groups and the determining of the target function demonstration segment clustering group from the multiple function demonstration segment clustering groups, with the sum of the evaluation indicator values corresponding to the video mainline segment clustering groups in multiple evaluation dimensions respectively and the evaluation indicator values corresponding to the function demonstration segment clustering groups in multiple evaluation dimensions respectively being closest to the target evaluation indicator values, includes: determining dimension weights corresponding to the multiple evaluation dimensions respectively; taking the sum of the products of the evaluation indicator values corresponding to the video mainline segment clustering groups in multiple evaluation dimensions respectively and the dimension weights and the products of the evaluation indicator values corresponding to the function demonstration segment clustering groups in multiple evaluation dimensions respectively and the dimension weights as evaluation indicator intermediate values; determining the target video mainline segment clustering group from the multiple video mainline segment clustering groups and determining the target function demonstration segment clustering group from the multiple function demonstration segment clustering groups, with the evaluation indicator intermediate values being closest to the target evaluation indicator values.

[0009] In another possible implementation, the method further includes: determining multiple promotion videos corresponding to the target product category corresponding to the target film and television shooting task in a preset historical time period, and obtaining average play quantity and / or peak play quantity of the multiple promotion videos on a first broadcast platform; determining multiple target promotion videos in a descending order of the average play quantity and / or the peak play quantity; obtaining evaluation indicator values corresponding to the multiple target promotion videos in multiple evaluation dimensions respectively as the target evaluation indicator values corresponding to the target film and television shooting task in multiple evaluation dimensions respectively.

[0010] In another possible implementation, the method further includes: determining, by the video effect evaluation model, a video mainline segment evaluation value corresponding to each target video mainline segment in the target video mainline script according to each target video mainline segment in the target video mainline script; determining, by the video effect evaluation model, a function demonstration segment evaluation value corresponding to each target function demonstration segment in the target function demonstration script according to each target function demonstration segment in the target function demonstration script; determining a sum of each video mainline segment evaluation value and a time-sequentially adjacent function demonstration segment evaluation value as a target script evaluation mean value; determining a difference between target script evaluation mean values corresponding to time-sequentially adjacent video mainline segments and function demonstration segments as an evaluation mean value difference; and rearranging the plurality of target video mainline segments and the plurality of target function demonstration segments with a minimum evaluation mean value difference between the target script evaluation mean values corresponding to the time-sequentially adjacent video mainline segments and function demonstration segments as a target.

[0011] In another possible implementation, the method further includes: when the evaluation mean value difference between the adjacent video mainline segments and function demonstration segments is greater than or equal to the first evaluation mean value difference, determining the target function demonstration script by the function demonstration generation model according to the target video mainline script, the video mainline segment evaluation value corresponding to each target video mainline segment in the target video mainline script, the function demonstration segments in the plurality of promotional videos, and the user viewing behavior data corresponding to the function demonstration segments.

[0012] In another possible implementation, the method further includes: when the evaluation mean value difference between the adjacent video mainline segments and function demonstration segments is greater than or equal to the second evaluation mean value difference, determining the target video mainline script corresponding to the target film shooting task by the video mainline generation model according to the video mainline segments in the plurality of promotional videos, the video mainline segment evaluation value corresponding to each target video mainline segment in the target video mainline script, and the user viewing behavior data corresponding to the video mainline segments; and wherein the second evaluation mean value difference is greater than the first evaluation mean value difference.

[0013] In another possible implementation, the method further includes: when the evaluation mean value difference between the adjacent video mainline segments and function demonstration segments is greater than or equal to the second evaluation mean value difference, rearranging the plurality of target video mainline segments and the plurality of target function demonstration segments with a maximum evaluation mean value difference between the target script evaluation mean values corresponding to the time-sequentially adjacent video mainline segments and function demonstration segments as a target.

[0014] The embodiments of the present application also provide a film shooting planning system based on deep learning, which includes units for executing the method according to any one of the above.

[0015] Compared with the prior art, the embodiments of the present application have the beneficial effects that:

[0016] The embodiment of the present application provides a kind of based on deep learning's film shooting planning method, this method includes: the target commodity category corresponding to target film shooting task is obtained, and the target commodity category corresponding multiple promotion videos are obtained, and the video main line segment and function demonstration segment corresponding respectively in multiple promotion videos are determined;The video main line segment and function demonstration segment corresponding respectively in each promotion video are obtained User viewing behavior data;Wherein, user viewing behavior data includes the video play operation data and comment data of user;By video main line generation model, according to the video main line segment in multiple promotion videos and the user viewing behavior data corresponding to video main line segment, the target video main line script corresponding to target film shooting task is determined;By function demonstration generation model, according to target video main line script, the function demonstration segment in multiple promotion videos and the user viewing behavior data corresponding to function demonstration segment, target function demonstration script is determined;Wherein, the time period corresponding to target function demonstration script is inserted into the time period corresponding to target video main line script.The film shooting planning method in the embodiment of the present application can be determined in turn target video main line script, target function demonstration script, can ensure that advertisement content is closer to audience interest, improve user participation and advertisement attraction, while avoiding subjective bias in traditional planning, so that model-driven script generation can reduce lengthy segment, improve advertisement conversion rate. BRIEF DESCRIPTION OF DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed to be used in the embodiments or prior art description will be briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0018] Figure 1 The flowchart of the first deep learning-based film shooting planning method provided by the embodiment of the present application is shown in the figure.

[0019] Figure 2 The workflow diagram of the first deep learning-based film shooting planning method provided by the embodiment of the present application is shown in the figure.

[0020] Figure 3 The flowchart of the second deep learning-based film shooting planning method provided by the embodiment of the present application is shown in the figure.

[0021] Figure 4 The flowchart of the third deep learning-based film shooting planning method provided by the embodiment of the present application is shown in the figure.

[0022] Figure 5A fourth deep learning-based film shooting planning method provided by an embodiment of the present application is shown in the flowchart.

[0023] Figure 6 A structure diagram of a target video script provided by an embodiment of the present application is shown in the flowchart.

[0024] Figure 7 A logic structure diagram of a deep learning-based film shooting planning system provided by an embodiment of the present application is shown in the flowchart. DETAILED DESCRIPTION

[0025] It should be understood that the term "comprises" as used in the specification and the appended claims indicates the presence of the described features, integers, steps, operations, elements, and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0026] It should also be understood that the term "and / or" as used in the specification and the appended claims, means any one or more of the associated listed items, as well as all possible combinations of the items.

[0027] As used in the specification and the appended claims, the term "if" can be interpreted as meaning "when" or "once" or "in response to a determination" or "in response to detecting" depending on the context. Similarly, the phrase "if determined" or "if detected [the described condition or event]" can be interpreted as meaning "once determined" or "in response to a determination" or "once detected [the described condition or event]" or "in response to detecting [the described condition or event]" depending on the context.

[0028] In addition, in the description of the specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish descriptions, and cannot be understood as indicating or implying relative importance.

[0029] The reference "one embodiment" or "some embodiments" and the like described in the present specification means that a particular feature, structure, or characteristic described in connection with the embodiment is included in one or more embodiments of the present application. Therefore, the statements "in one embodiment", "in some embodiments", "in other some embodiments", "in yet some embodiments", and the like appearing in different places in the specification are not necessarily all referring to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized. The terms "include", "contain", "have", and their variants mean "including but not limited to", unless otherwise specifically emphasized.

[0030] The subjective decision-making model of existing product video advertisements leads to a lack of objective basis for emphasizing the selling points of products in the advertisements, which may be misaligned with the real needs of the target audience.

[0031] Based on the above reasons, this application provides a deep learning-based film and television shooting planning method. This method includes: obtaining the target product category corresponding to the target film and television shooting task, obtaining multiple promotional videos corresponding to the target product category, and determining the main video segment and function demonstration segment corresponding to each of the multiple promotional videos; obtaining user viewing behavior data corresponding to the main video segment and function demonstration segment in each promotional video; wherein, the user viewing behavior data includes user video playback operation data and comment data; determining the target video mainline script corresponding to the target film and television shooting task through a video mainline generation model based on the main video segments in the multiple promotional videos and the user viewing behavior data corresponding to the main video segments; determining the target function demonstration script through a function demonstration generation model based on the target video mainline script, the function demonstration segments in the multiple promotional videos, and the user viewing behavior data corresponding to the function demonstration segments; wherein, the time period corresponding to the target function demonstration script is inserted into the time period corresponding to the target video mainline script. The film and television shooting planning method in this application embodiment can target the main video script and the target function demonstration script in sequence, which can ensure that the advertising content is closer to the audience's interests, improve user participation and advertising attractiveness, and at the same time avoid the subjective bias in traditional planning. This allows model-driven script generation to reduce lengthy segments and improve advertising conversion rate.

[0032] In some scenarios, the deep learning-based film and television shooting planning method of this application embodiment can be applied to the film and television shooting planning of video advertisements for various products, which can improve the efficiency and reliability of film and television shooting planning for video advertisements for products.

[0033] The following section provides a detailed explanation of a deep learning-based film and television shooting planning method provided in this application, using specific examples.

[0034] Figure 1 A flowchart illustrating the first deep learning-based film and television shooting planning method provided in this application embodiment is shown below. Figure 1 As shown, this deep learning-based film and television shooting planning method includes S110 to S120, and S110 to S120 will be explained in detail below.

[0035] S110, obtain the target commodity category corresponding to the target film and television shooting task, and obtain a plurality of promotion videos corresponding to the target commodity category, and determine the video main line segment and the function demonstration segment corresponding to each of the plurality of promotion videos. Obtain the user viewing behavior data corresponding to each of the video main line segment and the function demonstration segment in each promotion video. The user viewing behavior data includes video playback operation data and comment data of the user.

[0036] Figure 2 The working flow diagram of the first deep learning-based film and television shooting planning method provided by the embodiments of the present application is shown in Figure 2 In the deep learning-based film and television shooting planning method, the target commodity category corresponding to the target film and television shooting task can be obtained first, and then a plurality of promotion videos under the target commodity category can be obtained.

[0037] Exemplarily, the promotion videos corresponding to different commodity categories can be stored in a video database, facilitating subsequent analysis.

[0038] After obtaining the plurality of promotion videos under the target commodity category, the video main line segment and the function demonstration segment corresponding to each of the plurality of promotion videos can be further determined. The video main line segment mainly includes a video story line, a brand story narration or a video segment conveying product value, while the function demonstration segment focuses on real-time operation demonstration of product functions.

[0039] After obtaining the video main line segment and the function demonstration segment corresponding to each of the plurality of promotion videos, the user viewing behavior data corresponding to each of the video main line segment and the function demonstration segment in each promotion video can be obtained.

[0040] Exemplarily, the user viewing behavior data includes video playback operation data of the user, specifically including pause, fast forward and playback frequency of the user on the advertisement video, and comment data, specifically including real-time feedback and emotional analysis of the user on the content of the advertisement video.

[0041] Exemplarily, the user viewing behavior data can be extracted from the video platform background or the data analysis system, helping to understand the audience interest points. Taking a commodity video advertisement as an example, for example, for an advertisement series of a household cleaning product, the video main line segment may include a story narration of a family scene, and the function demonstration segment shows the cleaning effect of the product. The user viewing behavior data such as frequent playback of the user on the function demonstration part can reflect the audience's attention to the product performance.

[0042] S120, determine a target video main line script corresponding to the target film shooting task according to the video main line segments in the plurality of promotional videos and the user viewing behavior data corresponding to the video main line segments through a video main line generation model. Determine a target function demonstration script through a function demonstration generation model according to the target video main line script, the function demonstration segments in the plurality of promotional videos and the user viewing behavior data corresponding to the function demonstration segments. The time period corresponding to the target function demonstration script is inserted into the time period corresponding to the target video main line script.

[0043] After obtaining the data in S110, as shown in Figure 2 the target video main line script corresponding to the target film shooting task can be determined through a video main line generation model according to the video main line segments in the plurality of promotional videos and the user viewing behavior data corresponding to the video main line segments.

[0044] In determining the target video main line script, the video main line generation model can analyze the heat trend and audience behavior mode of the video main line segments in the plurality of promotional videos by using a deep learning algorithm, generate a target video main line script that meets the audience's preferences, and the target video main line script is used to determine the main line narrative structure of the advertising video.

[0045] For example, the video main line generation model can be fine-tuned based on a deep learning-based video understanding model. The video main line generation model is fine-tuned based on sample video main line segments in the plurality of promotional videos, sample user viewing behavior data corresponding to the sample video main line segments, and sample video main line scripts corresponding to sample film shooting tasks. Then the target video main line script can be determined through the video main line generation model.

[0046] For example, in promoting a product video advertisement of a smart phone, user comment data (such as high-frequency discussions of the audience on shooting experience) can be integrated to automatically generate a video main line script sequence that emphasizes the product camera performance.

[0047] After obtaining the target video main line script, a target function demonstration script can be determined through a function demonstration generation model according to the target video main line script, the function demonstration segments in the plurality of promotional videos and the user viewing behavior data corresponding to the function demonstration segments.

[0048] In determining the target function demonstration script, the function demonstration generation model takes the target video main line script as input, combines the user behavior data of the function segments, and automatically generates function demonstration details that are connected with the main line content.

[0049] For example, the function demonstration generation model can be a deep learning-based video understanding model, and the target function demonstration script can be determined through the function demonstration generation model.

[0050] Exemplarily, the function demonstration generation model can be obtained by fine-tuning a deep learning-based video understanding model, and the function demonstration generation model is fine-tuned based on a sample video mainline script, a sample function demonstration segment in a plurality of promotion videos, sample user viewing behavior data corresponding to the sample function demonstration segment, and a sample function demonstration script. Then, the target function demonstration script can be determined by using the function demonstration generation model.

[0051] When the target function demonstration script is determined, the time period corresponding to the target function demonstration script can be inserted into the time period corresponding to the target video mainline script, so as to ensure the overall coherence and rhythm.

[0052] Taking a commodity video advertisement as an example, for example, in an advertisement for promoting a smart kitchen device, the function demonstration generation model can insert product operation demonstration into the climax part of the cooking story according to the time point of the mainline script, so as to balance product information and emotional appeal.

[0053] The above-mentioned implementation manner has the beneficial effects that by obtaining user viewing behavior data and combining a deep learning model, the generation process of the video script can be optimized. This method ensures that the advertisement content is closer to the interests of the audience, improves user participation and advertisement appeal, and at the same time avoids subjective bias in traditional planning, so that the model-driven script generation can reduce lengthy segments and improve advertisement conversion rate.

[0054] The above-mentioned implementation manner also has the beneficial effects that the time period insertion mechanism of the target function demonstration script can enhance the fluency and integrity of the advertisement. In the integration process, the function demonstration content is seamlessly integrated with the video mainline, which not only optimizes the expression rhythm of the advertisement, but also effectively conveys the core function of the product, and helps to improve the visual impact and market competitiveness of the advertisement while ensuring the integrity of the information.

[0055] Figure 3 A flowchart of a second deep learning-based film and television shooting planning method provided by the embodiments of the present application is shown in FIG. 2. Figure 3 As shown in FIG. 2, the above-mentioned method further includes S130 to S140, which are specifically described as follows.

[0056] S130. Based on user viewing behavior data corresponding to the main video segments, cluster the main video segments of multiple promotional videos to obtain multiple main video segment cluster groups. Based on user viewing behavior data corresponding to the function demonstration segments of multiple promotional videos, obtain multiple function demonstration segment cluster groups. Determine the average similarity between the target product category and the product categories corresponding to the multiple main video segment cluster groups as the main video similarity, and determine the average similarity between the target product category and the product categories corresponding to the multiple function demonstration segment cluster groups as the function demonstration segment similarity. Determine the maximum main video similarity among the multiple main video similarities, and determine the maximum function demonstration similarity among the multiple function demonstration segment similarities.

[0057] In the deep learning-based film and television shooting planning method in this implementation, user viewing behavior data corresponding to the main video segments and function demonstration segments of multiple promotional videos can be further obtained.

[0058] For example, user viewing behavior data may include video playback operation data and comment data, which can be extracted from the video platform's backend or user feedback system.

[0059] After obtaining user viewing behavior data corresponding to the main video segments, cluster analysis can be performed on the main video segments of multiple promotional videos based on this data, resulting in multiple main video segment cluster groups. This clustering process can automatically group segments based on the similarity characteristics of the data, facilitating subsequent processing.

[0060] For example, in a product video ad promoting a smartphone, the main video segments may include product appearance introductions or user scenario story narrations. High user viewing behavior data, such as high pause frequency or high number of comments, indicates that the audience's interests are consistent, and clustering groups can group segments with similar themes into the same category.

[0061] Similarly, feature demonstration clips can be clustered based on user viewing behavior data corresponding to multiple promotional video feature demonstration clips, resulting in multiple feature demonstration clip cluster groups. Feature demonstration clips mainly focus on product operation demonstrations; for example, in smartphone advertisements, feature demonstration clips may demonstrate camera functions or screen operations. Clustering can be based on user data such as the number of times the video has been replayed to achieve grouping.

[0062] After obtaining the video mainline segment clustering groups, the average of the similarities between the target product category and the product categories corresponding to the plurality of video mainline segment clustering groups can be determined respectively as the video mainline similarity. The video mainline similarity is a statistical value of the average of the similarities of the video mainline segment clustering groups, and quantifies the overall matching degree of each video mainline segment clustering group to the target product.

[0063] For example, when determining the similarity between the target product category and the product categories corresponding to the plurality of video mainline segment clustering groups, the similarity can be calculated by a pre-defined classification model, and the similarity between the target product category and the product categories corresponding to the plurality of video mainline segment clustering groups represents the thematic relevance of the target product category and the product categories corresponding to the plurality of video mainline segment clustering groups.

[0064] Similarly, the average of the similarities between the target product category and the product categories corresponding to the plurality of function demonstration segment clustering groups can be determined respectively as the function demonstration segment similarity.

[0065] After obtaining the plurality of video mainline similarities, the maximum video mainline similarity in the plurality of video mainline similarities can be determined, which represents the highest relevance in the plurality of video mainline similarities. At the same time, the maximum function demonstration similarity in the plurality of function demonstration segment similarities can be determined. Further, the video mainline segment clustering group and the function demonstration segment clustering group that best match the target product can be used preferentially.

[0066] For example, in an advertisement promoting a smart phone, the target product category is “electronic device”, the video mainline clustering group can include “smart phone”, “smart wearable device”, and other categories, and the function demonstration clustering group can include “user experience”, “wearing experience”, and other categories. By determining the maximum video mainline similarity in the plurality of video mainline similarities and the maximum function demonstration similarity in the plurality of function demonstration segment similarities, the video mainline segment clustering group and the function demonstration segment clustering group that best match the target product category can be used preferentially.

[0067] S140, by the video mainline generation model, according to the video mainline segments in the plurality of promotional videos in the video mainline segment clustering group corresponding to the maximum video mainline similarity and the user viewing behavior data corresponding to the video mainline segments, a target video mainline script corresponding to the target film shooting task is determined. By the function demonstration generation model, according to the target video mainline script, the function demonstration segments in the plurality of promotional videos in the function demonstration segment clustering group corresponding to the maximum function demonstration similarity and the user viewing behavior data corresponding to the function demonstration segments, a target function demonstration script is determined.

[0068] After the maximum video mainline similarity and the maximum function demonstration similarity are determined, the video mainline generation model can be used to cluster the video mainline segments in the multiple promotional videos corresponding to the maximum video mainline similarity and the user viewing behavior data corresponding to the video mainline segments, i.e., based on the video mainline segment clustering group that best fits the target commodity category, to determine the target video mainline script corresponding to the target film and television shooting task.

[0069] Meanwhile, the function demonstration generation model can be used to cluster the function demonstration segments in the multiple promotional videos corresponding to the target video mainline script and the maximum function demonstration similarity and the user viewing behavior data corresponding to the function demonstration segments, i.e., based on the function demonstration segment clustering group that best fits the target commodity category, to determine the target function demonstration script.

[0070] The implementation manner described above has the beneficial effect that, by clustering and similarity comparison of video segments, a segment group highly relevant to the target commodity category can be selected, thereby improving the relevance and accuracy of script generation and avoiding blind selection of video segments, ensuring that the advertisement content is closer to the audience preference and improving user engagement and conversion rate.

[0071] The implementation manner described above also has the beneficial effect that, by using the video mainline segment clustering group corresponding to the maximum similarity, the function demonstration segment clustering group, and the deep learning model to generate the script, the video mainline and the function demonstration part can be seamlessly integrated, the overall smoothness and appeal of the advertisement are optimized, and the market adaptability and shooting efficiency of the advertisement are improved while ensuring the integrity of information communication.

[0072] Figure 4 A flowchart of a third deep learning-based film and television shooting planning method provided by the embodiments of the present application is shown in FIG. 6. Figure 4 As shown in FIG. 6, the method described above further includes S210 to S230, which are described in detail below.

[0073] S210, obtaining target evaluation index values corresponding to a target film and television shooting task in multiple evaluation dimensions respectively. Obtain evaluation index values corresponding to each video mainline segment clustering group in multiple evaluation dimensions respectively, and obtain evaluation index values corresponding to each function demonstration segment clustering group in multiple evaluation dimensions respectively. The multiple evaluation dimensions include sales data, play data, and user sentiment tendency data.

[0074] In this method, during the deep learning-based film and television shooting planning process, the target evaluation index values ​​corresponding to the target film and television shooting task under multiple evaluation dimensions can be further obtained. The target evaluation index values ​​can include multiple evaluation dimensions such as sales data, playback volume data, and user sentiment data. These multiple evaluation dimensions reflect the commercial conversion effect, dissemination breadth, and audience sentiment feedback of the advertisement, respectively.

[0075] For example, for a product video ad for a new smartphone, sales data can correspond to the pre-sale order growth rate, play count data can reflect the number of times the ad was exposed across the entire network, and user sentiment data can be obtained through a sentiment analysis model of the comment text.

[0076] At the same time, it is possible to obtain the evaluation index values ​​corresponding to each video main segment cluster group under multiple evaluation dimensions, and to obtain the evaluation index values ​​corresponding to each function demonstration segment cluster group under multiple evaluation dimensions. These index values ​​can be obtained through historical deployment data or simulation tests.

[0077] For example, in smartphone advertising cases, a video main story cluster (such as a group of narrative segments emphasizing camera performance) may have a higher user sentiment value (such as the frequency of positive keywords such as "clear image quality" in user comments), while a feature demonstration cluster (such as a night mode operation demonstration) may correspond to a better playback retention rate.

[0078] S220. Using the sum of the evaluation index values ​​corresponding to the main video segment cluster group under multiple evaluation dimensions and the evaluation index values ​​corresponding to the functional demonstration segment cluster group under multiple evaluation dimensions as the target, determine the target main video segment cluster group among multiple main video segment cluster groups, and determine the target functional demonstration segment cluster group among multiple functional demonstration segment cluster groups.

[0079] In this implementation, when determining the target video mainline segment cluster group and the target function demonstration segment cluster group for determining the target function demonstration script and the target video mainline script, the optimization objective can be to make the sum of the evaluation index values ​​of the video mainline segment cluster group and the sum of the evaluation index values ​​of the function demonstration segment cluster group under multiple evaluation dimensions closest to the target evaluation index value. This allows the target video mainline segment cluster group and the target function demonstration segment cluster group to be determined from multiple video mainline segment cluster groups, ensuring that the target video mainline segment cluster group and the target function demonstration segment cluster group are closest to the target evaluation index value of the target film and television shooting task, thereby improving the degree of conformity between the target function demonstration script, the target video mainline script and the target.

[0080] By determining the target video main line segment cluster group and the target function demonstration segment cluster group, when the target advertisement needs to consider both conversion rate and reputation, the method can select the segment group that best matches the comprehensive indicators (such as the segment set that simultaneously satisfies the predetermined playback threshold and the positive sentiment proportion).

[0081] In S230, a target video main line script corresponding to the target film shooting task is determined by a video main line generation model according to the video main line segments in the multiple promotional videos in the target video main line segment cluster group and the user viewing behavior data corresponding to the video main line segments. A target function demonstration script is determined by a function demonstration generation model according to the target video main line script, the function demonstration segments in the multiple promotional videos in the target function demonstration segment cluster group, and the user viewing behavior data corresponding to the function demonstration segments.

[0082] After determining the target video main line segment cluster group, a target video main line script corresponding to the target film shooting task can be generated by a video main line generation model according to the multiple promotional video segments in the target video main line segment cluster group and the user viewing behavior data thereof.

[0083] Similarly, after determining the target function demonstration segment cluster group, a target function demonstration script that matches the timeline of the main line script can be generated by a function demonstration generation model in combination with the target video main line script, the segments in the target function demonstration segment cluster group, and the user behavior data thereof.

[0084] The above-mentioned implementation manner has the beneficial effect that, through the target matching mechanism of multi-dimensional evaluation indicators, it can ensure that the selected segment combination simultaneously satisfies the comprehensive requirements of multi-dimensional evaluation indicators such as commercial communication and user acceptance. This method improves the systematicness of advertisement planning and the effectiveness of data support, and avoids the imbalance of communication caused by single-dimensional optimization.

[0085] The above-mentioned implementation manner also has the beneficial effect that, in the deep learning model, the multi-dimensional evaluation results and the association optimization of the target script are integrated, which can generate a film script with highly coordinated technical parameters and user experience. This helps to simultaneously improve the market adaptability and the effect of multi-dimensional evaluation indicators such as emotional resonance of the advertisement while controlling the production cost.

[0086] In some implementation manners, in S220, the sum of the evaluation indicator values respectively corresponding to the video main line segment cluster group in multiple evaluation dimensions and the evaluation indicator values respectively corresponding to the function demonstration segment cluster group in multiple evaluation dimensions is closest to the target evaluation indicator value, the target video main line segment cluster group is determined from the multiple video main line segment cluster groups, and the target function demonstration segment cluster group is determined from the multiple function demonstration segment cluster groups. The following describes S221 to S222.

[0087] S221, determine the dimension weight corresponding to each of the plurality of evaluation dimensions. The sum of the product of the evaluation index value corresponding to each of the plurality of evaluation dimensions of the video main line segment clustering group and the dimension weight and the product of the evaluation index value corresponding to each of the plurality of evaluation dimensions of the function demonstration segment clustering group and the dimension weight is taken as the evaluation index intermediate value.

[0088] In this implementation, in the process of implementing the video shooting plan based on deep learning, the screening of the segment clustering group can be optimized by introducing a differentiated weight mechanism of the evaluation dimension. First, the dimension weight corresponding to each of the plurality of evaluation dimensions can be determined, which can reflect the relative importance of different evaluation indexes in business decision-making.

[0089] For example, when formulating a video advertisement strategy for a smartphone product, the sales data dimension can be given a higher weight to highlight the market conversion target, and the weight configuration of the user sentiment tendency data can strengthen the consideration of user satisfaction.

[0090] After determining the dimension weight corresponding to each of the plurality of evaluation dimensions, after obtaining the evaluation index value of each video main line segment clustering group and function demonstration segment clustering group in the plurality of evaluation dimensions, the sum of the product of the video main line segment clustering group evaluation index value and the corresponding dimension weight can be calculated as the video main line evaluation index intermediate value; and the sum of the product of the function demonstration segment clustering group evaluation index value and the corresponding dimension weight is calculated as the function demonstration evaluation index intermediate value. This weighted calculation method can highlight the contribution value of the key dimension.

[0091] For example, when a certain sports camera advertisement project pays more attention to user reviews, the weight of the user sentiment tendency dimension will be increased, so that the segment group with a high positive comment ratio will be significantly improved.

[0092] S222, determine the target video main line segment clustering group in the plurality of video main line segment clustering groups and the target function demonstration segment clustering group in the plurality of function demonstration segment clustering groups, with the evaluation index intermediate value closest to the target evaluation index value as the target.

[0093] After obtaining the weighted calculation result in S221, the sum of the video main line evaluation index intermediate value and the function demonstration evaluation index intermediate value can be taken as the optimization target, and the final target video main line segment clustering group and the target function demonstration segment clustering group can be determined in the plurality of clustering groups. The mechanism in this implementation can achieve fine balance of multi-dimensional indexes. For example, in the case of an intelligent watch advertisement, even if the play data of a certain segment group is slightly low, but it performs well in the sales data dimension with a higher weight, it can still be selected preferentially.

[0094] The implementation manner has the beneficial effects that, by means of flexible configuration and weighting calculation of the dimension weight, the screening criteria can be customized for different business targets, the matching precision of the advertisement strategy and the core appeal of the enterprise is improved, and the decision deviation caused by equalization of the evaluation dimension importance in the traditional method is effectively avoided.

[0095] The implementation manner also has the beneficial effects that, the calculation manner of the evaluation index intermediate value fully respects the actual influence weight of each dimension index, so that the finally selected fragment cluster combination not only meets the overall requirement of the target evaluation index, but also mainly guarantees that the key dimension index achieves the expected effect, and the market adaptability of the film and television planning scheme is significantly enhanced.

[0096] Figure 5 A flowchart of a fourth film and television shooting planning method based on deep learning provided by the embodiments of the present application is shown in FIG. 6. Figure 5 As shown in FIG. 6, the method further includes S310 to S320, which are specifically described as follows.

[0097] S310, a plurality of promotional videos corresponding to a target product category corresponding to a target film and television shooting task in a preset historical time period are determined, and average play quantity and / or peak play quantity of the plurality of promotional videos on a first broadcast platform is obtained.

[0098] In the implementation process of the film and television shooting planning based on deep learning, the target evaluation index can be accurately set through the historical data backtracking mechanism. First, a plurality of promotional videos corresponding to a target product category corresponding to a target film and television shooting task in a preset historical time period can be determined.

[0099] For example, the plurality of promotional videos as historical video samples can be extracted from an advertisement database or a delivery platform, so as to ensure data timeliness and comparability. For example, when planning an advertisement for a new smart phone, the preset time period can be selected as a promotional period of three months after the release of a previous product of the same brand.

[0100] After obtaining the plurality of promotional videos corresponding to the target product category corresponding to the target film and television shooting task in the preset historical time period, the average play quantity and / or peak play quantity of the plurality of promotional videos on the first broadcast platform can be obtained as core reference data.

[0101] It should be noted that the average play quantity reflects the sustained attraction of the advertisement, and the peak play quantity reflects the explosive propagation effect. For example, the historical advertisement of a certain sports camera brand can show a high peak play quantity on an outdoor sports community platform, and a relatively stable average play quantity on a comprehensive platform.

[0102] S320. Identify multiple target promotional videos in descending order of average and / or peak play counts. Obtain the evaluation index values ​​corresponding to the multiple target promotional videos under multiple evaluation dimensions, and use them as the target evaluation index values ​​corresponding to the target film and television shooting task under multiple evaluation dimensions.

[0103] In this implementation, multiple target promotional videos can be filtered out in descending order of average and / or peak play counts based on play count data sorting logic. The sorting mechanism can automatically identify the best-performing ad samples in history; for example, in the promotion of smart home products, ad videos with the top 20% average play counts can be selected as benchmark cases.

[0104] After obtaining multiple target promotional videos, their performance across multiple evaluation dimensions, such as sales data, play count data, and user sentiment data, can be analyzed to extract evaluation index values, which can then be used as target evaluation index values ​​for the current film and television shooting task.

[0105] For example, sales growth rate data from multiple targeted promotional videos can be set as the target value for a new advertising campaign.

[0106] The beneficial effect of the above implementation method is that by screening the playback volume of historical high-quality advertising samples and extracting multi-dimensional indicators, an objective and quantitative target evaluation system can be established, which further avoids the bias of subjective experience-based decision-making and ensures that the planning scheme always benchmarks the successful models verified by the market.

[0107] The beneficial effect of the above implementation method is that the dynamic indicator setting mechanism based on real playback data enables the target film and television script to simultaneously take into account the needs of wide dissemination and real-time performance. This helps to improve the predictability and replicability of advertising effectiveness while maintaining a competitive advantage in the industry.

[0108] In some implementations, the above method also includes S410 to S420, which are described in detail below.

[0109] S410. Using the video effect evaluation model, determine the video mainline segment evaluation value corresponding to each target video mainline segment in the target video mainline script. Using the video effect evaluation model, determine the function demonstration segment evaluation value corresponding to each target function demonstration segment in the target function demonstration script.

[0110] In this implementation, the arrangement and combination of script segments can be further optimized during the implementation of deep learning-based film and television shooting planning to improve the effectiveness of film and television shooting planning.

[0111] In this implementation, the video effect evaluation model can automatically determine the corresponding video main line segment evaluation value according to the characteristics of each video main line segment in the target video main line script, and the video main line segment evaluation value can reflect the pros and cons of the estimated video narrative effect.

[0112] For example, when making an advertisement for a smart phone, the video effect evaluation model can analyze the rhythm of the product story introduction into the target video main line script and output a corresponding numerical score as the video main line segment evaluation value.

[0113] For example, the video effect evaluation model can be trained by sample video main line segment scripts and sample video main line segment evaluation values in historical data.

[0114] Similarly, the video effect evaluation model can determine the corresponding function demonstration segment evaluation value of the target function demonstration script according to the characteristics of each function demonstration segment in the target function demonstration script, and the target function demonstration script can reflect the strength of the product demonstration effect.

[0115] For example, for a function demonstration segment of a smart watch, the video effect evaluation model can generate a function demonstration segment evaluation value according to the clarity of user operation guidance and the interaction trigger rate.

[0116] S420, determine the sum of each video main line segment evaluation value and the time-sequentially adjacent function demonstration segment evaluation value as the target script evaluation mean. Determine the difference between the target script evaluation means corresponding to the time-sequentially adjacent video main line segment and function demonstration segment as the evaluation mean difference. Reorder the multiple target video main line segments and the multiple target function demonstration segments with the minimum evaluation mean difference corresponding to the time-sequentially adjacent video main line segment and function demonstration segment as the target.

[0117] In this implementation, after obtaining the double evaluation values including the video main line segment evaluation value and the function demonstration segment evaluation value, the sum of the time-sequentially adjacent video main line segment evaluation value and the function demonstration segment evaluation value can be calculated as the target script evaluation mean, and the target script evaluation mean reflects the comprehensive performance of the segment combination.

[0118] Figure 6 A structure diagram of a target video script provided by an embodiment of the present application is shown in FIG. 1. Figure 6As shown, the video mainline segments include X1, X2 and X3, the function demonstration segments include Y1, Y2 and Y3, and when the sum of the evaluation values of the time-sequentially adjacent video mainline segments and the function demonstration segments is calculated, the sum of the evaluation values of the time-sequentially adjacent video mainline segment X1 and the function demonstration segment Y1 can be calculated as the target script evaluation average Z1, and the target script evaluation average Z2 (corresponding to the video mainline segment X2 and the function demonstration segment Y2), the target script evaluation average Z3 (corresponding to the video mainline segment X3 and the function demonstration segment Y3) can be calculated in the same way.

[0119] When the target script evaluation average Z1, the target script evaluation average Z2 and the target script evaluation average Z3 are obtained, the difference between the evaluation value of the time-sequentially adjacent video mainline segment and the corresponding target script evaluation average can be calculated as the evaluation average difference value, and the evaluation average difference value can reflect the difference in performance between the segments.

[0120] Exemplarily, the evaluation average difference value can include the evaluation average difference value Z12 corresponding to the target script evaluation average Z1 and the target script evaluation average Z2, the evaluation average difference value Z23 corresponding to the target script evaluation average Z2 and the target script evaluation average Z3, and the like.

[0121] After obtaining all the evaluation average difference value data, the order of the plurality of video mainline segments and the function demonstration segments in the target video script can be rearranged with the minimum evaluation average difference value of the combination of the time-sequentially adjacent video mainline segment and the function demonstration segment as the optimization target.

[0122] The above-mentioned implementation manner has the beneficial effects that the precise matching of the narrative flow and the function demonstration flow can be realized through the double evaluation mechanism and the difference optimization algorithm, and the non-immediate segment flow alignment manner can significantly improve the information receiving coherence of the audience and enhance the effective communication of the brand information.

[0123] The above-mentioned implementation manner also has the beneficial effects that the dynamic arrangement mechanism of the segments of the target video script in the method respects the quality characteristics of the individual segments and optimizes the performance distribution of the overall script, which helps to maximize the propagation efficiency of the key product information and improve the comprehensive conversion effect of the advertisement placement on the premise of maintaining the integrity of the core content.

[0124] In some implementation manners, the above-mentioned method further includes: when the evaluation average difference value corresponding to the adjacent video mainline segment and the function demonstration segment is greater than or equal to the first evaluation average difference value, determining, by the function demonstration generation model, the target function demonstration script according to the target video mainline script, the video mainline segment evaluation value corresponding to each target video mainline segment in the target video mainline script, the function demonstration segment in the plurality of promotional videos and the user viewing behavior data corresponding to the function demonstration segment.

[0125] In the process of optimizing the shooting plan, an evaluation threshold triggering mechanism can be further set to dynamically adjust the key scripts. When it is detected that the evaluation mean difference corresponding to the adjacent video main line segments and the function demonstration segments in time sequence is greater than or equal to the preset first evaluation threshold, it indicates that the cooperation effect of the two groups of segments does not meet the expected standard.

[0126] For example, when the evaluation mean difference Z12 corresponding to the target script evaluation mean Z1 and the target script evaluation mean Z2, and the evaluation mean difference Z23 corresponding to the target script evaluation mean Z2 and the target script evaluation mean Z3 are greater than or equal to the preset first evaluation threshold, it indicates that the cooperation effect of the first group of video scripts (corresponding to the target script evaluation mean Z1 and the target script evaluation mean Z2) and the second group of video scripts (corresponding to the target script evaluation mean Z2 and the target script evaluation mean Z3) does not meet the expected standard.

[0127] When the cooperation effect of the two groups of segments does not meet the expected standard, the target function demonstration script can be re-optimized by the function demonstration generation model. The function demonstration generation model can comprehensively analyze the target video main line script content, the quality evaluation values of each video main line segment, and all function demonstration segments in the promotion video database and their corresponding user viewing behavior data, to generate a function demonstration scheme with higher matching degree.

[0128] Taking a commodity video advertisement as an example, when generating a function demonstration scheme with higher matching degree, when making an advertisement video for a new energy automobile, the function demonstration generation model can comprehensively analyze the target video main line script content, the quality evaluation values of each video main line segment, and all function demonstration segments in the promotion video database and their corresponding user viewing behavior data, to re-determine the fast charging technology demonstration segment with the highest user playback rate in the historical database to replace and generate.

[0129] The above-mentioned implementation manner has the beneficial effects that by establishing an automatic identification and script calibration mechanism for expressiveness difference, the display effect of the key nodes can be accurately improved while keeping the main line architecture stable. This closed-loop optimization mode effectively avoids the fragmentation of narrative flow and function demonstration flow, and guarantees the overall conveying efficiency of the advertisement information.

[0130] The above-mentioned implementation manner also has the beneficial effects that the function demonstration script is re-generated based on multi-dimensional data, which can realize local reinforcement without damaging the core narrative structure of the advertisement, which helps to solve the expressiveness short board problem in the segment combination and improve the overall quality of the final film under the premise of controlling the production cost.

[0131] In some implementations, the method further includes: when the evaluation mean difference values of the adjacent video mainline segments and the function demonstration segments are greater than or equal to a second evaluation mean difference value, determining, by the video mainline generation model, the target video mainline script corresponding to the target film shooting task according to the video mainline segments in the multiple promotional videos, the video mainline segment evaluation values corresponding to each target video mainline segment in the target video mainline script, and the user viewing behavior data corresponding to the video mainline segments, wherein the second evaluation mean difference value is greater than the first evaluation mean difference value.

[0132] In the optimization of the film shooting plan, a hierarchical triggering mechanism can be further established to deal with different degrees of segment matching problems. When it is detected that the evaluation mean difference values of the adjacent video mainline segments and the function demonstration segments in time sequence are greater than or equal to a preset second evaluation threshold value (the second evaluation mean difference value is greater than the first evaluation mean difference value), it indicates that there is a serious mismatch in the segment combination.

[0133] When there is a serious mismatch in the segment combination, the target video mainline script can be regenerated by the video mainline generation model according to the video mainline segments in the multiple promotional videos, the video mainline segment evaluation values corresponding to each target video mainline segment in the target video mainline script, and the user viewing behavior data corresponding to the video mainline segments, so as to re-determine the target video mainline script corresponding to the target film shooting task, thereby improving the overall effect of the target video mainline script.

[0134] The implementation manner has the beneficial effects that the script reconstruction mechanism for major matching problems can systematically solve the split problem of the core narrative logic and the function demonstration, and significantly improve the completeness and persuasiveness of the key information transmission.

[0135] The implementation manner also has the beneficial effects that the differential response strategy triggered by the hierarchical triggering can reasonably allocate computing resources, locally optimize when only the function segments are mismatched, and globally reconstruct when there is an overall script structural problem, which helps to improve the system running efficiency while ensuring the quality of the advertisement production.

[0136] In some implementations, the method further includes: when the evaluation mean difference values of the adjacent video mainline segments and the function demonstration segments are greater than or equal to a second evaluation mean difference value, rearranging the multiple target video mainline segments and the multiple target function demonstration segments with the maximum evaluation mean difference value of the adjacent video mainline segments and the function demonstration segments in time sequence as the target.

[0137] In the method, in the deepening optimization stage of the film and television shooting plan, a strong difference processing mechanism can be further set to solve the serious segment imbalance problem. When it is detected that the evaluation mean difference corresponding to the video main line segment and the function demonstration segment adjacent in time sequence is greater than or equal to a preset second evaluation threshold, a special segment recombination strategy can be implemented.

[0138] In the implementation of the special segment recombination strategy, the evaluation mean difference of the adjacent segment combination can be maximized as the optimization direction, the playing order of the plurality of video main line segments and function demonstration segments is rearranged, and then the position marker is actively found at the segment combination point with the most significant expressiveness difference, and then the playing sequence is reorganized to arrange the maximum difference point at the key node.

[0139] The beneficial effects brought by the above implementation manner are that when the evaluation mean difference corresponding to the adjacent video main line segment and function demonstration segment is greater than or equal to the second evaluation mean difference, the strategic layout of the expressiveness difference point is actively strengthened, the information gap tension can be created at the key propagation node, and the reverse difference activation mechanism can effectively improve the memory point of the audience to the core product characteristics and enhance the dramatic expression effect of the advertisement.

[0140] The beneficial effects brought by the above implementation manner are that the matching problem of the evaluation mean difference corresponding to the video main line segment and the function demonstration segment is solved, and at the same time, the potential deficiency is converted into a propagation advantage, and the overall expressiveness of the advertisement is improved.

[0141] The embodiment of the application also provides a film and television shooting planning system based on deep learning.

[0142] Figure 7 A logical structure schematic diagram of a film and television shooting planning system based on deep learning provided by the embodiment of the application is shown in FIG. 1. Figure 7 As shown in the figure, the system 1 of the embodiment includes a processing unit 11, a storage unit 12 and a transceiver unit 13, the processing unit 11 is used for processing data, the storage unit 12 is used for storing data, and the transceiver unit 13 is used for transceiving data. The processing unit 11, the storage unit 12 and the transceiver unit 13 cooperate with each other to realize the above method. The beneficial effects of the embodiment of the application have been described in the above method, and will not be described here.

[0143] It should be noted that the information interaction, execution process and the like between the above devices / units are based on the same concept as the method embodiments of the application, and the specific functions and the technical effects brought by them can be referred to the method embodiments part. Here, it will not be described again.

[0144] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above functional units and modules is taken as an example, and in actual application, the above functions can be completed by different functional units and modules according to needs, that is, the internal structure of the apparatus is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit, and the integrated unit can be realized in the form of hardware or in the form of software functional unit. In addition, the specific names of each functional unit and module are only for easy distinction, and do not limit the protection scope of the application. The specific working process of the units and modules in the system can refer to the corresponding process in the foregoing method embodiments, which will not be repeated here.

[0145] The integrated unit, if realized in the form of a software functional unit and sold or used as an independent product, can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the foregoing embodiment methods can be completed by a computer program instructing related hardware, and the computer program can be stored in a computer-readable storage medium. When the processor executes the computer program, the steps of each method embodiment can be implemented. The computer program includes computer program code, which can be in the form of source code, object code, executable file or some intermediate form. The computer-readable medium at least includes any entity or device that can carry the computer program code to the photographing device / terminal equipment, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal and software distribution medium. For example, U disk, mobile hard disk, magnetic disk or optical disk, etc. In some jurisdictions, according to legislation and patent practice, the computer-readable medium cannot be an electrical carrier signal and a telecommunication signal.

[0146] In the above embodiments, the description of each embodiment has its own emphasis, and the parts not described or recorded in detail in a certain embodiment can be referred to the relevant description of other embodiments.

[0147] Those skilled in the art can understand that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be realized in electronic hardware or a combination of computer software and electronic hardware. Whether the functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0148] In the embodiments provided in the present application, it should be understood that the disclosed apparatus and method can be implemented by other ways. For example, the apparatus embodiments described above are only schematic, and for example, the division of the modules or units is only a logical function division, and there can be another division in actual implementation, for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections between different parts can be indirect couplings or communication connections through some interfaces, devices or units, and can be electrical, mechanical or other forms.

[0149] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, that is, they can be located in one place, or can be distributed on multiple network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiments.

[0150] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that the technical solutions recorded in the foregoing embodiments can be modified, or some technical features can be replaced by equivalents; and these modifications or replacements do not make the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application.

Claims

1. A deep learning-based film shooting planning method, characterized in that, The method comprises: acquiring a target commodity category corresponding to a target film and television shooting task, and acquiring a plurality of promotional videos corresponding to the target commodity category, and determining video main line segments and function demonstration segments respectively corresponding in the plurality of promotional videos; acquiring user viewing behavior data respectively corresponding to the video main line segments and the function demonstration segments in each promotional video; wherein the user viewing behavior data comprises video playing operation data and comment data of a user; determining, by a video main line generation model, a target video main line script corresponding to the target film and television shooting task according to the video main line segments in the plurality of promotional videos and the user viewing behavior data corresponding to the video main line segments; determining, by a function demonstration generation model, a target function demonstration script according to the target video main line script, the function demonstration segments in the plurality of promotional videos and the user viewing behavior data corresponding to the function demonstration segments; wherein a time period corresponding to the target function demonstration script is inserted into a time period corresponding to the target video main line script; The method further comprises: respectively clustering the video main line segments of the plurality of promotional videos according to the user viewing behavior data respectively corresponding to the video main line segments, to obtain a plurality of video main line segment clustering groups; obtaining a plurality of function demonstration segment clustering groups according to the user viewing behavior data corresponding to the function demonstration segments of the plurality of promotional videos; determining, respectively, a mean value of the similarity of the commodity categories respectively corresponding to the target commodity category and the plurality of video main line segment clustering groups as a video main line similarity, and determining, respectively, a mean value of the similarity of the commodity categories respectively corresponding to the target commodity category and the plurality of function demonstration segment clustering groups as a function demonstration segment similarity; determining a maximum video main line similarity in the plurality of video main line similarities, and determining a maximum function demonstration similarity in the plurality of function demonstration segment similarities; determining, by the video main line generation model, the target video main line script corresponding to the target film and television shooting task according to the video main line segments in the plurality of promotional videos in the video main line segment clustering group corresponding to the maximum video main line similarity and the user viewing behavior data corresponding to the video main line segments; determining, by the function demonstration generation model, the target function demonstration script according to the target video main line script, the function demonstration segments in the plurality of promotional videos in the function demonstration segment clustering group corresponding to the maximum function demonstration similarity and the user viewing behavior data corresponding to the function demonstration segments.

2. The method of claim 1, wherein the step of creating a storyboard comprises the step of: The method further comprises: ​ acquiring target evaluation index values respectively corresponding to the target film and television shooting task under a plurality of evaluation dimensions; acquiring evaluation index values respectively corresponding to each video main line segment clustering group under the plurality of evaluation dimensions, and acquiring evaluation index values respectively corresponding to each function demonstration segment clustering group under the plurality of evaluation dimensions; wherein the plurality of evaluation dimensions comprise sales data, play count data and user sentiment tendency data; Determine a target video main line script corresponding to the target film shooting task through a video main line generation model according to the video main line segments in the multiple promotion videos in the target video main line segment cluster and the user viewing behavior data corresponding to the video main line segments; and determine a target function demonstration script through a function demonstration generation model according to the target video main line script, the function demonstration segments in the multiple promotion videos in the target function demonstration segment cluster, and the user viewing behavior data corresponding to the function demonstration segments. Determine a target video main line script corresponding to the target film shooting task through a video main line generation model according to the video main line segments in the multiple promotion videos in the target video main line segment cluster and the user viewing behavior data corresponding to the video main line segments; and determine a target function demonstration script through a function demonstration generation model according to the target video main line script, the function demonstration segments in the multiple promotion videos in the target function demonstration segment cluster, and the user viewing behavior data corresponding to the function demonstration segments.

3. The method of claim 2, wherein the step of creating a storyboard comprises the step of: creating a storyboard for the video production based on the selected theme. Determine a target video main line script corresponding to the target film shooting task through a video main line generation model according to the video main line segments in the multiple promotion videos in the target video main line segment cluster and the user viewing behavior data corresponding to the video main line segments; and determine a target function demonstration script through a function demonstration generation model according to the target video main line script, the function demonstration segments in the multiple promotion videos in the target function demonstration segment cluster, and the user viewing behavior data corresponding to the function demonstration segments. Determine a target video main line script corresponding to the target film shooting task through a video main line generation model according to the video main line segments in the multiple promotion videos in the target video main line segment cluster and the user viewing behavior data corresponding to the video main line segments; and determine a target function demonstration script through a function demonstration generation model according to the target video main line script, the function demonstration segments in the multiple promotion videos in the target function demonstration segment cluster, and the user viewing behavior data corresponding to the function demonstration segments. The method further comprises:

4. The method of claim 3, wherein the step of creating a storyboard comprises the step of: creating a storyboard for the video production. Determine multiple promotion videos of the corresponding target commodity category corresponding to the target film shooting task in a preset historical time period, and obtain the average and / or peak play quantity of the multiple promotion videos on the first broadcast platform; Determine multiple target promotion videos in order from high to low according to the average and / or peak play quantity, and obtain the evaluation index values corresponding to the multiple evaluation dimensions of the multiple target promotion videos as the target evaluation index values corresponding to the multiple evaluation dimensions of the target film shooting task. The method further comprises:

5. The method of claim 4, wherein the step of creating a storyboard comprises the step of: Determine a target video main line script corresponding to the target film shooting task through a video main line generation model according to the video main line segments in the multiple promotion videos in the target video main line segment cluster and the user viewing behavior data corresponding to the video main line segments; and determine a target function demonstration script through a function demonstration generation model according to the target video main line script, the function demonstration segments in the multiple promotion videos in the target function demonstration segment cluster, and the user viewing behavior data corresponding to the function demonstration segments. ​ Determine a target video main line script corresponding to the target film shooting task through a video main line generation model according to the video main line segments in the multiple promotion videos in the target video main line segment cluster and the user viewing behavior data corresponding to the video main line segments; and determine a target function demonstration script through a function demonstration generation model according to the target video main line script, the function demonstration segments in the multiple promotion videos in the target function demonstration segment cluster, and the user viewing behavior data corresponding to the function demonstration segments. Determine the sum of each video main line segment evaluation value and the time-sequentially adjacent function demonstration segment evaluation value as the target script evaluation mean value; determine the difference between the target script evaluation mean values corresponding to the time-sequentially adjacent video main line segment and function demonstration segment as the evaluation mean value difference; and rearrange the plurality of target video main line segments and the plurality of target function demonstration segments with the minimum evaluation mean value difference between the time-sequentially adjacent video main line segment and function demonstration segment as the target.

6. The method of claim 5, wherein the step of creating a storyboard comprises the step of: creating a storyboard for the video production based on the selected theme. The method further comprises: When the evaluation mean value difference between the adjacent video main line segment and function demonstration segment is greater than or equal to the first evaluation mean value difference, determining the target function demonstration script by the function demonstration generation model according to the target video main line script, the video main line segment evaluation value corresponding to each target video main line segment in the target video main line script, the function demonstration segment in the plurality of promoted videos, and the user viewing behavior data corresponding to the function demonstration segment.

7. The method of claim 6, wherein the step of creating a storyboard comprises the step of: The method further comprises: ​ When the evaluation mean value difference between the adjacent video main line segment and function demonstration segment is greater than or equal to the second evaluation mean value difference, determining the target video main line script corresponding to the target film shooting task by the video main line generation model according to the video main line segment in the plurality of promoted videos, the video main line segment evaluation value corresponding to each target video main line segment in the target video main line script, and the user viewing behavior data corresponding to the video main line segment; wherein the second evaluation mean value difference is greater than the first evaluation mean value difference.

8. The method of claim 7, wherein the step of creating a storyboard comprises the step of: The method further comprises: ​ When the evaluation mean value difference between the adjacent video main line segment and function demonstration segment is greater than or equal to the second evaluation mean value difference, rearranging the plurality of target video main line segments and the plurality of target function demonstration segments with the maximum evaluation mean value difference between the time-sequentially adjacent video main line segment and function demonstration segment as the target.

Citation Information

Patent Citations

  • Method for generating commodity materials and electronic equipment

    CN118569944A

  • Slice video generation method and device, storage medium and electronic equipment

    CN119893233A