Stroma advertisement video generation method and system, terminal and storage medium
By fusing user interaction data and multi-source narrative data across modalities, narrative advertising videos are generated, solving the problem of inaccurate user profiling in existing technologies. This achieves alignment between the narrative and user needs, as well as logical coherence in the video, thereby improving the effectiveness of advertising.
Patent Information
- Application Number
- CN202610058968.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-16
- Publication Date
- 2026-02-13
- Estimated Expiration
- 2046-01-16
AI Technical Summary
Existing technologies struggle to comprehensively collect multi-platform interaction data from target users when generating narrative advertising videos, and are unable to effectively integrate cross-modal semantics. This results in a lack of accuracy and completeness in user profiles, a disconnect between narrative design and user needs, and an inability for advertisements to resonate with the audience, thus limiting their dissemination effectiveness.
By acquiring native interaction data of target users across multiple platforms, cross-modal semantic fusion is performed to construct user cognitive profiles. Combined with multi-source heterogeneous narrative data from video playback platforms, a narrative knowledge graph is generated to drive plot development, optimize narrative scripts, and achieve multimodal consistent arrangement and automated rendering.
Accurately capture user preferences and emotional inclinations to ensure that the plot direction matches user needs, improve generation efficiency, guarantee the consistency of video style, emotion and logic, and enhance the viewing experience and information delivery effectiveness of advertisements.
Smart Images

Figure CN121531207A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of advertising services, and in particular to a plot advertisement video generation method and system, a terminal and a storage medium. BACKGROUND
[0002] In the prior art, when generating a plot advertisement video, it is difficult to comprehensively collect native interaction data of target users on multiple platforms, and it is also difficult to effectively perform cross-modal semantic fusion on different types of interaction data such as text, vision, and audio, resulting in a lack of accuracy and completeness of the user-related portrait constructed, and the inability to accurately capture the preference dimension and emotional tendency of the user, so that the subsequent plot design is disconnected from the user's needs, the advertisement cannot evoke the user's resonance, and the propagation effect is limited.
[0003] The prior art lacks a perfect narrative knowledge system support, cannot effectively integrate and dynamically optimize multi-source heterogeneous narrative related resources, and the generation of the plot script often relies on manual guidance, which not only has a complicated process and takes a long time, but also is difficult to naturally link the core functional attributes of the product to be promoted, the brand tone, and the plot, and the arrangement of multi-modal materials lacks logical consistency and emotional coherence, resulting in frequent problems such as abrupt transitions and style conflicts, which leads to a lack of watchability and professionalism of the advertisement video, and the inability to fully convey the product value and brand connotation, affecting the promotion effect, so how to improve the generation efficiency of the plot advertisement video has become a problem to be solved. SUMMARY
[0004] The present disclosure provides a plot advertisement video generation method, system, terminal and storage medium.
[0005] In a first aspect, the present disclosure provides a plot advertisement video generation method, comprising: S1, obtaining native interaction data of a target user on multiple platforms, and performing cross-modal semantic fusion on the native interaction data to obtain a user cognitive portrait of the target user; S2, taking multi-source heterogeneous classic narrative texts, commercial advertisement scripts and multi-modal materials in a video playback platform as basic data, taking deep semantic analysis and relationship mining as a construction means, taking user behavior feedback flow and external cultural semantic flow in the video playback platform as evolution basis, and taking incremental correlation learning as an evolution mechanism to obtain a narrative knowledge graph of the video playback platform; S3, based on the user cognitive portrait and the core functional attributes of a product to be promoted in the video playback platform, driving plot deduction of the narrative knowledge graph to obtain a primary narrative script of the video playback platform; S4, taking brand tone constraints and key information points of the product to be promoted as optimization targets, and strengthening plot anchors in the primary narrative script to obtain a branded narrative script of the video playback platform. S5, multi-modal consistency arrangement is carried out on the narrative nodes and emotional logic relations in the brand narrative script, and an executable narrative blueprint of the video playing platform is obtained; S6, according to the timing, superposition relation and transition logic in the executable narrative blueprint, the materials in the multimedia material resource pool of the video playing platform are automatically rendered, and a plot advertisement video of the video playing platform is obtained.
[0006] In a preferred embodiment, the native interaction data of the target user on multiple platforms is obtained, and cross-modal semantic fusion is performed on the native interaction data to obtain a user cognitive portrait of the target user, including: The text comment data, visual focus trajectory data and audio interaction segment data of the target user on multiple platforms are collected in parallel; The explicit interest keywords and implicit emotional tendencies in the text comment data are extracted to obtain a text semantic vector of the target user; The attention area of the interface elements in the visual focus trajectory data is analyzed, and a visual interest vector of the target user is obtained; The audio interaction segment data voice emotional keywords are screened out to obtain an audio semantic vector of the target user; Based on the preset simulation weight strategy, the text semantic vector, the visual interest vector and the audio semantic vector are weighted and spliced to obtain a user feature expression of the target user; Based on the user feature expression, the distribution intensity of the user feature expression in different interest dimensions is analyzed, and a user cognitive portrait of the target user is constructed.
[0007] In a preferred embodiment, the multi-source heterogeneous classic narrative text, commercial advertisement script and multi-modal material in the video playing platform are used as the basic data, deep semantic analysis and relationship mining are used as the construction means, the user behavior feedback flow and external cultural semantic flow in the video playing platform are used as the evolution basis, and incremental correlation learning is used as the evolution mechanism, and a narrative knowledge graph of the video playing platform is obtained, including: The plot structure of the classic narrative text and the commercial advertisement script in the video playing platform is disassembled, the character motivation is identified, and the emotional context is labeled, so as to extract the narrative entities and the logical association between the classic narrative text and the commercial advertisement script, and obtain the narrative logic of the video playing platform; The multi-modal material is aligned with the cross-media content, and the visual key frame, background audio segment and script element corresponding to the narrative logic are bound, and the multi-modal content binding relation of the video playing platform is obtained; The narrative logic is taken as a node, and logical association between the narrative logics and the binding relationship of the multi-modal content are taken as connecting edges, structured assembly is performed, and an initial knowledge graph of the video playing platform is obtained; User behavior feedback flow in the video playing platform is parsed as an optimization guide for a specific logical association path in the initial knowledge graph, and external cultural semantic flow in the video playing platform is parsed as an emerging narrative mode to be merged into the initial knowledge graph; According to the optimization guide, the association strength of the corresponding connecting edge in the initial knowledge graph is iteratively adjusted, and according to the emerging narrative mode, new nodes and connecting edges are extended in the initial knowledge graph, and through continuous iterative adjustment, the initial knowledge graph evolves into a narrative knowledge graph of the video playing platform. The evolution formula of the iterative adjustment of the association strength is as follows: ; In the formula, is a new association strength of the corresponding connecting edge in the initial knowledge graph, is a current association strength of the corresponding connecting edge in the initial knowledge graph, is a user behavior feedback coefficient of the user behavior feedback flow, is a cultural semantic injection coefficient of the external cultural semantic flow, is a preset user feedback adjustment coefficient, is a preset cultural injection adjustment coefficient, is a preset evolution smoothing factor.
[0008] In a preferred embodiment, based on the user cognitive portrait and the core function attribute of the product to be promoted in the video playing platform, the plot deduction of the narrative knowledge graph is driven, a primary narrative script of the video playing platform is obtained, which includes: The preference dimension and the strength weight in the user cognitive portrait are extracted, and the key function point and the value proposition in the core function attribute are extracted, the preference dimension, the strength weight, the key function point and the value proposition are cross-fused, and a joint query vector of the video playing platform is obtained; The joint query vector is mapped to the narrative knowledge graph as an initial probe, multi-step path exploration is performed on the narrative knowledge graph, and a candidate plot development path of the video playing platform is obtained; According to the emotional tendency weight in the user cognitive portrait and the logical completeness requirement of the core function attribute, the candidate plot development path is matched and sorted, and an optimal plot path of the video playing platform is obtained; The node sequence and relationships contained in the optimal plot path are filled in according to a preset script structure template to obtain the primary narrative script of the video playback platform.
[0009] In a preferred embodiment, the step of using the brand tone constraints and key information points of the product to be promoted as optimization targets to strengthen the plot anchors in the primary narrative script, thereby obtaining the branded narrative script of the video playback platform, includes: By deconstructing the brand tone constraints of the product to be promoted, the brand emotional tone and visual style guidelines of the product to be promoted are obtained. The key information points of the product to be promoted are logically sorted to obtain the information embedding priority of the video playback platform. Mark the narrative nodes in the primary narrative script that have a potential connection with the product feature demonstration or value communication as anchor points to be strengthened; The brand emotional tone, visual style guidelines, and information embedding priority are matched and mapped with the anchor points to be strengthened to obtain the brand embedding scheme of the video playback platform. Based on the branding integration scheme, the plot descriptions, character dialogues, or scene settings in the primary narrative script are reconstructed to obtain the branding narrative script of the video playback platform.
[0010] In a preferred embodiment, the step of performing multimodal consistency orchestration on the narrative nodes and emotional logic relationships in the branded narrative script to obtain an executable narrative blueprint for the video playback platform includes: The branded narrative script is structurally analyzed, and the emotional connection and logical progression between narrative nodes in the branded narrative script are marked. Based on the aforementioned emotional connection and logical progression, visual material items, background audio clips, and screen text elements of the narrative node are associated from the multimedia material resource pool of the video playback platform. Check whether the connection between the narrative node and the visual materials, background audio and screen text associated with the adjacent nodes is natural and smooth in terms of style, emotional intensity and rhythm, and mark the connection points in the branded narrative script that have abrupt jumps or contradictions. Based on the connection point, trace back to the emotional logic relationship in the branded narrative script, adjust the multimodal material association scheme of the connection point, and obtain the conflict elimination connection point of the video playback platform; Based on the conflict resolution connection point, the narrative nodes and the temporal and superposition relationships between the narrative nodes are encapsulated according to the logical order of the branded narrative script to obtain the executable narrative blueprint of the video playback platform.
[0011] In a preferred embodiment, the automatic rendering of the multimedia material in the video playing platform according to the timing, superimposition relationship and transition logic in the executable narrative blueprint, to obtain the plot advertisement video of the video playing platform, comprises: extracting the multimedia material identifier in the executable narrative blueprint, the start and end timing of the material in the video stream, and the superimposition level and transition logic between adjacent materials; According to the multimedia material identifier, the original visual segment, audio segment and graphic element of the video playing platform are dispatched from the multimedia material resource pool; According to the start and end timing, the original visual segment and the graphic element are aligned and synthesized according to the superimposition level to obtain the preliminary visual track of the video playing platform, and the corresponding audio segment is aligned and mixed to obtain the audio track of the video playing platform; According to the transition logic, transition effects are inserted between the visual segment in the preliminary visual track and the audio segment in the audio track to obtain the smooth synthesized audio and video track of the video playing platform; Encode and package the smooth synthesized audio and video track to obtain the plot advertisement video of the video playing platform.
[0012] Compared with the prior art, the present application has the following beneficial effects: 1. The present application collects multi-platform user original interactive data in parallel, constructs accurate user cognitive portrait through cross-modal semantic fusion, accurately captures user preference dimension, emotional tendency and intensity weight, and at the same time, based on multi-source heterogeneous narrative related data, combines user behavior feedback and external cultural semantic dynamic evolution narrative knowledge graph, makes the plot deduction have clear user orientation and time adaptation, and ensures that the plot direction is highly consistent with user demand and market trend.
[0013] 2. The present application generates a primary script through the cooperation of user cognitive portrait and product core function attribute, optimizes the plot anchor point with brand tone and key information point, realizes the natural implantation of product value and brand connotation, eliminates material connection conflict through multi-modal consistency arrangement, and completes audio and video synthesis with automatic rendering, which not only improves the generation efficiency of plot advertisement video, but also guarantees the coherence of video in style, emotion and logic, and strengthens the ornamental and information transmission effectiveness of the advertisement. BRIEF DESCRIPTION OF DRAWINGS
[0014] In the following, the present disclosure will be described in more detail based on the embodiments and with reference to the accompanying drawings: Figure 1 A work flow diagram of a plot advertisement video generation method according to an embodiment of the present application is shown; Figure 2 A functional module diagram of a plot advertisement video generation system according to Embodiment Two of the present application is shown; Figure 3 A structural diagram of a terminal for implementing the plot advertisement video generation method according to Embodiment Three of the present application is shown. DETAILED DESCRIPTION
[0015] In order to enable those skilled in the art to better understand the technical solutions of the present disclosure, and to understand the implementation process of how the present disclosure applies technical means to solve technical problems and achieve corresponding technical effects, the technical solutions in the embodiments of the present disclosure will be described clearly and completely below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, not all. The embodiments of the present disclosure and various features in the embodiments can be combined with each other without conflict, and the technical solutions formed thereby are all within the protection scope of the present disclosure. Based on the embodiments in the present disclosure, all other embodiments obtained by those skilled in the art without creative labor should be within the protection scope of the present disclosure.
[0016] It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described herein can be executed in an order different from that shown.
[0017] Embodiment One Figure 1 A flowchart of a plot advertisement video generation method according to an embodiment of the present disclosure is shown. As shown in the flowchart, a plot advertisement video generation method includes: Figure 1 S1, obtaining native interaction data of a target user on multiple platforms, and performing cross-modal semantic fusion on the native interaction data to obtain a user cognitive portrait of the target user; In the embodiment of the present application, the obtaining of the native interaction data of the target user on the multiple platforms and the cross-modal semantic fusion on the native interaction data to obtain the user cognitive portrait of the target user includes: parallelly collecting text comment data, visual focus trajectory data and audio interaction segment data of the target user on the multiple platforms; extracting explicit interest keywords and implicit sentiment tendencies in the text comment data to obtain a text semantic vector of the target user; performing attention area analysis on interface elements that are continuously stayed and quickly swept in the visual focus trajectory data to obtain a visual interest vector of the target user; Screening the audio interaction segment data voice emotion keywords to obtain the audio semantic vector of the target user; Based on the preset simulation weight strategy, the text semantic vector, the visual interest vector and the audio semantic vector are weighted and spliced to obtain the user feature expression of the target user. Based on the user feature expression, the distribution intensity of the user feature expression in different interest dimensions is analyzed, and the user cognitive portrait of the target user is constructed.
[0018] Through the data interface opened by each platform, the text comment content published by the target user on multiple platforms such as social media, video and shopping is synchronously acquired, and the position, duration data and element information of the interface region corresponding to the rapid swipe of the user's eyes when browsing the interface of each platform are recorded. In addition, complete audio segments generated by the user in the platform voice interaction scene are also collected, realizing the synchronous acquisition of text comment data, visual focus track data and audio interaction segment data.
[0019] The text comment data of the target user is analyzed sentence by sentence, the explicit interest keywords are determined by identifying the repeatedly appearing and semantic core words in the text, and the implicit emotional tendency behind the text is judged by analyzing the emotional color of the words and the tone intensity of the sentence. These explicit interest keywords and implicit emotional tendencies are converted into structured text semantic vectors according to the preset semantic coding rules.
[0020] The visual focus track data collected is analyzed frame by frame, the interface elements with a user stay duration exceeding the set standard are marked as continuous stay elements, and the interface elements with a user stay duration below the set standard are marked as rapid swipe elements. The content type, presentation form and functional attribute of these elements are analyzed in depth, and the visual features that the user focuses on are extracted, and these visual features are organized into visual interest vectors according to a unified dimension.
[0021] The audio interaction segment data is processed by speech-to-text, the words expressing emotions in the converted text are accurately identified, the redundant words without emotional direction and meaning are removed, the core voice emotion keywords reflecting the user's emotional state are screened out, and these keywords are converted into audio semantic vectors with semantic direction according to the classification standard of voice emotion.
[0022] According to the importance of text, visual and audio data in depicting the user's cognitive process, a fixed weight distribution ratio is set as a simulation weight strategy, the text semantic vector, the visual interest vector and the audio semantic vector are respectively given weights according to the ratio, and then the three vectors after weight assignment are spliced in order from text to visual to audio, and integrated to form a user feature expression containing multi-dimensional features of the user.
[0023] The user feature expression is mapped into a preset interest dimension system, the feature highlighting degree, i.e., distribution intensity, of the user feature expression on each interest dimension is analyzed one by one, the interest levels of the user in various fields are determined by comparing the distribution intensities of different dimensions, the core preference, potential demand and emotional appeal of the user are sorted in combination with the emotional tendency feature of the user, and finally a comprehensive and accurate user cognitive portrait is constructed.
[0024] The beneficial effects are that three types of original interactive data of text comments, visual focus trajectories and audio interaction fragments of users on multiple platforms are collected in parallel, multi-dimensional interactive information of users is comprehensively covered, explicit interest keywords and implicit emotional tendencies in the text are extracted, attention areas in the visual focus trajectories are analyzed, and voice emotional keywords in the audio are screened, corresponding semantic vectors and interest vectors are formed, a preset simulation weight strategy is combined to realize weighted splicing of the three types of vectors, cross-modal information is accurately fused to form complete user feature expressions, the distribution intensity of the feature expressions on different interest dimensions is analyzed to construct a user cognitive portrait with comprehensiveness and accuracy, and user preferences and emotional tendencies are fully captured to provide a solid user basis for accurate generation of subsequent plot advertisements.
[0025] S2, based on the multi-source heterogeneous classic narrative text, commercial advertisement script and multi-modal material in the video playing platform, taking deep semantic analysis and relationship mining as a construction means, taking the user behavior feedback flow and external cultural semantic flow in the video playing platform as evolution basis, taking incremental correlation learning as evolution mechanism, obtaining the narrative knowledge graph of the video playing platform; In the embodiment of the present application, the narrative knowledge graph of the video playing platform obtained based on the multi-source heterogeneous classic narrative text, commercial advertisement script and multi-modal material in the video playing platform, taking deep semantic analysis and relationship mining as a construction means, taking the user behavior feedback flow and external cultural semantic flow in the video playing platform as evolution basis, taking incremental correlation learning as evolution mechanism, includes: The classic narrative text and the commercial advertisement script in the video playing platform are subjected to plot structure disassembly, character motivation identification and emotional context labeling, so as to extract narrative entities and logical associations between the classic narrative text and the commercial advertisement script, and obtain the narrative logic of the video playing platform; The multi-modal material is subjected to cross-media content alignment, and the visual key frame, background audio segment and script element corresponding to the narrative logic are bound, and the multi-modal content binding relationship of the video playing platform is obtained; The narrative logic is taken as a node, and the logical association between the narrative logics and the multi-modal content binding relationship are taken as connection edges, and structured assembly is performed, and the initial knowledge graph of the video playing platform is obtained; user behavior feedback stream in the video playing platform is parsed as an optimization guide for a specific logical association path in the initial knowledge graph, and an external cultural semantic stream in the video playing platform is parsed as an emerging narrative mode to be integrated into the initial knowledge graph; According to the optimization guide, the association strength of the corresponding connection edge in the initial knowledge graph is iteratively adjusted, and according to the emerging narrative mode, new nodes and connection edges are extended in the initial knowledge graph, and through continuous iterative adjustment, the initial knowledge graph evolves into a narrative knowledge graph of the video playing platform. The evolution formula for iteratively adjusting the association strength is as follows: ; In the formula, is the new association strength of the corresponding connection edge in the initial knowledge graph, is the current association strength of the corresponding connection edge in the initial knowledge graph, is a user behavior feedback coefficient of the user behavior feedback stream, is a cultural semantic injection coefficient of the external cultural semantic stream, is a preset user feedback adjustment coefficient, is a preset cultural injection adjustment coefficient, is a preset evolution smoothing factor.
[0026] The classic narrative texts and commercial advertisement scripts in the video playing platform are comprehensively combed, and complete structures such as beginning, development, climax, and ending are split according to the natural context of plot development. The internal motivation of the characters is clarified by analyzing the behavior choices and target appeals of the characters in the plot, and the ups and downs of emotions in the texts and scripts are tracked and clearly labeled section by section. From these analysis results, core narrative entities such as characters, scenes, and events are extracted, and the causes and effects, progression, and parallelism of the associations between different texts and scripts in plot advancement, character setting, and emotional expression are combed. Finally, the narrative logic of the video playing platform is formed.
[0027] Image, audio, text and other multi-modal materials in the video playing platform are collected, cross-media content alignment is achieved by comparing the semantic content and time dimension association of different materials, and it is ensured that different types of materials corresponding to the same theme or plot can be accurately matched. Then, against the constructed narrative logic, visual key frames that fit each narrative logic segment are selected one by one. These key frames need to intuitively present the core information of the corresponding plot, background audio segments that meet the emotional tone of the narrative logic are selected, and copy elements that can strengthen the narrative expression are extracted. The selected visual key frames, background audio segments, and copy elements are one-to-one bound with the corresponding narrative logic to form the multi-modal content binding relationship of the video playing platform.
[0028] The determined narrative logic is taken as an independent node, the various logical associations between the narrative logics and the multi-modal content binding relationships between the narrative logics and the multi-modal materials are taken as connecting edges connecting different nodes, and the nodes are classified and the edge attributes are marked in a structured manner. All nodes and connecting edges are organized and assembled according to a unified structure specification to construct an initial knowledge graph containing narrative logic nodes, association edges and multi-modal material binding information, so that the initial knowledge graph can completely present the association relationships between various elements.
[0029] The behavior data of users in the video playing platform, such as watching time, interaction frequency, comment tendency and the like, are collected to form a user behavior feedback stream. The behavior data are classified and summarized to determine the preference degree of users for different narrative logic association paths. The preference degree is converted into a strengthening or weakening indication for a specific logic association path in the initial knowledge graph. Meanwhile, the popular cultural trends, mass aesthetic tendencies, language expression styles and the like in the society are collected to form an external cultural semantic stream. The external cultural semantic stream is deeply interpreted to extract emerging narrative modes conforming to the current communication environment, and the specific directions of the modes that can be integrated into the initial knowledge graph are determined.
[0030] According to the optimization indication converted from the user behavior feedback stream, the association strength of the connecting edges corresponding to the logic association paths in the initial knowledge graph is adjusted according to the user preference degree. The connecting edges corresponding to the paths with high user preference degree are strengthened, and the connecting edges corresponding to the paths with low user preference degree are appropriately weakened. Meanwhile, the emerging narrative modes obtained by analysis are converted into new narrative logic nodes, and corresponding multi-modal content binding relationships are added to the new nodes as connecting edges to be integrated into the initial knowledge graph. The operation of adjusting the association strength and expanding the nodes and edges is continuously repeated, so that the initial knowledge graph is continuously adapted to the user demand and the cultural trend, and finally evolves into a narrative knowledge graph of the video playing platform.
[0031] The current association strength of the connecting edges in the initial knowledge graph is the initial strength value of the connecting edges directly given when the initial knowledge graph is constructed according to the logical association tightness between the narrative logics and the matching degree of the multi-modal content binding relationships.
[0032] The user behavior feedback coefficient is obtained by analyzing the user behavior feedback stream in the video playing platform. The behavior data of users, such as watching time, interaction frequency, comment tendency and the like, generated by users for different narrative content corresponding to the logic association paths in the initial knowledge graph are collected. After the data are classified and summarized, specific numerical values reflecting the user preference degree are obtained.
[0033] The cultural semantic injection coefficient is derived from the interpretation of external cultural semantic flow, collecting popular cultural trends, mass aesthetic tendencies, language expression styles and other related information at the social level, and after deep analysis and refinement, it is converted into specific numerical values that reflect the integration of emerging narrative mode values.
[0034] The user feedback adjustment coefficient is a fixed value set in advance according to the actual needs of the narrative knowledge graph construction, combined with the influence experience of past user behavior feedback on the adjustment of association strength, used to control the degree of influence of the user behavior feedback coefficient on the adjustment of association strength.
[0035] The culture injection adjustment coefficient is a fixed value set in advance based on the rationality requirements of external cultural semantic flow integration, referring to the influence effect of historical cultural trends on the evolution of the narrative graph, used to regulate the intervention intensity of the cultural semantic injection coefficient on the adjustment of association strength.
[0036] The evolution smoothing factor is a fixed value set in advance to avoid large fluctuations in the adjustment process of association strength, according to the structural characteristics of the initial knowledge graph and the evolution stability requirements, used to balance the influence of the difference between user behavior feedback and external cultural semantic flow on association strength.
[0037] The significance of the formula is to provide a clear basis for adjusting the association strength of the connection edges in the initial knowledge graph, by integrating the dual influence of user behavior feedback and external cultural semantic flow, to achieve dynamic optimization of association strength.
[0038] The formula combines the current association strength with the adjustment amplitude corresponding to user behavior feedback and the adjustment amplitude corresponding to external cultural semantics, and then balances the difference between the two through smoothing processing, finally obtains the new association strength, ensuring that the adjusted association strength can reflect user preferences and adapt to emerging narrative patterns, allowing the initial knowledge graph to gradually evolve into a narrative knowledge graph that meets user needs and the trend of the times through continuous iterative adjustment, providing accurate knowledge support for subsequent plot development.
[0039] The beneficial effect is that based on multi-source heterogeneous classic narrative text, commercial advertisement script and multi-modal material, through operation such as plot structure disassembly, character motive recognition, emotion context annotation and cross-media content alignment, narrative entities, logical association and corresponding multi-modal material are accurately extracted, bound and assembled to form an initial knowledge graph with clear logic and complete content, and at the same time, the optimization guide of user behavior feedback circulation and the emerging narrative mode of external cultural semantic stream extraction are integrated, the initial knowledge graph is continuously adapted to user preferences and cultural trends by iteratively adjusting the connection edge association strength and expanding new nodes and new connection edges, and finally a narrative knowledge graph with narrative integrity, user adaptability and timeliness is formed, which provides comprehensive and dynamic knowledge support for the accurate deduction of subsequent plot advertisement video, and ensures that the plot development conforms to the narrative law and meets market demand.
[0040] S3, based on the user cognitive portrait and the core function attribute of the product to be promoted in the video playback platform, driving the plot deduction of the narrative knowledge graph to obtain the primary narrative script of the video playback platform; In the embodiment of the application, based on the user cognitive portrait and the core function attribute of the product to be promoted in the video playback platform, driving the plot deduction of the narrative knowledge graph to obtain the primary narrative script of the video playback platform, comprising: extracting the preference dimension and intensity weight in the user cognitive portrait, and extracting the key function point and value proposition in the core function attribute, cross-fusing the preference dimension, the intensity weight and the key function point, and the value proposition to obtain the joint query vector of the video playback platform; mapping the joint query vector to the narrative knowledge graph as an initial probe, performing multi-step path exploration on the narrative knowledge graph to obtain the candidate plot development path of the video playback platform; According to the emotional tendency weight in the user cognitive portrait and the logical completeness requirement of the core function attribute, the matching degree of the candidate plot development path is sorted to obtain the optimal plot path of the video playback platform; filling the node sequence and the relationship contained in the optimal plot path according to the preset script structured template to obtain the primary narrative script of the video playback platform.
[0041] From the constructed user cognitive portrait, the system sorts out the specific direction of the user in each interest field, and determines the corresponding intensity weight according to the prominence of each preference dimension in the portrait. Meanwhile, from the core functional attributes of the product to be promoted, the core action points that can solve the user's needs are extracted as key functional points, and the core benefits brought to the user by the product are summarized as value propositions. These preference dimensions, intensity weights, key functional points, and value propositions are matched according to semantic association. Through the integration of the core information of each feature, a joint query vector is formed that can reflect both user needs and product value.
[0042] According to the semantic characteristics and dimension distribution of the joint query vector, find the initial node that matches the semantics in the narrative knowledge graph. Take this initial node as the exploration starting point, i.e., the initial probe. From this node, explore the subsequent nodes associated with it along the connection edges between the narrative logics in the narrative knowledge graph, continuously expand the exploration range, traverse all possible plot development directions, collect each complete node connection sequence, and form multiple candidate plot development paths with logical coherence.
[0043] Extract the user's emphasis on different emotional types from the user cognitive portrait as emotional tendency weights, and explicitly define the core requirements that the core functional attributes of the product to be promoted must fully present, i.e., the logical completeness requirements. Analyze each candidate plot development path one by one, judge the degree of fit between the plot content and emotional tendency weights contained in the path, and check whether the path can fully cover the logical points of the product core functional attributes. According to the comprehensive evaluation of the matching degree of each path based on the degree of fit and logical coverage, arrange the candidate plot development paths in order from high to low. The path ranked first is the optimal plot path.
[0044] Pre-set a script structured template containing fixed modules such as scene description, character behavior, plot advancement, and logical transition. Arrange the narrative logic nodes in the optimal plot path in sequence as a node sequence, sort out the logical associations such as causality and progression between nodes, and fill in the corresponding plot content and logical relationships between nodes in the template according to the content requirements of each module in the template. Ensure that the filled content structure is clear and logically coherent, and finally form a primary narrative script.
[0045] The beneficial effect is that by extracting the preference dimension, intensity weight in the user cognitive portrait, the key function points and value proposition of the product to be promoted are cross-fused to form a joint query vector which can simultaneously consider user demand and product value. The vector provides accurate guidance for plot deduction of the narrative knowledge graph. Through multi-step path exploration, multiple candidate plot development paths with logical coherence are obtained. Then, according to the user emotional tendency weight and the logical completeness requirement of the product core function attribute, the matching degree is sorted to ensure that the optimal plot path selected not only fits the user preference but also fully presents the core value of the product. Finally, the node sequence and relationship of the optimal path are filled in the preset script structured template to generate a primary narrative script with structured specification and clear logic, which provides a solid foundation for subsequent brand optimization and ensures that the generated plot advertisement video meets user expectations and accurately delivers the core information of the product.
[0046] S4, taking the brand tone constraint and key information points of the product to be promoted as optimization targets, reinforcing the plot anchor points in the primary narrative script, to obtain a brand narrative script of the video playback platform; In the embodiment of the present application, taking the brand tone constraint and key information points of the product to be promoted as optimization targets, reinforcing the plot anchor points in the primary narrative script, to obtain a brand narrative script of the video playback platform, comprises: deconstructing the brand tone constraint of the product to be promoted to obtain the brand emotional tone and visual style guide of the product to be promoted; logically sorting the key information points of the product to be promoted to obtain the information implantation priority of the video playback platform; marking the narrative nodes in the primary narrative script that have potential correlation with product function display or value delivery as to-be-reinforced anchor points; mapping the brand emotional tone, the visual style guide and the information implantation priority to the to-be-reinforced anchor points to obtain a brand implantation scheme of the video playback platform; reconstructing the text of the plot description, role dialogue or scene setting in the primary narrative script according to the brand implantation scheme to obtain a brand narrative script of the video playback platform.
[0047] comprehensively disassembling the brand tone constraint of the product to be promoted, sorting out the core requirements for emotional expression, clarifying the core emotional tendency that the brand wants to convey to the user, forming the brand emotional tone of the product to be promoted, and analyzing the fixed specifications and preferences of the brand in visual presentation, including color matching, picture texture, element type and other requirements, to extract the visual style guide of the product to be promoted which can guide the visual presentation of the script.
[0048] Collect all key information points of the product to be promoted, judge the importance of each key information point to the user's understanding of the core advantages of the product and the internal correlation between the information points from the user's cognitive logic and the product value transmission logic, arrange these key information points in order of importance from high to low and correlation logic from main to secondary, determine the implantation sequence of each information point, and obtain the information implantation priority of the video playback platform.
[0049] Review each narrative node in the primary narrative script one by one, analyze the plot content, scene setting and function positioning of each node, and determine whether the node can provide scene support for the natural display of product functions or whether it can carry the transmission needs of product value. As long as there is a potential correlation between the node and the product function display or value transmission, the narrative node is marked as an anchor point to be strengthened.
[0050] For each anchor point to be strengthened, match the corresponding brand emotional tone based on its corresponding plot scene and function positioning, ensure that the emotional expression at the anchor point is consistent with the overall brand emotion, and determine the visual presentation requirements of the scene and picture at the anchor point according to the visual style guide. Then, assign each anchor point to be strengthened with the corresponding key information point according to the information implantation priority, and determine the specific way and presentation focus of information implantation. Integrate all matching results to obtain the brand implantation scheme of the video playback platform.
[0051] According to the specific requirements of the brand implantation scheme, optimize the plot description in the primary narrative script, adjust the language style to fit the brand emotional tone, supplement details to make the visual style guide land on the text level, adjust the character dialogue to naturally integrate the assigned key information points without damaging the coherence and authenticity of the dialogue, and perfect the scene setting according to the visual style guide to clearly determine the elements and colors in the scene. Through systematic text reconstruction of plot description, character dialogue or scene setting, the brand narrative script of the video playback platform is obtained.
[0052] The beneficial effects are that by deconstructing the brand tone constraints, determining the brand emotional tone and visual style guide, and logically sorting the key information points to determine the information implantation priority, a clear direction is provided for brand element integration. By marking the anchor points to be strengthened in the primary narrative script related to product function and value transmission, and forming a precise brand implantation scheme through adaptive mapping, text reconstruction is performed on plot description, character dialogue or scene setting according to the scheme, so that the brand narrative script not only maintains the logical coherence of the original plot, but also naturally carries the brand tone and product key information, achieving precise transmission of brand connotation and product value, improving the brand recognition and information transmission efficiency of the plot advertisement, and laying a solid foundation for subsequent multi-modal arrangement.
[0053] S5. Perform multimodal consistency arrangement of the narrative nodes and emotional logic relationships in the branded narrative script to obtain the executable narrative blueprint of the video playback platform; In this embodiment of the invention, the step of performing multimodal consistency orchestration on the narrative nodes and emotional logic relationships in the branded narrative script to obtain an executable narrative blueprint for the video playback platform includes: The branded narrative script is structurally analyzed, and the emotional connection and logical progression between narrative nodes in the branded narrative script are marked. Based on the aforementioned emotional connection and logical progression, visual material items, background audio clips, and screen text elements of the narrative node are associated from the multimedia material resource pool of the video playback platform. Check whether the connection between the narrative node and the visual materials, background audio and screen text associated with the adjacent nodes is natural and smooth in terms of style, emotional intensity and rhythm, and mark the connection points in the branded narrative script that have abrupt jumps or contradictions. Based on the connection point, trace back to the emotional logic relationship in the branded narrative script, adjust the multimodal material association scheme of the connection point, and obtain the conflict elimination connection point of the video playback platform; Based on the conflict resolution connection point, the narrative nodes and the temporal and superposition relationships between the narrative nodes are encapsulated according to the logical order of the branded narrative script to obtain the executable narrative blueprint of the video playback platform.
[0054] The overall structure of the branded narrative script is deconstructed, and different types of narrative nodes such as scene transitions, character interactions, and plot progression are divided. The emotional connection between each narrative node and the nodes before and after it is analyzed segment by segment. It is clarified whether the emotions are a continuation, a turning point, or a sublimation. At the same time, the logical connections between nodes in terms of cause and effect, parallelism, or progression in the plot development are sorted out. These emotional connections and logical progressions are marked one by one at the corresponding narrative node connection positions.
[0055] Based on the labeled emotional sequence and logical progression, a precise search is conducted in the multimedia material resource pool of the video playback platform. For each narrative node, visual material items that are consistent with the emotional tone of the node and meet the needs of logical progression are selected. These visual materials must be able to intuitively present the plot content corresponding to the node. Background audio clips that match the rhythm of emotional sequence are selected to ensure that the emotional direction of the audio is synchronized with the emotional changes in the plot. Screen text elements that can strengthen the core information of the node and fit the logical expression are extracted. A unique association is established between the three types of materials and the corresponding narrative nodes.
[0056] Compare and verify the visual material, background audio and screen text associated with each adjacent narrative node one by one, judge whether the visual material is naturally connected in color style and picture texture, whether the background audio is smoothly transitioned in tone, rhythm, whether the screen text is consistent in language style and information density, and whether the emotional intensity of the three types of materials is coordinated. If there is a style conflict, emotional abruptness or rhythm imbalance at a certain connection position, mark it as a connection point with abrupt changes or contradictions.
[0057] For the marked connection point, trace back to the corresponding emotional connection and logical progression relationship in the brand narrative script, determine the emotional trend and logical context that should be followed at this position, and re-examine the current associated multi-modal materials. If the conflict is caused by incompatible material style, replace the material that is consistent with the emotional logic. If the connection is abrupt due to improper use of material segments, extract the segment that meets the connection requirements from the material. By adjusting the material association scheme, the conflict problem of the connection point is completely solved, forming a conflict-eliminated connection point.
[0058] Based on the conflict-eliminated connection point, arrange all narrative nodes in the original logical order of the brand narrative script, determine the start and end order and time length ratio of each narrative node in the overall plot, determine the superposition relationship between the visual material, background audio and screen text corresponding to different narrative nodes, such as synchronous display of part of the text and picture, and use of specific audio with multiple consecutive nodes, etc. Systematically integrate and package these node sequences, time sequence relationships and superposition relationships to form an executable narrative blueprint that can be directly used for subsequent rendering.
[0059] The beneficial effects are that by structurally analyzing the brand narrative script and marking the emotional connection and logical progression relationship between the narrative nodes, a clear basis is provided for multi-modal material association. Based on this relationship, visual material items, background audio segments and screen text elements are accurately associated from a multi-media material resource pool. By checking the connection of adjacent node materials in style, emotional intensity and rhythm, and eliminating conflicts, the consistency and smoothness of multi-modal material arrangement are ensured. Finally, the narrative nodes and their time sequence and superposition relationships are packaged in the order of the script logic to form an executable narrative blueprint that not only retains the core connotation and logical context of the brand script, but also provides clear and accurate execution instructions for subsequent automated rendering, ensuring the coherence of the drama advertisement video in emotional expression, style presentation and logical progression, and improving the presentation quality of the final video.
[0060] S6、According to the time sequence, superposition relationship and transition logic in the executable narrative blueprint, automatically render the materials in the multi-media material resource pool of the video playback platform to obtain the drama advertisement video of the video playback platform.
[0061] In the embodiment of the present application, the video play platform's plot advertisement video is obtained by automatically rendering the multimedia material resource pool in the video play platform according to the timing, superposition relationship and transition logic in the executable narrative blueprint, comprising: extracting the multimedia material identifier in the executable narrative blueprint, the start and end timing of the material in the video stream, and the superposition level and transition logic between adjacent materials; According to the multimedia material identifier, the original visual segment, audio segment and graphic element of the video play platform are dispatched from the multimedia material resource pool; According to the start and end timing, the original visual segment and the graphic element are aligned and synthesized according to the superposition level to obtain the preliminary visual track of the video play platform, and the corresponding audio segment is aligned and mixed synchronously to obtain the audio track of the video play platform; According to the transition logic, transition effects are inserted between the visual segment in the preliminary visual track and the audio segment in the audio track to obtain the smooth synthesized audio and video track of the video play platform; The smooth synthesized audio and video track is encoded and packaged to obtain the plot advertisement video of the video play platform.
[0062] The structured packaging content of the executable narrative blueprint is parsed one by one, the unique identifier of each multimedia material corresponding to the narrative node is identified, the start and end presentation time of each material in the overall video stream is determined, the upper and lower level relationship of adjacent materials in picture display is combed, the transition type and execution mode between different material segments are determined, and the multimedia material identifier, material start and end timing, adjacent material superposition level and transition logic are completely extracted.
[0063] In the multimedia material resource pool of the video play platform, each original visual segment, audio segment and graphic element corresponds to a unique multimedia material identifier. The extracted material identifier is accurately matched with the identifier of the material in the resource pool, and the corresponding original visual segment, audio segment and graphic element are directly dispatched according to the matching result, so as to ensure that the dispatched material is completely consistent with the material specified in the executable narrative blueprint.
[0064] The start and end timing in the executable narrative blueprint is taken as the unified time reference, the dispatched original visual segment is arranged in time sequence, and the corresponding graphic element is placed in the specified picture level of the visual segment according to the superposition level requirement. The upper element covers the corresponding area of the lower element according to the setting, the alignment and synthesis of all visual materials are completed, and the preliminary visual track is formed. At the same time, the corresponding audio segment is aligned and synchronized with the time axis of the visual track according to the start and end timing, the volume intensity of each audio segment is adjusted, the sound is clear and the volume is balanced after the mixing of different audio segments, and the audio track is obtained.
[0065] Referring to the transition logic in the executable narrative blueprint, the position and type of each transition are clearly defined. At the junctions of adjacent visual segments in the initial visual track and adjacent audio segments in the audio track, corresponding transition effects are inserted respectively. The visual transition effects must ensure natural scene switching, and the audio transition effects must achieve smooth sound transition. At the same time, it is ensured that the visual transitions and audio transitions are completely synchronized in time, ultimately forming a smooth synthesized audio and video track.
[0066] Select an encoding format that meets the compatibility requirements of the video playback platform, convert the audio and video data in the smooth synthesized audio and video track, integrate the converted audio and video data according to a unified standard, add necessary video metadata information, and then package the integrated audio and video data into a complete video file through a packaging process to ensure that the video file can be loaded and played normally on the target video playback platform, and finally obtain the plot advertisement video of the video playback platform.
[0067] The beneficial effects are that by accurately extracting multimedia material identifiers, start and end sequences, overlay levels, and transition logic from the executable narrative blueprint, a clear basis is provided for material scheduling and compositing. This ensures that the original visual segments, audio segments, and graphic elements scheduled from the multimedia material resource pool are completely matched with the blueprint requirements. The initial visual and audio tracks are synthesized synchronously according to the sequence and overlay levels, ensuring the time synchronization and hierarchical rationality of the audio and video content. Then, transition effects are inserted according to the transition logic, making the audio and video segments connect naturally and smoothly. Finally, through encoding and encapsulation, a playable storyline advertisement video is formed. The entire process automates and standardizes material rendering, which not only improves the generation efficiency of storyline advertisement videos, but also ensures the integrity and smoothness of the video in terms of picture, sound, and transition presentation, ensuring the accurate delivery of the product's core information and brand connotation.
[0068] Example 2 like Figure 2 As shown in the figure, this embodiment also provides a functional module diagram of a story advertising video generation system.
[0069] The narrative advertising video generation system 100 described in this embodiment can be installed on a terminal. Depending on the functions implemented, the narrative advertising video generation system 100 may include a user profile construction module 101, a narrative graph construction module 102, a primary script deduction and generation module 103, a brand script optimization module 104, a narrative blueprint arrangement module 105, and an automated advertising video rendering module 106. The modules described in this invention can also be referred to as units, which are a series of computer program segments that can be executed by the terminal processor and perform fixed functions, stored in the terminal's memory.
[0070] In this embodiment, the functions of each module / unit are as follows: The user portrait construction module 101 is configured to acquire native interaction data of a target user on multiple platforms, and perform cross-modal semantic fusion on the native interaction data to obtain a user cognitive portrait of the target user. The narrative graph construction module 102 is configured to take, as basic data, a plurality of source heterogeneous classic narrative texts, commercial advertisement scripts and multi-modal materials in a video playing platform, take deep semantic analysis and relationship mining as a construction means, take user behavior feedback flow and external cultural semantic flow in the video playing platform as evolution basis, and take incremental correlation learning as an evolution mechanism to obtain a narrative knowledge graph of the video playing platform. The primary script deduction generation module 103 is configured to drive plot deduction of the narrative knowledge graph based on the user cognitive portrait and core function attributes of a product to be promoted in the video playing platform to obtain a primary narrative script of the video playing platform. The brand script optimization module 104 is configured to take brand tonality constraints and key information points of the product to be promoted as optimization targets, and strengthen plot anchors in the primary narrative script to obtain a branded narrative script of the video playing platform. The narrative blueprint arrangement module 105 is configured to perform multi-modal consistency arrangement on narrative nodes and emotional logic relationships in the branded narrative script to obtain an executable narrative blueprint of the video playing platform. The advertisement video automatic rendering module 106 is configured to automatically render materials in a multimedia material resource pool of the video playing platform according to timing, superposition relationships and transition logic in the executable narrative blueprint to obtain a plot advertisement video of the video playing platform.
[0071] In detail, each module in the plot advertisement video generation system 100 in the embodiment of the present application uses the same technical means as the plot advertisement video generation method in Embodiment One and Embodiment Two when in use, and can produce the same technical effects, which will not be repeated here.
[0072] Embodiment Three As shown in Figure 3 The present embodiment also provides a computer terminal, which can include a processor 10, a memory 11, a communication bus 12 and a communication interface 13, and can further include a computer program stored in the memory 11 and executable on the processor 10, such as a plot advertisement video generation program.
[0073] The processor 10 can be composed of integrated circuits in some embodiments, for example, can be composed of a single packaged integrated circuit, or can be composed of multiple packaged integrated circuits with the same function or different functions, including one or more combinations of central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The processor 10 is the control core (Control Unit) of the terminal, which connects various components of the terminal through various interfaces and lines, executes programs or modules stored in the memory 11 (for example, executes a plot advertisement video generation program, etc.), and calls data stored in the memory 11 to perform various functions of the terminal and process data.
[0074] The memory 11 includes at least one type of medium, including flash memory, mobile hard disk, multimedia card, card-type memory (for example, SD or DX memory, etc.), magnetic memory, disk, optical disk, etc. The memory 11 can be an internal storage unit of the terminal in some embodiments, for example, a mobile hard disk of the terminal. The memory 11 can also be an external storage terminal of the terminal in other embodiments, for example, a plug-in mobile hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the terminal. Further, the memory 11 can include both an internal storage unit and an external storage terminal of the terminal. The memory 11 can be used not only to store application software and various data installed on the terminal, such as the code of a plot advertisement video generation program, etc., but also to temporarily store data that has been output or will be output.
[0075] The communication bus 12 can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. The bus is configured to realize the connection and communication between the memory 11 and at least one processor 10, etc.
[0076] The communication interface 13 is used for communication between the terminal and other terminals, including a network interface and a user interface. Optionally, the network interface can include a wired interface and / or a wireless interface (such as a WI-FI interface, a Bluetooth interface, etc.), which is usually used to establish a communication connection between the terminal and other terminals. The user interface can be a display, an input unit (such as a keyboard), and optionally, the user interface can also be a standard wired interface, a wireless interface. Optionally, in some embodiments, the display can be an LED display, a liquid crystal display, a touch liquid crystal display, an OLED (Organic Light-Emitting Diode) touch, etc. The display can also be appropriately referred to as a display screen or a display unit, which is used to display information processed in the terminal and to display a visualized user interface.
[0077] Only terminals with components are shown in the figure, and those skilled in the art can understand that the structures shown in the figure do not constitute a limitation on the terminal, and can include fewer or more components than shown in the figure, or combine certain components, or different component arrangements.
[0078] For example, although not shown, the terminal can also include a power supply (such as a battery) for powering various components. Preferably, the power supply can be logically connected to the at least one processor 10 through a power management system, so as to realize functions such as charge management, discharge management, and power consumption management through the power management system. The power supply can also include one or more direct current or alternating current power supplies, a recharging system, a power supply fault detection circuit, a power supply converter or inverter, a power supply status indicator, and any other components. The terminal can also include various sensors, Bluetooth modules, Wi-Fi modules, etc., which are not described here.
[0079] It should be understood that the embodiments are only for illustration and are not limited in the scope of the patent application by the structure.
[0080] The scenario advertisement video generation program stored in the memory 11 in the terminal is a combination of multiple instructions, which, when running in the processor 10, can realize: S1, obtaining native interaction data of a target user on multiple platforms, and performing cross-modal semantic fusion on the native interaction data to obtain a user cognitive portrait of the target user; S2, taking multi-source heterogeneous classic narrative texts, commercial advertisement scripts, and multi-modal materials in a video playback platform as basic data, taking deep semantic analysis and relationship mining as construction means, taking user behavior feedback flow and external cultural semantic flow in the video playback platform as evolution basis, and taking incremental correlation learning as evolution mechanism, obtaining a narrative knowledge graph of the video playback platform; S3, driving plot deduction of the narrative knowledge graph based on the user cognitive portrait and core function attributes of the product to be promoted in the video playing platform, to obtain a primary narrative script of the video playing platform; S4, taking brand tone constraints and key information points of the product to be promoted as optimization targets, reinforcing scene anchors in the primary narrative script, to obtain a branded narrative script of the video playing platform; S5, performing multi-modal consistency arrangement on narrative nodes and emotional logic relationships in the branded narrative script, to obtain an executable narrative blueprint of the video playing platform; S6, automatically rendering materials in a multimedia material resource pool in the video playing platform according to timing, superposition relationships and transition logic in the executable narrative blueprint, to obtain a plot advertisement video of the video playing platform.
[0081] Specifically, the specific implementation method of the processor 10 on the above instructions can refer to the description of the related steps in the corresponding embodiments of the accompanying drawings, which will not be described here.
[0082] Further, the terminal integrated module / unit, if realized in the form of a software function unit and sold or used as an independent product, can be stored in a medium. The medium can be volatile or non-volatile. For example, the medium can include any entity or system that can carry the computer program code, a recording medium, a U disk, a mobile hard disk, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM, Read-Only Memory).
[0083] In several embodiments provided by the present application, it should be understood that the disclosed terminal, system and method can be implemented by other ways. For example, the above-mentioned system embodiments are only schematic, for example, the division of the modules is only a logical function division, and there can be another division way in actual implementation.
[0084] The modules described as separate components can or can not be physically separate, and the components displayed as modules can or can not be physical units, that is, they can be located in one place, or distributed on multiple network units. Part or all of the modules can be selected according to actual needs to achieve the purpose of the embodiment scheme.
[0085] In addition, each functional module in each embodiment of the present application can be integrated in one processing unit, or each unit can exist physically independently, or two or more units can be integrated in one unit. The above integrated unit can be realized in the form of hardware or in the form of hardware plus software function module.
[0086] It is apparent for a person skilled in the art that the present application is not limited to the details of the above-described exemplary embodiments, but that the present application can be implemented in other concrete forms without departing from the spirit or essential characteristics of the present application.
[0087] Embodiments of the present application can acquire and process related data based on artificial intelligence technology. Among them, artificial intelligence (Artificial Intelligence, AI) is to use digital computers or digital computer controlled machines to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. Theory, method, technology and application system.
[0088] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and are not limiting. Although the present application has been described in detail with reference to the preferred embodiments, it should be understood by those skilled in the art that the technical solutions of the present application can be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present application.
Claims
1. A method for generating narrative advertising videos, characterized in that, The method includes: S1. Obtain the target user's native interaction data on multiple platforms, and perform cross-modal semantic fusion on the native interaction data to obtain the target user's cognitive profile; S2. Based on the multi-source heterogeneous classic narrative texts, commercial advertising scripts and multimodal materials in the video playback platform, using deep semantic analysis and relationship mining as construction methods, using the user behavior feedback stream and external cultural semantic stream in the video playback platform as evolutionary basis, and using incremental association learning as the evolution mechanism, the narrative knowledge graph of the video playback platform is obtained. S3. Based on the user cognitive profile and the core functional attributes of the product to be promoted in the video playback platform, drive the plot deduction of the narrative knowledge graph to obtain the primary narrative script of the video playback platform; S4. Taking the brand tone constraints and key information points of the product to be promoted as optimization targets, strengthen the plot anchor points in the primary narrative script to obtain the branded narrative script of the video playback platform. S5. Perform multimodal consistency arrangement of the narrative nodes and emotional logic relationships in the branded narrative script to obtain the executable narrative blueprint of the video playback platform; S6. Based on the timing, overlay relationship and transition logic in the executable narrative blueprint, automatically render the materials in the multimedia material resource pool of the video playback platform to obtain the plot advertisement video of the video playback platform.
2. The method for generating a story-driven advertising video as described in claim 1, characterized in that, The process of acquiring native interaction data of the target user across multiple platforms and performing cross-modal semantic fusion on the native interaction data to obtain the user cognitive profile of the target user includes: Parallel acquisition of text comment data, visual focus trajectory data, and audio interaction segment data of target users on multiple platforms; Explicit interest keywords and implicit sentiment tendencies are extracted from the text comment data to obtain the text semantic vector of the target user; The visual interest vector of the target user is obtained by analyzing the attention regions of the interface elements that remain continuously or quickly pass through the visual focus trajectory data. The audio interaction segment data is filtered to extract voice emotion keywords, thus obtaining the audio semantic vector of the target user; Based on a preset simulated weighting strategy, the text semantic vector, the visual interest vector, and the audio semantic vector are weighted and concatenated to obtain the user feature expression of the target user. Based on the user feature expression, the distribution intensity of the user feature expression on different interest dimensions is analyzed to construct the user cognitive profile of the target user.
3. The method for generating a story-driven advertising video as described in claim 1, characterized in that, The narrative knowledge graph of the video playback platform is derived from multi-source heterogeneous classic narrative texts, commercial advertising scripts, and multimodal materials as its basic data. It employs deep semantic analysis and relationship mining as its construction methods, user behavior feedback streams and external cultural semantic streams within the video playback platform as its evolutionary basis, and incremental association learning as its evolutionary mechanism. This graph includes: The classic narrative texts and commercial advertising scripts in the video playback platform are deconstructed for plot structure, character motivation is identified, and emotional context is annotated in order to extract the narrative entities and the logical relationship between the classic narrative texts and the commercial advertising scripts, so as to obtain the narrative logic of the video playback platform. Cross-media content alignment is performed on the multimodal materials, and visual keyframes, background audio clips, and text elements corresponding to the narrative logic are bound to obtain the multimodal content binding relationship of the video playback platform; Using the narrative logic as nodes and the logical connections between the narrative logics and the binding relationships of the multimodal content as connecting edges, a structured assembly is performed to obtain the initial knowledge graph of the video playback platform. The user behavior feedback stream in the video playback platform is parsed as an optimization guide for specific logical association paths in the initial knowledge graph, and the external cultural semantic stream in the video playback platform is parsed as an emerging narrative mode to be integrated into the initial knowledge graph. According to the optimization guidelines, the association strength of corresponding connecting edges in the initial knowledge graph is iteratively adjusted. Simultaneously, based on the emerging narrative mode, new nodes and connecting edges are expanded in the initial knowledge graph. Through continuous iterative adjustments, the initial knowledge graph evolves into the narrative knowledge graph of the video playback platform. The evolution formula for iteratively adjusting the association strength is as follows: ; In the formula, The new association strength of the corresponding connecting edges in the initial knowledge graph. The current association strength of the corresponding connecting edge in the initial knowledge graph. The user behavior feedback coefficient of the user behavior feedback stream. The cultural semantic injection coefficients for the external cultural semantic stream. The preset user feedback adjustment coefficient, Injecting a moderating factor into the pre-set culture, This is the preset evolution smoothing factor.
4. The method for generating a story-driven advertising video as described in claim 1, characterized in that, The process of using the user cognitive profile and the core functional attributes of the product to be promoted on the video playback platform to drive the plot deduction of the narrative knowledge graph, thereby obtaining the initial narrative script of the video playback platform, includes: Extract the preference dimension and intensity weight from the user cognitive profile, and extract the key functional points and value propositions from the core functional attributes. Perform feature cross-fusion of the preference dimension, the intensity weight, the key functional points, and the value proposition to obtain the joint query vector of the video playback platform. The joint query vector is mapped to the narrative knowledge graph as an initial probe, and a multi-step path exploration is performed on the narrative knowledge graph to obtain the candidate plot development path of the video playback platform. Based on the emotional tendency weights in the user's cognitive profile and the logical completeness requirements of the core functional attributes, the candidate plot development paths are ranked according to their matching degree to obtain the optimal plot path of the video playback platform. The node sequence and relationships contained in the optimal plot path are filled in according to a preset script structure template to obtain the primary narrative script of the video playback platform.
5. The method for generating a story-driven advertising video as described in claim 1, characterized in that, The optimization target is to strengthen the plot anchors in the initial narrative script by using the brand tone constraints and key information points of the product to be promoted, thereby obtaining the branded narrative script for the video playback platform, including: By deconstructing the brand tone constraints of the product to be promoted, the brand emotional tone and visual style guidelines of the product to be promoted are obtained. The key information points of the product to be promoted are logically sorted to obtain the information embedding priority of the video playback platform. Mark the narrative nodes in the primary narrative script that have a potential connection with the product feature demonstration or value communication as anchor points to be strengthened; The brand emotional tone, visual style guidelines, and information embedding priority are matched and mapped with the anchor points to be strengthened to obtain the brand embedding scheme of the video playback platform. Based on the branding integration scheme, the plot descriptions, character dialogues, or scene settings in the primary narrative script are reconstructed to obtain the branding narrative script of the video playback platform.
6. The method for generating a story-driven advertising video as described in claim 1, characterized in that, The process of performing multimodal consistency orchestration on the narrative nodes and emotional logic relationships in the branded narrative script to obtain an executable narrative blueprint for the video playback platform includes: The branded narrative script is structurally analyzed, and the emotional connection and logical progression between narrative nodes in the branded narrative script are marked. Based on the aforementioned emotional connection and logical progression, visual material items, background audio clips, and screen text elements of the narrative node are associated from the multimedia material resource pool of the video playback platform. Check whether the connection between the narrative node and the visual materials, background audio and screen text associated with the adjacent nodes is natural and smooth in terms of style, emotional intensity and rhythm, and mark the connection points in the branded narrative script that have abrupt jumps or contradictions. Based on the connection point, trace back to the emotional logic relationship in the branded narrative script, adjust the multimodal material association scheme of the connection point, and obtain the conflict elimination connection point of the video playback platform; Based on the conflict resolution connection point, the narrative nodes and the temporal and superposition relationships between the narrative nodes are encapsulated according to the logical order of the branded narrative script to obtain the executable narrative blueprint of the video playback platform.
7. The method for generating a story-driven advertising video as described in claim 1, characterized in that, The process of automatically rendering materials from the multimedia resource pool of the video playback platform based on the timing, overlay relationships, and transition logic in the executable narrative blueprint to obtain the story-driven advertising video of the video playback platform includes: Extract the multimedia material identifiers, the start and end times of the materials in the video stream, and the overlay levels and transition logic between adjacent materials from the executable narrative blueprint; Based on the multimedia material identifier, the original visual segments, audio segments, and graphic elements of the video playback platform are retrieved from the multimedia material resource pool; According to the start and end sequence, the original visual segments and the graphic elements are aligned and synthesized according to the overlay level to obtain the preliminary visual track of the video playback platform. Simultaneously, the corresponding audio segments are aligned and mixed to obtain the audio track of the video playback platform. Based on the transition logic, a transition effect is inserted between the visual segments in the initial visual track and the audio segments in the audio track to obtain a smooth synthesized audio and video track of the video playback platform; The smooth synthesized audio and video track is encoded and encapsulated to obtain the plot advertisement video of the video playback platform.
8. A narrative advertising video generation system, characterized in that, include: The user profile building module is used to acquire the target user's native interaction data on multiple platforms and perform cross-modal semantic fusion on the native interaction data to obtain the target user's cognitive profile. The narrative graph construction module is used to obtain the narrative knowledge graph of the video playback platform by using multi-source heterogeneous classic narrative texts, commercial advertising scripts and multimodal materials as basic data, deep semantic analysis and relationship mining as construction methods, user behavior feedback stream and external cultural semantic stream in the video playback platform as evolution basis, and incremental association learning as evolution mechanism. The primary script deduction and generation module is used to drive the plot deduction of the narrative knowledge graph based on the user cognitive profile and the core functional attributes of the product to be promoted in the video playback platform, so as to obtain the primary narrative script of the video playback platform. The brand script optimization module is used to take the brand tone constraints and key information points of the product to be promoted as optimization targets, strengthen the plot anchors in the primary narrative script, and obtain the branded narrative script of the video playback platform. The narrative blueprint arrangement module is used to perform multimodal consistency arrangement of narrative nodes and emotional logic relationships in the branded narrative script to obtain an executable narrative blueprint for the video playback platform. The automated rendering module for advertising videos is used to automatically render the materials in the multimedia material resource pool of the video playback platform according to the timing, overlay relationship and transition logic in the executable narrative blueprint, so as to obtain the plot advertising video of the video playback platform.
9. A terminal, comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 7.
10. A storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the steps of the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
High-quality video content automatic generation method and related equipment
CN120050487A
Commodity video intelligent generation method based on multi-modal analysis and dynamic narrative architecture
CN120711259A
Advertisement generation method, device and equipment and readable storage medium
CN120746645A
Multi-agent collaborative commodity plot advertisement script generation system and method
CN120893577A
Advertisement creativity dynamic generation and adaptation method, device, equipment and medium
CN121120149A