Video plot generation and scene synthesis method and system based on natural language processing
By building a multi-layer scene map and feature pyramid, combining self-attention mechanism and style transfer technology, the problem of insufficient understanding of scene semantics in video generation is solved, and high-quality videos that are highly consistent with the script are generated, improving video production efficiency and artistic expression.
Patent Information
- Application Number
- CN202510819613.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-19
AI Technical Summary
The existing technology lacks the understanding of scene semantics in video generation, and it is difficult to accurately reflect the narrative logic and emotional expression of the script, the visual feature reconstruction is inaccurate, and the lack of an effective feature fusion mechanism, resulting in large differences between the generated video scenes and the original script, unnatural visual effects, inconsistent transitions and style transfers, making it difficult to achieve high-quality video production.
By constructing a multi-layer scene map, extracting keyword sequences and calculating semantic correlations, performing feature decomposition and reconstruction, combining feature pyramids and adaptive feature fusion, applying a self-attention mechanism to perform video segmentation and transition effect processing, and finally performing style transfer to generate high-quality video finished products.
It realizes video scene generation that is highly consistent with the original script, improves narrative coherence and visual expression, improves video editing efficiency and finished product quality, and provides intelligent and automated film and television creation support.
Smart Images

Figure CN120339919A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video generation, and in particular, to a method and system for video plot generation and scene synthesis based on natural language processing. Background Art
[0002] With the rapid development of artificial intelligence and computer vision technologies, the technology of converting text content into video scenes based on natural language processing has gradually become a research hotspot. By parsing text descriptions, extracting semantic features, and using corresponding visual generation models, the automatic conversion from script text to video content is realized. The traditional video production process usually requires a large amount of manpower and material resources for shooting, editing, and post-processing. However, the video generation technology based on natural language processing provides an efficient and automated solution, which can significantly reduce the video production cost and improve the production efficiency. However, the existing technologies still have deficiencies in the depth of scene semantic understanding and limited ability to grasp complex plots and multi-level narrative structures, resulting in a large difference in semantic expression between the generated video scenes and the original script text, and being unable to accurately reflect the narrative logic and emotional expression of the script. The reconstruction and representation of visual features are not precise enough, lacking an effective feature fusion mechanism, and it is difficult to generate video content with natural visual effects and unified styles. In particular, it performs poorly in dealing with scene transitions and visual coherence, and lacks a systematic method in video sequence combination and style processing, being unable to achieve intelligent segmentation and transition effect generation based on content semantics. At the same time, the style transfer process lacks consideration for the global consistency of the video, and it is difficult to ensure the artistic expressiveness and professional quality of the final video works. Therefore, there is an urgent need for a solution to solve the problems existing in the prior art. Summary of the Invention
[0003] Embodiments of the present invention provide a method and system for video plot generation and scene synthesis based on natural language processing, which can at least solve some of the problems existing in the prior art.
[0004] In a first aspect of the embodiments of the present invention, there is provided a method for video plot generation and scene synthesis based on natural language processing, including: Receiving the script text input by a user and extracting scene semantic features, and constructing the scene semantic features into a multi-layer scene graph; Extract keyword sequences based on the multi-layer scene graph, calculate the semantic relevance of the keyword sequences to obtain a feature vector matrix, perform semantic feature decomposition and visual feature reconstruction on the feature vector matrix respectively to obtain a scene description vector and a scene structure vector, alternately update the feature representation according to the scene description vector and the scene structure vector to obtain a scene synthesis vector, calculate the similarity between the scene synthesis vector and the reference samples in the video sample library to obtain a distance metric value, and when the distance metric value is less than a preset threshold, generate a target video scene based on the current scene synthesis vector; Based on the pre-constructed feature pyramid and the target video scene, combined with the inter-layer correlation relationship in the multi-layer scene graph, extract spatio-temporal features at each layer to obtain a sequence of feature maps, perform adaptive feature fusion on the sequence of feature maps to obtain a scene feature representation and calculate an inter-frame similarity matrix, apply a self-attention mechanism to the similarity matrix to obtain scene key frame indices and segment the video content, and insert a depth map-based transition effect between adjacent segments to obtain a combined video sequence; Perform style transfer on the combined video sequence and output the finished video.
[0005] In an alternative embodiment, Receive the script text input by the user and extract the scene semantic features, and construct the multi-layer scene graph from the scene semantic features, including: Receive the script text input by the user, perform semantic word segmentation on the script text to obtain a word sequence, and encode the word sequence based on a pre-trained language model to obtain initial semantic features; Input the initial semantic features into a multi-head attention network, calculate the feature representations of different attention heads in parallel, and perform hierarchical clustering on the feature representations to obtain scene semantic features, where each attention head focuses on the semantic information of character interaction, scene attributes, and plot development respectively; Construct a multi-layer scene graph based on the scene semantic features, calculate the neighborhood relationship of the nodes in the graph using a graph attention network, dynamically aggregate the nodes with similar semantic features into a character relationship layer, a scene attribute layer, and a plot development layer, and determine the inter-layer correlation relationship.
[0006] In an alternative embodiment, Extract keyword sequences based on the multi-layer scene graph, calculate the semantic relevance of the keyword sequences to obtain a feature vector matrix, perform semantic feature decomposition and visual feature reconstruction on the feature vector matrix respectively to obtain a scene description vector and a scene structure vector, and alternately update the feature representation according to the scene description vector and the scene structure vector to obtain a scene synthesis vector, including: Extract node features from the multi-layer scene graph, calculate the information entropy between each hierarchical node and its adjacent nodes to obtain node weights, perform hierarchical clustering based on the node weights, and select the clustering center points to form a keyword sequence; Calculate the semantic relevance of word pairs in the keyword sequence to obtain a relevance matrix. The semantic relevance is obtained by multiplying the cosine distance of word vectors by the edge weights between nodes in the graph. The relevance matrix is subjected to weight mapping and non-linear transformation to obtain a feature vector matrix; Perform semantic feature decomposition on the feature vector matrix to obtain a scene description vector, and perform visual feature reconstruction to obtain a scene structure vector; Alternately iterate and optimize the scene description vector and the scene structure vector to obtain a candidate feature representation. Calculate the joint optimization objective of the semantic consistency loss, the visual reconstruction loss, and the temporal continuity loss based on the candidate feature representation. Update the candidate feature representation until convergence by minimizing the joint optimization objective, and determine the converged feature representation as the scene synthesis vector.
[0007] In an alternative embodiment, Calculate the similarity between the scene synthesis vector and the reference samples in the video sample library to obtain a distance metric value. When the distance metric value is less than a preset threshold, generating a target video scene based on the current scene synthesis vector includes: Construct a scene semantic gradient map, calculate the information entropy based on the normalized eigenvalue of the scene elements in the scene semantic gradient map, obtain the importance distribution of the scene elements, divide the scene into a key area and a transition area according to the information entropy and the importance distribution, extract the first scene feature vector corresponding to the key area and the second scene feature vector corresponding to the transition area, and calculate the semantic association strength between the key area and the transition area based on the Euclidean distance between the first scene feature vector and the second scene feature vector; Extract the scene synthesis vector and the reference samples in the video sample library, perform cosine distance measurement on the key area, perform Mahalanobis distance measurement based on the feature covariance matrix on the transition area, weight the cosine distance measurement value of the key area and the Mahalanobis distance measurement value of the transition area by a preset regional weight coefficient, and weight the semantic association strength by a pre-acquired semantic association coefficient, and combine the weighted values to obtain the distance metric value; Calculate the distance gradient of the distance metric value, superimpose the product of the distance gradient and the update step size on the original regional boundary to obtain an updated regional boundary, and judge whether the distance metric value is less than the preset threshold based on the updated regional boundary. If it is less, generate the target video scene based on the scene synthesis vector.
[0008] In an alternative embodiment, Based on a pre-constructed feature pyramid and the target video scene, combining the inter-layer correlation relationships in the multi-layer scene graph, spatio-temporal features are extracted at each layer to obtain a sequence of feature maps. Adaptive feature fusion is performed on the sequence of feature maps to obtain a scene feature representation and calculate an inter-frame similarity matrix. The self-attention mechanism is applied to the similarity matrix to obtain scene key frame indices and segment the video content. A depth map-based transition effect is inserted between adjacent segments to obtain a combined video sequence, including: Based on the target video scene, feature maps of each layer in the pre-constructed feature pyramid are obtained. Combining the inter-layer correlation relationships in the multi-layer scene graph, spatio-temporal features are extracted from the feature maps through multi-scale convolution. Convolution operations are performed on the feature maps using convolution kernel parameters corresponding to each layer and bias terms are added to obtain the sequence of feature maps; Perform matrix multiplication on the query matrix of the current layer of the feature pyramid and the key matrix of the adjacent layer. After normalizing the calculation result, multiply it with the value matrix of the adjacent layer to obtain feature correlation data. The feature correlation data is fused with the sequence of feature maps to obtain an association-enhanced feature sequence; Perform average pooling and multi-layer perceptron processing on the association-enhanced feature sequence to obtain channel attention weights. The channel attention weights are weighted and summed with the association-enhanced feature sequence to obtain the scene feature representation. Based on the scene feature representation, calculate the inter-frame similarity matrix and apply the self-attention mechanism to extract the scene key frame indices; Segment the video content according to the scene key frame indices to obtain video segments. Generate a transition effect based on the time-varying interpolation coefficient calculated by the cosine function and the depth maps of adjacent video segments. Combine the video segments with the transition effect to obtain the combined video sequence.
[0009] In an alternative embodiment, Generate a transition effect based on the time-varying interpolation coefficient calculated by the cosine function and the depth maps of adjacent video segments. Combine the video segments with the transition effect to obtain the combined video sequence, including: For the transition time period between adjacent video segments, calculate the time-varying interpolation coefficient based on the ratio of the current time point to the transition duration using the cosine function; Collect the scene information of the adjacent video segments. Perform position encoding on the spatial position and viewing direction of any point in the current scene through a multi-layer perceptron to obtain encoded features. Based on the encoded features, construct a neural radiance field representation including spatial density features and color features; Obtain the depth maps of the adjacent video segments, calculate the depth consistency loss between the spatial density integral of the camera rays and the depth maps, optimize the neural radiance field representation based on the depth consistency loss, apply the time-varying interpolation coefficients to the optimized neural radiance field representation for weighted fusion, perform volume rendering on the fused features, and generate an intermediate view sequence by calculating the integral of the cumulative transmittance and the spatial density and color features; Calculate the gradient difference of the depth maps to obtain the depth consistency weights, apply the depth consistency weights to the intermediate view sequence and the depth maps weighted by the time-varying interpolation coefficients for adaptive fusion to obtain the transition effect, and combine the adjacent video segments and the transition effect in chronological order to obtain the combined video sequence.
[0010] In an alternative embodiment, Perform style transfer on the combined video sequence, and the output video product includes: Obtain the combined video sequence, and extract the spatio-temporal features and the sequence of feature maps corresponding to the combined video sequence; Perform mapping and matching on the spatio-temporal features and the sequence of feature maps corresponding to the combined video sequence with reference to the feature space of the target style template to obtain the fused features after style transfer; Reconstruct and decode the fused features after style transfer to obtain a video product with a unified style.
[0011] In the second aspect of the embodiments of the present invention, there is provided a video plot generation and scene synthesis system based on natural language processing, including: A first unit for receiving the script text input by the user and extracting the scene semantic features, and constructing the scene semantic features into a multi-layer scene graph; A second unit for extracting a keyword sequence based on the multi-layer scene graph, calculating the semantic relevance of the keyword sequence to obtain a feature vector matrix, respectively performing semantic feature decomposition and visual feature reconstruction on the feature vector matrix to obtain a scene description vector and a scene structure vector, alternately updating the feature representation according to the scene description vector and the scene structure vector to obtain a scene synthesis vector, calculating the similarity between the scene synthesis vector and the reference samples in the video sample library to obtain a distance metric value, and when the distance metric value is less than a preset threshold, generating a target video scene based on the current scene synthesis vector; The third unit is used to extract spatio-temporal features at each layer based on a pre-constructed feature pyramid and the target video scene, combine the inter-layer correlation relationships in the multi-layer scene graph, obtain a sequence of feature maps, perform adaptive feature fusion on the sequence of feature maps, obtain a scene feature representation, calculate an inter-frame similarity matrix, apply a self-attention mechanism to the similarity matrix to obtain scene key frame indices, segment the video content, and insert depth map-based transition effects between adjacent segments to obtain a combined video sequence; The fourth unit is used to perform style transfer on the combined video sequence and output a finished video.
[0012] In a third aspect of the embodiments of the present invention, an electronic device is provided, including: A processor and a memory for storing processor-executable instructions, wherein the processor is configured to call the instructions stored in the memory to execute the method described above.
[0013] In a fourth aspect of the embodiments of the present invention, a computer-readable storage medium is provided, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.
[0014] In the present invention, by constructing a multi-layer scene graph to extract script semantic features, the accurate conversion from natural language to visual content is realized, effectively solving the problems of semantic information loss and inconsistent visual performance in traditional methods. Through the similarity matching mechanism between the scene synthesis vector and the reference sample, a video scene highly consistent with the original script can be generated, while retaining the intention expression of the creator, improving the narrative coherence and visual expressiveness of the video content. Based on the processing method of the feature pyramid and adaptive feature fusion, combined with the depth map transition effect and style transfer technology, the generation of high-quality video sequences is realized, significantly improving the video editing efficiency and the quality of the finished product, providing intelligent and automated technical support for film and television creation. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 It is a schematic flowchart of a method for video plot generation and scene synthesis based on natural language processing according to an embodiment of the present invention; Figure 2 It is a schematic architecture diagram of script text analysis and multi-layer scene graph construction; Figure 3 It is a schematic diagram of inter-layer correlation and adaptive feature fusion of the feature pyramid; Figure 4 It is a schematic diagram of performance comparison of video transition effects based on neural radiance fields. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0016] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0017] The technical solutions of the present invention will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.
[0018] Figure 1 The following is a schematic flowchart of a method for video plot generation and scene synthesis based on natural language processing according to an embodiment of the present invention, as Figure 1 shown, the method includes: Receiving the script text input by the user and extracting the scene semantic features, and constructing the scene semantic features into a multi-layer scene graph; Extracting a keyword sequence based on the multi-layer scene graph, calculating the semantic relevance of the keyword sequence to obtain a feature vector matrix, respectively performing semantic feature decomposition and visual feature reconstruction on the feature vector matrix to obtain a scene description vector and a scene structure vector, alternately updating the feature representation according to the scene description vector and the scene structure vector to obtain a scene synthesis vector, calculating the similarity between the scene synthesis vector and a reference sample in the video sample library to obtain a distance metric value, and when the distance metric value is less than a preset threshold, generating a target video scene based on the current scene synthesis vector; Based on a pre-constructed feature pyramid and the target video scene, combining the inter-layer association relationships in the multi-layer scene graph to extract spatio-temporal features at each layer to obtain a sequence of feature maps, performing adaptive feature fusion on the sequence of feature maps to obtain a scene feature representation and calculating an inter-frame similarity matrix, applying a self-attention mechanism to the similarity matrix to obtain scene key frame indices and segmenting the video content, and inserting a depth map-based transition effect between adjacent segments to obtain a combined video sequence; Performing style transfer on the combined video sequence and outputting a finished video.
[0019] In an alternative embodiment, Receiving the script text input by the user and extracting the scene semantic features, and constructing the scene semantic features into a multi-layer scene graph includes: Receiving the script text input by the user, performing semantic word segmentation on the script text to obtain a word sequence, and encoding the word sequence based on a pre-trained language model to obtain initial semantic features; Input the initial semantic features into a multi - head attention network. By parallelly calculating the feature representations of different attention heads, hierarchical clustering is performed on the feature representations to obtain scene semantic features, where each attention head respectively focuses on the semantic information of character interaction, scene attributes, and plot development; Construct a multi - layer scene graph based on the scene semantic features. Use a graph attention network to calculate the neighborhood relationships of the nodes in the graph, dynamically aggregate the nodes with similar semantic features into a character relationship layer, a scene attribute layer, and a plot development layer, and determine the inter - layer association relationships.
[0020] Figure 2 It is a schematic diagram of the system architecture for drama text analysis and multi - layer scene graph construction. As Figure 2 shown, receive the drama text input by the user, and perform semantic word segmentation on the received drama text to obtain a word sequence. During the semantic word segmentation process, a word segmentation model based on a bidirectional long - short - term memory network is used. By encoding the characters in the text, the word boundaries are identified, thereby generating a word sequence. For example, for the input drama fragment "On a rainy night, Li Ming pushed open the door of the coffee shop and looked around for the friend who had an appointment to meet", it is segmented into ["On a rainy night", ",", "Li Ming", "pushed open", "the coffee shop", "of", "the door", ",", "looked around", "searched for", "appointment", "to meet", "of", "the friend"].
[0021] After obtaining the word sequence, encode the word sequence based on a pre - trained language model to obtain initial semantic features. A pre - trained language model with a 12 - layer Transformer encoder structure is used, and the model has been pre - trained on a large - scale text corpus. For each input word, it is converted into a 512 - dimensional word embedding vector, and then feature extraction is performed through the multi - layer encoder of the language model to obtain an initial semantic feature vector representing the semantics of the word. Each word corresponds to a 768 - dimensional vector. Taking the above - mentioned word segmentation result as an example, an initial semantic feature set containing 16 768 - dimensional vectors will be generated.
[0022] Input the initial semantic features into a multi - head attention network, and parallelly calculate the feature representations of different attention heads. The multi - head attention network contains 8 attention heads, where 3 attention heads specifically focus on character interaction information, 3 attention heads focus on scene attribute information, and 2 attention heads focus on plot development information. Each attention head uses a self - attention mechanism to capture the association relationships between words by calculating the similarities between query vectors, key vectors, and value vectors. Taking the attention head that focuses on character interaction as an example, higher attention weights are assigned to the character names and the words describing the character's behaviors. In the example, "Li Ming", "pushed open", "looked around", "searched for", "friend" obtain higher weights in the character interaction attention head, forming a feature representation focused on character interaction.
[0023] When performing hierarchical clustering on the feature representation to obtain the scene semantic features, a bottom-up hierarchical clustering algorithm is used to calculate the cosine similarity between the word feature representations, and a similarity matrix is constructed. Initially, each word feature is regarded as an independent cluster, and the system gradually merges the clusters with the highest similarity until the preset number of clusters is reached. For the character interaction features, the similarity threshold is set to 0.75 to form semantic clusters such as "character - action - object"; for the scene attribute features, the similarity threshold is set to 0.7 to form semantic clusters such as "scene - description - state"; for the plot development features, the similarity threshold is set to 0.8 to form semantic clusters representing plot changes. In the example, "Li Ming - push open - the door" and "Li Ming - look for - friends" are generated as character interaction semantic clusters, "rainy night - coffee shop" as a scene attribute semantic cluster, and "push open the door - look around - look for friends" as a plot development semantic cluster.
[0024] When constructing a multi-layer scene graph based on the scene semantic features, the graph attention network is used to calculate the neighborhood relationship of the nodes in the graph. Each semantic cluster is regarded as a node in the graph, and the node feature is obtained by the weighted average of all word features in the cluster. The graph attention network contains 2 layers of graph convolutional layers, and the number of hidden units in each layer is 256. Through the graph attention mechanism, the relationship strength between each node and its neighborhood nodes is calculated. The relationship strength is determined by the similarity of the node features. The attention coefficient is calculated for each pair of nodes, and the attention coefficient is obtained by the inner product of the linearly transformed node features and is normalized by the SoftMax function.
[0025] When dynamically aggregating nodes with similar semantic features into the character relationship layer, the scene attribute layer, and the plot development layer, hierarchical partitioning is performed based on the node representations learned by the graph attention network. For the character relationship layer, the nodes that contain the character name and focus on character interaction are aggregated into this layer; for the scene attribute layer, the nodes that describe scene elements such as environment, time, and space are aggregated into this layer; for the plot development layer, the nodes that describe the event development and plot changes are aggregated into this layer. Exemplarily, the nodes of "Li Ming - push open - the door" and "Li Ming - look for - friends" are aggregated into the character relationship layer, the node of "rainy night - coffee shop" is aggregated into the scene attribute layer, and the node of "push open the door - look around - look for friends" is aggregated into the plot development layer.
[0026] When determining the inter-layer association relationship, the cross-attention scores between nodes in different layers are calculated. The cross-attention scores are obtained by the dot product between node features, and weak association relationships are filtered using a threshold. Node pairs with cross-attention scores greater than 0.6 are established with inter-layer connections, thus forming a complete multi-layer scene graph. In the constructed multi-layer scene graph, the node "Rainy Night - Café" in the scene attribute layer is associated with the node "Li Ming - Push - Door" in the character relationship layer, and at the same time is associated with the node "Push the door - Look around - Look for friends" in the plot development layer, thus forming a multi-level semantic representation describing a complete scene.
[0027] In this embodiment, through semantic processing of the script text and constructing a multi-layer scene graph, the deep semantic understanding and structured expression of the script content are realized, and the key semantic information such as character interaction, scene attributes, and plot development in the script are automatically recognized and distinguished. The original flat text content is transformed into a hierarchical semantic structure, which not only retains the core semantic content of the original script, but also clarifies the association relationship between elements through the graph structure, providing a richer and more accurate semantic basis for subsequent video scene generation, so as to be able to generate video scenes highly matching the script content, improving the accuracy and content consistency of the text-to-video conversion.
[0028] In an alternative embodiment, Based on the multi-layer scene graph, extract a keyword sequence, calculate the semantic relevance of the keyword sequence to obtain a feature vector matrix, perform semantic feature decomposition and visual feature reconstruction on the feature vector matrix respectively, obtain a scene description vector and a scene structure vector, and alternately update the feature representation according to the scene description vector and the scene structure vector to obtain a scene synthesis vector, including: Extract node features from the multi-layer scene graph, calculate the information entropy of each layer node and its adjacent nodes to obtain node weights, perform hierarchical clustering based on the node weights, and select the cluster center points to form a keyword sequence; Calculate the semantic relevance of the word pairs in the keyword sequence to obtain a relevance matrix. The semantic relevance is obtained by multiplying the cosine distance of the word vectors by the edge weights between nodes in the graph. The relevance matrix is subjected to weight mapping and non-linear transformation to obtain a feature vector matrix; Perform semantic feature decomposition on the feature vector matrix to obtain a scene description vector, and perform visual feature reconstruction to obtain a scene structure vector; Alternately iterate and optimize the scene description vector and the scene structure vector to obtain a candidate feature representation. Based on the candidate feature representation, calculate the joint optimization objective of semantic consistency loss, visual reconstruction loss, and temporal continuity loss. Update the candidate feature representation by minimizing the joint optimization objective until convergence, and determine the converged feature representation as the scene synthesis vector.
[0029] Construct a multi - layer scene graph, which includes object nodes, relationship nodes and their hierarchical structures. Extract keyword sequences from the graph and calculate semantic relevance, and finally generate a scene synthesis vector for subsequent scene generation tasks.
[0030] Calculate the information entropy for each node to determine its importance. For node vi in the graph, calculate the information transfer amount between it and its adjacent node vj. Obtain the feature representation fi of node vi, including node type, attributes and context information. For object nodes, features include object category, location, size, etc.; for relationship nodes, features include relationship type and associated object information. Taking the scene "A person is sitting on a chair reading a book" as an example, the features of the node "person" include the category "person", location coordinates (x1, y1), attribute "adult male", etc.
[0031] For node vi and all its adjacent nodes vj, calculate the information entropy Hi of node vi, considering the connection strength wij between nodes, which represents the degree of association between two nodes. For example, the connection strength between "person" and "sit" may be 0.85, while the connection strength between "person" and "read" is 0.78. Based on the calculated information entropy Hi, assign a weight wi to each node. The higher the weight value, the more important the node is in scene understanding. In the example scene, the weights of nodes "person", "chair", "book", "sit", "read" may be 0.92, 0.87, 0.83, 0.79, 0.76 respectively.
[0032] Perform hierarchical clustering based on node weights. The clustering process is bottom - up. Initially, each node is regarded as a separate cluster, and clusters with similarity higher than a pre - set similarity threshold are gradually merged. The similarity calculation comprehensively considers the cosine similarity of node feature vectors and the distance between nodes in the graph. When the clustering reaches a preset stop condition (such as the number of clusters or the within - cluster variance threshold), select the node with the highest weight from each cluster as a representative to form a keyword sequence. In the above example, the possible keyword sequence may be ["person", "chair", "book", "read"].
[0033] After obtaining the keyword sequence, calculate the semantic relevance between word pairs in the sequence. For each pair of words wi and wj in the sequence, obtain their pre - trained word vectors vi and vj, and calculate the cosine similarity of the word vectors sim(vi, vj). At the same time, find the edge weight edge(wi, wj) between the corresponding nodes in the graph. The semantic relevance is obtained by multiplying these two factors: rel(wi, wj)=sim(vi, vj)×edge(wi, wj). In actual calculation, if the cosine similarity of the word vectors of "person" and "read" is 0.65 and the edge weight in the graph is 0.78, then their semantic relevance is 0.507.
[0034] After computing for all word pairs, an n×n relevance matrix R is constructed, where n is the length of the keyword sequence. To enhance the matrix's expressive power, the weight mapping function is applied to adjust R, and through non-linear transformations such as ReLU, the feature vector matrix F is obtained. For a sequence of 4 keywords, F may be a 4×128 matrix, with each row representing the feature representation of a keyword.
[0035] Semantic feature decomposition is performed on the feature vector matrix F to extract the semantic core information. The main semantic information is retained and noise is filtered through dimensionality reduction techniques. F is projected into the semantic space to obtain a scene description vector of dimension d. For the example scene, the scene description vector is a 128-dimensional vector that captures the core semantics of "a person reading a book on a chair".
[0036] Visual feature reconstruction is performed by mapping F into the visual representation space, considering the spatial relationships and visual attributes between objects, and generating a scene structure vector. The scene structure vector encodes the spatial layout and visual relationships of the objects in the scene, such as structural information like "the person is above the chair" and "the book is in front of the person".
[0037] Taking the scene description vector and the scene structure vector as the initial inputs, the feature representation is updated through an alternating iterative optimization method. In each iteration, the system updates the scene structure vector based on the current scene description vector, and then updates the scene description vector based on the updated scene structure vector, forming a mutually promoting optimization process. The optimization objectives include three parts: the semantic consistency loss ensures that the feature representation is consistent with the original semantics; the visual reconstruction loss ensures that the features can correctly reconstruct the visual scene; the temporal continuity loss ensures a smooth transition in time when processing consecutive scenes.
[0038] Taking the above scene as an example, the initial scene description vector may focus on the semantic representation of the "reading" behavior, while the initial scene structure vector may not fully capture the spatial relationship between the "person" and the "chair". Through iterative optimization, the final scene description vector will contain richer semantic information, and the scene structure vector will more accurately represent the spatial structure, and the two together form a comprehensive understanding of the scene.
[0039] When the combined optimization objective value is lower than a preset threshold (e.g., 0.001) or reaches the maximum number of iterations (e.g., 50 times), it is considered that the optimization has converged, and the final feature representation is determined as the scene synthesis vector. The scene synthesis vector combines the semantic content and structural information of the scene, providing a comprehensive scene representation for subsequent scene generation or understanding tasks.
[0040] In this embodiment, the node weights are calculated based on information entropy and hierarchical clustering is performed, so that the extracted keyword sequence can accurately reflect the core semantics of the multi-layer scene graph. The semantic relevance is calculated by combining the cosine distance of word vectors and the edge weights of the graph structure, fully integrating the information in the semantic space and the graph structure space, enhancing the expression ability of the relevance matrix. The eigenvector matrix is respectively subjected to semantic feature decomposition and visual feature reconstruction to obtain complementary scene description vectors and scene structure vectors, solving the problem that traditional methods cannot simultaneously take into account semantic expression and visual structure; In the prior art, simple keyword matching or single feature extraction methods are usually adopted, which are difficult to accurately capture the complex semantic relationships and visual structure information in the script, resulting in the generated scene not matching the semantics of the original script and insufficient visual expressiveness; This embodiment significantly improves the semantic expression ability and visual structured representation quality of the scene synthesis vector, provides a more accurate feature basis for subsequent video scene generation, and effectively solves the technical problems such as insufficient semantic-visual conversion, single feature expression, and one-sided optimization objective in traditional methods.
[0041] In an alternative embodiment, Calculating the similarity between the scene synthesis vector and the reference sample in the video sample library to obtain a distance metric value. When the distance metric value is less than a preset threshold, generating a target video scene based on the current scene synthesis vector includes: Constructing a scene semantic gradient map, calculating the information entropy based on the normalized eigenvalue of the scene elements in the scene semantic gradient map, obtaining the importance distribution of the scene elements, dividing the scene into a key region and a transition region according to the information entropy and the importance distribution, extracting the first scene feature vector corresponding to the key region and the second scene feature vector corresponding to the transition region, and calculating the semantic association strength between the key region and the transition region based on the Euclidean distance between the first scene feature vector and the second scene feature vector; Extracting the scene synthesis vector and the reference sample in the video sample library, performing cosine distance measurement on the key region, performing Mahalanobis distance measurement based on the feature covariance matrix on the transition region, weighting the cosine distance measurement value of the key region and the Mahalanobis distance measurement value of the transition region by a preset regional weight coefficient, weighting the semantic association strength by a pre-obtained semantic correlation coefficient, and combining the weighted values to obtain the distance metric value; Calculating the distance gradient of the distance metric value, adding the product of the distance gradient and the update step size to the original regional boundary to obtain an updated regional boundary, and judging whether the distance metric value is less than a preset threshold based on the updated regional boundary. If it is less, generating the target video scene based on the scene synthesis vector.
[0042] Analyze the feature distribution of each pixel point in the original scene, filter the scene image using a convolution kernel, and extract the edge features of scene elements. For a pixel point (x, y), its gradient intensity value can be obtained by taking the square root of the sum of the squares of the horizontal and vertical gradients. The gradient direction is determined by the arctangent value of the horizontal gradient and the vertical gradient. For example, for a scene containing people, buildings, and natural environments, after processing with a 5×5 Gaussian convolution kernel, the edge gradient value distribution of each element can be obtained, thus constructing a complete scene semantic gradient map.
[0043] Calculate the information entropy and importance distribution based on the scene semantic gradient map, normalize the feature values of each element in the scene so that all feature values are distributed in the interval [0, 1]. For the normalized feature values, calculate the scene information entropy. Assume there are n elements in the scene, and the normalized feature value of each element i is pi, then the scene information entropy can be calculated by taking the logarithm product of the normalized feature values of each element. In practical applications, for a video frame containing 8 main scene elements, the normalized feature values are [0.15, 0.22, 0.08, 0.12, 0.18, 0.09, 0.06, 0.10] respectively, and the calculated information entropy is approximately 2.87. Through information entropy analysis, determine the importance distribution of scene elements. The elements with higher importance values occupy more important positions in visual perception.
[0044] Divide the scene into key regions and transition regions according to the information entropy and importance distribution. Set the importance threshold to 0.15. The regions where the elements are above this threshold are marked as key regions, such as the human face, key action regions, etc.; the regions below this threshold are marked as transition regions, such as the background environment, non-focus objects, etc. In actual division, a video frame with a resolution of 1920×1080 may divide the 800×600 pixel range in the central region into key regions, and the rest as transition regions.
[0045] Extract the first scene feature vector of the key region and the second scene feature vector of the transition region, and use a pre-trained deep neural network model to extract features from the key region and the transition region respectively. The key region adopts a high-resolution processing mode, and the dimension of the extracted feature vector is 512; the transition region adopts a standard-resolution processing mode, and the dimension of the extracted feature vector is 256. For example, the first five dimensions of the feature vector in the key region may be [0.78, 0.45, 0.23, 0.67, 0.39], and the first five dimensions of the feature vector in the transition region may be [0.32, 0.56, 0.41, 0.29, 0.52].
[0046] Calculate the semantic association strength between the key region and the transition region based on the Euclidean distance between the first scene feature vector and the second scene feature vector. Calculate the square root of the sum of the squares of the differences of the corresponding position elements of the two feature vectors to obtain the Euclidean distance value. To process feature vectors of different dimensions, zero-padding is used to extend the shorter-dimensional vector to the longer dimension. In actual calculation, the Euclidean distance value that may be obtained for the feature vectors of the two regions is 8.74, and then this distance value is mapped to the interval [0, 1] to obtain a semantic association strength of 0.64.
[0047] Execute different distance measurement strategies for the key region and the transition region. Extract the scene synthesis vector and the reference samples in the video sample library, perform cosine distance measurement on the key region, and calculate the cosine value of the included angle between the feature vectors. For example, the cosine distance between the key region feature vector and the reference sample feature vector may be 0.13. Perform Mahalanobis distance measurement based on the feature covariance matrix on the transition region, considering the correlation between features. Construct a feature covariance matrix with a size of 256×256, and the obtained Mahalanobis distance value may be 3.28.
[0048] Perform weighted combination through the regional weight coefficient. Preset the key region weight coefficient to 0.7 and the transition region weight coefficient to 0.3, weight the distance measurement values of the two regions, and the pre-obtained semantic association coefficient is 0.5, and weight the semantic association strength of 0.64 to get 0.32. Combine the weighted key region distance measurement value, the transition region distance measurement value, and the semantic association strength to obtain a final distance measurement value of 1.46.
[0049] Calculate the distance gradient of the distance measurement value and update the region boundary. During the iteration process, calculate the difference between the current distance measurement value and the distance measurement value of the previous iteration to obtain the distance gradient. Assume that the distance measurement value of the previous iteration is 1.52 and the current value is 1.46, then the distance gradient is -0.06. Set the update step size to 5 pixels, and superimpose the product of the distance gradient and the update step size, -0.3 pixels, on the original region boundary. For example, the original key region boundary is a rectangle [400, 300, 1200, 900], and the updated boundary is [400.3, 300.3, 1199.7, 899.7].
[0050] Judge whether the distance measurement value is less than the preset threshold based on the updated region boundary. The preset threshold is 1.5, and the current distance measurement value of 1.46 is less than the threshold. Therefore, generate the target video scene based on the scene synthesis vector. During the generation process, use a deep generation model, take the scene synthesis vector as the conditional input, and generate a target video scene with a resolution of 1920×1080 and a frame rate of 30fps. The scene content meets the original semantic requirements and has high-quality visual effects.
[0051] In this embodiment, by constructing a scene semantic gradient map and partitioning the scene based on information entropy, the refined processing and accurate matching of the video scene are realized. By dividing the scene into key regions and transition regions and adopting different distance measurement methods for different regions, the core semantic information and secondary transition information in the scene can be captured more accurately, and the semantic coherence between the key region and the transition region is fully considered, ensuring that the generated video scene not only maintains the unity of the overall style in visual performance but also accurately expresses the semantic transition between different regions. Through the calculation of the distance gradient and the dynamic update of the region boundary, the region partitioning can be adaptively adjusted, making the finally generated video scene more conform to the semantic structure of the original script.
[0052] In an alternative embodiment, Based on the pre-constructed feature pyramid and the target video scene, spatio-temporal features are extracted layer by layer from each layer by combining the inter-layer correlation relationships in the multi-layer scene graph to obtain a sequence of feature maps. Adaptive feature fusion is performed on the sequence of feature maps to obtain a scene feature representation and calculate an inter-frame similarity matrix. The self-attention mechanism is applied to the similarity matrix to obtain scene key frame indices and segment the video content. A depth-map-based transition effect is inserted between adjacent segments to obtain a combined video sequence, including: Obtain the feature maps of each layer in the pre-constructed feature pyramid based on the target video scene. Combine the inter-layer correlation relationships in the multi-layer scene graph, extract spatio-temporal features from the feature maps through multi-scale convolution, and perform convolution operations on the feature maps using convolution kernel parameters corresponding to each layer and superimpose bias terms to obtain the sequence of feature maps; Perform a matrix multiplication operation on the query matrix of the current layer of the feature pyramid and the key matrix of the adjacent layer. After normalizing the calculation result, multiply it with the value matrix of the adjacent layer to obtain feature correlation data. Fuse the feature correlation data with the sequence of feature maps to obtain an association-enhanced feature sequence; Perform average pooling and multi-layer perceptron processing on the association-enhanced feature sequence to obtain channel attention weights. Weighted sum the channel attention weights with the association-enhanced feature sequence to obtain the scene feature representation. Calculate the inter-frame similarity matrix based on the scene feature representation and apply the self-attention mechanism to extract the scene key frame indices; Segment the video content according to the scene key frame indices to obtain video segments. Generate a transition effect based on the time-varying interpolation coefficient calculated by the cosine function and the depth maps of adjacent video segments. Combine the video segments with the transition effect to obtain the combined video sequence.
[0053] Figure 3 It is a schematic diagram of the inter-layer correlation and adaptive feature fusion of the feature pyramid, as Figure 3As shown, obtain the target video scene and the pre - constructed feature pyramid. The feature pyramid contains multiple scale levels, for example, it can be set to a three - layer structure, corresponding to high, medium, and low resolution features respectively. In practical applications, for a video scene with a resolution of 1920×1080, the first layer can maintain the original resolution, the second layer is downsampled to 960×540, and the third layer is further downsampled to 480×270.
[0054] For the feature extraction process, use multi - scale convolution operations on each layer of the feature pyramid to extract spatio - temporal features. Taking the first layer as an example, convolution kernels with sizes of 3×3, 5×5, and 7×7 are used for feature extraction, and the convolution kernel parameters are set to {W 11 , W 12 , W 13}, and the corresponding bias terms are {b 11 , b 12 , b 13}. For the input feature map F1, the output feature F'1 = W 11 *F1+W 12 *F1+W 13 *F1+b 11 +b 12 +b 13 is obtained through convolution operation, where "*" represents the convolution operation. Similarly, similar operations are performed on the second and third layers respectively to obtain F'2 and F'3. In actual implementation, the number of convolution kernels can be set to 64, the number of feature channels is 128, and the ReLU activation function is used.
[0055] Enhance the feature representation by combining the inter - layer correlation relationships in the multi - layer scene graph, and perform correlation calculations between the query matrix of the current layer of the feature pyramid and the key matrices of adjacent layers. Taking the second layer as an example, a query matrix Q2 is generated from F'2, and key matrices K1, K3 and value matrices V1, V3 are generated from the first layer F'1 and the third layer F'3 respectively. Calculate the matrix product of Q2 and K1 to obtain the attention score S 12 , divide S 12 by 8 (the square root of the feature dimension) for normalization, and then multiply it by V1 to obtain the feature correlation data A 12 . Similarly, A 23 is calculated. A 12 and A 23 are weighted and fused into F'2 to obtain the correlation - enhanced feature sequence E2 = F'2+0.5×A 12 +0.5×A 23 . The same processing method is adopted for other levels.
[0056] To achieve adaptive feature fusion, a channel attention mechanism is performed on the feature sequence with enhanced correlation. Taking E2 as an example, global average pooling is performed on E2 to obtain the feature vector G2, which is then processed through a two-layer fully connected network: G'2 = W 22 (ReLU(W 21 (G2))), where W 21 and W 22 are the parameters of the fully connected layers. The channel attention weight α2 is obtained by normalizing G'2 to the range of 0 - 1 through the Sigmoid function. Multiply α2 and E2 along the channel dimension to get the weighted feature E"2. Similarly, E"1 and E"3 are obtained for other levels. Combine the weighted features of all levels according to the weight coefficients: FS = 0.5×E"1 + 0.3×E"2 + 0.2×E"3 to obtain the scene feature representation.
[0057] Calculate the inter-frame similarity matrix based on the scene feature representation. For any two frames i and j in the video, extract the corresponding feature vectors FS_i and FS_j, and calculate the cosine similarity: Sim(i, j) = (FS_i·FS_j) / (||FS_i||×||FS_j||) to obtain the inter-frame similarity matrix M. For example, for a video clip containing 100 frames, a 100×100 similarity matrix is generated, where M[50, 51] = 0.92 indicates a high similarity between the 50th frame and the 51st frame.
[0058] Apply the self-attention mechanism to the similarity matrix to extract key frames. Calculate the average similarity of each frame with all other frames, select the frames with significant similarity changes as scene transition points, calculate the first derivative of the similarity, and when the absolute value of the derivative exceeds a preset threshold (such as 0.15), mark this frame as a key frame. For example, in a 30-second video, the system may mark scene transition points at the 3rd second, 12th second, and 25th second, thus dividing the video into four segments.
[0059] Generate a depth map-based transition effect between adjacent video segments. Extract the depth maps D1 and D2 from the end frame and start frame of adjacent segments, and normalize the depth value range to 0 - 1. Based on the time-varying interpolation coefficient λ(t) = 0.5 - 0.5×cos(πt) (t varies from 0 to 1), calculate the depth map of the transition frame: D(t) = (1 - λ(t))×D1 + λ(t)×D2. Use the depth map to control pixel blending to generate a natural transition effect. In practical applications, for the transition between two scenes, the transition duration can be set to 1 second (corresponding to 30 frames), and the depth blending weight is calculated frame by frame and applied to the original pixels.
[0060] Combine the video segments and the generated transition effects in chronological order to obtain a combined video sequence with a coherent transition effect.
[0061] In this embodiment, through the feature fusion mechanism of multi-scale feature extraction and correlation enhancement, high-quality video scene segmentation and the generation of natural and smooth transition effects are achieved. By combining the inter-layer correlation relationship between the feature pyramid and the multi-layer scene graph, the spatio-temporal features of the video scene can be captured at different scales, and the detail information and overall structure of the scene are retained. The query-key-value matrix operation mechanism between the layers of the feature pyramid enhances the information exchange between features at different scales, effectively solving the problem of information islands caused by the independent processing of features at different scales in traditional methods. Based on the enhanced scene feature representation, the inter-frame similarity can be accurately calculated and the key frames can be intelligently identified using the self-attention mechanism, thus realizing the natural segmentation of the video content and significantly improving the visual quality and narrative fluency of the generated video.
[0062] In an alternative embodiment, Generating a transition effect based on the time-varying interpolation coefficient calculated by the cosine function and the depth maps of adjacent video segments, and combining the video segments with the transition effect to obtain the combined video sequence includes: For the transition time period between adjacent video segments, calculate the time-varying interpolation coefficient based on the ratio of the current time point to the transition duration using the cosine function; Collect the scene information of the adjacent video segments, perform position encoding on the spatial position and viewing direction of any point in the current scene through a multi-layer perceptron to obtain the encoded features, and construct a neural radiance field representation containing spatial density features and color features based on the encoded features; Obtain the depth maps of the adjacent video segments, calculate the depth consistency loss between the spatial density integral of the camera ray and the depth maps, optimize the neural radiance field representation based on the depth consistency loss, apply the time-varying interpolation coefficient to the optimized neural radiance field representation for weighted fusion, perform volume rendering on the fused features, and generate an intermediate view sequence by calculating the cumulative transmittance and the integral of the spatial density and color features; Calculate the gradient difference of the depth maps to obtain the depth consistency weight, apply the depth consistency weight to the intermediate view sequence and the depth maps weighted by the time-varying interpolation coefficient for adaptive fusion to obtain the transition effect, and combine the adjacent video segments and the transition effect in chronological order to obtain the combined video sequence.
[0063] For the generation of transition effects between adjacent video segments, each time point within the transition time period is processed. Assuming the transition duration is T seconds and the current time point is t seconds, the system calculates a time-varying interpolation coefficient α, where the calculation of the α value is based on the cosine function. Specifically, α = 0.5 * (1 - cos(π * t / T)). For example, when t = 0, α = 0; when t = T / 2, α = 0.5; when t = T, α = 1. This smoothly varying interpolation coefficient ensures a gradual change effect during the transition, avoiding visual discomfort caused by sudden changes.
[0064] During the process of obtaining the scene information of adjacent video segments, the image data and camera parameters in the adjacent video segments are collected. For any three-dimensional spatial point P(x, y, z) and viewing direction D(θ, φ) in the scene, a multi-layer perceptron is used for position encoding to generate a high-dimensional feature vector. The multi-layer perceptron contains 4 fully connected layers, with each layer containing 256 neurons, and the ReLU activation function is used. The input layer receives a 6-dimensional vector of the spatial position (x, y, z) and the viewing direction (θ, φ), which is extended to 96-dimensional features through position encoding. The output layer produces a neural radiance field representation containing spatial density features and RGB color features. For example, for the point P(1.5, 2.3, 0.8) and the viewing direction D(0.7, 1.2) in the scene, after position encoding and processing by the multi-layer perceptron, the density feature σ = 0.85 and the color feature c = (0.7, 0.6, 0.5) are obtained.
[0065] The depth maps of adjacent video segments are obtained through existing depth estimation methods. To optimize the neural radiance field representation, the depth consistency loss between the spatial density integral of the camera ray and the depth map is calculated. Assuming the depth value of a certain point in the depth map is d_gt = 5.2 meters, and the expected depth value calculated through the neural radiance field is d_pred = 5.5 meters, then the depth consistency loss can be expressed as |d_gt - d_pred| = 0.3 meters. During the training process, the Adam optimizer is used, with the initial learning rate set to 0.001, and 2000 iterations of optimization are performed. The learning rate is reduced to 0.5 times the original every 500 iterations until the depth consistency loss converges to less than 0.05 meters. The optimized neural radiance field can more accurately represent the geometric structure of the scene.
[0066] For the generation of intermediate views during the transition process, the time-varying interpolation coefficient α is applied to the neural radiance field representations of two adjacent video segments for weighted fusion. For example, for the same point in space, if the density feature σ1 = 0.8 and the color feature c1 = (0.7, 0.6, 0.5) in the first video segment, and the density feature σ2 = 0.6 and the color feature c2 = (0.5, 0.4, 0.6) in the second video segment, when α = 0.3, the fused density feature σ = 0.3 * 0.6 + (1 - 0.3) * 0.8 = 0.74, and the color feature c = 0.3 * (0.5, 0.4, 0.6) + (1 - 0.3) * (0.7, 0.6, 0.5) = (0.64, 0.54, 0.53).
[0067] Through volume rendering technology, 128 points are sampled along the camera ray, and the integral of the cumulative transmittance with the spatial density and color features is calculated to generate the intermediate view. For a video frame with a resolution of 1920×1080, the sampling interval for each ray is 0.05 meters, the near-plane distance is 0.1 meters, and the far-plane distance is 10 meters. The sequence of intermediate views generated by the volume rendering process smoothly transitions during the transition process, but there may be artifacts caused by depth inconsistencies.
[0068] To solve the depth inconsistency problem, the gradient difference of the depth map is calculated to obtain the depth consistency weight. The depth gradient is calculated by the Sobel operator. For example, if the depth values of adjacent pixels are 5.2 meters and 5.3 meters respectively, the gradient value is 0.1 meter. For the depth maps of two video segments, if the gradient difference at the corresponding positions is large, it indicates that there is depth inconsistency in that area. The system sets the threshold to 0.2 meters. When the gradient difference is greater than the threshold, the depth consistency weight w is set to 0.8, otherwise it is set to 0.2.
[0069] The depth consistency weight w is applied to the intermediate view sequence and the depth map weighted by the time-varying interpolation coefficient for adaptive fusion to obtain the transition effect. For example, for the pixel point (800, 600), if the color value of the intermediate view is (0.64, 0.54, 0.53), the depth consistency weight w = 0.8, the time-varying interpolation coefficient α = 0.3, the color value of the first video segment is (0.7, 0.6, 0.5), and the color value of the second video segment is (0.5, 0.4, 0.6), then the fused color value is 0.8 * (0.64, 0.54, 0.53) + (1 - 0.8) * (0.3 * (0.5, 0.4, 0.6) + (1 - 0.3) * (0.7, 0.6, 0.5)) = (0.628, 0.532, 0.534).
[0070] After generating the transition effect, segment adjacent videos and combine them with the transition effect in chronological order to obtain the final combined video sequence. For example, for two video segments each lasting 5 seconds, if the transition duration is 1 second, the final combined video sequence consists of the first 4.5 seconds of the first video segment, 1 second of the transition effect, and the last 4.5 seconds of the second video segment, with a total length of 10 seconds.
[0071] In this embodiment, through the neural radiance field technology and depth consistency optimization, a high-quality and natural video transition effect is achieved. The time-varying interpolation coefficient calculated based on the cosine function provides a smooth time transition characteristic, avoiding the mechanical feeling that may be brought by traditional linear interpolation, making the transition process present a more natural acceleration-deceleration rhythm. The scene representation method based on the neural radiance field breaks through the limitations of traditional transition technologies based on pixels or optical flow, and can accurately model the geometric structure and appearance characteristics of the scene in three-dimensional space. By performing time-varying weighted fusion on the optimized neural radiance field representation, smooth transitions between different perspectives and scenes are achieved, avoiding common geometric distortions and flickering artifacts in traditional transition effects.
[0072] Figure 4 It is a schematic diagram for performance comparison of video transition effects based on the neural radiance field. In terms of visual quality score, the neural radiance field method obtains 8.7 points, far higher than the optical flow transition (6.3 points) and traditional fade-in / fade-out (4.5 points). In terms of the geometric consistency index, the neural radiance field method performs particularly outstandingly, reaching 9.0 points, indicating that this technology can more accurately maintain the three-dimensional geometric structure of the scene and avoid the geometric distortion problems common in traditional methods.
[0073] Temporal smoothness is the index with the smallest gap among the methods, but the neural radiance field method still leads with 8.5 points, while the optical flow transition and traditional fade-in / fade-out are 7.1 points and 7.3 points respectively. This indicates that all methods have good performance in temporal coherence, but the time-varying interpolation of the cosine function of the neural radiance field provides a more natural transition effect.
[0074] The user experience score reflects the subjective feelings of the audience. The neural radiance field method obtains the highest 9.1 points, far exceeding the baseline level (5.0 points), indicating that this transition technology can significantly improve the viewing experience.
[0075] Generally speaking, the transition technology based on the neural radiance field has achieved obvious breakthroughs in aspects such as visual quality, geometric consistency, and user experience through advanced algorithms such as depth consistency optimization and adaptive fusion, providing a new technical path for high-quality video content production.
[0076] In an alternative embodiment, Perform style transfer on the combined video sequence, and the output video product includes: Obtain the combined video sequence, and extract the spatio-temporal features and the sequence of feature maps corresponding to the combined video sequence; Perform mapping and matching on the spatio-temporal features and the sequence of feature maps corresponding to the combined video sequence with reference to the feature space of the target style template to obtain the fused features after style transfer; Reconstruct and decode the fused features after style transfer to obtain a video product with a unified style.
[0077] Obtain a combined video sequence, which can be composed of multiple video clips shot in different scenes, with different shooting devices, or at different times. Taking an actual application scenario as an example, assume that a user has three videos: the first is a clip of a person walking shot outdoors in sunlight, with a duration of 5 seconds and a resolution of 1920×1080; the second is a close-up of a person shot indoors under a light, with a duration of 3 seconds and a resolution of 1280×720; the third is a landscape shot taken in the evening, with a duration of 4 seconds and a resolution of 3840×2160. There are obvious differences in hue, light, and atmosphere among the three videos, and style unification processing is required.
[0078] After obtaining the combined video sequence, use a deep convolutional neural network to analyze the video sequence and extract the spatio-temporal features and the sequence of feature maps. Spatio-temporal features refer to the feature representations in the time dimension between video frames and the spatial dimension within each frame, while the sequence of feature maps is a set of multi-level visual features extracted for each frame in the video. A pre-trained video feature extraction network is used, and the video feature extraction network contains multiple convolutional layers, pooling layers, and non-linear activation functions. For each input frame image, low-level features such as edges and textures are extracted through shallow convolution, intermediate-level features such as object parts are extracted through intermediate convolution, and high-level features at the semantic level are extracted through deep convolution. The extraction of spatio-temporal features is achieved through 3D convolution or by combining 2D convolution with long short-term memory networks to capture the temporal correlation between video frames.
[0079] Process the three videos separately. For the first outdoor video, the extracted features show characteristics of high brightness, high contrast, and a warm color tone; for the second indoor video, the extracted features show medium brightness, low contrast, and a cold color tone; for the third evening video, the extracted features show low brightness, medium contrast, and a magenta color tone. The extracted features are encoded into multi-dimensional feature vectors and feature maps, and the dimension depends on the neural network architecture used, usually between 512 and 2048.
[0080] Select a target style template as a reference. The target style template can be a video, an image, or a predefined style descriptor specified by the user. Assume that the user selects a movie clip as the target style template. The clip has a typical movie color style, characterized by medium brightness, high contrast, warm tones but without loss of details. Use the aforementioned feature extraction network to analyze the target style template and obtain its feature space representation.
[0081] In the mapping and matching stage, the style transfer algorithm is used to map the spatiotemporal features and feature map sequence of the combined video sequence to the feature space of the target style template. While maintaining the original video content, its visual style is adjusted to match the target template. The adaptive instance normalization technology is used to achieve style transfer by adjusting the mean and variance of the feature map. For the feature map of each video clip, its mean μ and standard deviation σ are calculated, and then the original feature map is normalized and readjusted according to the mean μ_target and standard deviation σ_target of the target style feature map to obtain the feature map after style transfer.
[0082] In order to maintain the coherence of the video sequence in the temporal dimension, the features between adjacent frames are subjected to temporal consistency constraints, the differences in features between adjacent frames are calculated, and a temporal smoothing term is introduced to make the style transfer process smoother in the temporal dimension and avoid flickering or incoherence.
[0083] For example, the high-brightness feature of the first outdoor video is adjusted to medium brightness, but the high contrast is maintained; the cool tones of the second indoor video are adjusted to warm tones, while the contrast is improved; the low brightness of the third evening video is appropriately increased, and the purple-red tones are adjusted to warm tones that are more in line with the movie style.
[0084] After completing the feature map matching, the fused features after style transfer are obtained. The fused features contain both the content information of the original video and the style features of the target style template. The generative network is used. The network consists of multiple deconvolution layers, upsampling layers and residual connections. The fused features are input into the generative network. Through layer-by-layer decoding and upsampling, the video frames with the target style are reconstructed. In order to ensure the quality of reconstruction, a multi-scale feature fusion mechanism is also introduced in the decoding process to combine features at different levels to retain more detailed information.
[0085] After processing, the three videos of different styles were unified into a movie-style video sequence, showing a medium brightness, high contrast, and warm-toned visual effect, while maintaining the content and narrative integrity of the original video. The final output video has a unified resolution of 1920×1080, a total length of 12 seconds, a frame rate of 30fps, an encoding format of H.264, and a file size of about 35MB.
[0086] In this embodiment, through the style transfer and fusion technology, the style unification of the video finished product is realized, which significantly improves the artistic expressiveness and visual consistency of the video. By extracting spatio-temporal features and multi-level feature maps from the combined video sequence, the structural information and temporal coherence of the video content are retained, providing a complete feature basis for subsequent style transfer. In the way of mapping and matching with the target style template as a reference, the style transfer process can not only accurately capture the feature distribution of the target style, but also maintain the content integrity of the original video, avoiding the common problems of content distortion and style imbalance in traditional style transfer methods. While ensuring the style effect, it can also maintain high-quality temporal coherence, avoiding problems such as style flickering and temporal instability, and greatly improving the overall quality and artistic expressiveness of the video finished product.
[0087] In the second aspect of the embodiment of the present invention, a video plot generation and scene synthesis system based on natural language processing is provided, including: The first unit is used to receive the script text input by the user, extract the scene semantic features, and construct the multi-level scene map with the scene semantic features; The second unit is used to extract the keyword sequence based on the multi-level scene map, calculate the semantic correlation degree of the keyword sequence to obtain the feature vector matrix, respectively perform semantic feature decomposition and visual feature reconstruction on the feature vector matrix to obtain the scene description vector and the scene structure vector, alternately update the feature representation according to the scene description vector and the scene structure vector to obtain the scene synthesis vector, calculate the similarity between the scene synthesis vector and the reference sample in the video sample library to obtain the distance metric value, and when the distance metric value is less than the preset threshold, generate the target video scene based on the current scene synthesis vector; The third unit is used to extract spatio-temporal features at each layer based on the pre-constructed feature pyramid and the target video scene, combine the inter-layer correlation relationship in the multi-level scene map to obtain the feature map sequence, perform adaptive feature fusion on the feature map sequence to obtain the scene feature representation and calculate the inter-frame similarity matrix, apply the self-attention mechanism to the similarity matrix to obtain the scene key frame index and segment the video content, and insert the depth map-based transition effect between adjacent segments to obtain the combined video sequence; The fourth unit is used to perform style transfer on the combined video sequence and output the video finished product.
[0088] In the third aspect of the embodiment of the present invention, an electronic device is provided, including: A processor and a memory for storing processor-executable instructions, wherein the processor is configured to call the instructions stored in the memory to execute the method described above.
[0089] In a fourth aspect of the embodiments of the present invention, there is provided a computer-readable storage medium having computer program instructions stored thereon, and when the computer program instructions are executed by a processor, the foregoing method is implemented.
[0090] The present invention may be a method, an apparatus, a system, and / or a computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions for performing various aspects of the present invention loaded thereon.
[0091] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for video plot generation and scene synthesis based on natural language processing, characterized in that It includes: Receiving the script text input by the user, extracting the scene semantic features, and constructing the scene semantic features into a multi-layer scene graph; Extracting a keyword sequence based on the multi-layer scene graph, calculating the semantic relevance of the keyword sequence to obtain a feature vector matrix, respectively performing semantic feature decomposition and visual feature reconstruction on the feature vector matrix to obtain a scene description vector and a scene structure vector, alternately updating the feature representation according to the scene description vector and the scene structure vector to obtain a scene synthesis vector, calculating the similarity between the scene synthesis vector and the reference sample in the video sample library to obtain a distance metric value, and when the distance metric value is less than a preset threshold, generating a target video scene based on the current scene synthesis vector; Based on the pre-constructed feature pyramid and the target video scene, combining the inter-layer association relationships in the multi-layer scene graph to extract spatio-temporal features at each layer to obtain a sequence of feature maps, performing adaptive feature fusion on the sequence of feature maps to obtain a scene feature representation and calculating an inter-frame similarity matrix, applying a self-attention mechanism to the similarity matrix to obtain scene key frame indices and segmenting the video content, and inserting a depth map-based transition effect between adjacent segments to obtain a combined video sequence; Performing style transfer on the combined video sequence and outputting the finished video.
2. The method according to claim 1, wherein Receiving the script text input by the user and extracting the scene semantic features, and constructing the scene semantic features into a multi-layer scene graph includes: Receiving the script text input by the user, performing semantic word segmentation on the script text to obtain a word sequence, and encoding the word sequence based on a pre-trained language model to obtain initial semantic features; Inputting the initial semantic features into a multi-head attention network, calculating the feature representations of different attention heads in parallel, and performing hierarchical clustering on the feature representations to obtain scene semantic features, where each attention head respectively focuses on the semantic information of character interaction, scene attributes, and plot development; Constructing a multi-layer scene graph based on the scene semantic features, calculating the neighborhood relationship of nodes in the graph using a graph attention network, dynamically aggregating nodes with similar semantic features into a character relationship layer, a scene attribute layer, and a plot development layer, and determining the inter-layer association relationships.
3. The method according to claim 1, wherein Extracting a keyword sequence based on the multi-layer scene graph, calculating the semantic relevance of the keyword sequence to obtain a feature vector matrix, respectively performing semantic feature decomposition and visual feature reconstruction on the feature vector matrix to obtain a scene description vector and a scene structure vector, and alternately updating the feature representation according to the scene description vector and the scene structure vector to obtain a scene synthesis vector includes: Extracting node features from the multi-layer scene graph, calculating the information entropy between each hierarchical node and its adjacent nodes to obtain node weights, performing hierarchical clustering based on the node weights, and selecting the cluster center points to form a keyword sequence; Calculating the semantic relevance of word pairs in the keyword sequence to obtain a relevance matrix, where the semantic relevance is obtained by multiplying the cosine distance of word vectors by the edge weights between nodes in the graph, and obtaining a feature vector matrix by performing weight mapping and non-linear transformation on the relevance matrix; Perform semantic feature decomposition on the feature vector matrix to obtain a scene description vector, and perform visual feature reconstruction to obtain a scene structure vector; Alternately iterate and optimize the scene description vector and the scene structure vector to obtain a candidate feature representation. Calculate the joint optimization objective of semantic consistency loss, visual reconstruction loss, and temporal continuity loss based on the candidate feature representation. Update the candidate feature representation by minimizing the joint optimization objective until convergence, and determine the converged feature representation as the scene synthesis vector.
4. The method according to claim 1, characterized in that Calculate the similarity between the scene synthesis vector and the reference samples in the video sample library to obtain a distance metric value. When the distance metric value is less than a preset threshold, generating a target video scene based on the current scene synthesis vector includes: Construct a scene semantic gradient map, calculate the information entropy based on the normalized feature values of the scene elements in the scene semantic gradient map, obtain the importance distribution of the scene elements, divide the scene into key regions and transition regions according to the information entropy and the importance distribution, extract the first scene feature vector corresponding to the key region and the second scene feature vector corresponding to the transition region, and calculate the semantic association strength between the key region and the transition region based on the Euclidean distance between the first scene feature vector and the second scene feature vector; Extract the scene synthesis vector and the reference samples in the video sample library, perform cosine distance measurement on the key region, perform Mahalanobis distance measurement based on the feature covariance matrix on the transition region, weight the cosine distance measurement value of the key region and the Mahalanobis distance measurement value of the transition region through a preset regional weight coefficient, weight the semantic association strength by combining a pre-obtained semantic correlation coefficient, and combine the weighted values to obtain the distance metric value; Calculate the distance gradient of the distance metric value, superimpose the product of the distance gradient and the update step size on the original region boundary to obtain an updated region boundary, and determine whether the distance metric value is less than a preset threshold based on the updated region boundary. If it is less, generate the target video scene based on the scene synthesis vector.
5. The method according to claim 1, wherein Based on a pre-constructed feature pyramid and the target video scene, combined with the inter-layer association relationship in the multi-layer scene graph, extract spatio-temporal features at each layer to obtain a sequence of feature maps, perform adaptive feature fusion on the sequence of feature maps to obtain a scene feature representation and calculate an inter-frame similarity matrix, apply a self-attention mechanism to the similarity matrix to obtain scene key frame indices and segment the video content, and insert a depth map-based transition effect between adjacent segments to obtain a combined video sequence including: Obtain the feature maps of each layer in the pre-constructed feature pyramid based on the target video scene, combined with the inter-layer association relationship in the multi-layer scene graph, extract spatio-temporal features from the feature maps through multi-scale convolution, and perform convolution operations on the feature maps with convolution kernel parameters corresponding to each layer and superimpose bias terms to obtain the sequence of feature maps; Perform a matrix multiplication operation on the query matrix of the current layer of the feature pyramid and the key matrix of the adjacent layer, multiply the calculation result after normalization with the value matrix of the adjacent layer to obtain feature correlation data, and fuse the feature correlation data with the feature map sequence to obtain an associated enhanced feature sequence; Perform average pooling and multi-layer perceptron processing on the associated enhanced feature sequence to obtain channel attention weights, perform weighted summation of the channel attention weights and the associated enhanced feature sequence to obtain the scene feature representation, calculate the inter-frame similarity matrix based on the scene feature representation, and apply the self-attention mechanism to extract the scene key frame index; Segment the video content according to the scene key frame index to obtain video segments, generate a transition effect based on the time-varying interpolation coefficient calculated by the cosine function and the depth maps of adjacent video segments, and combine the video segments and the transition effect to obtain the combined video sequence.
6. The method according to claim 5, wherein Generating a transition effect based on the time-varying interpolation coefficient calculated by the cosine function and the depth maps of adjacent video segments, and combining the video segments and the transition effect to obtain the combined video sequence includes: For the transition time period between adjacent video segments, calculate the time-varying interpolation coefficient based on the ratio of the current time point to the transition duration using the cosine function; Collect the scene information of the adjacent video segments, perform position encoding on the spatial position and viewing direction of any point in the current scene through a multi-layer perceptron to obtain encoded features, and construct a neural radiance field representation including spatial density features and color features based on the encoded features; Obtain the depth maps of the adjacent video segments, calculate the depth consistency loss between the spatial density integral of the camera rays and the depth maps, optimize the neural radiance field representation based on the depth consistency loss, apply the time-varying interpolation coefficient to the optimized neural radiance field representation for weighted fusion, perform volume rendering on the fused features, and generate an intermediate view sequence by calculating the integral of the cumulative transmittance with the spatial density and color features; Calculate the gradient difference of the depth maps to obtain the depth consistency weight, apply the depth consistency weight to the intermediate view sequence and the depth maps weighted by the time-varying interpolation coefficient for adaptive fusion to obtain the transition effect, and combine the adjacent video segments and the transition effect in chronological order to obtain the combined video sequence.
7. The method according to claim 1, characterized in that, Performing style transfer on the combined video sequence and outputting the video finished product includes: Obtain the combined video sequence, and extract the spatio-temporal features and feature map sequence corresponding to the combined video sequence; Perform mapping and matching on the spatio-temporal features and feature map sequence corresponding to the combined video sequence with reference to the feature space of the target style template to obtain the fused features after style transfer; Reconstruct and decode the fused features after style transfer to obtain a video finished product with unified style.
8. A video plot generation and scene synthesis system based on natural language processing, which is used to implement the method described in any one of the foregoing claims 1-7, characterized in that, Including: The first unit is used to receive the script text input by the user and extract the scene semantic features, and construct the scene semantic features into a multi-layer scene graph; A second unit, configured to extract a keyword sequence based on the multi-layer scene graph, calculate the semantic relevance of the keyword sequence to obtain a feature vector matrix, perform semantic feature decomposition and visual feature reconstruction on the feature vector matrix respectively to obtain a scene description vector and a scene structure vector, alternately update the feature representation according to the scene description vector and the scene structure vector to obtain a scene synthesis vector, calculate the similarity between the scene synthesis vector and a reference sample in the video sample library to obtain a distance metric value, and when the distance metric value is less than a preset threshold, generate a target video scene based on the current scene synthesis vector; A third unit, configured to extract spatio-temporal features at each layer based on a pre-constructed feature pyramid and the target video scene, in combination with the inter-layer correlation relationship in the multi-layer scene graph, to obtain a sequence of feature maps, perform adaptive feature fusion on the sequence of feature maps to obtain a scene feature representation and calculate an inter-frame similarity matrix, apply a self-attention mechanism to the similarity matrix to obtain scene key frame indices and segment the video content, and insert a depth-map-based transition effect between adjacent segments to obtain a combined video sequence; A fourth unit, configured to perform style transfer on the combined video sequence and output a finished video.
9. An electronic device, characterized in that, Comprising: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Environment modeling method and system for realizing intelligent security and protection
CN119863588A
Multimodal heterogeneous feature fusion-based compact video event description method
WO2023050295A1
Cited By
Film and television script overall planning and paging method and system based on natural language processing
CN120744137A
A method and system for pagination of film and television scripts based on natural language processing
CN120744137B
Context semantic modulation method and system for video subtitle generation
CN120997741A
A method and system for context semantic modulation for video caption generation
CN120997741B
Automatic video generation system and method based on AI Agent multi-mode cooperative control
CN121126084A