Video understanding method and device, computer device and storage medium

By employing multimodal data processing and a self-multi-head attention mechanism, and combining video, text, and audio data, the problem of accuracy in understanding film and television plots was solved, enabling in-depth understanding and multi-dimensional analysis of the plot.

CN120726542BActive Publication Date: 2025-11-11ASPIRE TECH (SHENZHEN) LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511211470.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-28
Publication Date
2025-11-11
Estimated Expiration
2045-08-28

AI Technical Summary

Technical Problem

Existing technologies have low accuracy in understanding film and television plots, making it difficult to fully grasp complex internal relationships and deeper meanings, and unable to deeply understand character relationships, plot development, and thematic ideas.

Method used

We employ a multimodal data processing approach, combining self-attention and multi-head attention mechanisms, to achieve deep understanding using video, text, and audio data through feature extraction, enhancement, fusion, and weight adjustment.

Benefits of technology

It improves the accuracy of understanding film and television plots, enabling in-depth exploration of the deeper meaning of the plot and providing multi-dimensional plot analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120726542B_ABST
    Figure CN120726542B_ABST
Patent Text Reader

Abstract

This invention discloses a video understanding method, comprising: acquiring multimodal data corresponding to the video to be analyzed; performing feature extraction processing based on the multimodal data to obtain modal features corresponding to each type of multimodal data; enhancing the modal features through a self-attention mechanism to obtain enhanced modal features; performing feature fusion processing on the enhanced modal features through a multi-head attention mechanism to obtain initial fused features; adjusting the weights of each enhanced modal feature in the initial fused features based on the similarity between modal features to obtain target fused features; and performing inference based on the target fused features to obtain the understanding result of the video to be analyzed. By combining multimodal feature fusion and dynamic weight adjustment mechanisms with self-attention and multi-head attention mechanisms to achieve cross-modal information complementarity, this method can fully utilize the complementarity of multimodal data, improve the accuracy of video understanding, and delve deeper into the deeper meaning of the plot.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing, and more particularly to a video understanding method, apparatus, computer device, and storage medium. Background Technology

[0002] In the field of video understanding, such as film and television plots, traditional methods primarily rely on manual analysis. Professionals watch the content and interpret and summarize the plot based on their knowledge and experience. Some early AI technologies attempted to be applied to film and television plot understanding, but these typically employed single-modal analysis methods. For example, they might only perform image recognition on video frames, analyzing visual information such as scenes and character movements; or only perform speech recognition and sentiment analysis on audio, extracting dialogue content and emotional features of speech; or simply perform semantic analysis on subtitle text to understand the surface-level textual meaning of the plot.

[0003] Traditional manual methods of plot understanding are not only extremely inefficient, but also suffer from significant discrepancies in interpretation among different individuals due to differences in knowledge background, comprehension ability, and subjective feelings. This makes it difficult to meet the demands of rapid and accurate processing of large-scale film and television content. Existing AI-based plot understanding technologies, limited to a single modality or only performing superficial semantic analysis, cannot fully grasp the complex internal relationships and deeper meanings within a plot. For example, they struggle to accurately capture subtle emotional changes between characters, accurately interpret metaphors and symbolic meanings, and fail to provide a comprehensive and in-depth understanding of character relationships, plot development, and thematic themes. Therefore, existing solutions for video understanding exhibit low accuracy.

[0004] Therefore, there is an urgent need to propose a video understanding method that can effectively improve the accuracy of video understanding. Summary of the Invention

[0005] Therefore, it is necessary to provide a video understanding method, apparatus, computer device, and storage medium to address the aforementioned technical problems and solve the problem of low video understanding accuracy in traditional video understanding methods.

[0006] A video understanding method, the method comprising:

[0007] Obtain multimodal data corresponding to the video to be parsed, wherein the multimodal data includes at least two of the following: video data, text data, and audio data;

[0008] Based on the multimodal data, feature extraction processing is performed to obtain the modal features corresponding to each type of multimodal data;

[0009] The modal features are enhanced by using a self-attention mechanism to obtain enhanced modal features;

[0010] The enhanced modal features are fused using a multi-head attention mechanism to obtain initial fused features.

[0011] Based on the similarity between the modal features, the weights of each enhanced modal feature in the initial fusion feature are adjusted to obtain the target fusion feature;

[0012] Based on the target fusion features, inference is performed to obtain the understanding result of the video to be parsed.

[0013] Optionally, before obtaining the multimodal data corresponding to the video to be parsed, the method further includes:

[0014] Obtain the initial multimodal data corresponding to the video to be parsed, wherein the initial multimodal data includes at least two of the following: initial video data, initial text data, and initial audio data;

[0015] The initial multimodal data is preprocessed to obtain the multimodal data. When the initial multimodal data includes the initial video data, the preprocessing includes normalization and noise reduction. When the initial multimodal data includes the initial text data, the preprocessing includes word segmentation and part-of-speech tagging. When the initial multimodal data includes the initial audio data, the preprocessing includes noise reduction and segmentation.

[0016] Optionally, the step of adjusting the weights of each enhanced modal feature in the initial fusion feature based on the similarity between the modal features to obtain the target fusion feature includes:

[0017] The similarity between the modal features is input into a preset normalization function to calculate the target weight of each enhanced modal feature;

[0018] The weights of each enhanced modal feature in the initial fusion feature are adjusted to the corresponding target weights to obtain the target fusion feature.

[0019] Optionally, the step of inputting the similarity between the modal features into a preset normalization function to calculate the target weight of each enhanced modal feature includes:

[0020] An exponential operation is performed based on the similarity between the modal features to obtain the exponential value of each of the enhanced modal features;

[0021] The normalization factor is obtained by summing the exponent values.

[0022] The target weight of each enhanced modal feature is calculated based on the exponential value of each enhanced modal feature and the normalization factor.

[0023] Optionally, the step of reasoning based on the target fusion features to obtain the understanding result of the video to be parsed includes:

[0024] The target fusion features and the preset first prompt words are provided to a preset large language model so that the large language model can construct a knowledge graph of the video to be parsed.

[0025] The knowledge graph and the preset second prompt words are provided to the large language model so that the large language model can reason about the video to be parsed based on the knowledge graph and obtain the understanding result.

[0026] Optionally, after inferring the understanding result of the video to be parsed based on the target fusion features, the method further includes:

[0027] Obtain user feedback on the understanding results;

[0028] Based on the feedback information, the learnable parameters in the feature extraction process, the enhancement process, the feature fusion process, and the adjustment process are updated using the gradient descent algorithm to reduce the deviation between the understanding result and the expected result.

[0029] Optionally, the method further includes:

[0030] Construct multiple story understanding sub-models to implement the feature extraction process, the enhancement process, the feature fusion process, and the adjustment process;

[0031] Obtain the training set for each of the plot understanding sub-models. The training set includes sample multimodal data and result labels corresponding to the sample multimodal data. The sample multimodal data in the training sets of different plot understanding sub-models are different.

[0032] Based on the training set, each of the plot understanding sub-models is trained to obtain multiple trained plot understanding sub-models.

[0033] The multimodal data corresponding to the video to be analyzed is input into each of the trained story understanding sub-models to obtain the story understanding results output by the multiple story understanding sub-models;

[0034] The final understanding result of the video to be analyzed is obtained by weighted averaging and fusion calculation based on the plot understanding results output by the multiple plot understanding sub-models.

[0035] A video understanding device, the device comprising:

[0036] The first acquisition module is used to acquire multimodal data corresponding to the video to be parsed, wherein the multimodal data includes at least two of the following: video data, text data, and audio data.

[0037] The feature extraction module is used to perform feature extraction processing based on the multimodal data to obtain the modal features corresponding to each type of multimodal data;

[0038] The feature enhancement module is used to enhance the modal features through a self-attention mechanism to obtain enhanced modal features;

[0039] The feature fusion module is used to perform feature fusion processing on the enhanced modal features through a multi-head attention mechanism to obtain initial fused features;

[0040] The feature adjustment module is used to adjust the weight of each enhanced modal feature in the initial fusion feature based on the similarity between the modal features, so as to obtain the target fusion feature;

[0041] The reasoning module is used to perform reasoning based on the target fusion features to obtain the understanding result of the video to be parsed.

[0042] A computer device includes a memory, a processor, and computer-readable instructions stored in the memory and executable on the processor, wherein the processor implements the video understanding method described above when executing the computer-readable instructions.

[0043] A readable storage medium having computer-readable instructions stored thereon, which, when executed by a processor, implement the video understanding method.

[0044] The aforementioned video understanding method, apparatus, computer equipment, and storage medium acquire multimodal data corresponding to the video to be analyzed. The multimodal data includes at least two of video data, text data, and audio data. Feature extraction processing is performed based on the multimodal data to obtain modal features corresponding to each type of multimodal data. The modal features are enhanced using a self-attention mechanism to obtain enhanced modal features. The enhanced modal features are fused using a multi-head attention mechanism to obtain initial fused features. Based on the similarity between the modal features, the weights of each enhanced modal feature in the initial fused features are adjusted to obtain target fused features. Inference is performed based on the target fused features to obtain the understanding result of the video to be analyzed. By combining multimodal feature fusion and dynamic weight adjustment mechanisms with self-attention and multi-head attention mechanisms to achieve cross-modal information complementarity, the problem of one-sided single-modal analysis and insufficient feature fusion is solved. This approach fully utilizes the complementarity of multimodal data, improves the accuracy of video understanding, and allows for deeper exploration of the deeper meaning of the plot. Attached Figure Description

[0045] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0046] Figure 1 This is a flowchart illustrating a video understanding method according to an embodiment of the present invention;

[0047] Figure 2 This is a schematic diagram of the structure of a video understanding device in one embodiment of the present invention;

[0048] Figure 3 This is a schematic diagram of a computer device according to an embodiment of the present invention. Detailed Implementation

[0049] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0050] In one embodiment, such as Figure 1 As shown, a video understanding method is provided, including the following steps:

[0051] 101. Obtain the multimodal data corresponding to the video to be parsed.

[0052] In this embodiment of the invention, the video understanding method described above can be applied to a video understanding platform. The video understanding platform can be constructed from a server or server cluster. The server or server cluster can be any electronic device with functions such as multimodal data processing, multimodal data analysis, data storage, and data transmission. Correspondingly, the video understanding platform can also deploy a large language model or an application interface with a large language model. Through the application interface or the large language model directly deployed on the video understanding platform, reasoning about the target fusion features can be realized, thereby obtaining the understanding result.

[0053] The aforementioned multimodal data can include at least two of the following: video data, text data, and audio data. Specifically, video data, text data, and audio data can belong to the same film or television work. Within the same film or television work, video frames, corresponding audio segments, and subtitle text should be linked by timestamp alignment.

[0054] The video to be parsed can be uploaded by the user. Video data, audio data, and text data are extracted from the uploaded data to form the multimodal data.

[0055] 102. Perform feature extraction processing based on multimodal data to obtain the modal features corresponding to each type of multimodal data.

[0056] In this embodiment of the invention, the above-mentioned feature extraction process may specifically include video frame feature extraction, audio feature extraction, and text feature extraction, etc. The corresponding feature extraction process can be selected according to the actual data included in the above-mentioned multimodal data.

[0057] Specifically, video frame feature extraction can be performed using convolutional neural networks (CNNs), such as EfficientNet, to extract features from each frame of a video. EfficientNet balances network depth, width, and resolution through a composite scaling method, enabling the extraction of visual features with rich semantic information with relatively few computational resources. For example, it can identify information such as the categories and locations of people, scenes, and objects in video frames.

[0058] Audio feature extraction: Mel-frequency cepstral coefficients (MFCC) and short-time Fourier transform (STFT) can be used to process audio in video. MFCC can simulate the auditory characteristics of the human ear to extract the spectral features of audio; STFT can convert the audio signal to the time-frequency domain to obtain the time-frequency features of the audio. By combining the two methods, information such as pitch, timbre, and rhythm of audio can be comprehensively described, for example, distinguishing different types of audio such as dialogue, music, and sound effects.

[0059] Text feature extraction: For subtitles and speech-to-text in videos, pre-trained word embedding models, such as BERT (Bidirectional Encoder Representations from Transformers), can be used to convert the text into vector representations. Through large-scale unsupervised learning, BERT can capture semantic and syntactic information in the text, providing rich text features for subsequent plot understanding.

[0060] 103. Enhance modal features by using a self-attention mechanism to obtain enhanced modal features.

[0061] In this embodiment of the invention, the self-attention mechanism allows the model to automatically focus on the correlation between features of different modalities. For example, when analyzing a dialogue scene, the model can automatically focus on the correlation between the facial expressions and actions of people in video frames, the tone of voice in audio, and the dialogue content in text.

[0062] Specifically, the input modal features can be linearly transformed to obtain the corresponding query vector (Q), key vector (K), and value vector (V); the similarity between the query vector and the key vector can be calculated, and the similarity can be normalized to obtain the attention weight; the attention weight can be applied to the value vector to obtain the weighted enhanced features, thereby completing the modal feature enhancement process.

[0063] 104. Through a multi-head attention mechanism, the enhanced modal features are fused to obtain the initial fused features.

[0064] In this embodiment of the invention, the multi-head attention mechanism can learn feature representations from multiple different subspaces, further improving the effect of feature fusion. Compared with a single attention mechanism, the multi-head attention mechanism can capture the correlation between modal features from multiple perspectives, thereby improving the expressive power and robustness of the fused features.

[0065] Specifically, the enhanced modal features can be input into multiple independent attention heads. Each attention head performs a linear mapping on the input features to obtain different query vectors (Q), key vectors (K), and value vectors (V). Each attention head independently calculates its corresponding attention weights and outputs a feature representation in a subspace. Finally, the outputs of multiple attention heads are concatenated and mapped to a unified feature space through a linear transformation to obtain the initial fused features.

[0066] 105. Based on the similarity between modal features, the weights of each enhanced modal feature in the initial fusion feature are adjusted to obtain the target fusion feature.

[0067] In this embodiment of the invention, the similarity of modal features can refer to the cosine distance between different modal feature vectors in the embedding space, which can be implemented using vector dot product operations to quantify the semantic association strength between different modalities. Weight adjustment processing can refer to dynamically allocating the contribution ratio of each modality in the fused features based on similarity, which can be implemented using a normalized exponential function to focus the model on highly correlated modal combinations.

[0068] Specifically, the similarity scores between the enhanced modal features in the initial fusion features can be calculated; these similarity scores are then input into a weighting function, such as the normalized exponential function (Softmax), to obtain the weight coefficients corresponding to each modality; finally, the initial fusion features are weighted and combined based on these weight coefficients to obtain the target fusion features. Alternatively, the Sigmoid independent normalization function or other normalization functions can be used to calculate the aforementioned weight coefficients.

[0069] More specifically, the aforementioned similarity can be cosine similarity. Assuming we have a visual feature vector v, an audio feature vector a, and a text feature vector t, the similarity between the modal features of the video data and the modal features of the audio data can be calculated using the following formula:

[0070]

[0071] Where v represents the visual feature vector (i.e., the modal features corresponding to the video data), a represents the audio feature vector (i.e., the modal features corresponding to the audio data), and t represents the text feature vector (i.e., the modal features corresponding to the text data). Represented as the dot product of vectors v and a, the specific calculation formula can be:

[0072]

[0073] and Let v and a represent the moduli respectively. The specific calculation formula can be:

[0074]

[0075] Similarly, the similarity Svt between visual features and text features, and the similarity Sat between audio features and text features can be calculated by simply replacing the corresponding feature vectors in the above formulas. To avoid repetition and redundancy, this will not be elaborated further here.

[0076] 106. Based on the target fusion features, reasoning is performed to obtain the understanding results of the video to be analyzed.

[0077] In this embodiment of the invention, the aforementioned target fusion features can be provided to a preset large language model for reasoning, thereby obtaining the aforementioned understanding results.

[0078] In one possible embodiment, reasoning can also be based on a pre-trained narrative understanding model, which performs nonlinear mapping and semantic decoding of target fusion features through multiple network layers to generate a structured understanding of the video narrative.

[0079] Specifically, the reasoning process may include: inputting the target fusion features into at least one structure, such as a fully connected layer, a convolutional neural network layer, a recurrent neural network layer, or a Transformer layer, to extract high-level semantic representations; then inputting the semantic representations into a classifier or generator to obtain the understanding results corresponding to the video to be analyzed. The understanding results may include at least one of the following: plot classification tags, event time sequence, character relationship graph, semantic summary text, or scene sentiment distribution, thereby achieving multi-dimensional analysis of the video plot.

[0080] In this embodiment of the invention, multimodal data corresponding to the video to be analyzed is acquired, including at least two of video data, text data, and audio data. Feature extraction processing is performed based on the multimodal data to obtain modal features corresponding to each type of multimodal data. The modal features are enhanced using a self-attention mechanism to obtain enhanced modal features. The enhanced modal features are then fused using a multi-head attention mechanism to obtain initial fused features. Based on the similarity between the modal features, the weights of each enhanced modal feature in the initial fused features are adjusted to obtain target fused features. Inference is then performed based on the target fused features to obtain the understanding result of the video to be analyzed. By combining multimodal feature fusion and dynamic weight adjustment mechanisms with self-attention and multi-head attention mechanisms to achieve cross-modal information complementarity, the problems of one-sided single-modal analysis and insufficient feature fusion are solved. This approach fully utilizes the complementarity of multimodal data, improves the accuracy of video understanding, and allows for deeper exploration of the deeper meaning of the plot.

[0081] Optionally, before obtaining the multimodal data corresponding to the video to be parsed, initial multimodal data corresponding to the video to be parsed can also be obtained. The initial multimodal data includes at least two of the following: initial video data, initial text data, and initial audio data. The initial multimodal data is preprocessed to obtain multimodal data. When the initial multimodal data includes initial video data, the preprocessing includes normalization and noise reduction. When the initial multimodal data includes initial text data, the preprocessing includes word segmentation and part-of-speech tagging. When the initial multimodal data includes initial audio data, the preprocessing includes noise reduction and segmentation.

[0082] In this embodiment of the invention, normalization processing can refer to adjusting the pixel values ​​of video data to a uniform numerical range. Specifically, it can be achieved using the Min-Max normalization method. By eliminating the dimensional differences between video data from different sources, subsequent feature extraction becomes consistent.

[0083] Denoising can refer to eliminating noise interference in video data. Specifically, it can be achieved by using Gaussian filtering or median filtering algorithms to suppress high-frequency noise components and retain effective visual information.

[0084] Word segmentation refers to dividing continuous text into independent semantic units. Specifically, it can be achieved using sub-word segmentation methods based on the BERT model. By decomposing long text into processable word sequences, it provides a foundation for semantic analysis.

[0085] Part-of-speech tagging refers to assigning grammatical category labels to segmented words. Specifically, it can be achieved by combining a conditional random field model with a pre-trained language model. By identifying the grammatical attributes of words, the accuracy of text semantic representation can be enhanced.

[0086] Noise reduction processing refers to eliminating environmental noise in audio data. Specifically, it can be achieved using spectral subtraction or deep neural network noise reduction models. By suppressing non-speech audio signals, the clarity of speech content can be improved.

[0087] Segmentation can refer to cutting continuous audio into segments with independent semantics. Specifically, it can be achieved using endpoint detection algorithms based on silence detection. By dividing the boundaries of speech segments, it is easier to extract subsequent temporal features.

[0088] Specifically, the initial multimodal data may contain various interference factors during the acquisition process. For example, video data may have uneven lighting or sensor noise, text data may contain unsegmented continuous sentences, and audio data may have environmental noise or reverberation effects. Normalization processes map video data from different devices to a unified numerical space, eliminating the impact of brightness differences on feature extraction. Denoising processes filter out salt-and-pepper noise and Gaussian noise in the video while retaining key motion information. For text data, word segmentation transforms the original sentences into a sequence of words, and part-of-speech tagging further reveals the grammatical functions of the words, such as distinguishing between verbs and nouns, providing support for subsequent analysis of character behavior and event relationships. Noise reduction processing removes background music interference from audio data, and segmentation processing divides dialogue segments according to silence intervals, facilitating the extraction of speech features related to the plot.

[0089] Optionally, in the step of adjusting the weights of each enhanced modal feature in the initial fusion features based on the similarity between modal features to obtain the target fusion feature, the similarity between modal features can be input into a preset normalization function to calculate the target weight of each enhanced modal feature; the weights of each enhanced modal feature in the initial fusion features can be adjusted to the corresponding target weights to obtain the target fusion feature.

[0090] In this embodiment of the invention, the normalization function can be a function that converts similarity into a probability distribution. Specifically, it can be implemented using the Softmax function. The Softmax function maps similarity to weight values ​​through exponential operations and normalization, ensuring that the sum of the weights of each modal feature is 1. The target weight refers to the contribution ratio of each modal feature in the fusion process determined after normalization. Specifically, it can be achieved by adjusting the weight allocation strategy, for example, dynamically allocating weights according to the degree of correlation between modalities.

[0091] Specifically, when calculating the target weights, the similarity between modal features is first exponentially calculated to obtain the exponential value of each modal feature. Then, all exponential values ​​are summed to obtain a normalization factor. Finally, the exponential value of each modal feature is divided by the normalization factor to obtain the corresponding target weight. By adjusting the weights of each modal feature in the initial fusion features to the calculated target weights, the fusion ratio can be dynamically adjusted according to the inherent correlation between modalities. For example, when video visuals and audio emotional features have high similarity, their weights are synchronously increased during the fusion process, thereby strengthening the expression of relevant information.

[0092] Optionally, in the step of inputting the similarity between modal features into a preset normalization function to calculate the target weight of each enhanced modal feature, an exponential operation can be performed based on the similarity between modal features to obtain the exponential value of each enhanced modal feature; the exponential value is summed to obtain the normalization factor; and the target weight of each enhanced modal feature is calculated based on the exponential value and the normalization factor.

[0093] In this embodiment of the invention, the exponential operation can refer to the power operation of the similarity between modal features with the natural constant e as the base. Specifically, it can be implemented using the exponential calculation module in the mathematical function library. Its function is to amplify the similarity difference between different modal features, so that the highly similar modal features can obtain more significant distinguishability in the weight allocation.

[0094] The normalization factor can be the sum of the values ​​of all modal features after exponential operation. Specifically, it can be achieved by summing the exponential values ​​of each modal feature using an accumulator. Its function is to map the values ​​after exponential operation to the probability distribution space, ensuring that the sum after weight adjustment is 1.

[0095] The target weight can refer to the normalized exponential value, which can be obtained by dividing the exponential value of a single modality feature by the normalization factor. Its function is to dynamically adjust the influence of different modality features based on their similarity contribution during the fusion process through a probabilistic weight allocation mechanism.

[0096] Specifically, the target weights of the enhanced modal features corresponding to the aforementioned video data can be calculated using the following formula:

[0097]

[0098] in, This represents the target weight of the enhanced modal features corresponding to the video data. This is represented as the similarity between the video feature vector and the audio feature vector. This is represented as the similarity between video feature vectors and text feature vectors. This represents the similarity between the text feature vector and the audio feature vector, where e is a natural constant, approximately equal to 2.71828.

[0099] Accordingly, the target weights of the enhanced modal features corresponding to the aforementioned audio and text data can be calculated using the following formula:

[0100]

[0101] in, This represents the target weight of the enhanced modal features corresponding to the audio data. This represents the target weight of the enhanced modal features corresponding to the text data. This is represented as the similarity between the video feature vector and the audio feature vector. This is represented as the similarity between video feature vectors and text feature vectors. This represents the similarity between the text feature vector and the audio feature vector. e is a natural constant, approximately equal to 2.71828. It can be understood that, through an exponential function and normalization, weights can be dynamically allocated based on the similarity between different modal features; modal features with higher similarity receive greater weight during the fusion process.

[0102] After calculating the target weights, the weights in the initial fusion features can be adjusted to the target weights. Specifically, the target fusion features can also be further calculated using the following formula:

[0103]

[0104] in, The above-mentioned target fusion features are represented as follows: This represents the target weight of the enhanced modal features corresponding to the video data. This represents the target weight of the enhanced modal features corresponding to the audio data. Let V represent the target weight of the enhanced modal feature corresponding to the text data, V represent the enhanced modal feature corresponding to the video data, a represent the enhanced modal feature corresponding to the audio data, and t represent the enhanced modal feature corresponding to the text data.

[0105] Optionally, in the step of reasoning based on target fusion features to obtain the understanding result of the video to be parsed, the target fusion features and the preset first prompt words can be provided to the preset large language model so that the large language model can construct a knowledge graph of the video to be parsed; the knowledge graph and the preset second prompt words can be provided to the large language model so that the large language model can reason about the video to be parsed based on the knowledge graph to obtain the understanding result.

[0106] In this embodiment of the invention, the first prompt word may refer to a text instruction template used to guide the construction of a knowledge graph. Specifically, it may be implemented using framework constraint statements described in natural language, such as guiding text containing entity relationship definitions and event type classifications.

[0107] A knowledge graph can be a graph data structure that represents the relationships between elements of video content. Specifically, it can be implemented in the form of a graph database where nodes represent entities and edges represent relationships, and is used to store the semantic relationships between roles, events, and scenes.

[0108] The second prompt can refer to a text instruction template used to drive the reasoning process. Specifically, it can be implemented using natural language statements that include requirements for logical reasoning paths, such as guiding text that asks for the analysis of causal relationships of events or changes in a character's emotions.

[0109] Specifically, when the target fusion features are input into the large language model, the first prompt word guides the model to transform visual objects, vocal emotions, and textual semantics into nodes and edges of the knowledge graph by defining entity extraction rules and relation definition templates. During the knowledge graph construction phase, the model automatically identifies core elements in the video, such as character identities, scene attributes, and event types, and establishes semantic relationships such as spatiotemporal associations and emotional connections. Subsequently, the second prompt word drives the large language model to perform multi-hop reasoning based on the knowledge graph by setting the reasoning task type and logical constraints. For example, it can infer the plot development trend by traversing character interaction paths or deduce the story's structure by analyzing the temporal relationships of events.

[0110] For example, knowledge graph construction can involve inputting fused multimodal features into a large language model. The large language model then utilizes its rich linguistic knowledge and reasoning capabilities to construct a knowledge graph of the plot. The knowledge graph uses entities (such as characters, scenes, and items) as nodes and relationships between entities (such as relationships between characters, and associations between characters and scenes) as edges, comprehensively describing the semantic information of the plot.

[0111] Deep plot reasoning can be based on constructed knowledge graphs and large language models to understand the logic of plot development, predict subsequent plot directions, and uncover metaphors and symbolic meanings within the plot. For example, by analyzing dialogue and behavior between characters, large language models can infer the characters' emotional changes and their impact on plot development. Furthermore, it can combine common plot patterns and narrative structures to evaluate the completeness and plausibility of the plot.

[0112] Optionally, after the step of reasoning based on the target fusion features to obtain the understanding result of the video to be parsed, user feedback information on the understanding result can also be obtained; based on the feedback information, the gradient descent algorithm is used to update the learnable parameters in the feature extraction processing, enhancement processing, feature fusion processing, and adjustment processing to reduce the deviation between the understanding result and the expected result.

[0113] In this embodiment of the invention, feedback information can refer to the user's evaluation data on the model output results. Specifically, it can be implemented using a scoring mechanism or labeled correction data, which is used to quantify the difference between the model output and the actual situation.

[0114] Gradient descent can be an optimization method that iteratively updates parameter values ​​by calculating the partial derivatives of the loss function with respect to the learnable parameters. Specifically, it can be implemented using stochastic gradient descent or adaptive moment estimation algorithms, and the network weights are adjusted through backpropagation.

[0115] Learnable parameters refer to the weight matrix and bias vector that need to be optimized during model training. They are specifically found in the feature extraction network, attention mechanism module, and feature fusion layer. Updating these parameters can change how the model processes multimodal data. Specifically, learnable parameters can include the convolutional kernel weights of the feature extraction network, the query matrix of the self-attention mechanism, the key-value matrix of the multi-head attention module, and the parameters of the fully connected layers in the weight adjustment module.

[0116] Specifically, when a user rates or labels a video understanding result incorrectly, the feedback information is converted into a supervisory signal to construct a loss function. The gradient of the loss function with respect to the weights of the feature extraction network convolutional kernels, the query matrix of the self-attention mechanism, the multi-head attention key-value matrix, and the parameters of the fully connected layer of the weight adjustment module is calculated, and the parameters are updated along the negative gradient direction according to a preset learning rate. Backpropagation of errors is achieved through the chain rule, enabling the model to more accurately capture intermodal correlation features when processing similar videos, gradually reducing the deviation between the inference result and the user's expectations.

[0117] For example, the updating of the aforementioned learnable parameters can be further illustrated by the following feedback learning mechanism:

[0118] Model parameter update: In feedback learning mechanisms, the model parameters are adjusted based on user feedback. Assume the model's loss function is L(θ), where θ is the model's parameter vector (i.e., the learnable parameters mentioned above). In the gradient descent algorithm, the parameter update formula is:

[0119]

[0120] in, These are the updated parameters. These are the parameters before the update. It's the learning rate, which controls the step size for updating parameters. It is the gradient of the loss function with respect to the parameter θ.

[0121] The performance of the plot understanding algorithm can be evaluated using accuracy as the evaluation metric for the aforementioned feedback learning mechanism. Assuming there are N plot understanding samples in the test set, and Ncorrectly, the number of samples correctly understood by the algorithm is Ncorrect, then the formula for calculating the plot understanding accuracy is:

[0122]

[0123] Optionally, the above method can also construct multiple story understanding sub-models for implementing feature extraction, enhancement, feature fusion, and adjustment processing; obtain a training set for each story understanding sub-model, the training set including sample multimodal data and corresponding result labels, the sample multimodal data being different for different story understanding sub-models; train each story understanding sub-model based on the training set to obtain multiple trained story understanding sub-models; input the multimodal data corresponding to the video to be analyzed into each trained story understanding sub-model to obtain story understanding results output by multiple story understanding sub-models; perform weighted average fusion calculation based on the story understanding results output by multiple story understanding sub-models to obtain the final understanding result of the video to be analyzed.

[0124] In this embodiment of the invention, the plot understanding sub-model can refer to an independent neural network module optimized for different modal combinations or data distributions. Specifically, it can be implemented using a combination structure of convolutional neural networks and long short-term memory networks, used to extract deep correlation features from specific types of multimodal data.

[0125] The training set can refer to a collection containing samples of specific modal combinations and their annotation results. Specifically, it can be implemented using a data structure that pairs video clips with corresponding plot summaries. Different sub-models enhance the model's generalization ability through differentiated training data.

[0126] Weighted average fusion calculation refers to an integration method that integrates the output results of different sub-models by weighting them with confidence. Specifically, it can be implemented by using a linear combination method that assigns weights based on the accuracy of the model validation set, thereby improving the reliability of the final result by suppressing the interference of low-quality models.

[0127] Specifically, multiple plot understanding sub-models are trained in parallel, with each sub-model receiving sample data with different modal combinations. For example, the first sub-model can be trained on a combination of video and audio, while the second sub-model can be trained on a combination of video and text.

[0128] During training, each sub-model independently learns the feature association patterns of its corresponding modality combination. In the inference phase, the multimodal data of the video to be parsed is simultaneously input into all trained sub-models, and each model generates a preliminary understanding result based on its own learning experience. Finally, a weighted averaging algorithm, for example, assigning weights of 0.3, 0.5, and 0.2 to each model based on its performance on the validation set, merges the multiple results into a unified output.

[0129] For example, suppose there are M different plot comprehension models, and the output of the i-th model is... The corresponding weight is (Weights can be determined by the model's performance on the validation set; for example, higher accuracy means higher weights.) The output after multi-model fusion is then... The calculation formula is:

[0130]

[0131] in, This is to ensure that the fused output is within a reasonable range.

[0132] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0133] In one embodiment, a video understanding device is provided, which corresponds one-to-one with the video understanding methods described in the above embodiments. For example... Figure 2 As shown, the video understanding device includes a first acquisition module 201, a feature extraction module 202, a feature enhancement module 203, a feature fusion module 204, a feature adjustment module 205, and an inference module 206. Detailed descriptions of each functional module are as follows:

[0134] The first acquisition module 201 is used to acquire multimodal data corresponding to the video to be parsed, wherein the multimodal data includes at least two of the following: video data, text data, and audio data.

[0135] The feature extraction module 202 is used to perform feature extraction processing based on the multimodal data to obtain modal features corresponding to each type of multimodal data;

[0136] The feature enhancement module 203 is used to enhance the modal features through a self-attention mechanism to obtain enhanced modal features;

[0137] The feature fusion module 204 is used to perform feature fusion processing on the enhanced modal features through a multi-head attention mechanism to obtain initial fused features;

[0138] The feature adjustment module 205 is used to adjust the weight of each enhanced modal feature in the initial fusion feature based on the similarity between the modal features, so as to obtain the target fusion feature;

[0139] The reasoning module 206 is used to perform reasoning based on the target fusion features to obtain the understanding result of the video to be parsed.

[0140] Optionally, the device further includes:

[0141] The second acquisition module is used to acquire the initial multimodal data corresponding to the video to be parsed. The initial multimodal data includes at least two of the following: initial video data, initial text data, and initial audio data.

[0142] A preprocessing module is used to preprocess the initial multimodal data to obtain the multimodal data. When the initial multimodal data includes the initial video data, the preprocessing includes normalization processing and noise reduction processing. When the initial multimodal data includes the initial text data, the preprocessing includes word segmentation processing and part-of-speech tagging processing. When the initial multimodal data includes the initial audio data, the preprocessing includes noise reduction processing and segmentation processing.

[0143] Optionally, the feature adjustment module 205 is further configured to:

[0144] The similarity between the modal features is input into a preset normalization function to calculate the target weight of each enhanced modal feature;

[0145] The weights of each enhanced modal feature in the initial fusion feature are adjusted to the corresponding target weights to obtain the target fusion feature.

[0146] Optionally, the feature adjustment module 205 is further configured to:

[0147] An exponential operation is performed based on the similarity between the modal features to obtain the exponential value of each of the enhanced modal features;

[0148] The normalization factor is obtained by summing the exponent values.

[0149] The target weight of each enhanced modal feature is calculated based on the exponential value of each enhanced modal feature and the normalization factor.

[0150] Optionally, the inference module 206 is further configured to:

[0151] The target fusion features and the preset first prompt words are provided to a preset large language model so that the large language model can construct a knowledge graph of the video to be parsed.

[0152] The knowledge graph and the preset second prompt words are provided to the large language model so that the large language model can reason about the video to be parsed based on the knowledge graph and obtain the understanding result.

[0153] Optionally, the device further includes:

[0154] The third acquisition module is used to acquire user feedback information on the understanding results;

[0155] An update module is used to update the learnable parameters in the feature extraction process, the enhancement process, the feature fusion process, and the adjustment process based on the feedback information using a gradient descent algorithm, so as to reduce the deviation between the understanding result and the expected result.

[0156] Optionally, the device further includes:

[0157] The construction module is used to build multiple story understanding sub-models for implementing the feature extraction process, the enhancement process, the feature fusion process, and the adjustment process;

[0158] The fourth acquisition module is used to acquire the training set of each of the plot understanding sub-models. The training set includes sample multimodal data and result labels corresponding to the sample multimodal data. The sample multimodal data of the training sets of different plot understanding sub-models are different.

[0159] The training module is used to train each of the plot understanding sub-models based on the training set to obtain multiple trained plot understanding sub-models.

[0160] The input module is used to input the multimodal data corresponding to the video to be parsed into each of the trained plot understanding sub-models to obtain the plot understanding results output by the multiple plot understanding sub-models;

[0161] The calculation module is used to perform a weighted average fusion calculation based on the plot understanding results output by the multiple plot understanding sub-models to obtain the final understanding result of the video to be analyzed.

[0162] For specific limitations regarding the video understanding device, please refer to the limitations of the video understanding method above, which will not be repeated here. Each module in the aforementioned video understanding device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.

[0163] In one embodiment, a computer device is provided, which may be a terminal device, and its internal structure diagram may be as follows: Figure 3As shown, the computer device includes a processor, memory, and a network interface connected via a system bus. The processor provides computing and control capabilities. The memory includes a readable storage medium storing computer-readable instructions. The network interface communicates with external terminals via a network connection. When executed by the processor, the computer-readable instructions implement a video understanding method. The readable storage medium provided in this embodiment includes both non-volatile and volatile readable storage media.

[0164] In this application embodiment, a computer device is provided, including a memory, a processor, and computer-readable instructions stored in the memory and executable on the processor. When the processor executes the computer-readable instructions, it implements the steps of the video understanding method described above.

[0165] In one embodiment of the application, a readable storage medium is provided, which stores computer-readable instructions. When the computer-readable instructions are executed by a processor, they implement the steps of the video understanding method described above.

[0166] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by instructing related hardware with computer-readable instructions. These computer-readable instructions can be stored in a non-volatile readable storage medium or a volatile readable storage medium. When executed, these computer-readable instructions can include the processes of the embodiments of the methods described above. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0167] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0168] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A video understanding method, characterized in that, The method includes: Obtain multimodal data corresponding to the video to be parsed, wherein the multimodal data includes at least two of the following: video data, text data, and audio data; Based on the multimodal data, feature extraction processing is performed to obtain the modal features corresponding to each type of multimodal data; The modal features are enhanced by using a self-attention mechanism to obtain enhanced modal features; The enhanced modal features are fused using a multi-head attention mechanism to obtain initial fused features. Based on the similarity between the modal features, the weights of each enhanced modal feature in the initial fusion feature are adjusted to obtain the target fusion feature; Based on the target fusion features, reasoning is performed to obtain the understanding result of the video to be parsed; The step of adjusting the weights of each enhanced modal feature in the initial fusion feature based on the similarity between the modal features to obtain the target fusion feature includes: The similarity between the modal features is input into a preset normalization function to calculate the target weight of each enhanced modal feature; The weights of each enhanced modal feature in the initial fusion feature are adjusted to the corresponding target weights to obtain the target fusion feature; The step of inputting the similarity between the modal features into a preset normalization function to calculate the target weight of each enhanced modal feature includes: An exponential operation is performed based on the similarity between the modal features to obtain the exponential value for each of the enhanced modal features; The normalization factor is obtained by summing the exponent values. The target weight of each enhanced modal feature is calculated based on the exponential value of each enhanced modal feature and the normalization factor.

2. The video understanding method as described in claim 1, characterized in that, Before acquiring the multimodal data corresponding to the video to be parsed, the method further includes: Obtain the initial multimodal data corresponding to the video to be parsed, wherein the initial multimodal data includes at least two of the following: initial video data, initial text data, and initial audio data; The initial multimodal data is preprocessed to obtain the multimodal data. When the initial multimodal data includes the initial video data, the preprocessing includes normalization and noise reduction. When the initial multimodal data includes the initial text data, the preprocessing includes word segmentation and part-of-speech tagging. When the initial multimodal data includes the initial audio data, the preprocessing includes noise reduction and segmentation.

3. The video understanding method as described in claim 1, characterized in that, The reasoning based on the target fusion features to obtain the understanding result of the video to be parsed includes: The target fusion features and the preset first prompt words are provided to a preset large language model so that the large language model can construct a knowledge graph of the video to be parsed. The knowledge graph and the preset second prompt words are provided to the large language model so that the large language model can reason about the video to be parsed based on the knowledge graph and obtain the understanding result.

4. The video understanding method as described in claim 1, characterized in that, After obtaining the understanding result of the video to be parsed by inference based on the target fusion features, the method further includes: Obtain user feedback on the understanding results; Based on the feedback information, the learnable parameters in the feature extraction process, the enhancement process, the feature fusion process, and the adjustment process are updated using the gradient descent algorithm to reduce the deviation between the understanding result and the expected result.

5. The video understanding method as described in claim 1, characterized in that, The method further includes: Construct multiple story understanding sub-models to implement the feature extraction process, the enhancement process, the feature fusion process, and the adjustment process; Obtain the training set for each of the plot understanding sub-models. The training set includes sample multimodal data and result labels corresponding to the sample multimodal data. The sample multimodal data in the training sets of different plot understanding sub-models are different. Based on the training set, each of the plot understanding sub-models is trained to obtain multiple trained plot understanding sub-models. The multimodal data corresponding to the video to be analyzed is input into each of the trained story understanding sub-models to obtain the story understanding results output by the multiple story understanding sub-models; The final understanding result of the video to be analyzed is obtained by weighted averaging and fusion calculation based on the plot understanding results output by the multiple plot understanding sub-models.

6. A video understanding device, characterized in that, The device includes: The first acquisition module is used to acquire multimodal data corresponding to the video to be parsed, wherein the multimodal data includes at least two of the following: video data, text data, and audio data. The feature extraction module is used to perform feature extraction processing based on the multimodal data to obtain the modal features corresponding to each type of multimodal data; The feature enhancement module is used to enhance the modal features through a self-attention mechanism to obtain enhanced modal features; The feature fusion module is used to perform feature fusion processing on the enhanced modal features through a multi-head attention mechanism to obtain initial fused features; The feature adjustment module is used to adjust the weight of each enhanced modal feature in the initial fusion feature based on the similarity between the modal features, so as to obtain the target fusion feature; The reasoning module is used to perform reasoning based on the target fusion features to obtain the understanding result of the video to be parsed; The feature adjustment module is also used for: The similarity between the modal features is input into a preset normalization function to calculate the target weight of each enhanced modal feature; The weights of each enhanced modal feature in the initial fusion feature are adjusted to the corresponding target weights to obtain the target fusion feature; The feature adjustment module is also used for: An exponential operation is performed based on the similarity between the modal features to obtain the exponential value for each of the enhanced modal features; The normalization factor is obtained by summing the exponent values. The target weight of each enhanced modal feature is calculated based on the exponential value of each enhanced modal feature and the normalization factor.

7. A computer device comprising a memory, a processor, and computer-readable instructions stored in the memory and running on the processor, characterized in that, When the processor executes the computer-readable instructions, it implements the video understanding method as described in any one of claims 1 to 5.

8. A readable storage medium having computer-readable instructions stored thereon, characterized in that, When the computer-readable instructions are executed by a processor, they implement the video understanding method as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Multi-modal emotion recognition method

    CN119475252A

  • Method and system for detecting irony in dialogue scene

    CN120470518A