Video engagement prediction method, system, device and medium based on graph learning

Through the graph-based learning method, multimodal fully connected graphs and fusion of multimodal data, the low accuracy problem caused by ignoring multimodal correlation in the prior art is solved, and more efficient video user participation prediction is achieved.

CN115187893BActive Publication Date: 2025-06-03ZHEJIANG NORMAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210620916.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-02
Publication Date
2025-06-03
Estimated Expiration
2042-06-02

AI Technical Summary

Technical Problem

The prior art ignores the correlation between multimodal data in video user engagement prediction, resulting in low accuracy.

Method used

Using a graph-based learning method, multimodal fully connected graphs are constructed to extract emotional features by obtaining video content and extracting text, audio and video features, and multimodal data is fused with attention mechanism to improve prediction accuracy.

Benefits of technology

By considering the interrelationship between multimodal data, the accuracy of video user participation prediction is improved, and the problem of low credibility of single-modal feature prediction results is overcome.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115187893B_ABST
    Figure CN115187893B_ABST
Patent Text Reader

Abstract

The present invention discloses a video engagement prediction method, system, device and medium based on graph learning, which relates to the field of computer technology. This application combines an emotion feature matrix based on multimodal analysis and text feature matrices, audio feature matrices and video feature matrices based on unimodal analysis to predict user engagement. It not only considers the influence of unimodal data on the prediction result, but also considers the mutual correlation between multimodal data through sentiment analysis, thereby improving the accuracy of video user engagement prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular, to a method, system, device and medium for predicting video engagement based on graph learning. Background Art

[0002] With the rapid development of mobile Internet technology, creators can share videos on different platforms. An effective way to evaluate the quality of a video is user engagement, such as the number of clicks, likes, and comments on the video. If the quality of a video is evaluated by real user engagement after the video is released, although the result is relatively accurate, there is a lag, which is not conducive to creators optimizing the video content in advance. Therefore, it is necessary to predict the user engagement of the created videos in advance.

[0003] Currently, video user engagement prediction models the various features of a single modality such as text to predict engagement. The results obtained by only considering the features of a single modality are not highly reliable, and the correlation between modalities is also ignored, resulting in low accuracy. Summary of the Invention

[0004] The present invention aims to solve at least one of the technical problems existing in the prior art. For this purpose, the present invention provides a method, system, device and storage medium for predicting video engagement based on graph learning, which can improve the accuracy of predicting video user engagement.

[0005] On the one hand, an embodiment of the present invention provides a method for predicting video engagement based on graph learning, including the following steps:

[0006] Obtain video content;

[0007] Extract modal features from the video content to obtain text data, audio data, and video data;

[0008] Extract emotional features through graph learning based on the text data, the audio data, and the video data to obtain an emotional feature matrix;

[0009] Extract keyword features from the text data to obtain a text feature matrix;

[0010] Extract target object features from the video data to obtain a video feature matrix;

[0011] Extract spectrogram features from the audio data to obtain an audio feature matrix;

[0012] Input the emotional feature matrix, the text feature matrix, the video feature matrix, and the audio feature matrix into a user engagement prediction model to obtain a user engagement prediction result.

[0013] According to some embodiments of the present invention, the extraction of emotional features through graph learning based on the text data, the audio data, and the video data to obtain an emotional feature matrix includes the following steps:

[0014] Input the text data into a text feed-forward neural network for feature encoding to obtain a text embedding sequence, wherein the text embedding sequence includes a plurality of text embeddings, and the text embedding includes a first position identifier for characterizing the position of the text embedding in the text embedding sequence;

[0015] Input the audio data into an audio feed-forward neural network for feature encoding to obtain an audio embedding sequence, wherein the audio embedding sequence includes a plurality of audio embeddings, and the audio embedding includes a second position identifier for characterizing the position of the audio embedding in the audio embedding sequence;

[0016] Input the video data into a video feed-forward neural network for feature encoding to obtain a video embedding sequence, wherein the video embedding sequence includes a plurality of video embeddings, and the video embedding includes a third position identifier for characterizing the position of the video embedding in the video embedding sequence;

[0017] Determine the temporal relationship between two embeddings based on the first position identifier, the second position identifier, and the third position identifier, wherein the two embeddings include at least one of two text embeddings, two video embeddings, two audio embeddings, a text embedding and a video embedding, a text embedding and an audio embedding, and an audio embedding and a video embedding;

[0018] Construct a multimodal fully-connected graph according to the text embeddings, the audio embeddings, the video embeddings, and the temporal relationship between two embeddings, wherein the text embeddings, the audio embeddings, and the video embeddings are all used as nodes of the multimodal fully-connected graph, and the temporal relationship between two embeddings is used as an edge of the multimodal fully-connected graph;

[0019] Input the multimodal fully-connected graph into a graph neural network for emotional feature extraction to obtain the emotional feature matrix.

[0020] According to some embodiments of the present invention, the determination of the temporal relationship between two embeddings based on the first position identifier, the second position identifier, and the third position identifier includes the following steps:

[0021] When the two embeddings are embeddings of the same modality, determine the temporal relationship between the two embeddings according to the position identifier corresponding to the modality;

[0022] When the two embeddings are embeddings of different modalities, the convolution kernel and convolution step size are set according to the lengths of the two embedding sequences where the two embeddings are located, and an alignment operation is performed on the two embedding sequences according to the convolution kernel and the convolution step size to determine the first embedding and the second embedding that are aligned with each other in the two embedding sequences. Based on the position identifier of the first embedding, the temporal relationship between each embedding in the embedding sequence where the first embedding is located and the second embedding is determined.

[0023] According to some embodiments of the present invention, the step of extracting the emotion feature matrix through graph learning based on the text data, the audio data, and the video data further includes the following steps:

[0024] Determine the attention weights of the corresponding edges according to the embeddings of adjacent nodes in the multimodal fully connected graph;

[0025] Perform information fusion on adjacent nodes according to the attention weights to obtain new embeddings of each node;

[0026] Determine the similarity between adjacent nodes according to the embeddings of the nodes;

[0027] When the similarity of adjacent nodes is greater than the similarity threshold, delete the edges of the adjacent nodes;

[0028] Delete the isolated nodes in the multimodal fully connected graph that have no edges connected.

[0029] According to some embodiments of the present invention, the step of extracting the text feature matrix by performing feature extraction on the text data for keywords includes the following steps:

[0030] Extract the video text, title text, and various parts of speech from the text data;

[0031] Calculate the video text length, title text length, and part-of-speech ratio and perform representation learning using a multi-layer perceptron to obtain the text feature matrix.

[0032] According to some embodiments of the present invention, the step of extracting the audio feature matrix by performing feature extraction on the audio data for spectrograms includes the following steps:

[0033] Extract the Mel spectrogram from the audio data;

[0034] Input the Mel spectrogram into a recurrent autoencoder for feature representation to obtain the audio feature matrix.

[0035] According to some embodiments of the present invention, the step of extracting the video feature matrix by performing feature extraction on the video data for target objects includes the following steps:

[0036] Divide the video data into several frame segments;

[0037] Respectively input several of the frame segments into the trained YOLO v3 model for target object recognition, and obtain a video feature matrix for characterizing the appearance time of the target object in the video.

[0038] On the other hand, an embodiment of the present invention further provides a video engagement prediction system based on graph learning, including:

[0039] A first module for obtaining video content;

[0040] A second module for performing modal feature extraction on the video content to obtain text data, audio data, and video data;

[0041] A third module for performing emotion feature extraction through graph learning based on the text data, the audio data, and the video data to obtain an emotion feature matrix;

[0042] A fourth module for performing keyword feature extraction on the text data to obtain a text feature matrix;

[0043] A fifth module for performing target object feature extraction on the video data to obtain a video feature matrix;

[0044] A sixth module for performing spectrogram feature extraction on the audio data to obtain an audio feature matrix;

[0045] A seventh module for inputting the emotion feature matrix, the text feature matrix, the video feature matrix, and the audio feature matrix into a user engagement prediction model to obtain a user engagement prediction result.

[0046] On the other hand, an embodiment of the present invention further provides a video engagement prediction device based on graph learning, including:

[0047] At least one processor;

[0048] At least one memory for storing at least one program;

[0049] When the at least one program is executed by the at least one processor, the at least one processor implements the video engagement prediction method based on graph learning as described above.

[0050] On the other hand, an embodiment of the present invention further provides a computer-readable storage medium storing computer-executable instructions for causing a computer to execute the video engagement prediction method based on graph learning as described above.

[0051] At least one of the above technical solutions of the present invention has the following advantages or beneficial effects: This application combines the emotional feature matrix based on multimodal analysis and the text feature matrix, audio feature matrix, and video feature matrix based on unimodal analysis to predict user engagement. It not only considers the impact of unimodal data on the prediction result but also takes into account the mutual correlation between multimodal data through sentiment analysis, thereby improving the accuracy of video user engagement prediction. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 is a flowchart of a video engagement prediction method based on graph learning provided by an embodiment of the present invention;

[0053] Figure 2 is a schematic diagram of a video engagement prediction device based on graph learning provided by an embodiment of the present invention;

[0054] Figure 3 is a schematic diagram of a video engagement prediction process based on graph learning provided by an embodiment of the present invention;

[0055] Figure 4 is a schematic diagram of a node alignment operation provided by an embodiment of the present invention;

[0056] Figure 5 is a schematic diagram of a node alignment operation provided by another embodiment of the present invention;

[0057] Figure 6 is a schematic diagram of a prediction model processing flow provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0058] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention and should not be construed as a limitation of the present invention.

[0059] In the description of the present invention, it should be understood that the orientation descriptions, such as up, down, left, right, etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention.

[0060] In the description of the present invention, if the first, second, etc. are described, it is only for the purpose of distinguishing technical features and should not be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features or implicitly indicating the sequence of the indicated technical features.

[0061] An embodiment of the present invention provides a method for predicting video engagement based on graph learning. Referring to Figure 1 , the method for predicting video engagement based on graph learning according to the embodiment of the present invention includes but is not limited to steps S110, S120, S130, S140, S150, S160, and S170.

[0062] Step S110, obtain video content;

[0063] Step S120, perform modal feature extraction on the video content to obtain text data, audio data, and video data;

[0064] Step S130, perform sentiment feature extraction through graph learning based on the text data, audio data, and video data to obtain a sentiment feature matrix;

[0065] Step S140, perform feature extraction of keywords on the text data to obtain a text feature matrix;

[0066] Step S150, perform feature extraction of target objects on the video data to obtain a video feature matrix;

[0067] Step S160, perform feature extraction of spectrograms on the audio data to obtain an audio feature matrix;

[0068] Step S170, input the sentiment feature matrix, text feature matrix, video feature matrix, and audio feature matrix into the user engagement prediction model to obtain the user engagement prediction result.

[0069] The embodiment of the present invention combines the sentiment feature matrix based on multimodal analysis and the text feature matrix, audio feature matrix, and video feature matrix based on unimodal analysis to predict user engagement, and obtains the user engagement prediction result of the video. It not only considers the influence of unimodal data on the prediction result, but also considers the mutual correlation between multimodal data through sentiment analysis, thereby improving the accuracy of video user engagement prediction.

[0070] In some embodiments, the prediction result can be an overall evaluation score, which reflects to a certain extent the user's willingness to participate in the video. The prediction result can also be the predicted number of clicks, likes, and evaluation tendency values of the video. The embodiment of the present invention does not make specific limitations. The manifestation form of the prediction result is related to the labels used for training the prediction model. If the data labels in the training set are manually labeled scores, the prediction result is the overall evaluation score. If the data labels are the actually obtained number of clicks, the prediction result is the number of clicks.

[0071] Next, please refer to Figure 3 , and elaborate on the detailed solution of the embodiment of the present invention.

[0072] According to some specific embodiments of the present invention, step S130 includes but is not limited to the following steps:

[0073] Step S210, input the text data into a text feedforward neural network for feature encoding to obtain a text embedding sequence, where the text embedding sequence includes multiple text embeddings, and each text embedding includes a first position identifier for characterizing the position of the text embedding in the text embedding sequence;

[0074] Step S220, input the audio data into an audio feedforward neural network for feature encoding to obtain an audio embedding sequence, where the audio embedding sequence includes multiple audio embeddings, and each audio embedding includes a second position identifier for characterizing the position of the audio embedding in the audio embedding sequence;

[0075] Step S230, input the video data into a video feedforward neural network for feature encoding to obtain a video embedding sequence, where the video embedding sequence includes multiple video embeddings, and each video embedding includes a third position identifier for characterizing the position of the video embedding in the video embedding sequence;

[0076] Step S240, determine the temporal relationship between two embeddings based on the first position identifier, the second position identifier, and the third position identifier, where the two embeddings include at least one of two text embeddings, two video embeddings, two audio embeddings, a text embedding and a video embedding, a text embedding and an audio embedding, and an audio embedding and a video embedding;

[0077] Step S250, construct a multimodal fully connected graph according to the text embeddings, audio embeddings, and video embeddings, and the temporal relationship between the two embeddings, where the text embeddings, audio embeddings, and video embeddings serve as nodes of the multimodal fully connected graph, and the temporal relationship between the two embeddings serves as the edge of the multimodal fully connected graph;

[0078] Step S260, input the multimodal fully connected graph into a graph neural network for emotion feature extraction to obtain an emotion feature matrix.

[0079] In some embodiments, in order to associate the features of each modality data, it is necessary to fuse each modality data for emotion analysis to improve the accuracy of subsequent user engagement prediction. After extracting different modalities of video content to obtain video data, audio data, and text data, input the text data into a text feedforward neural network for feature encoding to obtain a text embedding sequence T, input the audio data into an audio feedforward neural network for feature encoding to obtain an audio embedding sequence A, input the video data into a video feedforward neural network for feature encoding to obtain a video embedding sequence S. The text embedding sequence includes multiple text embeddings e t , the audio embedding sequence includes multiple audio embeddings e a, the video embedding sequence includes multiple video embeddings e of the same size s .

[0080] Add a position identifier to each embedding to represent the position of the embedding in the embedding sequence where it is located. Exemplarily, the text embedding includes a first position identifier, and the first position identifier represents the position of the text embedding in the text embedding sequence. The audio embedding includes a second position identifier, and the second position identifier represents the position of the audio embedding in the audio embedding sequence. The video embedding includes a third position identifier, and the third position identifier represents the position of the video embedding in the video embedding sequence. The position identifiers are set according to the chronological order, so that the embeddings of each modality can be combined into an embedding sequence in chronological order. The position identifiers can be determined according to formulas (1) and (2):

[0081]

[0082]

[0083] where i is the dimension of the mapping, d emb is the dimension of the embedding, pos is the position to be calculated, PE (pos,2i) and PE (pos,2i+1) are both position encodings.

[0084] Furthermore, the temporal relationship between embeddings of each modality will affect sentiment analysis, and different sequence orders may produce different sentiment recognition results. Therefore, the temporal relationship between all pairs of embeddings can be determined based on the position identifiers (the first position identifier, the second position identifier, and the third position identifier) of the embeddings of each modality.

[0085] Take each embedding as a node, and connect two embeddings according to the temporal relationship between the two embeddings to form an edge, obtaining a multi-modal fully connected graph. That is, the multi-modal fully connected graph G=(v, E) includes a node set ν and an edge set E. The node set v includes a video node set S, an audio node set A, and a text node set T, as follows:

[0086] S = {s 1 , s 2 ,..., s i}

[0087] A = {a 1 , a 2 ,..., a j}

[0088] T = {t 1 , t 2 ,..., t k}

[0089] Use the modality identifier π to label the modality of the node:

[0090] π ∈ {S, A, T}

[0091] The set of nodes v is represented as:

[0092] ν = S ∪ A ∪ T

[0093] The set of edges E = {(s, a)(a, s)(s, t)(t, s)(a, t)(t, a)(a, a)(s, s)(t, t) | s ∈ S, a ∈ A, t ∈ T}, which represents the existence relationship between any two of the video nodes, audio nodes, and text nodes. On this basis, the direction identifier ω and the time feature φ are added to characterize the temporal relationship between the nodes.

[0094] For example, a directed edge is represented as (s′, a′), where the source node is s′ and the target node is a′. The direction identifier ω is used to identify the set of edges:

[0095] ω ∈ {(a′, a′), (a′, s′), (a′, t′), (s′, a′), (s′, s′), (s′, t′), (t′, a′), (t′, s′), (t′, t′)}

[0096] Among them, a′ represents an audio node, s′ represents a video node, and t′ represents a text node.

[0097] Time feature An edge represented by the eigenvalue p indicates that its source node is a past node relative to the target node. An edge represented by the eigenvalue n indicates that its source node is the current node relative to the target node. An edge represented by the eigenvalue f indicates that its source node is a future node relative to the target node.

[0098] In some embodiments, the multi-modal fully connected graph is input into the graph representation layer of the graph neural network for feature encoding to obtain the graph vector Vec e , as shown in formula (3), the graph vector is input into the MLP classifier for sentiment recognition to obtain the sentiment representation matrix V p , as shown in formula (4), the sentiment representation matrix is normalized to obtain the sentiment feature matrix X 1 , exemplarily, the sentiment recognition results can be joy, sadness, anger, fear, calmness, etc. The sentiment feature matrix is the encoded representation of the sentiment recognition results.

[0099] V p = MLP e (Vec e ) (3)

[0100] X 1 = W 1 V p (4)

[0101] Among them, W 1 is the normalization coefficient.

[0102] According to some specific embodiments of the present invention, step S240 includes the following steps:

[0103] Step S310, when the two embeddings are embeddings of the same modality, determine the temporal relationship between the two embeddings according to the position identifier corresponding to the modality;

[0104] Step S320, when the two embeddings are embeddings of different modalities, set the convolution kernel and convolution step according to the lengths of the two embedding sequences where the two embeddings are located, and perform an alignment operation on the two embedding sequences according to the convolution kernel and convolution step to determine the first embedding and the second embedding that are aligned with each other in the two embedding sequences, and determine the temporal relationship between each embedding in the embedding sequence where the first embedding is located and the second embedding based on the position identifier of the first embedding.

[0105] In some embodiments, for the temporal relationship between two embeddings, or two nodes, if they are embeddings of the same modality, the temporal relationship between them can be determined according to the position identifier corresponding to the modality. Taking text embeddings as an example, if the first position identifiers of the first text embedding and the second text embedding in the text embedding sequence are 1 and 2 respectively, it can be determined that the first text embedding is earlier than the second text embedding.

[0106] In some embodiments, if they are embeddings of different modalities, or nodes, an alignment operation needs to be performed on the embeddings where the two embeddings are located, and then the temporal relationship between the embeddings of different modalities is determined based on the alignment result.

[0107] Taking the audio embedding sequence and the text embedding sequence as an example, the alignment operation process is as follows:

[0108] Take the audio embedding sequence and the text embedding sequence as the input and output of a one-dimensional convolution operation. It can be that the audio embedding sequence is the input sequence and the text embedding sequence is the output sequence, or the audio embedding sequence is the output sequence and the text embedding sequence is the input sequence. Generally, take the sequence with the longer length as the input sequence and the sequence with the shorter length as the output sequence, so that the embeddings in the input sequence and the embeddings in the output sequence can be aligned in a one-to-one or many-to-one form.

[0109] Set the convolution kernel and convolution step according to the following formula (5):

[0110]

[0111] Among them, M represents the length of the input sequence, N represents the length of the output sequence, W represents the convolution kernel, and S represents the convolution step.

[0112] In the output sequence (the longer sequence), convolution is performed using the convolution kernel calculated by formula (5). Starting from the first embedding in the output sequence, the embedding in the convolution kernel is aligned with the first embedding in the input sequence (the shorter sequence). The convolution kernel is moved according to the convolution stride calculated by formula (5), and then the embedding in the convolution kernel is continued to be aligned with the second embedding in the input sequence, and so on. If the number of remaining unaligned embeddings in the output sequence is the same as the number of remaining unaligned embeddings in the input sequence, then both the convolution kernel and the convolution stride are adjusted to 1 and convolution is continued for one-to-one alignment. Generally speaking, if the convolution kernel size calculated according to formula (5) is n, then there are n nodes including the current node and its subsequent nodes in the input sequence (the longer one) determined to correspond to the current node in the output sequence (the shorter one) as co-existing nodes.

[0113] As Figure 4 shown, the input sequence is a text node sequence with the number of nodes M = 8, and the output sequence is an audio node sequence with the number of nodes N = 4, satisfying N > M / 2. According to formula (5), the convolution kernel size W = 2 and the stride S = 2 are set, so a 1 corresponding current node is t 1 , t 2 , a 2 corresponding current node is t 3 , t 4 , a 3 corresponding current node is t 5 , t 6 . To ensure that each audio node has at least one corresponding co-existing current text node, therefore, the convolution kernel size for a 4 and each of its subsequent nodes is set to 1, and the stride is also set to 1 accordingly. For the audio nodes a 3 and a 5 , there is only one corresponding co-existing text node. For the audio node a 2 , the text nodes t 1 , t 2 are past nodes, t 3 , t 4 are current nodes, and t 5 ~t 8 are future nodes.

[0114] As Figure 5 shown, the input sequence is a text node sequence with the number of nodes M = 8, and the output sequence is a video node sequence with the number of nodes N = 4, satisfying N = M / 2. According to formula (5), the convolution kernel size W = 2 and the stride S = 2 are set. Since the input sequence is long enough to ensure that each video node corresponds to two text nodes, therefore, the convolution kernel and the convolution stride do not need to be adjusted to 1. s 1 corresponding current node is t1 and t 2 ,s 2 The corresponding current node is t 3 and t 4 ,s 3 The corresponding current node is t 5 and t 6 ,s 4 The corresponding current node is t 7 and t 8 。For the audio node s 2 ,the text nodes t 1 and t 2 are past nodes, t 3 and t 4 are current nodes, t 5 ~t 8 are future nodes.

[0115] In some embodiments, the derivation process of formula (5) is as follows:

[0116] The convolution operation based on the input sequence M, output sequence N, convolution kernel W, and convolution step S can be expressed by formula (6):

[0117]

[0118] Formula (7) is obtained from formula (6):

[0119] W = M - (N - 1) * S (7)

[0120] The step S is at least 1 and at most M / (N - 1). When N ≤ M / 2, the alignment algorithm takes the average of the maximum step and the minimum step as the step S, and takes M - (N - 1) * S as the size of the convolution kernel W; when N > M / 2, the alignment algorithm sets the size of the convolution kernel W to 2 and the step S to 2, thus obtaining formula (5).

[0121] In addition, on the premise of ensuring that all nodes have current nodes, find the maximum number of nodes with a convolution kernel W size of 2 from the output sequence N, and set the convolution kernel W size of the remaining nodes to 1, and the step is also set to 1 accordingly.

[0122] According to some specific embodiments of the present invention, step S130 further includes but is not limited to the following steps:

[0123] Step S410, determining the attention weight of the corresponding edge according to the embeddings of adjacent nodes in the multimodal fully connected graph;

[0124] Step S420, performing information fusion on adjacent nodes according to the attention weight to obtain the new embedding of each node;

[0125] Step S430, determine the similarity between adjacent nodes according to the embeddings of the nodes;

[0126] Step S440, when the similarity between adjacent nodes is greater than the similarity threshold, delete the edge between the adjacent nodes;

[0127] Step S450, delete the isolated nodes in the multimodal fully-connected graph that have no edges connected.

[0128] In some embodiments, before calculating the attention weights of the edges, first use a simple linear transformation to transform the features of all nodes into a common feature space. The modality identifier π of the nodes is used to distinguish different modality nodes, so as to ensure that there are different linear transformation coefficients for different types of nodes. The linear transformation is shown in formula (8):

[0129] x i =M π x i (8)

[0130] π ∈ {S, A, T}

[0131] where S identifies the video modality, A identifies the audio modality, T identifies the text modality, M π is the linear transformation coefficient, x i is the node embedding, and x' i is the transformed node embedding.

[0132] Then calculate the original attention weights of the edges between adjacent nodes through formula (9):

[0133]

[0134] where β i,j is the original attention weight of the edge between node i and node j, is an attention vector corresponding to each type of edge. The type of the edge is determined by the tuple ω is the edge direction identifier, is the time feature of the edge.

[0135] Then use the SoftMax function to normalize the original attention scores of the neighbor nodes of all nodes to obtain the attention weight α ij , to guide the information transmission between heterogeneous nodes and maintain the scale of the node features in the graph, as shown in formula (10):

[0136]

[0137] For node i, according to the normalized attention weight α ij , aggregate the information from the linearly transformed neighbor node j according to formula (11) to obtain zi ,z i as the new embedding of node i.

[0138]

[0139] After aggregation, node i changes from a node including unimodal data to a node integrating multimodal data, containing rich information from adjacent nodes. Since the multimodal fully-connected graph is completely connected, node i can collect information from all modalities on the time-series edges, thus enabling the modeling of rich and complex cross-modal and temporal information.

[0140] Considering that not every edge and node in the graph is meaningful, and too many edges will lead to an overly large computational graph, causing an excessive burden on the processor. Therefore, it is necessary to prune the edges in the graph and remove isolated nodes, as follows:

[0141] Calculate the embedding x of node i on edge (i, j) according to formula (12) i and the embedding x of node j j The feature smoothness es(i, j) between them:

[0142]

[0143] where d represents the dimensionality of the feature space of the node.

[0144] If es(i, j) is less than the defined threshold, such as 0.2, it means that nodes i and j are highly similar, and there is not much different information that can be aggregated from each other. Therefore, prune edge (i, j).

[0145] Repeatedly execute the two steps of aggregating the information of adjacent nodes and dynamically pruning redundant edges several times alternately, so that the degree of information aggregation is higher and the number of redundant edges is also greatly reduced. Then, delete the isolated nodes without edge connections to obtain a mixed multimodal graph G′(V, E) with an appropriate size and rich information for each node, so that a graph convolutional neural network can be used for representation learning to obtain the graph vector Vec e .

[0146] In this embodiment, calculate the attention score according to the time and modal features of the edge, and guide the information transfer between heterogeneous nodes through the normalized attention score, effectively fusing the information of multimodal nodes. By calculating the feature smoothness between node feature vectors, dynamically prune the edges between two nodes with lower feature smoothness, and at the same time remove the isolated nodes without edge connections. This processing removes the redundant edges and nodes that have little impact on the prediction result, effectively reducing the size of the computational graph and thus improving the computational efficiency.

[0147] According to some specific embodiments of the present invention, step S140 includes but is not limited to the following steps:

[0148] Step S510: Extract the text data to obtain video text, title text, and various parts of speech.

[0149] Step S520: Calculate the video text length, title text length, and part-of-speech ratio, and represent them using a multi-layer perceptron to obtain a text feature matrix.

[0150] In some embodiments, feature extraction is performed on the text data included in the video to obtain video text, title text, and various parts of speech. The parts of speech include stop words, prepositions, auxiliary verbs, conjunctions, nouns, pronouns, etc. Then, calculate eight text features γ, namely the text length Len (total number of words in all video text), title length Tittle (number of words in the video title), stop word ratio Stop (number of stop words / text length), preposition ratio Prep (number of prepositions / text length), auxiliary verb ratio Auxi (number of auxiliary verbs / text length), conjunction ratio Conj (number of conjunctions / text length), noun ratio Noun (number of nouns / text length), and pronoun ratio Pro (number of pronouns / text length). Input the text features into a multi-layer perceptron (MLP t ) to obtain a text representation matrix Vec t , and perform normalization processing on the text representation matrix to obtain a text feature matrix X 2 .

[0151] γ = {Len, Tittle, Stop, Prep, Auxi, Conj, Noun, Pro}

[0152] Vec t = MLP t (γ) (13)

[0153] X 2 = W 2 Vec t (14)

[0154] According to some specific embodiments of the present invention, step S150 includes but is not limited to the following steps:

[0155] Step S610: Extract the Mel spectrogram from the audio data;

[0156] Step S620: Input the Mel spectrogram into a recurrent autoencoder for feature representation to obtain the audio feature matrix.

[0157] In some embodiments, as audio is an integral part of the video and also has an important impact on video engagement, therefore, the embodiments of the present invention analyze audio unimodal features and use them in the prediction model.

[0158] According to formulas (15) to (17), the Mel-frequency cepstral coefficients (MFCC) are extracted from the original audio file. Then, a recurrent sequence-to-sequence autoencoder (STSAEncoder) is trained on the spectrogram, and the learned spectrogram representation is extracted as the feature vector for the corresponding audio instance. Finally, the generated feature vectors are concatenated to obtain the audio representation matrix Vec. a The audio representation matrix is normalized to obtain the audio feature matrix X. 2 。

[0159]

[0160]

[0161] X 4 =W 4 Vec a (17)

[0162] According to some specific embodiments of the present invention, step S160 includes but is not limited to the following steps:

[0163] Step S710, dividing the video data into several frame segments;

[0164] Step S720, respectively inputting the several frame segments into the trained YOLO v3 model for target object recognition to obtain a video feature matrix for characterizing the appearance time of the target object in the video.

[0165] In some embodiments, for an instructional video, if animations are used, more objects will be generated compared to traditional instructional content, and the rich objects will affect video engagement. Therefore, the YOLO model is introduced to capture the objects contained in the video frame, and the model output r represents the time when special objects appear in the video. The target detection uses the YOLO v3 model, which can use logistic regression to calculate the target scores, thereby giving the scores of all targets in each bounding box. YOLOv3 uses a logistic classifier for each class and can give multi-label classification. The darknet-53 used by YOLO v3 has 53 convolutional layers. Compared with the darknet 19 used in YOLO v2, the 53 convolutional layers can learn more deeply. Darknet-53 mainly contains 3x3 and 1x1 filters and bypass links. According to formulas (18) and (19), the final output of the YOLO v3 model is r, which characterizes the time when special objects appear in the video. r is normalized to obtain the video feature matrix X. 3 。

[0166] r=YOLO v3 (video) (18)

[0167] X 3 = W 3 r (19)

[0168] In some other embodiments, the frame images of each frame in the video data can be subjected to similarity calculation based on pixel bits, and several frame images with higher similarity are divided into a group of frame segments. Inputting the frame segments to be analyzed into the YOLO v3 model can reduce the recognition difficulty of the YOLO v3 model and improve the calculation efficiency compared with the method of inputting the entire video data. Exemplarily, in an instructional video, blackboard images, animation images, and PPT images will appear alternately. The differences in pixel bits among these three types of images are large, while the differences in pixel bits within one type of image are small. Therefore, these three types of images can be divided based on the overall similarity of pixel bits. The content of the PPT images and blackboard images overlaps with the content of the audio data and text data. Therefore, only the animation images can be input into the YOLO v3 model for recognition.

[0169] According to some specific embodiments of the present invention, the processing process of the prediction model is as Figure 6 shown. The emotion feature matrix X 1 , the text feature matrix X 2 , the video feature matrix X 3 , and the text feature matrix X 4 are processed by the attention mechanism to obtain their respective original attention scores s i . After normalizing the original attention scores using the SoftMax function, α i is obtained. Using α i as the weight to perform weighted sum on the four outputs, the user engagement prediction result

[0170] s i = X i W i ′ (20)

[0171]

[0172]

[0173] Furthermore, to improve the accuracy of the prediction model, sentiment analysis, audio analysis, video analysis, and text analysis models involved in the present application, after each user finishes watching the video, data such as the user's viewing duration, pause times, user rating, and positive / negative comments of the user are collected to calculate the true user engagement value y. The loss function is used to compare the differences between the two engagement scores to calculate the loss, where μ is a hyperparameter.

[0174]

[0175] Through error backpropagation, optimize the parameters of the model involved above. Perform a certain number of iterations of backpropagation until the model parameters converge. The finally trained model parameters can make the prediction of the engagement score obtain an accurate value close to the actual value, so as to obtain a reliable user engagement prediction result.

[0176] On the other hand, an embodiment of the present invention also provides a video engagement prediction system based on graph learning, including:

[0177] The first module is used to obtain video content;

[0178] The second module is used to extract modal features from the video content to obtain text data, audio data, and video data;

[0179] The third module is used to extract emotional features through graph learning based on the text data, audio data, and video data to obtain an emotional feature matrix;

[0180] The fourth module is used to extract keyword features from the text data to obtain a text feature matrix;

[0181] The fifth module is used to extract features of the target object from the video data to obtain a video feature matrix;

[0182] The sixth module is used to extract spectrogram features from the audio data to obtain an audio feature matrix;

[0183] The seventh module is used to input the emotional feature matrix, text feature matrix, video feature matrix, and audio feature matrix into the user engagement prediction model to obtain the user engagement prediction result.

[0184] It can be understood that the content in the above embodiments of the video engagement prediction method based on graph learning is applicable to the embodiments of this system. The functions specifically implemented by the embodiments of this system are the same as those of the above embodiments of the video engagement prediction method based on graph learning, and the beneficial effects achieved are also the same as those of the above embodiments of the video engagement prediction method based on graph learning.

[0185] Refer to Figure 2 , Figure 2 is a schematic diagram of a video engagement prediction device based on graph learning provided by an embodiment of the present invention. The video engagement prediction device according to the embodiment of the present invention includes one or more control processors and memories. Figure 2 In

[0186] Take one control processor and one memory as an example. Figure 2Take the bus connection as an example.

[0187] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory optionally includes a memory remotely provided relative to the control processor, and these remote memories can be connected to the graph learning-based video engagement prediction device through a network. Examples of the above networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0188] Those skilled in the art can understand that Figure 2 the device structure shown in the figure does not constitute a limitation on the graph learning-based video engagement prediction device, and may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.

[0189] The non-transitory software programs and instructions required to implement the graph learning-based video engagement prediction method applied to the graph learning-based video engagement prediction device in the above embodiments are stored in the memory. When executed by the control processor, the graph learning-based video engagement prediction method applied to the graph learning-based video engagement prediction device in the above embodiments is executed.

[0190] In addition, an embodiment of the present invention further provides a computer-readable storage medium storing computer-executable instructions, and when the computer-executable instructions are executed by one or more control processors, the one or more control processors can be caused to execute the graph learning-based video engagement prediction method in the above method embodiments.

[0191] Those of ordinary skill in the art will understand that all or some of the steps and systems disclosed in the above methods can be implemented as software, firmware, hardware, and their appropriate combinations. Some physical components or all physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or can be implemented as hardware, or can be implemented as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include a computer storage medium (or non-transitory medium) and a communication medium (or transitory medium). As is well known to those of ordinary skill in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disk (DVD) or other optical disk storage, magnetic cassettes, tapes, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, it is well known to those of ordinary skill in the art that communication media typically contains computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and can include any information delivery medium.

[0192] The embodiments of the present invention have been described in detail above in conjunction with the accompanying drawings. However, the present invention is not limited to the above embodiments, and various changes can be made without departing from the spirit of the present invention within the scope of knowledge possessed by those of ordinary skill in the relevant art.

Claims

1. A video engagement prediction method based on graph learning, characterized in that, it includes the following steps: Obtain video content; Extract modal features from the video content to obtain text data, audio data, and video data; Extract emotional features through graph learning based on the text data, the audio data, and the video data to obtain an emotional feature matrix; Extract features of keywords from the text data to obtain a text feature matrix; Extract features of target objects from the video data to obtain a video feature matrix; Extract features of spectrograms from the audio data to obtain an audio feature matrix; Input the emotional feature matrix, the text feature matrix, the video feature matrix, and the audio feature matrix into a user engagement prediction model to obtain a user engagement prediction result; Among them, the step of extracting emotional features through graph learning based on the text data, the audio data, and the video data to obtain an emotional feature matrix includes the following steps: Input the text data into a text feed-forward neural network for feature encoding to obtain a text embedding sequence, where the text embedding sequence includes multiple text embeddings, and each text embedding includes a first position identifier for characterizing the position of the text embedding in the text embedding sequence; Input the audio data into an audio feed-forward neural network for feature encoding to obtain an audio embedding sequence, where the audio embedding sequence includes multiple audio embeddings, and each audio embedding includes a second position identifier for characterizing the position of the audio embedding in the audio embedding sequence; Input the video data into a video feed-forward neural network for feature encoding to obtain a video embedding sequence, where the video embedding sequence includes multiple video embeddings, and each video embedding includes a third position identifier for characterizing the position of the video embedding in the video embedding sequence; Determine the temporal relationship between two embeddings based on the first position identifier, the second position identifier, and the third position identifier, where the two embeddings include at least one of two text embeddings, two video embeddings, two audio embeddings, a text embedding and a video embedding, a text embedding and an audio embedding, and an audio embedding and a video embedding; Construct a multimodal fully connected graph according to the text embeddings, the audio embeddings, the video embeddings, and the temporal relationship between two embeddings, where the text embeddings, the audio embeddings, and the video embeddings are all used as nodes of the multimodal fully connected graph, and the temporal relationship between two embeddings is used as the edge of the multimodal fully connected graph; Input the multimodal fully connected graph into a graph neural network for emotional feature extraction to obtain the emotional feature matrix.

2. The video engagement prediction method based on graph learning according to claim 1, characterized in that, the step of determining the temporal relationship between two embeddings based on the first position identifier, the second position identifier, and the third position identifier includes the following steps: When the two embeddings are embeddings of the same modality, determine the temporal relationship between the two embeddings according to the position identifier corresponding to the modality; When the two embeddings are embeddings of different modalities, the convolution kernel and convolution step size are set according to the lengths of the two embedding sequences where the two embeddings are located, and the two embedding sequences are aligned according to the convolution kernel and convolution step size to determine the first embedding and the second embedding that are aligned with each other in the two embedding sequences. Based on the position identifier of the first embedding, the temporal relationship between each embedding in the embedding sequence where the first embedding is located and the second embedding is determined.

3. The method for predicting video engagement based on graph learning according to claim 1, wherein, the step of extracting the sentiment feature matrix through graph learning according to the text data, the audio data, and the video data further includes the following steps: Determine the attention weights of the corresponding edges according to the embeddings of adjacent nodes in the multi-modal fully connected graph; Perform information fusion on adjacent nodes according to the attention weights to obtain new embeddings of each node; Determine the similarity between adjacent nodes according to the embeddings of the nodes; When the similarity of adjacent nodes is greater than the similarity threshold, delete the edges of the adjacent nodes; Delete the isolated nodes in the multi-modal fully connected graph that have no edges connected.

4. The method for predicting video engagement based on graph learning according to claim 1, wherein, the step of extracting the feature matrix of keywords from the text data includes the following steps: Extract the video text, the title text, and various parts of speech from the text data; Calculate the video text length, the title text length, and the part-of-speech ratio, and perform representation learning using a multi-layer perceptron to obtain the text feature matrix.

5. The method for predicting video engagement based on graph learning according to claim 1, wherein, the step of extracting the feature matrix of spectrograms from the audio data includes the following steps: Extract the Mel spectrogram from the audio data; Input the Mel spectrogram into a recurrent autoencoder for feature representation to obtain the audio feature matrix.

6. The method for predicting video engagement based on graph learning according to claim 1, wherein, the step of extracting the feature matrix of target objects from the video data includes the following steps: Divide the video data into several frame segments; Input the several frame segments into the trained YOLO v3 model for target object recognition respectively to obtain the video feature matrix for characterizing the appearance time of the target object in the video.

7. A system for predicting video engagement based on graph learning, wherein, it includes: A first module for obtaining video content; A second module for performing modal feature extraction on the video content to obtain text data, audio data, and video data; A third module for extracting sentiment feature matrix through graph learning according to the text data, the audio data, and the video data; A fourth module for extracting the feature matrix of keywords from the text data; A fifth module for extracting the feature matrix of target objects from the video data; The sixth module is configured to perform feature extraction on the spectrogram of the audio data to obtain an audio feature matrix; The seventh module is configured to input the emotion feature matrix, the text feature matrix, the video feature matrix, and the audio feature matrix into a user engagement prediction model to obtain a user engagement prediction result; Among them, the third module is specifically configured to perform the following steps: Input the text data into a text feed-forward neural network for feature encoding to obtain a text embedding sequence. Among them, the text embedding sequence includes multiple text embeddings, and the text embedding includes a first position identifier, and the first position identifier is used to represent the position of the text embedding in the text embedding sequence; Input the audio data into an audio feed-forward neural network for feature encoding to obtain an audio embedding sequence. Among them, the audio embedding sequence includes multiple audio embeddings, and the audio embedding includes a second position identifier, and the second position identifier is used to represent the position of the audio embedding in the audio embedding sequence; Input the video data into a video feed-forward neural network for feature encoding to obtain a video embedding sequence. Among them, the video embedding sequence includes multiple video embeddings, and the video embedding includes a third position identifier, and the third position identifier is used to represent the position of the video embedding in the video embedding sequence; Determine the temporal relationship between two embeddings based on the first position identifier, the second position identifier, and the third position identifier. Among them, the two embeddings include at least one of two text embeddings, two video embeddings, two audio embeddings, a text embedding and a video embedding, a text embedding and an audio embedding, and an audio embedding and a video embedding; Construct a multi-modal fully connected graph according to the text embedding, the audio embedding, the video embedding, and the temporal relationship between two embeddings. Among them, the text embedding, the audio embedding, and the video embedding are all used as nodes of the multi-modal fully connected graph, and the temporal relationship between the two embeddings is used as an edge of the multi-modal fully connected graph; Input the multi-modal fully connected graph into a graph neural network for emotion feature extraction to obtain the emotion feature matrix.

8. A video engagement prediction device based on graph learning, Characterized in that, It includes: At least one processor; At least one memory for storing at least one program; When the at least one program is executed by the at least one processor, at least one of the processors implements the graph learning-based video engagement prediction method according to any one of claims 1 to 6.

9. A computer-readable storage medium, in which a program executable by a processor is stored, Characterized in that, The program executable by the processor is used to implement the graph learning-based video engagement prediction method according to any one of claims 1 to 6 when executed by the processor.

Citation Information

Patent Citations

  • Systems and methods for interactive content generation

    CA2769994A1

  • Video recommendation method based on multi-modal video content and multi-task learning

    CN111246256A