A Human-Machine Collaboration Video Summarization Method Based on Pupil Size
Through a human-computer collaboration method based on pupil size, GRU is used to build an Encoder-Decoder structure to enhance global and local attention mechanisms, solving the problem of the inability to accurately express the audience's dynamic perceptual response in the existing technology, and achieving more efficient video digest generation.
Patent Information
- Application Number
- CN202211231244.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-09
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-10-09
AI Technical Summary
Existing deep learning-based video summary methods fail to accurately express the complex and dynamic perceptual responses of human audiences. Artificial annotations relying on public data sets are susceptible to individual audience memory bias and subjective selection, and cannot provide the audience's real-time interest when watching videos.
By modeling the subjects' real-time pupil size data, combining GRU to build a video summary model with Encoder-Decoder structure, introducing pupil dilation information to enhance global and local attention mechanisms, learning the relationship between video frames and audience attention, and generating attention-driven video summary.
Effectively extracting more attractive parts of the video, the average F-score is about 8% higher than the random score-generated video summary method, providing more accurate indicators of audience interest.
Smart Images

Figure CN115658963B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of digital video processing, and particularly relates to a human-computer collaboration video summarization method based on pupil size. Background Art
[0002] With the development of Internet technology and video devices in recent years, the data volume of digital videos on various platforms has shown an explosive growth trend. According to OMNICORE statistics and industry insights, more than 500 hours of content are uploaded to YouTube every minute, and the total daily viewing time of people exceeds 1 billion hours of videos. Therefore, how to help users quickly and effectively select the content they are interested in from a vast amount of videos has become an increasingly challenging problem.
[0003] Video summarization technology is to solve the above problems. Video summarization, also known as video summary, is a brief summary of video content; video summarization technology captures important and representative information in the original video to generate a series of video frames or video segments, providing users with a fast and relatively comprehensive video browsing method. In recent years, the attention mechanism has been widely applied in computer vision tasks and achieved many important breakthroughs; essentially, human viewers are the ultimate object and consumer of video summarization, and the attention mechanism and viewing habits of viewers are of great significance for capturing key and interesting information in the original video. Based on the above two points, it is natural and has great potential to introduce the attention mechanism into the field of video summarization.
[0004] Obtaining the importance score of each frame in the video is a key link in video summarization technology. The importance score represents the importance and representativeness of the information contained in the video frame and is the basis for the subsequent key frame selection link. Most current mainstream deep learning-based video summarization algorithms use GRU to model the temporal information between video frames to obtain the long temporal depth features of video frames and use them to regress the importance score of each frame. Introducing the attention mechanism of human viewers into the video summarization method, assigning different importance weights to video frames, and thus adjusting the importance scores of each frame, such an internal connection is established between the video key frames selected by the summarization method and the viewers' interest in video content, making the generated summary effectively improve the viewing experience of viewers.
[0005] At present, many scholars have tried to introduce attention mechanisms in the field of video summarization and achieved fruitful results. For example, in the literature of Ma et al., "A user attention model for video summarization[C] / / Proceedings of the tenth ACM international conference on Multimedia. 2002:533-542", a video summarization technology based on a computational audiovisual attention model was proposed. By using a computational human attention model, the need for complex heuristic rules in video summarization was filled. Specifically, this method modeled how the attention of viewers was attracted by actions, objects, sounds, and languages when watching video programs, and designed a video summarization technology based on an audiovisual modeling method. Another example is that Ji and his colleagues first proposed an attention-based video summarization method using an encoder-decoder network in the literature "Video summarization with attention-based encoder–decoder networks. IEEE Transactions on Circuits and Systems for Video Technology. 30(6):1709-17(2019)". This method introduced the attention mechanism in the field of natural language processing and imitated the way humans select key frames by inserting an intermediate attention layer in the video summarization framework. Fajtl et al. proposed a video summarization method with the attention mechanism as the core of computational analysis in the literature "Monekosso D, Remagnino P.: Summarizing Videos with Attention. In: Computer Vision–ACCV2018 Workshops, pp. 39-54. Springer International Publishing. (2019)". This method avoided the use of computationally demanding LSTM by applying a soft self-attention mechanism with simple concepts and high computational efficiency. Its video summarization network consisted of only a self-attention mechanism structure and a two-layer fully connected network for regressing frame importance scores.
[0006] However, it is not difficult to find that most current video summarization methods based on deep learning that introduce attention mechanisms often adopt simplified computational models, rather than learning the true visual attention mechanism from human perceptual responses. These computational models cannot accurately express the complex and dynamic perceptual responses of human viewers. More importantly, most of these methods rely on manual annotations of public datasets, which are vulnerable to the memory biases and subjective choices of individual viewers and cannot provide the real-time interests of viewers when watching videos. Summary of the Invention
[0007] In view of the above, the present invention provides a human-computer collaborative video summarization method based on pupil size, which can model the attention mechanism of human viewers. The model will learn which parts of the video are more likely to arouse the interest of viewers, and then extract a video summary from the moments when attention is concentrated.
[0008] A human-computer collaborative video summarization method based on pupil size includes the following steps:
[0009] (1) Conduct an eye movement tracking experiment on the subject to obtain the video file watched by the subject, and record the real-time pupil size data of the subject during the viewing process;
[0010] (2) Decompose the video file into a video frame sequence, and use a pre-trained convolutional neural network to perform deep feature extraction on the video frame sequence to obtain a video frame deep feature sequence X;
[0011] (3) Calculate an attention score sequence AS and a pupil dilation information sequence PD based on the real-time pupil size data;
[0012] (4) Conduct multiple tests on different subjects according to the above steps to obtain multiple groups of samples, and divide all samples into a training set and a test set. Each group of samples includes a video frame deep feature sequence X, an attention score sequence AS, and a pupil dilation information sequence PD;
[0013] (5) Use GRU (Gated Recurrent Units) to build a video summarization model based on the Encoder-Decoder structure, which includes an encoder, a decoder, and an attention mechanism module. The encoder is used to encode the input video frame deep feature sequence X and output an encoding result E; the attention mechanism module enhances local attention with video position encoding information and enhances global attention with the pupil dilation information sequence PD, and outputs an attention weight score Attention; the decoder uses the result Z after adding E and Attention as the input to learn the dependence relationship between video frames and attention scores, so as to predict an attention score sequence Y corresponding to the video frame sequence.
[0014] (6) Use X and PD in the training set samples as the model inputs, and AS as the label to train the video summarization model;
[0015] (7) Input X and PD in the test set samples into the trained video summarization model, then the corresponding attention score sequence Y can be predicted, and then key shots can be selected according to this sequence and synthesized into a video summary.
[0016] Further, the specific implementation of step (3) is as follows: First, convert the real-time pupil size data of the subject during the viewing process into a pupil size sequence P = {p1, p2,..., p t ,…, p N} corresponding to the video frame sequence; then calculate the corresponding attention score sequence AS = {AS1, AS2,..., AS t ,…, AS N} and pupil dilation information sequence PD = {PD1, PD2,..., PD t ,…, PD N} through the following formula;
[0017]
[0018] PD t =p t -p t-1
[0019] where: p max and p min are the maximum and minimum values in the pupil size sequence P respectively, and p t represents the pupil size when the subject sees the t-th frame image, t is a natural number and 1 ≤ t ≤ N, and N is the total number of frames of the video file.
[0020] Further, the attention mechanism module includes two different processing mechanisms. One is the global multi-head attention mechanism guided by eye movement data, and the other is the local multi-head attention mechanism guided by video position encoding information. The sum of the outputs of the two processing mechanisms is the attention weight score Attention.
[0021] Further, the global multi-head attention mechanism uses multiple queries to parallelly calculate multiple pieces of information selected from the input sequence X. Each parallel global attention structure focuses on vectors in different subspaces of the sequence X, and then they are concatenated to finally obtain the global attention weight score Multi-Attention Global , and the overall calculation process is as follows:
[0022]
[0023]
[0024]
[0025] Wherein: respectively represent the query vector, key vector, and value vector corresponding to the i-th subspace of the global space, respectively represent the query vector weight, key vector weight, and value vector weight corresponding to the i-th subspace of the global space, is the attention weight score output for the i-th subspace, is the global dot product weight, d k is the number of hidden units of the GRU, * is the scalar multiplication symbol, T is the transpose symbol, Concact() is the concatenation operation, softmax() is the Softmax function, i is a natural number and 1 ≤ i ≤ num head , num head is the number of subspaces of the multi-head attention mechanism.
[0026] Furthermore, the local multi-head attention mechanism first divides the input sequence X into M small segments, each small segment containing the depth feature information of N / M video frames, uses the video position encoding matrix PE to enhance the dot product process of the query vector and the key vector, and performs the concatenation of the local attention weight scores of each video segment after the concatenation of the multi-head attention mechanism, and finally obtains the local attention weight score Multi-Attention Local , and the overall calculation process is as follows:
[0027]
[0028]
[0029]
[0030] Multi-Attention Local
[0031] = Concact(Multi-Attention 1-th ,..., Multi-Attention M-th )
[0032] Wherein: respectively represent the query vector, key vector, and value vector corresponding to the i-th subspace on the j-th segment of the video, respectively represent the query vector weight, key vector weight, and value vector weight corresponding to the i-th subspace on the j-th segment of the video, X j represents the j-th segment of the sequence X, is the attention weight score outputted by the i-th subspace on the j-th segment, is the dot product weight of the jth fragment, Multi-Attention j-th is the attention weight score of the jth segment, d k is the number of hidden units of GRU, T is the transposition symbol, Concact() is the concatenation operation, softmax() is the Softmax function, i is a natural number and 1≤i≤num head , num head is the number of subspaces of the multi-head attention mechanism, j is a natural number and 1≤j≤M, N is the total number of frames in the video file, and M is a natural number greater than 1.
[0033] Furthermore, the dimension of the video position encoding matrix PE is The expressions for the values of each element in the matrix are as follows:
[0034]
[0035]
[0036] Among them: PE pos,2r Represents the element value at the posth row and 2rth column in the matrix PE, PE pos,2r+1 represents the element value of the posth row and 2r+1th column in the matrix PE, where pos represents the absolute position of the current frame in the video to which it belongs and pos is a natural number and pos∈[0,N), d model =d k / num_head, r is a natural number.
[0037] Furthermore, the process of training the network model in step (6) is as follows:
[0038] 6.1 Initialize model parameters, including bias and weight, learning rate, optimization method, and maximum number of iterations;
[0039] 6.2 Input X and PD in the training set samples into the model, forward propagate the model output corresponding prediction results, and calculate the loss function L between the prediction results and the labels;
[0040] 6.3 Based on the loss function L, the model parameters are continuously updated using the gradient descent method until the loss function converges or the maximum number of iterations is reached, and the training is completed.
[0041] Furthermore, the expression of the loss function L is as follows:
[0042]
[0043] Where: y tis the attention score corresponding to the t-th frame image in the attention score sequence Y predicted by the model, AS t is the attention score corresponding to the t-th frame image in the label sequence.
[0044] Further, in the step (7), for any video file, visually continuous frames are combined into shots, and then the predicted attention score sequence Y = {y1, y2, …, y i , …, y N} is converted into a shot importance score sequence, that is, the importance score of the shot can be obtained by averaging the attention scores of the frames in the shot; furthermore, according to the shot importance score sequence and the shot length, the 0 / 1 knapsack algorithm is used to select key shots, and the length of the key shots does not exceed 15% of the entire video length. Finally, all key shots are synthesized into a video summary.
[0045] The method of the present invention is based on the theory that there is a close connection between the spontaneous non-verbal reactions of the audience and their real-time attention changes when watching videos. Utilizing the characteristics that the pupillary light response (pupil size) can be used to indicate the more attractive parts in the video and the data is easy to obtain, a perception-driven video dataset is made, providing a basis for the video summary model to learn the real-time and dynamic attention mechanism of the audience.
[0046] In addition, the present invention uses a human-machine collaborative video summary framework composed of an encoder-decoder module, an attention mechanism module, and a key frame selection module. It can supervised learn the relationship between video features and the audience's attention to the video, and finally obtain an attention-driven video summary model that can automatically generate a summary according to the original video. Experiments show that the present invention can effectively extract the more attractive parts in the video. Compared with the video summary generated using random scores, the average F-score of the present invention is increased by about 8%. Description of the Drawings
[0047] Figure 1 is a schematic flow chart of the video summary method of the present invention.
[0048] Figure 2 is a schematic diagram of the overall process of the eye movement data acquisition experiment.
[0049] Figure 3 is a schematic diagram of the structure of the video summary model of the present invention.
[0050] Figure 4 is a schematic diagram of the multi-head attention structure.
[0051] Figure 5 is a schematic diagram of the module structure of the global attention mechanism and the local attention mechanism.
[0052] Figure 6It is a schematic diagram of the bidirectional GRU encoder structure. Detailed implementation manners
[0053] To describe the present invention more specifically, the technical solutions of the present invention will be described in detail below in conjunction with the accompanying drawings and specific implementation manners.
[0054] As Figure 1 shown, the human-computer collaborative video summarization method based on pupil size of the present invention includes the following steps:
[0055] S1. Design and conduct an experiment using eye tracking technology, record the real-time pupil responses of the subjects, and establish a perception-driven video dataset based on the collected data.
[0056] The present invention uses an Eyelink Portable Duo eye tracker to obtain the pupil size information of the subjects. The eye tracker is set to the monocular "pupil-corneal reflect" tracking mode and records data at a sampling rate of 500 Hz. In the experiment, a 23-inch liquid crystal display with a refresh rate of 60 Hz and a resolution of 1920*1080 is used to play videos. As Figure 2 shown, the process of the data acquisition experiment is as follows:
[0057] (1) At the beginning of each experiment, 13-point calibration and verification are applied, and drift correction is performed at the start of video playback. If the drift error exceeds 0.5°, a new calibration is carried out.
[0058] (2) During the experiment, the participants are required to freely watch the videos, and their pupil size data are recorded by the eye tracker.
[0059] (3) After half of the videos (10 videos) are played, forced calibration is performed; to avoid fatigue, there is a short break between two consecutive videos.
[0060] (4) During the experiment, a test of the degree of attention concentration is randomly carried out after a video is played: two pictures are presented, and the participants are required to select the picture that just appeared in the video. Such tests are carried out 4 times in total when a single subject conducts the experiment.
[0061] The obtained original pupil size data are processed through preprocessing steps such as screening out abnormal extreme values, filling in missing values, smoothing, normalizing, and downsampling to obtain a frame-level (30 fps) pupil size data sequence P = {p1, p2, …, p t , …, p N}. Since attention is closely related to the pupil light response, an attention score AS = {AS1, AS2, …, AS t , …, AS N, as a measure of the participants' interest in and attention to the video content, is shown as follows:
[0062]
[0063] Note that the calculation formula of the pupil dilation information PD used in the attention model is as follows:
[0064] PD t =Δp t =p t -p t-1 i = 2, 3, …, N
[0065] PD1 = PD2
[0066] The perception-driven dataset contains 20 videos from the TVSum video dataset with a total duration of about 1 hour. Each video is in the standard frame rate of 30fps. The comparison of the consistency between the video information in the dataset and the importance annotation of each part in the video is shown in Table 1, where Pair. (Proposed dataset) and Pair. (TVSum) are the paired F-scores of the automatic annotation of the videos in the perception-driven dataset and the manual annotation of the same videos in the TVSum dataset respectively, which are used to indicate the consistency of the annotations of different viewers.
[0067] Table 1
[0068]
[0069]
[0070] The automatic annotation based on the pupil size signal used in the perception-driven dataset has the advantages of better real-time performance, insensitivity to the memory errors and subjective choices of the subjects, high consistency of the information provided by each subject, and saving of labor and time costs compared with the original manual annotation, which has positive significance for video summarization research.
[0071] S2. Convert the video file into a set of video frame pictures, and then use the pre-trained convolutional neural network to extract the depth features X of the video frames.
[0072] In this example, the original video is first converted into a set of video frame pictures at a sampling rate of 2fps, and then the GoogLeNet network pre-trained on the large-scale image dataset ImageNet is used to extract the depth features of the frame pictures, obtaining the depth feature sequence X = {x1, x2, …, x t , …, x N}, where N represents the video length, and the dimension of the depth feature corresponding to each video frame is 1024.
[0073] S3. Design and build an abstract model based on the Encoder-Decoder structure, as Figure 3 shown. It includes an image extraction module using CNN, a Bi-GRU encoder, a GRU decoder, and an attention mechanism module guided by external attention for eye movement data representation.
[0074] The deep feature sequence X advances along two different processing paths in the attention layer. One of the paths is mainly an eye movement data-guided global multi-head attention mechanism, as Figure 4 and Figure 5 shown. First, through linear transformation, the representations of the deep feature sequence X in different vector spaces, namely Q (Query, query vector), K (Key, key vector), and V (Value, value vector), are obtained. For the single-head attention mechanism, there are:
[0075]
[0076]
[0077]
[0078] According to the Scaled Dot-Product Attention used in "Attention is All You Need", the dot product of Q Global and K Global generates a weight matrix, and then the weight is multiplied by V Global , normalized by the softmax function to obtain the global video frame attention weight score. To prevent the value input to the softmax function from being too large, the dot product result is divided by a value (d k is the dimension of the hidden layer, which is 1024 in this example), as shown in the following formula:
[0079]
[0080] Introduce the pupil dilation information PD to enhance the global attention mechanism (* represents scalar multiplication), then there is:
[0081]
[0082] The multi-head attention mechanism uses multiple queries to parallelly calculate multiple pieces of information selected from the input information X. Each parallel global attention structure focuses on different parts of X (vectors in different subspaces), and then they are concatenated. Finally, the global attention weight score in the case of multi-head attention is obtained. The overall calculation process is as follows:
[0083]
[0084]
[0085]
[0086] where: i = 1, 2, 3, .. num_head, respectively represent the query vector, key vector, and value vector corresponding to the i-th subspace of the global, is the attention weight score output by the i-th subspace, is the global dot product weight, Multi-Attention Global i.e., the final, concatenated global attention weight score.
[0087] The main body of another processing path is multiple local attention mechanism units, such as Figure 3 shown, the depth feature sequence X first undergoes a segmentation step and is subdivided into M small segments, each segment containing the information of N / M video frames, with:
[0088]
[0089] Similarly, the multi-head attention mechanism is also used on the local attention processing path. Compared with the global attention processing path, there are mainly two differences between the two: ① The local attention processing path uses the video position encoding information PE to enhance the dot product process of the query vector Q and the key vector K; ② The local attention processing path will also perform the concatenation of the local attention weight scores of each segment of the video after the concatenation step of the multi-head attention mechanism; the overall calculation process is as follows:
[0090]
[0091]
[0092]
[0093] Multi-Attention Local
[0094] = Concact(Multi-Attention 1-th ,..., Multi-Attention M-th )
[0095] where: j = 1, 2, 3,..., M, respectively represent the query vector, key vector, and value vector corresponding to the i-th subspace on the j-th segment of the video, is the attention weight score output by the i-th subspace on the j-th segment, is the dot product weight of the j-th segment, Multi-Attention j-th is the attention weight score of the j-th segment, Multi-Attention Local That is, the final, concatenated local attention weight score.
[0096] The video position encoding matrix PE used in the present invention provides information on the position of each video frame. Specifically, the absolute position of the frame is encoded by sine and cosine functions of different frequencies, as shown in the following formula:
[0097]
[0098]
[0099] where: pos is the absolute position of the current frame in the j-th segment of the video to which it belongs, pos ∈ [0, N), d model is the feature dimension under the current multi-head attention mechanism. For the single-head attention mechanism, d model takes the value of the number of GRU hidden units, 1024; for the multi-head attention mechanism, d model = 1024 / num_head, and the dimension of the calculated two-dimensional matrix PE is
[0100] In summary, the calculation formula for the video weight score Attention output by the attention module is as follows:
[0101] Attention = Multi-Attention Global + Multi-Attention Local
[0102] As Figure 6 shown, the Bi-GRU-based encoder reads the depth feature sequence X, and then the forward GRU calculates the forward hidden state from it Similarly, the backward GRU calculates the backward hidden state The encoder connects the two states and calculates the encoded result E of X = {e1, e2,..., e t ,..., e N}.
[0103]
[0104] The decoder is implemented by a single-layer GRU. As Figure 3 shown, it receives the sum Z of the weight score Attention and the encoder output E, and calculates the hidden state h = {h1, h2,..., d t ,..., h N}, the output decoded result D = {d1, d2, ..., d t , ..., d N}.
[0105] Z = Attention + E
[0106]
[0107] The activation function area consists of a double - layer tanh activation function, a dropout layer, a relu activation function, a sigmoid activation function, and multiple linear transformation layers. Its function is to process the decoded result D to obtain the attention scores Y = {y1, y2, …, y t , …, y N}.
[0108] S4. Conduct model parameter training. In this process, note that the model learns the attention mechanism of human viewers through a perception - driven dataset.
[0109] The perception - driven dataset only has 20 videos, and the dataset size is small, which is suitable for training and testing using the cross - validation method to make full use of the limited data. The present invention selects 80% of the data as the training set and 20% of the data as the test set, and uses the ten - fold cross - validation method. In each partitioning method, the number of test sets is 16 and the number of validation sets is 4. During the experiment, the mean squared error loss (MSE Loss) is used to evaluate the gap between the model - predicted attention scores and the annotations in the dataset, as shown in the following formula, where (y1, …, y N ) are the model - predicted attention scores, and (a1, …, a N ) are the attention scores annotated in the dataset.
[0110]
[0111] S5. Use the video summarization model to process the videos in the test set, generate summaries, and evaluate the summary effects.
[0112] Combine visually continuous frames into shots, and at the same time convert the frame - level attention scores to shot - level attention scores. Use the 0 / 1 knapsack algorithm to select key shots. The length of the key shots needs to be limited to 15% of the original video length. The combination of key shots is the video summary.
[0113] To verify the effectiveness of the method of the present invention, the results of video summaries generated based on random scores are compared. Generating video summaries based on random scores means replacing the sequence of attention scores predicted by the model with a sequence of random values (the range of random values is 0 - 1) of the same length and alignment, and using the same key - frame selection method to obtain the corresponding video summaries.
[0114] We use precision Pr, recall Re, and F-measure as evaluation metrics, and their calculation methods are as follows:
[0115]
[0116]
[0117]
[0118] Where: A is the true video summary provided by the dataset, B is the summary generated by the method of the present invention, and the indicators of the obtained summary and its comparison with the summary generated by the random score are shown in Table 2:
[0119] Table 2
[0120]
[0121] The above description of the embodiments is to facilitate the understanding and application of the present invention by ordinary technical personnel in the technical field. Those who are familiar with the technology in this field can obviously make various modifications to the above embodiments easily, and apply the general principles described here to other embodiments without creative labor. Therefore, the present invention is not limited to the above embodiments, and all improvements and modifications made by those skilled in the art based on the disclosure of the present invention should be within the protection scope of the present invention.
Claims
1. A human-machine collaboration video summarization method based on pupil size, comprising the following steps: (1) Conduct an eye-tracking experiment on the subject to obtain the video file watched by the subject, and record the real-time pupil size data of the subject during the viewing process; (2) Decompose the video file into a video frame sequence, and use a pre-trained convolutional neural network to extract deep features from the video frame sequence to obtain a video frame deep feature sequence X; (3) Calculate the attention score sequence AS and the pupil dilation information sequence PD based on the real-time pupil size data; (4) Conduct multiple tests on different subjects according to the above steps to obtain multiple groups of samples, and divide all samples into a training set and a test set. Each group of samples includes the video frame deep feature sequence X, the attention score sequence AS, and the pupil dilation information sequence PD; (5) Use GRU to build a video summarization model based on the Encoder-Decoder structure, which includes an encoder, a decoder, and an attention mechanism module. The encoder is used to encode the input video frame deep feature sequence X and output an encoded result E; the attention mechanism module enhances local attention with video position encoding information and enhances global attention with the pupil dilation information sequence PD, and outputs an attention weight score Attention; the decoder uses the result Z after adding E and Attention as the input to learn the dependence relationship between video frames and attention scores, so as to predict the attention score sequence Y corresponding to the video frame sequence; (6) Use X and PD in the training set samples as model inputs and AS as labels to train the video summarization model; (7) Input X and PD in the test set samples into the trained video summarization model, and the corresponding attention score sequence Y can be predicted, and then key shots can be selected according to this sequence and synthesized into a video summary.
2. The human-machine collaboration video summarization method according to claim 1, wherein: The specific implementation of step (3) is as follows: First, convert the real-time pupil size data of the subject during the viewing process into a pupil size sequence P = {p1, p2, …, p t , …, p N}; corresponding attention score sequence AS = {AS1, AS2, …, AS t , …, AS N} and pupil dilation information sequence PD = {PD1, PD2, …, PD t , …, PD N} are then calculated through the following formula; PD t = p t -p t-1 where: p max and p min are the maximum value and the minimum value in the pupil size sequence P respectively, and p t represents the pupil size when the subject sees the t-th frame image, where t is a natural number and 1 ≤ t ≤ N, and N is the total number of frames of the video file.
3. The human-machine collaborative video summarization method according to claim 1, wherein: The attention mechanism module contains two different processing mechanisms. One is a global multi-head attention mechanism guided by eye movement data, and the other is a local multi-head attention mechanism guided by video position encoding information. The outputs of the two processing mechanisms are added together to obtain the attention weight score Attention.
4. The human-machine collaboration video summarization method according to claim 3, wherein: The global multi-head attention mechanism uses multiple queries to calculate multiple pieces of information selected from the input sequence X in parallel. Each parallel global attention structure focuses on vectors in different subspaces of the sequence X, and then they are concatenated. Finally, the global attention weight scores in the case of multi-head attention, namely Multi-Attention, are obtained. Global , and the overall calculation process is as follows: Wherein: respectively represent the query vector, key vector, and value vector corresponding to the i-th subspace of the global space, respectively represent the query vector weight, key vector weight, and value vector weight corresponding to the i-th subspace of the global space, is the attention weight score output by the i-th subspace, is the global dot product weight, d k is the number of hidden units of the GRU, * is the scalar multiplication symbol, T is the transpose symbol, Concact() is the concatenation operation, softmax() is the Softmax function, i is a natural number and 1 ≤ i ≤ num head , num head is the number of subspaces of the multi-head attention mechanism.
5. The human-computer collaborative video summarization method according to claim 3, wherein: The local multi-head attention mechanism first divides the input sequence X into M small segments, each small segment containing the depth feature information of n / M video frames. It uses the video position encoding matrix PE to enhance the dot product process of the query vector and the key vector, and splices the local attention weight scores of each segment of video after the splicing of the multi-head attention mechanism, and finally obtains the local attention weight score Multi-Attention in the case of multi-head attention Local , and the overall calculation process is as follows: Multi-Attention Local = Concact(Multi - Attention 1-th ,…,Multi - Attention M-th ) Wherein: respectively represent the query vector, key vector, and value vector corresponding to the $i$-th subspace on the $j$-th segment of the video, respectively represent the query vector weight, key vector weight, and value vector weight corresponding to the $i$-th subspace on the $j$-th segment of the video, $X$ j represents the $j$-th segment of the sequence $X$, is the attention weight score output by the $i$-th subspace on the $j$-th segment, is the dot product weight of the $j$-th segment, Multi-Attention j-th is the attention weight score of the $j$-th segment, $d$ k is the number of hidden units of the GRU, $T$ is the transpose symbol, Concact() is the concatenation operation, softmax() is the Softmax function, $i$ is a natural number and $1\leq i\leq num$ head , $num$ head is the number of subspaces of the multi-head attention mechanism, $j$ is a natural number and $1\leq j\leq M$, $N$ is the total number of frames of the video file, and $M$ is a natural number greater than 1.
6. The human-computer collaborative video summarization method according to claim 5, wherein: The dimension of the video position encoding matrix PE is The expression of each element value in the matrix is as follows: Where: PE pos,2r represents the element value at the 2r-th column of the pos-th row in the matrix PE, and PE pos,2r+1 represents the element value at the (2r + 1)-th column of the pos-th row in the matrix PE. pos represents the absolute position of the current frame in the affiliated video, and pos is a natural number and pos ∈ [0, N). d model = d k / num_head, and r is a natural number.
7. The method for human-computer collaborative video summarization according to claim 1, wherein: The process of training the network model in step (6) is as follows: 6.1 Initialize the model parameters, including biases and weights, learning rate, optimization method, number of GRU hidden units, and maximum number of iterations; 6.2 Input X and PD in the training set samples into the model, and the model propagates forward to output the corresponding prediction results, and calculate the loss function L between the prediction results and the labels; 6.3 Continuously iterate and update the model parameters according to the loss function L using the gradient descent method until the loss function converges or reaches the maximum number of iterations, and the training is completed.
8. The human-machine collaborative video summarization method according to claim 7, wherein: The expression of the loss function L is as follows: where: y t is the attention score corresponding to the t-th frame image in the attention score sequence Y predicted by the model, AS t is the attention score corresponding to the t-th frame image in the label sequence.
9. The human-machine collaborative video summarization method according to claim 1, wherein: In step (7), for any video file, visually continuous frames are combined into shots, and then the predicted attention score sequence Y = {y1, y2, …, y i , …, y N} is converted into a shot importance score sequence, that is, the importance score of a shot can be obtained by averaging the attention scores of the frames in the shot; furthermore, according to the shot importance score sequence and the shot length, the 0 / 1 knapsack algorithm is used to select key shots, the length of the key shots does not exceed 15% of the entire video length, and finally all the key shots are synthesized into a video summary.
Citation Information
Patent Citations
Supervised video abstract extraction method utilizing visual attention mechanism
CN108024158A
Attention-assisted unsupervised video abstraction system
CN112560760A