Video Multi-Cue Social Relationship Extraction Method and Device Based on Knowledge Distillation

Through a knowledge distillation method, combined with scene recognition and semantic analysis models, a video multi-cue social relationship extraction framework is constructed, which solves the efficiency and accuracy of social relationship extraction of video characters in unconstrained scenarios, reduces the dependence on manual annotation, and improves the efficiency and accuracy of social relationship extraction of videos.

CN114972841BActive Publication Date: 2025-07-04BEIJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210426677.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-21
Publication Date
2025-07-04
Estimated Expiration
2042-04-21

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently extract the social relationships of video characters under video data in unconstrained scenarios, and the existing methods rely on manual labeling information and computing resources to consume a lot, ignoring the changes in complex relationships between multiple characters in the video.

Method used

Using a knowledge distillation-based method, soft targets are extracted by pre-training the teacher model, combined with scene recognition and semantic analysis models, scene features and semantic features are synchronized using cosine loss function, multi-layer attention network and graph convolution network are constructed, and a video multi-cues social relationship extraction framework is generated.

Benefits of technology

Provide richer semantic clues in unconstrained scenarios, enhance visual information, reduce dependence on manual annotation, and improve the efficiency and accuracy of social relationship extraction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114972841B_ABST
    Figure CN114972841B_ABST
Patent Text Reader

Abstract

The present invention proposes a method and device for video multi-cue social relationship extraction based on knowledge distillation. The method includes: obtaining a video frame sequence of an unconstrained scene video to be trained; preprocessing the video frame sequence through a pre-trained teacher model to extract soft targets; inputting the video frame sequence into a student model to obtain scene features and semantic features, and simultaneously performing synchronous training through a cosine loss function to narrow the distance between the scene features and semantic features and the soft targets. Among them, the student model includes a scene recognition model and a semantic analysis model; the scene features and semantic features are subjected to feature extraction and fusion through a multi-layer attention network, a convolutional layer, and a pooling layer, and the fused features, scene features, and semantic features are segmented and used as three types of nodes for graph construction; the node features after graph construction are aggregated through a graph convolutional network and classified through a classifier to generate a video multi-cue social relationship extraction framework.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer applications. Background Art

[0002] With the booming development of online social platforms and multimedia technologies, rich video content has attracted a large number of users to watch. Videos are gradually becoming the mainstream mode for people to record and spread their lives, and the quantity of interactive video data has thus increased massively. In this day and age, dynamic multimedia represented by videos has occupied the dominant position in Internet traffic. Video semantic analysis and content understanding are in urgent demand in practical applications and have thus gradually become a research hotspot in the field of computer applications. Video data provides richer time series and multi-modal clues. Compared with static pictures, multimedia data is heterogeneous in form and interrelated semantically, posing new challenges to artificial intelligence and deep learning algorithms. How to extract the relationships between video characters and understand video content has become a direction for promoting the intelligent development of society and one of the hotspots that researchers focus on.

[0003] As a key issue in multimedia content understanding, the task of extracting the social relationships of characters in videos is crucial for further analysis of character relationships, such as character behavior and emotion analysis, etc. It has great social and commercial value in the fields of public security monitoring, video content understanding, social network analysis, and visual quality assurance. Therefore, how to efficiently extract social relationships in videos is a very crucial issue.

[0004] We humans can relatively easily identify characters or infer their social relationships through some comprehensive clues, such as their appearance, interactions, conversations, clothing styles, and backgrounds. However, for artificial intelligence, automatically capturing the social relationships of characters by learning numerous clues in videos is still a challenging task. To solve this difficult task, people have done a great deal of work in relation extraction, with different motivations. Most of the work focuses on the spatio-temporal features of videos and pays attention to different video semantic information through the fusion of features. The effect of relation extraction depends on the quality of the features. There is also some work that models the connections between characters through composition, but most of them only focus on single-level visual clues, and the processing process and the overall model are very complex and not very friendly for applications.

[0005] In addition, current work mainly focuses on extracting social relationships from entire video clips. For example, some work aims to label general relationships in video clips, treating the relationships of many characters in a clip as the same. In this case, different relationships or complex relationships that change over time among multiple characters in the video may be overlooked. The video semantic information most relevant to social relationships usually requires the help of manually annotated relevant information and large-scale models to achieve the best results. However, most videos may have frequently changing characters and scenes as well as complex relationship description forms, which is an extremely heavy burden on manual preprocessing and computing resources. Therefore, there is an urgent need for a more general and simple solution for video data in unconstrained scenarios to provide richer semantic clues to enhance visual information, so as to better realize real-world applications. Summary of the Invention

[0006] The present invention aims to solve at least one of the technical problems in the related art to some extent.

[0007] To this end, the first object of the present invention is to propose a method for extracting multi-cue social relationships in videos based on knowledge distillation, which is used to provide richer semantic clues to enhance visual information for video data in unconstrained scenarios.

[0008] The second object of the present invention is to propose a device for extracting multi-cue social relationships in videos based on knowledge distillation.

[0009] To achieve the above object, an embodiment of the first aspect of the present invention proposes a method for extracting multi-cue social relationships in videos based on knowledge distillation, including:

[0010] Obtaining a video frame sequence of a video to be trained in an unconstrained scenario;

[0011] Preprocessing the video frame sequence through a pre-trained teacher model to extract soft targets;

[0012] Inputting the video frame sequence into a student model to obtain scene features and semantic features, and simultaneously performing synchronous training through a cosine loss function to narrow the distance between the scene features and semantic features and the soft targets; wherein, the student model includes a scene recognition model and a semantic analysis model;

[0013] Extracting and fusing the scene features and semantic features through a multi-layer attention network, a convolutional layer, and a pooling layer, segmenting the fused features, the scene features, and the semantic features, and using them as three types of nodes for graph construction;

[0014] Aggregating the node features after graph construction through a graph convolutional network and classifying them through a classifier to generate a framework for extracting multi-cue social relationships in videos.

[0015] In addition, the video multi-thread social relationship extraction method based on knowledge distillation according to the above embodiment of the present invention may also have the following additional technical features:

[0016] Furthermore, in one embodiment of the present invention, after generating a video multi-thread social relationship extraction framework, the method further includes:

[0017] Obtain a video frame sequence of an unconstrained scene video to be analyzed;

[0018] Inputting the video frame sequence of the unconstrained scene video to be analyzed into the multi-cue social relationship extraction framework of the video to be analyzed;

[0019] The social relations in the unconstrained scene video are extracted based on a video multi-cue social relation extraction framework.

[0020] Further, in one embodiment of the present invention, the synchronous training by the cosine loss function to shorten the distance between the scene feature and the semantic feature and the soft target includes:

[0021] The scene features and semantic features are mapped to the same feature space as the soft target output by the teacher model through pooling, and then the cosine loss function is used to shorten the distance between the soft target and the scene features and semantic features.

[0022] Furthermore, in one embodiment of the present invention, the fused features, the scene features, and the semantic features are segmented and composed as three types of nodes, including:

[0023] The scene features and semantic features are extracted through a multi-layer attention network, a convolution layer, and a pooling layer to adjust the features of their own weights and are mapped to obtain a feature sequence corresponding to the entire video frame. The first half, the middle half, and the second half of the features are selected as three nodes respectively, and then the fused features, the semantic features, and the scene features are used as three types of nodes. The fused feature nodes are fully connected with the semantic feature nodes and the scene feature nodes to perform composition.

[0024] Furthermore, in one embodiment of the present invention, after the classification is performed by the classifier, the method further comprises:

[0025] The student model is trained by weighted fusion of the cosine loss function of the scene feature and semantic feature and the classification loss function.

[0026] To achieve the above-mentioned purpose, the second embodiment of the present invention proposes a video multi-thread social relationship extraction device based on knowledge distillation, comprising:

[0027] An acquisition module, configured to acquire a video frame sequence of an unconstrained scenario video to be trained;

[0028] A preprocessing module, configured to preprocess the video frame sequence through a pre-trained teacher model to extract soft targets;

[0029] A distillation module, configured to input the video frame sequence into a student model to obtain scene features and semantic features, and simultaneously perform synchronous training through a cosine loss function to narrow the distance between the scene features and semantic features and the soft targets; wherein, the student model includes a scene recognition model and a semantic analysis model;

[0030] A composition module, configured to extract and fuse the scene features and semantic features through a multi-layer attention network, a convolutional layer, and a pooling layer, segment the fused features, the scene features, and the semantic features, and use them as three types of nodes for composition;

[0031] A generation module, configured to aggregate the node features after composition through a graph convolutional network and perform classification through a classifier to generate a video multi-cue social relationship extraction framework.

[0032] Further, in an embodiment of the present invention, it further includes an extraction module, configured to:

[0033] Acquire a video frame sequence of an unconstrained scenario video to be analyzed;

[0034] Input the video frame sequence of the unconstrained scenario video to be analyzed into the video multi-cue social relationship extraction framework to be analyzed;

[0035] Extract the social relationships in the unconstrained scenario video based on the video multi-cue social relationship extraction framework.

[0036] Further, in an embodiment of the present invention, the distillation module is further configured to:

[0037] Map the scene features and semantic features and the soft targets output by the teacher model to the same feature space through pooling, and then use the cosine loss function to narrow the distance between the soft targets and the scene features and semantic features.

[0038] Further, in an embodiment of the present invention, the composition module is further configured to:

[0039] The scene features and semantic features are used to extract features that adjust their own weights through a multi-layer attention network, a convolutional layer, and a pooling layer, and after mapping, a feature sequence corresponding to the entire video frame is obtained. The first half, the middle half, and the second half of the features are selected as three nodes respectively. Then, the fused features, the semantic features, and the scene features are used as three types of nodes, and the fused feature nodes are fully connected to the semantic feature nodes and the scene feature nodes to perform composition.

[0040] Further, in an embodiment of the present invention, the generation module further includes a training unit for:

[0041] After classification by the classifier, the student model is trained by weighted fusion of the cosine loss function and the classification loss function of the scene features and semantic features.

[0042] The method for extracting video multi-cue social relationships based on knowledge distillation proposed in the embodiments of the present invention: First, compress the model through the idea of knowledge distillation, so that the student model can learn as much knowledge as possible from the teacher network; Second, use the pre-trained teacher network to extract soft targets and learn knowledge related to social relationships, such as scenes and semantic objects in the video, without additional manual annotations; Third, construct ATCG to fuse features of multiple cues and capture temporal information; Fourth, experiments on the MovieGraphs dataset verify the effectiveness of our framework, indicating that our method is superior to the current state-of-the-art methods using compressed models. Description of the Drawings

[0043] The above and / or additional aspects and advantages of the present invention will become obvious and easy to understand from the following description of the embodiments in conjunction with the drawings, where:

[0044] Figure 1 It is a schematic flowchart of a method for extracting video multi-cue social relationships based on knowledge distillation provided by an embodiment of the present invention.

[0045] Figure 2 It is a schematic flowchart of a device for extracting video multi-cue social relationships based on knowledge distillation provided by an embodiment of the present invention.

[0046] Figure 3 It is a schematic diagram of a multi-cue social relationship extraction framework provided by an embodiment of the present invention. Detailed Embodiments

[0047] Embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present invention, but should not be construed as limiting the present invention.

[0048] The method and apparatus for extracting multi - clue social relationships in videos based on knowledge distillation according to embodiments of the present invention will be described below with reference to the accompanying drawings.

[0049] The present invention proposes a multi - clue social relationship extraction framework based on multi - teacher knowledge distillation to extract social relationships in videos of unconstrained scenarios. Specifically, first, different pre - trained teacher networks are used to extract semantic objects and scene features as soft targets respectively, and then the knowledge of the teachers is transferred to the compressed student to ensure the knowledge consistency between the teacher and student models, that is, to reduce the distance between the feature representation of the student and the soft target of the teacher. Secondly, a method of fusing multiple clue features and constructing a temporal clue graph based on the attention mechanism for the entire video is designed to capture temporal information and the connection between different clues. Specifically, we map multiple features to a new vector space through a multi - layer perceptron, then take the first half, middle half, and second half of the video as different nodes, and the features of multiple clues are regarded as representatives of different types of nodes, and then use GNN to learn and fuse the temporal information and the connection between different clues.

[0050] Figure 1 It is a schematic flowchart of a method for extracting multi - clue social relationships in videos based on knowledge distillation provided by an embodiment of the present invention.

[0051] As Figure 1 shown, the method for extracting multi - clue social relationships in videos based on knowledge distillation includes the following steps:

[0052] S1: Obtain a video frame sequence of an unconstrained scenario video to be trained;

[0053] S2: Pre - process the video frame sequence through a pre - trained teacher model to extract soft targets;

[0054] S3: Input the video frame sequence into the student model to obtain scene features and semantic features, and simultaneously perform synchronous training through a cosine loss function to narrow the distance between the scene features and semantic features and the soft targets; wherein, the student model includes a scene recognition model and a semantic analysis model;

[0055] S4: Extract and fuse the scene features and semantic features through a multi - layer attention network, as well as a convolutional layer and a pooling layer, segment the fused features, scene features, and semantic features and use them as three types of nodes to construct a graph;

[0056] S5: Aggregate the node features after composition through the graph convolutional network and classify them through the classifier to generate a video multi-cue social relationship extraction framework.

[0057] Furthermore, in one embodiment of the present invention, after generating a video multi-thread social relationship extraction framework, the method further includes:

[0058] Obtain a video frame sequence of an unconstrained scene video to be analyzed;

[0059] Inputting a video frame sequence of the unconstrained scene video to be analyzed into the multi-cue social relationship extraction framework of the video to be analyzed;

[0060] Extracting social relations in unconstrained scene videos based on video multi-cue social relation extraction framework.

[0061] Further, in one embodiment of the present invention, synchronous training is performed through a cosine loss function to shorten the distance between scene features and semantic features and soft targets, including:

[0062] Through pooling, the scene features and semantic features are mapped to the same feature space as the soft targets output by the teacher model, and then the cosine loss function is used to shorten the distance between the soft targets and the scene features and semantic features.

[0063] The scene recognition branch uses the pre-trained ResNet152 as the teacher model, and the relatively simple ResNet50 model as the student model. After that, we map the features of the student output and the soft targets of the teacher output to the same feature space through pooling. Then we use cosine similarity as the loss function to ensure the consistency of the scene information of the teacher-student model. The semantic analysis branch uses Deeplabv3plus pre-trained on the VOC2012 dataset as the teacher model. The student model adopts the spatial pyramid pooling idea of ​​Deeplabv3 and detects convolution features at multiple scales by applying dilated convolution.

[0064] Furthermore, in one embodiment of the present invention, the fused features, scene features, and semantic features are segmented and composed as three types of nodes, including:

[0065] The scene features and semantic features are extracted through a multi-layer attention network, convolutional layers, and pooling layers to adjust the features of their own weights and are mapped to obtain a feature sequence corresponding to the entire video frame. The first half, middle half, and second half of the features are selected as three nodes respectively, and then the fused features, semantic features, and scene features are used as three types of nodes. The fused feature nodes are fully connected with the semantic feature nodes and scene feature nodes to perform composition.

[0066] Among them, after obtaining the feature embeddings corresponding to multiple clues, the features for adjusting its own weights are extracted by using a multi-layer perceptron (Multi-Layer Attention) and mapped to obtain a feature sequence corresponding to the entire video frame. The first half, the middle half, and the second half of the features are selected as the three nodes for graph construction respectively. Similarly, the corresponding three parts of the features after semantic parsing and scene recognition are selected as nodes. Then, the fused features, semantic features, and scene features are used as three types of nodes, and there are edges connecting each pair of them, thereby constructing an ATCG. The feature representation of the entire network is learned through GCN to obtain a prediction result, and a cosine loss function is used to calculate the loss of relation classification.

[0067] Further, in an embodiment of the present invention, after classification by the classifier, it further includes:

[0068] The student model is trained by weighted fusion of the cosine loss function of the scene features and semantic features and the classification loss function.

[0069] Specifically, the knowledge distillation technology can not only compress the spatial network model, but also promote the performance of different student networks through corresponding multi-teacher knowledge distillation. Under the same input type and loss function, the comprehensive knowledge of multiple teachers is imparted to different students. In order to enable the students to learn the teacher model, we calculate the cosine similarity loss in the teacher-student models related to semantic objects and scenes respectively, and perform weighted training with the loss function fused and passed through the ATCG.

[0070] Figure 3 It is a schematic diagram of the multi-clue social relationship extraction framework of the present invention. First, the overall model reads the video frame sequence. The scene features and semantic features of the frame sequence are extracted and processed by the Teacher model in advance and used as the target soft target for the Student model to learn. The frame sequence is loaded one by one in batches and input into the Studnet model. The Student model is divided into two branches, which are small models corresponding to scene recognition and semantic analysis respectively. The specific structure is the small structure in the figure. Then, through the loss function, the features output by the Student model are respectively pulled closer to the features of the corresponding Teacher model pre-processed before, so as to achieve the effect of distillation.

[0071] Meanwhile, the features output by the two branches are subjected to feature extraction through a multi-layer Attention network, convolutional layers, and pooling layers, and used as nodes corresponding to the colors in the composition. The first half, middle half, and second half of the frame sequences are respectively used as three different aggregated nodes, and the same applies to the nodes for semantic analysis and scene recognition. After that, the features after composition are aggregated through a GCN (Graph Convolutional Network) and classified by a classifier. The classification loss is generated by comparing with the true labels. Finally, the three losses are weighted and fused for training simultaneously.

[0072] The method for extracting multi-cue social relationships in videos based on knowledge distillation proposed in the embodiments of the present invention: First, compress the model through the idea of knowledge distillation, enabling the student model to learn as much knowledge as possible from the teacher network; Second, use a pre-trained teacher network to extract soft targets and simultaneously learn knowledge related to social relationships, such as scenes and semantic objects in the video, without additional manual annotations; Third, construct an ATCG to fuse the features of multiple cues and capture temporal information; Fourth, experiments on the MovieGraphs dataset verify the effectiveness of our framework, indicating that our method is superior to the current state-of-the-art methods using compressed models.

[0073] To implement the above embodiments, the present invention also proposes a device for extracting multi-cue social relationships in videos based on knowledge distillation.

[0074] Figure 2 It is a schematic structural diagram of a device for extracting multi-cue social relationships in videos based on knowledge distillation provided by the embodiments of the present invention.

[0075] As Figure 2 shown, the device for extracting multi-cue social relationships in videos based on knowledge distillation includes: an acquisition module 10, a preprocessing module 20, a distillation module 30, a composition module 40, and a generation module 50.

[0076] Among them, the acquisition module is used to acquire the video frame sequence of the unconstrained scene video to be trained;

[0077] The preprocessing module is used to preprocess the video frame sequence through a pre-trained teacher model to extract soft targets;

[0078] The distillation module is used to input the video frame sequence into the student model to obtain scene features and semantic features, and simultaneously perform synchronous training through a cosine loss function to narrow the distance between the scene features and semantic features and the soft targets; wherein, the student model includes a scene recognition model and a semantic analysis model;

[0079] The composition module is used to extract and fuse scene features and semantic features through a multi-layer attention network, convolutional layer, and pooling layer, and then segment the fused features, scene features, and semantic features and compose them as three types of nodes;

[0080] The generation module is used to aggregate the node features after composition through the graph convolutional network and classify them through the classifier to generate a video multi-threaded social relationship extraction framework.

[0081] Furthermore, in one embodiment of the present invention, an extraction module is also included, which is used to:

[0082] Obtain a video frame sequence of an unconstrained scene video to be analyzed;

[0083] Inputting the video frame sequence of the unconstrained scene video to be analyzed into the multi-cue social relationship extraction framework of the video to be analyzed;

[0084] Extracting social relations in unconstrained scene videos based on video multi-cue social relation extraction framework.

[0085] Furthermore, in one embodiment of the present invention, the distillation module is also used for:

[0086] Through pooling, the scene features and semantic features are mapped to the same feature space as the soft targets output by the teacher model, and then the cosine loss function is used to shorten the distance between the soft targets and the scene features and semantic features.

[0087] Furthermore, in one embodiment of the present invention, the composition module is also used to:

[0088] The scene features and semantic features are extracted through a multi-layer attention network, convolutional layers, and pooling layers to adjust the features of their own weights and are mapped to obtain a feature sequence corresponding to the entire video frame. The first half, middle half, and second half of the features are selected as three nodes respectively, and then the fused features, semantic features, and scene features are used as three types of nodes. The fused feature nodes are fully connected with the semantic feature nodes and scene feature nodes to perform composition.

[0089] Furthermore, in one embodiment of the present invention, the generating module further includes a training unit, which is used to:

[0090] After classification by the classifier, the student model is trained by weighted fusion of the cosine loss function of scene features and semantic features and the classification loss function.

[0091] In the description of this specification, the description referring to terms such as "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0092] In addition, the terms "first" and "second" are used only for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of the present invention, "a plurality" means at least two, such as two, three, etc., unless otherwise specifically defined.

[0093] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.

Claims

1. A method for extracting multi-cue social relationships in videos based on knowledge distillation, characterized in that, It includes the following steps: Obtain the video frame sequence of the unconstrained scenario video to be trained; Preprocess the video frame sequence through a pre-trained teacher model to extract soft targets, where the soft targets include semantic objects and scene features extracted using different pre-trained teacher networks respectively; Input the video frame sequence into the student model to obtain scene features and semantic features, and simultaneously perform synchronous training through a cosine loss function to narrow the distance between the scene features and semantic features and the soft targets; where the student model includes a scene recognition model and a semantic analysis model, and the scene features and semantic features are mapped to the same feature space as the soft targets output by the teacher model through pooling, and then the cosine loss function is used to narrow the distance between the soft targets and the scene features and semantic features; Extract and fuse the scene features and semantic features through a multi-layer attention network, a convolutional layer, and a pooling layer, segment the fused features, the scene features, and the semantic features, and use them as three types of nodes for graph construction; Aggregate the node features after graph construction through a graph convolutional network, and classify them through a classifier to generate a video multi-cue social relationship extraction framework.

2. The method according to claim 1, wherein After generating the video multi-cue social relationship extraction framework, it further includes: Obtain the video frame sequence of the unconstrained scenario video to be analyzed; Input the video frame sequence of the unconstrained scenario video to be analyzed into the video multi-cue social relationship extraction framework to be analyzed; Extract the social relationships in the unconstrained scenario video based on the video multi-cue social relationship extraction framework.

3. The method according to claim 1, wherein The step of segmenting the fused features, the scene features, and the semantic features and using them as three types of nodes for graph construction includes: Extract the features with adjusted self-weights for the scene features and semantic features through a multi-layer attention network, a convolutional layer, and a pooling layer, and after mapping, obtain the feature sequence corresponding to the entire video frame. Select the first half, the middle half, and the second half of the features as three nodes respectively, and then use the fused features, the semantic features, and the scene features as three types of nodes. The fused feature node is fully connected to the semantic feature node and the scene feature node to perform graph construction.

4. The method according to claim 1, characterized in that After classification through the classifier, it further includes: Train the student model by weighted fusion of the cosine loss function and the classification loss function of the scene features and semantic features.

5. A video multi-cue social relationship extraction device based on knowledge distillation, characterized in that, It includes the following steps: An acquisition module for obtaining the video frame sequence of the unconstrained scenario video to be trained; A preprocessing module for preprocessing the video frame sequence through a pre-trained teacher model to extract soft targets, where the soft targets include semantic objects and scene features extracted using different pre-trained teacher networks respectively; A distillation module for inputting the video frame sequence into the student model to obtain scene features and semantic features, and simultaneously performing synchronous training through a cosine loss function to narrow the distance between the scene features and semantic features and the soft targets; where the student model includes a scene recognition model and a semantic analysis model; A composition module, used to extract and fuse the scene features and semantic features through a multi-layer attention network, a convolution layer and a pooling layer, and to segment the fused features, the scene features and the semantic features and compose them as three types of nodes; The generation module is used to aggregate the node features after composition through the graph convolution network and classify them through the classifier to generate a video multi-threaded social relationship extraction framework; The distillation module is also used for: The scene features and semantic features are mapped to the same feature space as the soft target output by the teacher model through pooling, and then the cosine loss function is used to shorten the distance between the soft target and the scene features and semantic features.

6. The device according to claim 5, characterized in that, Also includes extraction modules for: Obtain a video frame sequence of an unconstrained scene video to be analyzed; Inputting the video frame sequence of the unconstrained scene video to be analyzed into the multi-cue social relationship extraction framework of the video to be analyzed; The social relations in the unconstrained scene video are extracted based on a video multi-cue social relation extraction framework.

7. The device according to claim 5, characterized in that The composition module is further used for: The scene features and semantic features are extracted through a multi-layer attention network, a convolution layer, and a pooling layer to adjust the features of their own weights and are mapped to obtain a feature sequence corresponding to the entire video frame. The first half, the middle half, and the second half of the features are selected as three nodes respectively, and then the fused features, the semantic features, and the scene features are used as three types of nodes. The fused feature nodes are fully connected with the semantic feature nodes and the scene feature nodes to perform composition.

8. The device according to claim 5, characterized in that The generation module further includes a training unit, which is used to: After classification by the classifier, the student model is trained by weighted fusion of the cosine loss function of the scene feature and semantic feature and the classification loss function.

Citation Information

Patent Citations

  • Target object social relation identification method based on video semantics

    CN104778224A

  • Video figure behavior semantic meaning recognition method

    CN108509880A