A sketch-based scene-level fine-grained video retrieval method and system

By using scene-level sketch compression technology and adaptive frame sampling strategy in video retrieval, combined with graph convolutional neural network, the problem of fine-grained video retrieval in multiple objects in complex scenarios is solved, achieving wider application and more efficient feature extraction.

CN114969430BActive Publication Date: 2025-05-06INST OF SOFTWARE - CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210429989.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-04-28
Filing Date
2022-04-22
Publication Date
2025-05-06
Estimated Expiration
2042-04-22

AI Technical Summary

Technical Problem

Existing sketch-based video retrieval methods are difficult to effectively retrieve fine-grained videos containing multiple objects in complex scenarios.

Method used

A scene-level fine-grained video retrieval method based on sketches is proposed. By compressing objects from different time periods into the same sketch, scene-level video content is summarized, and using technologies such as adaptive frame sampling strategies and graph convolution neural networks, a sketch-video association relationship model and search model are constructed to realize scene-level fine-grained retrieval of videos.

Benefits of technology

It realizes effective retrieval of multiple object actions and background information in complex scenarios, expands the application scenarios of sketch search, and improves the efficiency of video feature extraction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114969430B_ABST
    Figure CN114969430B_ABST
Patent Text Reader

Abstract

The present invention discloses a scene-level fine-grained video retrieval method and system based on sketches, which belongs to the field of computer vision. The method extracts sketch features to construct a sketch space structure diagram based on appearance features and category features; extracts video features to construct a video spatiotemporal structure diagram based on video timing information, appearance features and category features, and uses an adaptive frame sampling strategy to sample video frames, first sparsely sample video frames, and then use a sketch-video association model to screen video frames, and use a sketch-video retrieval model to complete video fine-grained retrieval. The present invention compresses objects that appear in different time periods into the same sketch, performs scene-level video content summarization, and retrieves videos that are consistent with the sketch background elements, object appearance features and action types.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision, and in particular relates to a sketch-based scene-level fine-grained video retrieval method and system. Background Art

[0002] Sketches are hand-drawn images that can abstractly display the appearance features and location information of objects through simple line combinations. They are widely used in sketch-based image research, such as image completion, image retrieval, image generation and other tasks. Sketches can summarize objects and corresponding actions in different time periods in a sketch across temporal information. The superiority of combining temporal information and spatial position has gradually made them applied to video-related tasks. Sketch-based video retrieval (SBVR) tasks have made great research progress. It can be widely used in vision and multimedia fields, such as video browsing, video query and editing. However, most of the existing sketch-based video retrieval work uses traditional methods, where sketches only draw simple line outlines for category-level video retrieval; while instance-level video retrieval using deep learning methods only targets single objects and does not contain background information.

[0003] Sketches can describe the appearance characteristics and movement process of objects. Combining spatial position and temporal information can effectively improve the video retrieval effect. The traditional method (reference: JP Collosse, G. McNeill, and Y. Qian, "Storyboard sketches for content based video retrieval," in 2009 IEEE 12th International Conference on Computer Vision. IEEE, 2009, pp. 245–252.) is to first divide the video frame into color regions, extract regional features: area, color, foreground objects, etc., and then represent the camera motion through a single mapping between frames, thereby representing the motion trajectory of people, and then perform video retrieval by adding sketches to sketch symbols. Peng Xu (reference: P.Xu, K.Liu, T.Xiang, T.M Hospedales, Z.Ma, J.Guo, and Y.-Z.Song, "Fine-grained image-level sketch-based video retrieval," arXiv preprint arXiv:2002.09461, 2020.) uses a deep learning method and a triplet network to extract features into appearance feature streams and action feature streams. Finally, the features are fused and the triplet loss function is used for video retrieval tasks. Scene-level SBIR (reference: F.Liu, C.Zou, X.Deng, R.Zuo, Y.-K.Lai, C.Ma, Y.-J.Liu, and H.Wang, "Scenesketcher:Fine-grained image retrieval with scene sketches," 2020.) implements scene sketch retrieval images, which cannot capture the action information in the video.

[0004] From the above, we can see that the existing sketch-based video retrieval methods cannot solve the problem of fine-grained video retrieval containing multiple objects in complex scenes. Summary of the invention

[0005] The purpose of the present invention is to propose a scene-level fine-grained video retrieval method (Scene-level Video Retrieval with Sketches) and system based on sketches. The present invention performs sketch-based video retrieval at the scene level (i.e., including multiple foreground objects and background information), compresses objects appearing in different time periods into the same sketch, performs scene-level video content summary, and retrieves videos that are consistent with the sketch background elements, object appearance features and action types.

[0006] To achieve the above object, the technical solution adopted by the present invention is as follows:

[0007] A sketch-based scene-level fine-grained video retrieval method includes the following steps:

[0008] For a drawn scene sketch, sketch features are obtained, including overall appearance features, and appearance features, category features, and position features of instances on the sketch; a sketch space structure graph is constructed based on these sketch features, and the sketch space structure graph includes a sketch appearance structure graph and a sketch category structure graph, wherein the sketch appearance structure graph is composed of an instance node representing the appearance features of the instance, a scene node representing the overall appearance features of the sketch, and an edge representing the distance calculated based on the position features, and the sketch category structure graph is composed of an instance node representing the type features of the instance and an edge representing the distance calculated based on the position features;

[0009] According to the scene sketch, an adaptive frame sampling strategy is used to sample the video. That is, the video frames are first sparsely sampled to obtain candidate video frames, and then the candidate video frames are screened using the trained sketch-video association model to screen out the video frames most relevant to the scene sketch and encode them into videos.

[0010] For the above-mentioned encoded video, video features and timing information are obtained, wherein the video features include overall appearance features in the video image, appearance features of each instance, type features and position features; a video spatiotemporal structure graph is constructed according to these video features and timing information, wherein the video spatiotemporal structure graph includes a video space structure graph and a video timing structure graph, wherein the video space structure graph includes a video appearance structure graph and a video category structure graph, wherein the video appearance structure graph is composed of an instance node representing the appearance features of the instance, a scene node representing the overall appearance features of the image and an edge representing the distance calculated according to the position features, wherein the video category structure graph is composed of an instance node representing the type features of the instance and an edge representing the distance calculated according to the position features; wherein the video timing structure graph is constructed according to the timing information, the instance nodes and the scene nodes;

[0011] The sketch features and video features are input into the trained sketch-video retrieval model for video retrieval. The sketch-video retrieval model includes an appearance branch and a category branch. The appearance branch generates video retrieval results based on the sketch appearance structure graph and the video appearance structure graph. The category branch generates video retrieval results based on the sketch category structure graph and the video category structure graph. The two retrieval results are fused with appearance features and category features to obtain the final video retrieval results.

[0012] Furthermore, for the drawn scene sketch, the pre-trained GoogLeNet Inception-V3 is used to extract the appearance features of each instance in the sketch, the Bert model is used to encode the category features of each instance, the relative position processing method mentioned in Transformer is adopted and the sine and cosine functions are used to obtain the position features, and Distance-IOU is used to calculate the distance between instances based on the position features.

[0013] Furthermore, a two-layer GCN network is used to update the sketch features. The GCN network performs feature fusion on local instance nodes by adding an SE module.

[0014] Furthermore, for the above-encoded video, ResNet-152 is used to extract the appearance features of each instance in the video frame, the Bert model is used to encode the category features of each instance, the relative position processing method mentioned in Transformer is used and the sine and cosine functions are used to obtain the position features, and Distance-IOU is used to calculate the distance between instances based on the position features.

[0015] Furthermore, a two-layer GCN network combined with the SE module is used to update the features of the video spatial structure graph and the video temporal structure graph respectively.

[0016] Furthermore, the sketch-video association relationship model is constructed based on the triplet network, and its training method is as follows: utilizing the matching relationship between sketches and video frames in the training set, and forming triplet matching pairs consisting of sketches, video frame positive samples and video frame negative samples to train the sketch-video association relationship model, and learning the semantic and visual association relationship between sketches and video images through training.

[0017] Furthermore, the sketch-video retrieval model is constructed based on a triplet network, and its training method is as follows: using the sketch features and video features to be retrieved in the training set, constructing a triplet matching pair consisting of sketch features, video positive sample features and video negative sample features, training the sketch-video retrieval model, calculating the final loss function, and completing the training by adjusting the model parameters to minimize the loss.

[0018] Furthermore, the final loss function is obtained by averaging the loss functions of multiple batches, and the loss function of each batch is the difference between the distance between the sketch and the positive sample of the video and the distance between the sketch and the negative sample, plus the interval between the positive and negative samples themselves, and is maximized.

[0019] Furthermore, the retrieval results of the two branches of the sketch-video retrieval model are fused with appearance features and category features, which means fusing the Euclidean distances between the sketch and the video obtained by the category features and the appearance features respectively.

[0020] A sketch-based scene-level fine-grained video retrieval system, comprising:

[0021] The interactive interface for retrieving videos from sketches includes a user input interface and a video display interface. The user input interface is used to provide scene sketch drawing tools and a panel for drawing scene sketches. The video display interface is used to display the retrieved videos.

[0022] A sketch feature acquisition module is used to acquire the overall appearance features of the sketch, as well as the appearance features, category features, and position features of instances on the sketch;

[0023] The sketch feature update module uses a two-layer GCN network combined with an SE module to update the sketch features and perform feature fusion on local instance nodes;

[0024] The video feature acquisition module is used to acquire the overall appearance features of the video image, as well as the appearance features, category features and position features of the instances on the image, as well as the time sequence information;

[0025] The video feature update module is used to use a two-layer GCN network combined with an SE module to update the features of the video spatial structure graph and the temporal structure graph respectively;

[0026] The adaptive frame sampling module is used to sample video frames according to the scene sketch using an adaptive frame sampling strategy, that is, firstly obtain candidate video frames through sparse sampling, then screen the candidate video frames through a sketch-video association model built based on a triplet network and trained to select the video frames most relevant to the scene sketch and encode them into videos;

[0027] The sketch-video retrieval model is built based on a triplet network and completed through training. It includes an appearance branch and a category branch, and is used to perform fine-grained video retrieval based on the input sketch features and video features. The appearance branch generates video retrieval results based on the sketch appearance structure graph and the video appearance structure graph, and the category branch generates video retrieval results based on the sketch category structure graph and the video category structure graph. The two retrieval results are fused by the appearance features and the category features to obtain the final video retrieval results.

[0028] The present invention supports a single scene sketch input mode. After inputting a sketch of the corresponding form, a video clip that is consistent with the object shape, category, action and background elements in the sketch is retrieved. The present invention supports fine-grained video retrieval based on scene sketches, and supports two types of fine-grained changes in scenes: one is the change of foreground objects, and the other is the change of background elements. Fine-grained video retrieval based on scene sketches is achieved by combining a graph convolutional neural network with a triplet network. Among them, the sketch and video are updated by constructing a graph network structure, and the cross-modal features are mapped to the same common subspace. Then, a triplet network is constructed to calculate the feature distance between the sketch and the video to realize the retrieval function.

[0029] Compared with the prior art, the present invention has the following beneficial effects:

[0030] 1. Compared with the single-object, background-free SBVR, the scene-level video retrieval of the method of the present invention is more in line with common video types and expands the practical application scenarios of sketch retrieval.

[0031] 2. The present invention proposes an adaptive video frame sampling strategy. By pre-learning the scene-level sketch-image correspondence, multiple video frames that are most similar to the query-side sketch can be efficiently selected from the redundant video content, and their features are updated to represent the overall video content. This strategy achieves content alignment between the video and the sketch, and can further improve the efficiency of video feature extraction and reduce the amount of model calculation.

[0032] 3. In order to realize fine-grained video retrieval based on scene sketches, the present invention constructs a branch model to update the features of sketches and videos from the category branch and the appearance branch respectively, and learns the matching relationship between sketches and videos from the semantic and visual levels. Each branch contains a spatial structure graph of the scene sketch and a spatiotemporal structure graph of the video, and uses a graph convolutional neural network and a spatiotemporal graph convolutional neural network for feature update respectively. The two branches are trained separately, and the retrieval results of the two branches are combined in the retrieval stage. This method effectively realizes scene-level fine-grained retrieval of videos and broadens the application scope of sketch retrieval in the field of video. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 is a diagram of a sketch-based scene-level fine-grained video retrieval network structure according to an embodiment of the present invention;

[0034] Figure 2 is an adaptive video frame sampling strategy diagram of an embodiment of the present invention;

[0035] Figure 3 is a diagram of a scene-level fine-grained video retrieval result based on a sketch according to an embodiment of the present invention;

[0036] Figure 4.Interactive interface diagram of sketch retrieval video in an embodiment of the present invention (including user input interface and system display interface). DETAILED DESCRIPTION

[0037] In order to enable those skilled in the art to better understand the present invention, the technical solution of the present invention is further described in detail below in conjunction with the accompanying drawings, but does not constitute a limitation to the present invention.

[0038] This embodiment proposes a scene-level fine-grained video retrieval method based on sketches, and proposes to establish an adaptive frame sampling strategy to achieve efficient alignment between sketches and video content; proposes a branch model to establish a graph structure for sketches and videos from the category and appearance levels; based on the graph convolution network (GCN) and the spatial-temporal graph convolution network (ST-GCN, i.e. GCN updates features in space and time), the graph features of sketches and videos are updated respectively; and a triple network model for matching sketch and video features is established. The main content modules are divided into three modules: sketch feature encoding, video feature encoding, and sketch and video feature matching. Figure 1 The sketch-based scene-level fine-grained video retrieval framework is shown.

[0039] 1. Sketch feature encoding

[0040] 1) Sketch space structure diagram construction

[0041] For the scene sketch S, create a sketch space structure diagram: Sketch category structure diagram G s,c And sketch appearance structure diagram G s,a , each graph structure contains n instance nodes Represents instance-level information. Each object with motion features in the sketch represents an instance. The distance between instances is calculated using the Distance-IOU method as the connecting edge of the instance. In addition, the sketch appearance structure diagram also contains an independent scene node The sketch as a whole is treated as a node, not connected to other instance nodes, to represent the overall appearance characteristics.

[0042] The feature of each node in the sketch category structure diagram is a combination of category feature and position feature, and the feature of each node in the sketch appearance structure diagram is a combination of appearance feature and position feature. The pre-trained GoogLeNetInception-V3 is used to extract the appearance feature of each instance in the sketch, with a feature dimension of 2048 dimensions; the Bert model is used to encode the category feature of each instance, with a feature dimension of 768 dimensions; the relative position processing method mentioned in Transformer is used to obtain the absolute position feature using sine and cosine functions.

[0043] 2) Sketch feature update

[0044] After establishing the sketch space structure graph, a two-layer GCN is used to update the features. s,c And sketch appearance structure diagram G s,a Features in The update formulas are:

[0045] F s l+1 =ReLu(AW l F s l )

[0046] Where A is the adjacency matrix, W l is the training parameter of the lth layer, and ReLU is the activation function used.

[0047] Among them, after the feature update through GCN, the SE module (reference: J.Hu, L.Shen, and G.Sun, "Squeeze-and-excitation networks," in Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp.7132–7141.) is added as an attention mechanism to perform feature fusion on local instance nodes:

[0048]

[0049] Where σ is the Sigmoid activation function, W se is the weight vector of the SE module, F s I is the characteristic of the instance node.

[0050] Among them, for the sketch appearance structure diagram, after obtaining the local instance node features After that, the characteristics of the scene node Perform feature splicing to obtain the final sketch appearance feature Fs,a :

[0051]

[0052] 2. Video feature encoding

[0053] 1) Adaptive frame sampling strategy

[0054] In the process of video feature encoding, the biggest problem is that due to the temporal redundancy of the video, its key content is hidden in the video. In order to reduce the huge computational loss caused by processing the video frame by frame and align the content between the sketch and the video, an adaptive video frame sampling strategy is proposed to sample multiple video frames that are most similar to the input sketch. Figure 2 A flowchart showing the adaptive frame sampling strategy.

[0055] By utilizing the matching relationship between sketches and video frames in the dataset, that is, the sketches, positive video frame samples and negative video frame samples form triplet matching pairs, a sketch-video association relationship model based on the triplet network is constructed, and the semantic and visual association relationship between sketches and video images is learned through training.

[0056] During the retrieval process, the video is first sparsely sampled as candidate video frames, and then the sketch-image association model is applied to further screen the k video frames most relevant to the input sketch to represent the video for video encoding. This method quickly screens out the video with the highest relevance to the sketch.

[0057] 2) Construction of video spatiotemporal structure graph

[0058] Given the k selected video frames, a video spatiotemporal structure graph is established to extract the features of the video frames from the perspective of space and time. The video spatiotemporal structure graph includes a video spatial structure graph and a video time sequence structure graph. The spatial structure graph of each video frame in the video spatial structure graph is the same as the grass Figure 1 A consistent spatial structure diagram.

[0059] The appearance features of each instance in the video frame are extracted by ResNet-152. The extraction method of category features and position features is consistent with the sketch. The connection edges of the instances are obtained by calculating the distance between instances using Distance-IOU.

[0060] The video spatiotemporal structure graph contains two branches, namely the video appearance structure graph and the video category structure graph. For the video appearance structure graph, after the node features of all instances in the video spatial structure graph pass through the SE module, they are output as an instance node in the temporal structure graph. Furthermore, two temporal structure graphs are established. and It contains the scene nodes and instance nodes of each video frame after the spatial structure graph features are updated.

[0061] For the video category structure diagram, since the diagram only contains instance nodes and does not contain global scene nodes, only a timing structure diagram containing timing instance nodes is established.

[0062] 3) Video feature update

[0063] The same method as the sketch feature update is used to update the video spatial structure graph, and the two-layer GCN combined with the SE module is used again to update the video temporal structure graph. Finally, the video features F are obtained from the category level and appearance level. v,c and F v,a .

[0064] 3. Sketch-Video Feature Matching

[0065] After obtaining the sketch feature F s and video feature F v After that, construct the sketch features to be retrieved Video positive sample features And video negative sample features The three-tuple matching pairs are mapped to the same common subspace and input into a sketch-video retrieval model based on a triplet network, where the positive video sample is the correct video in the dataset that is consistent with the sketch content description, and the negative video sample is obtained by traversing all videos and calculating the video with the farthest distance from the positive sample, that is:

[0066]

[0067] Where argmin means that the objective function The variable value when the minimum value is taken, i represents the sequence number of the positive video sample, and j represents the sequence number of the negative video sample.

[0068] The sketch-video retrieval model is trained. Specifically, the sketch-video retrieval model contains two branches, namely the appearance branch and the category branch. The appearance branch is used to retrieve the video appearance structure diagram in the video spatiotemporal structure diagram according to the sketch appearance structure diagram in the sketch spatial structure diagram. The category branch is used to retrieve the video category structure diagram in the video spatiotemporal structure diagram according to the sketch category structure diagram in the sketch spatial structure diagram. By training these two branches separately, fine-grained video retrieval of a single sketch is achieved. The specific loss function is defined as maximizing the difference between the distance between the sketch and the video positive sample and the distance between the sketch and the negative sample, plus the interval between the positive and negative samples themselves:

[0069]

[0070] Where d is the Euclidean distance, Δ is the interval between positive and negative samples, and 0 represents The value is positive and the minimum is 0.

[0071] The results of multiple batches are averaged to obtain the final loss function:

[0072]

[0073] b is the batch size, and each batch contains a batch of sketches and corresponding video matching pairs.

[0074] By adjusting the model parameters to minimize the above loss, a trained sketch-video retrieval model is obtained.

[0075] In the normal use phase (including the testing phase), the trained sketch-video association model and sketch-video retrieval model are used to perform video retrieval. The specific steps are as follows:

[0076] 1. The user enters the sketch retrieval video system and forms his own scene sketch on the canvas. The system provides two ways to generate scene sketches: sketch drawing and sketch material splicing. In this example, the user uses a brush to draw the background elements of the current scene on the canvas, and further specifies the category of the foreground object in the material bar. By sliding the material bar, drag the desired object to the canvas to input the scene sketch, and use the zoom and move operations to adjust the object's direction and proportion to form the input scene sketch (such as Figure 4 left column).

[0077] 2. Use the pre-trained GoogLeNet Inception-V3 to extract the appearance features of each instance in the sketch, use the Bert model to encode the category features of each instance, use the relative position processing method mentioned in Transformer to use sine and cosine functions to obtain absolute position features, and use the Distance-IOU method to calculate the distance between instances as the connection edge of the instance. The appearance features and position features form the sketch appearance structure diagram, and the category features and position features form the sketch category structure diagram, and the sketch space structure diagram is established. Then use a two-layer GCN network to update the above features.

[0078] 3. The videos to be retrieved in the database are sampled using an adaptive frame sampling strategy. First, sparse sampling is performed to obtain candidate video frames. Then, the candidate video frames are further screened using the trained sketch-video association model to obtain the video with the highest relevance to the sketch.

[0079] 4. For the screened video frames, construct a video spatiotemporal structure diagram, which includes a video spatial structure diagram and a video temporal structure diagram. ResNet-152 is used to extract the appearance features of the instances in the video frames, and the Bert model is used to encode the category features of each instance. The relative position processing method mentioned in Transformer is used to obtain the absolute position features using sine and cosine functions, and the Distance-IOU method is used to calculate the distance between instances as the connecting edges of the instances. The video appearance structure diagram is composed of appearance features and position features, and the video category structure diagram is composed of category features and position features, to obtain the video spatial structure diagram. The above features are then updated using a two-layer GCN network.

[0080] The trained sketch-video retrieval model is used to retrieve the video, and the results are output through the appearance branch and the type branch respectively. The Euclidean distance D between the sketch and the video is obtained for the category features and appearance features in the two results. c and D a The fusion formula is as follows:

[0081] D=αD c +(1-α)D a

[0082] The specific value of α is determined during the experiment.

[0083] After fusing the sketch-video distances obtained from the category level and the appearance feature level, the videos to be retrieved are sorted to obtain the final sketch retrieval video results. The sketch retrieval video results are shown in the figure below: Figure 3 As shown, the effectiveness of the retrieval method is verified.

[0084] 5. The model outputs the top 10 videos with the highest similarity to the input sketch and displays them in the search results column. Users can slide in the search results column, click to enlarge and play the retrieved videos, such as Figure 4 Shown in the right column.

[0085] The above is a detailed description of the scene-level fine-grained video retrieval method and device based on sketches of the present invention, but it is obvious that the specific implementation form of the present invention is not limited thereto. For those skilled in the art, various obvious changes to the method of the present invention without departing from the spirit and scope of the claims are within the scope of protection of the present invention.

Claims

1. A sketch-based scene-level fine-grained video retrieval method, characterized in that: The following steps are involved: For the drawn scene sketch, the sketch features are obtained, including the overall appearance features, and the appearance features, category features, and position features of the instances on the sketch; A sketch space structure graph is constructed according to the sketch features, the sketch space structure graph including a sketch appearance structure graph and a sketch category structure graph, the sketch appearance structure graph is composed of an instance node representing the appearance feature of the instance, a scene node representing the overall appearance feature of the sketch, and an edge representing the distance calculated according to the position feature, and the sketch category structure graph is composed of an instance node representing the type feature of the instance and an edge representing the distance calculated according to the position feature; According to the scene sketch, an adaptive frame sampling strategy is used to sample the video. That is, the video frames are first sparsely sampled to obtain candidate video frames, and then the candidate video frames are screened using the trained sketch-video association model to screen out the video frames most relevant to the scene sketch and encode them into videos. For the encoded video, video features and timing information are obtained, wherein the video features include overall appearance features, appearance features of each instance, type features, and position features in the video image; A video spatiotemporal structure graph is constructed according to these video features and timing information, wherein the video spatiotemporal structure graph includes a video space structure graph and a video timing structure graph, wherein the video space structure graph includes a video appearance structure graph and a video category structure graph, wherein the video appearance structure graph is composed of an instance node representing the appearance feature of an instance, a scene node representing the overall appearance feature of an image, and an edge representing a distance calculated according to a position feature, and the video category structure graph is composed of an instance node representing the type feature of an instance and an edge representing a distance calculated according to a position feature; and the video timing structure graph is constructed according to the timing information, the instance node, and the scene node; The sketch features and video features are input into the trained sketch-video retrieval model for video retrieval. The sketch-video retrieval model includes an appearance branch and a category branch. The appearance branch generates video retrieval results based on the sketch appearance structure graph and the video appearance structure graph. The category branch generates video retrieval results based on the sketch category structure graph and the video category structure graph. The two retrieval results are fused with appearance features and category features to obtain the final video retrieval results.

2. The method according to claim 1, characterized in that For the drawn scene sketch, the pre-trained GoogLeNet Inception-V3 is used to extract the appearance features of each instance in the sketch, the Bert model is used to encode the category features of each instance, the relative position processing method mentioned in Transformer is used and the sine and cosine functions are used to obtain the position features, and Distance-IOU is used to calculate the distance between instances based on the position features.

3. The method according to claim 1, characterized in that A two-layer GCN network is used to update sketch features. The GCN network fuses features of local instance nodes by adding SE modules.

4. The method according to claim 1, characterized in that For the above-encoded video, ResNet-152 is used to extract the appearance features of each instance in the video frame, the Bert model is used to encode the category features of each instance, the relative position processing method mentioned in Transformer is used, and the sine and cosine functions are used to obtain the position features, and Distance-IOU is used to calculate the distance between instances based on the position features.

5. The method according to claim 1, characterized in that A two-layer GCN network is used in combination with the SE module to update the features of the video spatial structure graph and the video temporal structure graph respectively.

6. The method according to claim 1, characterized in that The sketch-video association model is built based on a triplet network. Its training method is as follows: using the matching relationship between sketches and video frames in the training set, a triplet matching pair consisting of sketches, video frame positive samples and video frame negative samples is used to train the sketch-video association model. The semantic and visual association relationship between sketches and video images is learned through training.

7. The method according to claim 1, characterized in that The sketch-video retrieval model is built based on a triplet network. Its training method is as follows: using the sketch features and video features to be retrieved in the training set, constructing a triplet matching pair consisting of sketch features, video positive sample features and video negative sample features, training the sketch-video retrieval model, calculating the final loss function, and completing the training by adjusting the model parameters to minimize the loss.

8. The method according to claim 7, characterized in that The final loss function is obtained by averaging the loss functions of multiple batches. The loss function of each batch is the difference between the distance between the sketch and the positive sample of the video and the distance between the sketch and the negative sample, plus the interval between the positive and negative samples themselves, and is maximized.

9. The method according to claim 1, characterized in that The retrieval results of the two branches of the sketch-video retrieval model are fused with appearance features and category features, which means fusing the Euclidean distances between the sketch and the video obtained by the category features and the appearance features respectively.

10. A sketch-based scene-level fine-grained video retrieval system, characterized in that: include: The interactive interface for retrieving videos from sketches includes a user input interface and a video display interface. The user input interface is used to provide scene sketch drawing tools and a panel for drawing scene sketches. The video display interface is used to display the retrieved videos. The sketch feature acquisition module is used to obtain the overall appearance features of the sketch, as well as the appearance features, category features and position features of the instances on the sketch; Sketch feature update module uses a two-layer GCN network combined with an SE module to update sketch features and perform feature fusion on local instance nodes; The video feature acquisition module is used to acquire the overall appearance features of the video image, as well as the appearance features, category features, position features, and timing information of the instances on the image; The video feature update module is used to use a two-layer GCN network combined with an SE module to update the features of the video spatial structure graph and the video temporal structure graph respectively; The adaptive frame sampling module is used to sample video frames according to the scene sketch using an adaptive frame sampling strategy, that is, firstly obtain candidate video frames through sparse sampling, then screen the candidate video frames through a sketch-video association model built based on a triplet network and trained to select the video frames most relevant to the scene sketch and encode them into videos; The sketch-video retrieval model is built based on a triplet network and completed through training. It includes an appearance branch and a category branch, and is used to perform fine-grained video retrieval based on the input sketch features and video features. The appearance branch generates video retrieval results based on the sketch appearance structure graph and the video appearance structure graph, and the category branch generates video retrieval results based on the sketch category structure graph and the video category structure graph. The two retrieval results are fused by the appearance features and the category features to obtain the final video retrieval results.