A Self-Supervised Video Scene Boundary Detection Method Based on Scene Montage
Through self-supervised learning and visual dual encoder to generate pseudo-boundary boundaries from unlabeled videos, the problem of high data labeling and upper pseudo-boundary hit rate in video scene boundary detection is solved, and efficient video scene boundary detection is achieved.
Patent Information
- Application Number
- CN202311107013.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-30
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2043-08-30
AI Technical Summary
In the prior art, in video scene boundary detection, model performance is impaired due to the high cost of data annotation and the upper hit rate limit of the pseudo-boundary generated by the split-based method.
A self-supervised video scene boundary detection method based on scene montage is used to generate pseudo-boundaries from labeled long videos through self-supervised learning, lens-level features are extracted using visual dual encoder, and contextual relationship modeling and scene boundary judgment modules are combined, and pseudo-boundary prediction is used as agent task to train neural networks.
It improves the accuracy of video scene boundary detection and the robustness of the model, reduces the dependence on labeled data, generates high-quality pseudo-boundaries for training, and significantly improves the performance of the detection model.
Smart Images

Figure CN117058593B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of deep learning in artificial intelligence, and particularly to a self-supervised video scene boundary detection method based on scene montage. Background Art
[0002] Video scene boundary detection refers to dividing the shot sequence in a video into semantically coherent story segments according to the content described by the shots.
[0003] Currently, the method of training a model to segment videos based on manually annotated data is severely limited by the high cost of data annotation. The methods that obtain pseudo-scene boundaries by splitting video segments into two pseudo-scenes are collectively referred to as splitting-based methods. Although the splitting-based methods have achieved certain results in the video scene boundary detection task, the generated pseudo-boundaries have an upper limit on the hit rate. The method of obtaining pseudo-boundaries by splitting video segments is very likely to over-segment the scene, resulting in a single scene being divided into two pseudo-scenes, disrupting the inter-shot relationship in the video segment, and thus damaging the performance of the video scene boundary detection model. Summary of the Invention
[0004] Aiming at the above problems existing in the prior art, the technical problem to be solved by the present invention is: how to accurately detect video scene boundaries.
[0005] To solve the above technical problem, the present invention adopts the following technical solution: a self-supervised video scene boundary detection method based on scene montage, including the following steps:
[0006] S1: Extraction of shot sequence: Obtain an input long video S video =[s1,…s N , obtain the shots therein, and sample key frames from the shots.
[0007] S2: Construct and train a video scene montage network model VSM.
[0008] S21: For each shot s video in S i encode it into a feature vector x i using a visual editor, and S video =[s1,…s N is encoded into a sequence X video =[x1,…,x b ,…,x N of feature vectors, where 1≤γ≤n - 1, 1≤α≤N - γ + 1 and β1≤β≤N - γ + 1, N represents the number of shots in S video and n represents the number of shots in the video segment generated by VSM.
[0009] S22: Determine the lengths of two feature subsequences and their starting positions through random parameters γ, α, β, and obtain the synthetic feature sequence X through splicing operations. syn , corresponding to the synthetic shot sequence S syn .
[0010] S23: Generate multiple video segments as training data by the method of S22, and use pseudo-boundary prediction as a proxy task to train the context encoder and the scene boundary judgment module.
[0011] S24: Use the real-annotated scene boundary information to fine-tune the pre-trained VSM obtained in S25 to obtain the final VSM.
[0012] S3: Detection. Pass the video segment to be detected through step S1 to obtain the shot sequence S', then obtain the corresponding feature sequence X' through step S21, input X' into the final VSM, and output the confidence that the middle shot of the sequence is the scene boundary.
[0013] Specifically, for each shot s in S in S21 video is encoded into a d-dimensional feature vector x i by the following specific steps: i :
[0014] x i = concat(FGE(s i ), BGE(s i ))
[0015] where concat(·) represents the splicing of feature vectors, FGE(·) and BGE(·) respectively represent the foreground encoder and the background encoder, s i represents the i-th shot, and x i represents the feature of the i-th shot.
[0016] Preferably, the specific steps for synthesizing the feature sequence X syn and S syn are as follows:
[0017] First, select a random positive integer γ such that one of the two video segments to be intercepted contains γ shots and the other contains n - γ shots. Then, select two random positive integers α and β as the positions of the starting shots of the two video segments to be intercepted in the long video S video . Thus, two video segments can be intercepted from S video as two pseudo-scenes, denoted as S left = [s α , s α+1 …, s α+γ+1 and S right = [s β , sβ+1 …, s β+n-γ-1 , there are n shots in two pseudo-scenes. Concatenate S left and S right together in the time dimension to form a synthesized video segment R syn = [s α , … s α+γ+1 , s β , …, s β+n-γ-1 :
[0018] S syn = splice(S left , S right )
[0019] where splice(·) represents concatenation in the time dimension. α, β, and γ are all unfixed random numbers, and these random numbers are reselected each time a video segment is synthesized.
[0020] When intercepting the corresponding part of S from X
[0021] = [x1, …, x video , we can get X N = [x syn , … x syn , x α , …, x α+γ+1 , x β , …, x β+n-γ-1 .
[0022] Preferably, the specific steps for training the context relationship modeling module and the scene boundary judgment module in S23 are as follows:
[0023] Concatenate the feature vector sequence P = [p1, …, p n to X syn to supplement the position information of each shot, and finally feed it into the context relationship modeling module to obtain the feature vector sequence R syn = [r α , … r α+γ+1 , r β , …, r β+n-γ-1 :
[0024] R syn = Context(concat(P, X syn ))
[0025] where Context(·) is the context relationship modeling module and concat(·) is the vector concatenation operation.
[0026] For the video segment S syn , the pseudo-scene S rightThe starting shot s β is regarded as a positive sample shot, and then a shot s syn is randomly selected from S i as a negative sample shot. Finally, the feature vectors {r β , r i} corresponding to the positive and negative samples are input into the scene boundary judgment module, and the context encoder is pre-trained by minimizing the binary cross-entropy loss:
[0027]
[0028] where h p (·) is the scene boundary judgment module, and the regularization term is used to combat overfitting, where λ represents the coefficient.
[0029] Preferably, it is fine-tuned using the training data set with scene boundary annotations in ImageNet. If the central shot s c in the data set is a shot at the scene boundary, the label y c = 1, otherwise y c = 0, and the loss is calculated by the following formula:
[0030] L sbd = -y c log(h sbd (r c )) + (1 - y c ) log(1 - h sbd (r c ))
[0031] where L sbd represents the loss in the fine-tuning stage, y c represents the label of the central shot s c in the shot sequence during fine-tuning, r c represents the context relationship feature of the central shot s c , and h sbd (·) represents the scene boundary judgment module.
[0032] Compared with the prior art, the present invention has at least the following advantages: BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 represents the data flow of the dual-shot feature encoder.
[0034] Figure 2 represents the schematic diagram of pseudo-boundary synthesis.
[0035] Figure 3 represents the working process of the visual dual encoder.
[0036] Figure 4 Shows the workflow of the VSM method.
[0037] Figure 5 Shows the process of the pre-trained context editor.
[0038] Figure 6 Shows the process of fine-tuning the context editor.
[0039] Figure 7 Shows the diagrams of the SC relationship, SI relationship, WC relationship, and WI relationship between shots. Detailed implementation mode
[0040] The present invention will be further described in detail below.
[0041] The method of the present invention uses a visual encoder pre-trained on a large-scale image dataset to extract shot-level features. The semantics within a video shot can be roughly divided into foreground and background. The foreground includes characters and various props, and ResNet50 pre-trained on ImageNet is a commonly used object detection model suitable for extracting foreground features, so this model is used as the foreground encoder; it often contains information about locations or environments, etc., and ResNet50 pre-trained on Places365 is good at extracting information such as environments and locations and is suitable for extracting background features, so this model is used as the background encoder. Decoupling the visual information of the shots makes the extracted semantics diverse, and inheriting the visual knowledge in the large dataset makes the extracted semantics richer, thus laying a solid foundation for modeling the relationships between shots.
[0042] The present invention method proposes a self-supervised video scene boundary detection method VSMBD (Video Scene Montage or Boundary Detection) based on scene montage, which decouples visual information to obtain rich semantics and uses reliable pseudo-boundaries to learn the relationships between shots. This method takes into account both the diversity of visual semantics in the shots and the complexity of the relationships between shots, and thus achieves excellent performance on public datasets.
[0043] This method aims to model the relationships between shots, explores a reliable pseudo-label generation method, and uses pseudo-boundary prediction as a proxy task to pre-train a neural network model. Modeling the relationships between shots requires the premise of extracting rich visual semantics. Therefore, this method inherits the knowledge of large-scale image datasets and decouples the visual semantics of the shots.
[0044] The present invention classifies the relationships between shots into the following four types: 1. Weak Inconsistency (WI) between shot semantics across scenes; 2. Strong Inconsistency (SI) between shot semantics across scenes; 3. Weak Consistency (WC) between shot semantics from the same scene; 4. Strong Consistency (SC) between shot semantics from the same scene. These four relationships are defined for both the relationships between shots within a scene and between scenes. The Video Scene Montage (VSM) method proposed by the present invention is used to generate video segments with artificial boundaries. As Figure 4 shown, the VSM method randomly selects two video segments in a long video as two artificial scenes, and by splicing the two artificial scenes sequentially, a video segment containing an artificial scene boundary can be obtained. The feasibility of the VSM method relies on the simulation of the relationships between shots. Multiple video segments in the same long video have a probability of sharing some visual elements, so they may have a WC relationship; since different video segments often describe different stories, they generally have an SI relationship; for randomly extracted video segments, there is almost always an SC relationship within them, because the vast majority of scenes contain locations composed of multiple shots; when the length of all video segments is n = 13 (including 13 shots), nearly half of the video segments have scene boundaries, and when two video segments are spliced together, the probability of having a scene boundary in the longer video segment formed is even greater. Therefore, there is generally a WI relationship in the spliced video segment. Based on the above analysis, the VSM method cleverly completes the simulation of the four relationships between shots through simple operations.
[0045] A self-supervised video scene boundary detection method based on scene montage, comprising the following steps:
[0046] S1: Extraction of the shot sequence: Obtain an input long video S video =[s1,…s N , obtain the shots therein, and sample key frames from the shots.
[0047] S2: Construct and train the video scene montage network model VSM.
[0048] First, perform shot segmentation and frame sampling on the long video, then use a pre-trained visual encoder to extract frame features, and finally apply a max pooling operation to the frame features of each shot so that each shot can be represented by a feature vector. The difference is that in the present invention, a foreground encoder and a background encoder are connected in parallel as a visual dual encoder to extract shot-level visual features. First, use these two visual encoders to extract 2048-dimensional frame-level visual feature vectors respectively, then perform max pooling on the frame-level feature vectors within each shot, and finally concatenate them into a 4096-dimensional shot-level feature vector. Figure 3 Shows the working process of the visual dual encoder for a single shot.
[0049] S21: For each shot s video in S i encode it into a -dimensional feature vector x i , S video = [s1,…s N is encoded into a sequence X video = [x1,…,x b ,…,x N of feature vectors, where 1 ≤ γ ≤ n - 1, 1 ≤ α ≤ N - γ + 1 and β1 ≤ β ≤ N - γ + 1, and N represents the number of shots in S video , and n represents the number of shots in the video segments generated by the VSM.
[0050] S22: Determine the lengths of two feature subsequences and their starting positions through random parameters γ, α, β, and the synthesized feature sequence X syn obtained through the concatenation operation corresponds to the synthesized shot sequence S syn .
[0051] S23: Generate multiple video segments as training data through the method of S22, and use pseudo-boundary prediction as a proxy task to train a context encoder [the context encoder can be constructed based on a two-layer Transformer encoder] and a scene boundary judgment module [a multi-layer perceptron (MLP) can be used].
[0052] S24: Use the real annotated scene boundary information to fine-tune the pre-trained VSM obtained in S25 to obtain the final VSM.
[0053] S3: Detection. Pass the video segment to be detected through the shot sequence S' obtained in step S1, then obtain the corresponding feature sequence X' through step S21, and input X' into the final VSM to output the confidence that the middle shot of the sequence is the scene boundary.
[0054] Specifically, in S21, each shot s video in S i is encoded into a -dimensional feature vector xi The specific steps are as follows:
[0055] x i = concat(FGE(s i ), BHE(s i ))
[0056] where concat(·) represents the concatenation of feature vectors, and FGE(·) and BGE(·) represent the foreground encoder and the background encoder respectively, and s i represents the i-th shot, and x i represents the feature of the i-th shot.
[0057] The two pre-trained models can include a foreground encoder FFE to extract foreground features and a background encoder FBE to extract background features. The FFE can use the ResNet50 model pre-trained on the ImageNet dataset, and the FBE can use the ResNet50 model pre-trained on the Places365 dataset.
[0058] Specifically, the specific steps for synthesizing the feature sequences X syn and S syn are as follows:
[0059] Given an unannotated long video S video = [s1, …, s N , VSM first determines the length n of the video segment to be generated, and then intercepts two video segments from S video and splices them to form a video segment containing n shots, with the splicing point as the pseudo-scene boundary. Specifically, to generate a video segment of length n, first select a random positive integer γ such that one of the two video segments to be intercepted contains γ shots and the other contains n - γ shots. Then, select two random positive integers α and β as the positions of the starting shots of the two video segments to be intercepted in the long video S video . Thus, two video segments can be intercepted from S video as two pseudo-scenes, denoted as S left = [s α , s α+1 …, s α+γ+1 and S right = [s β , s β+1 …, s β+n-γ-1 . The two pseudo-scenes have a total of n shots. S left and S right are spliced together in the time dimension to form a synthesized video segment S syn = [s α , …, s α+γ+1 , s β,…,s β+n-γ-1 :
[0060] S syn =splice(S left ,S right )
[0061] where splice(·) means splicing in the time dimension. α, β, and γ are all unfixed random numbers, and these random numbers are reselected each time a video segment is synthesized. The (γ + 1)-th shot s
[0062] in the video segment S syn is the starting shot of the pseudo-scene S β , that is, the boundary of the pseudo-scene S right is equivalent to the left time boundary of the shot s syn . β
[0063] When intercepting the corresponding part from X video =[x1,…,x N , X syn =[x syn ,…x α ,x α+γ+1 ,…,x β ,…,x β+n-γ-1 can be obtained.
[0064] The present invention controls the number of video segments generated by the VSM by controlling the random variable α. For a given long video S video , the variable α is set to a value starting from 1 and increasing. Each time the VSM generates a new video segment, the variable α is incremented by 1 until the shots in S video are traversed. The specific execution process of the VSM is as follows:
[0065]
[0066] where random(a, b) means randomly selecting an integer within the range [a, b]. The return value Chunks is a set of video segments with pseudo-boundaries generated, and each element is a video segment containing n shots. By controlling the variable α, the number of generated video segments can be controlled. In this way, each shot is applied to the pre-training process, thus maximizing the diversity of the shot semantics accessed by the context editor.
[0067] Specifically, the specific steps for training the context relationship modeling module and the scene boundary judgment module in S23 are as follows:
[0068] Splice the feature vector sequence P = [p1,…,p n to X syn To supplement the position information of each shot, and finally feed it into the context relationship modeling module to obtain a sequence of feature vectors R containing context information syn =[r α ,…r α+γ+1 ,r β ,…,r β+n-γ-1 :
[0069] R syn =XContext(concat(P,X syn ))
[0070] where Context(·) is the context relationship modeling module, concat(·) is the vector concatenation operation. To enable the context encoder to learn the relationships between shots, the pseudo-boundary prediction task is used as a proxy task. The pseudo-boundary prediction task is essentially a binary classification task for shots, that is, to determine whether the left temporal boundary of a shot is a pseudo-boundary generated by the VSM algorithm. The pre-training process of the context encoder is as Figure 5 shown.
[0071] For the video segment S syn , the starting shot s right of the pseudo-scene S β is regarded as a positive sample shot. To ensure the balance of the number of positive and negative samples, another shot s syn is randomly selected from S i as a negative sample shot. Finally, the feature vectors {r β ,r i} corresponding to the positive and negative samples are input into the scene boundary judgment module, and the context encoder is pre-trained by minimizing the binary classification cross-entropy loss:
[0072]
[0073] where h p (·) is the scene boundary judgment module, and the regularization term is used to combat overfitting, and λ represents the coefficient, which is set to 0.5 in the experiment
[0074] Although VSM can use a small amount of unlabeled long videos to generate a large number of video segments with pseudo-boundaries, there are differences in the visual semantics of shots in different long videos. Using more long videos to generate video segments can produce richer semantic relationships, thereby training a more robust context encoder. The pre-trained context encoder has learned a large number of semantic relationships between shots valuable for the video scene boundary detection task, laying a solid foundation for the subsequent fine-tuning. After completing the pre-training, the context encoder has initially mastered the ability to understand the relationships between shots. Next, the context encoder will be fine-tuned using long video data with real scene annotations to make the model adapt to the video scene boundary detection task.
[0075] Specifically, fine-tuning is performed using a training dataset with scene boundary annotations in ImageNet. If the central shot s in the dataset c is a shot at the scene boundary, the label y c = 1; otherwise, y c = 0. The loss is calculated using the following formula:
[0076] L sbd = -y c log(h sbd (r c )) + (1 - y c ) log(1 - h sbd (r c ))
[0077] where L sbd represents the loss in the fine-tuning stage, y c represents the label of the central shot s in the shot sequence during fine-tuning c , r c represents the context relationship feature of the central shot s c , and h sbd (γ) represents the scene boundary judgment module.
[0078] Example:
[0079] First, the original video is processed to obtain a sequence in units of shots. For each shot, the shot information and three key frames of each shot are directly provided in the MovieNet dataset.
[0080] After obtaining the shot information and key frame information, it is necessary to initialize the visual encoder and load the pre-trained model. The parameters of the model come from the ResNet50 model pre-trained on the ImageNet dataset and the ResNet50 model pre-trained on the Places365 dataset respectively.
[0081] First, these two visual encoders are used to extract 2048-dimensional frame-level visual feature vectors respectively. After stacking the features of the three key frames, a 6144-dimensional feature is obtained, and then a max-pooling operation is performed on it to obtain a 2048-dimensional feature. Subsequently, the features of the two encoders are concatenated into a 4096-dimensional shot-level feature.
[0082] After all the data processing is completed, it is necessary to sample and splice the feature sequences within a single video. Specifically, two random numbers γ (1 ≤ γ ≤ n - 1) can be generated to determine the lengths of the two feature sequences. Subsequently, two more random numbers α (1 ≤ α ≤ N - γ + 1) and β (1 ≤ β ≤ N - γ + 1) are generated as the starting positions of the two feature sequences, and then they are spliced to obtain a feature sequence X based on the shot sequence S syn =[s α ,…,s α+γ-1 ,s β ,…,s β+n-γ-1 , and the construction of all feature sequences is completed through multiple repeated processes. syn =[x α ,…,x α+γ-1 ,x β ,…,x β+n-γ-1 .
[0083] After the construction of the feature sequences containing pseudo-scene boundaries is completed, a context semantic encoder consisting of two layers of Transformer Encoder and two two-layer multi-layer perceptrons are constructed, serving as the scene boundary judgment model during pre-training and the scene boundary judgment model during fine-tuning respectively.
[0084] In the pre-training stage, multiple feature sequences are input into the network model. After passing through the context semantic encoder, the context features where the pseudo-scene boundaries in a sequence are at non-pseudo-scene boundary positions are selected for use in judging the scene boundaries, and the binary cross-entropy loss is used as the loss function.
[0085] In the fine-tuning stage, data of the same dimension is constructed on top of the features as the input. At this time, instead of data synthesis, sequential sampling is performed using a sliding window. For the shot at the central position of the sequence, if the true annotation data is a scene boundary, the label "1" is assigned, otherwise the label "0" is assigned. In this stage, only the prediction result at this position is used to calculate the loss through binary cross-entropy and then update the model parameters.
[0086] By comparing with existing methods on the MovieNet evaluation dataset, it can be found that the present invention has comprehensively surpassed in multiple evaluation criteria.
[0087] Table 1. Precision comparison between the present invention and existing methods on the MovieNet evaluation dataset.
[0088] Method AP mIoU AUC-ROC F1 BaSSL 57.4 50.69 90.54 47.02 TranS4mer 60.78 51.91 91.89 48.36 VSMBD (Ours) 63.65 56.35 92.73 55.30
[0089] The present invention proposes a self-supervised learning method VSMBD (Video Scene Montage for Boundary Detection) for video scene boundary detection. This method takes into account both the rich visual semantics contained in the shots and the complex and diverse semantic relationships between the shots. First, the visual semantics in the shots are decoupled into foreground and background, and a visual dual encoder is constructed based on the transfer learning method. Then, a video scene montage method (VSM) is designed to synthesize a large number of video segments with pseudo-boundaries using unannotated long videos. Finally, the pseudo-boundary prediction is used as a proxy task to guide the model to learn the semantic relationships between the shots, and the model is fine-tuned using real data. The experimental results show that VSMBD can extract rich and diverse semantic features and efficiently learn the relationships between the shots, thus demonstrating excellent performance on public datasets. This proves the effectiveness of the visual semantic decoupling method and also verifies the significance of comprehensively mastering the four types of semantic relationships between the shots for the video scene boundary detection task.
[0090] Experimental Analysis
[0091] Table 2, Transfer test results of this method across datasets (Average Precision: AP)
[0092]
[0093] Self-supervised methods including VSMBD have strong generalization ability on the OVSD dataset and the BBC dataset. This comparative experiment uses the model trained on the MovieNet dataset and does not perform additional fine-tuning on the OVSD and BBC datasets. The results show that VSMBD is superior to other self-supervised methods on the OVSD dataset, indicating its strong generalization ability. Although the long videos in the BBC dataset are all documentaries and there are significant differences from the movie data used during training, VSMBD still achieves competitive performance, indicating that the model trained by the method in this chapter has strong generalization ability.
[0094] Table 3, Performance comparison of unsupervised methods
[0095]
[0096] The experimental data in Table 3 show that VSMBD (U.) outperforms the existing unsupervised methods in all metrics. By comparing VSMBD (U.) with BaSSL (U.) and TranS4mer (U.) which also use the pseudo-boundary prediction algorithm, the superiority of the visual semantic decoupling scheme and the VSM algorithm proposed in the present invention in video scene boundary detection can be proven.
[0097] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the purpose and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.
Claims
1. A self-supervised video scene boundary detection method based on scene montage, characterized in that: Including the following steps: S1: Extraction of the shot sequence: Obtain an input long video S video =[s1,...s N , obtain the shots therein, sample key frames from the shots, and N represents the number of shots in S video ; S2: Construct and train the video scene montage network model VSM; S21: For each shot s video in S i encode it into a high-dimensional feature vector x i using a visual editor, where S video = [s1,..., s N is encoded into a sequence X video = [x1,..., x b ,..., x N ; S22: Determine the lengths of two feature subsequences and their starting positions through random parameters γ, α, β, and the synthesized feature sequence X through splicing operation syn , corresponding to the synthesized shot sequence S syn , where 1 ≤ γ ≤ n - 1, 1 ≤ α ≤ N - γ + 1, and 1 ≤ β ≤ N - γ + 1, and n represents the number of shots in the video segments generated by VSM; Feature sequence X syn and S syn The specific synthesis steps are as follows: Select a random positive integer γ such that among the two video segments to be intercepted, one contains γ shots and the other contains n - γ shots. Then select two random positive integers α and β as the positions of the starting shots of the two video segments to be intercepted in the long video S video so that two video segments can be intercepted from S video as two pseudo-scenes, denoted as S left = [s α , s α+1 …, s α+γ+1 and S right = [s β , s β+1 …, s β+n-γ-1 . The two pseudo-scenes have a total of n shots. Concatenate S left and S right together in the time dimension to form a synthesized video segment S syn = [s α , …s α+γ+1 , s β , …, s β+n-γ-1 : S syn = splice(S left , S right ) where splice(·) represents splicing in the time dimension; α, β, and γ are all unfixed random numbers, and these random numbers are reselected each time a video segment is synthesized; From X video = [x1,..., x N , intercepting the corresponding part of S syn yields X syn = [x α ,..., x α+γ+1 , x β ,..., x β+n-γ-1 ; S23: Generate multiple video segments as training data by the method of S22, and use pseudo-boundary prediction as a proxy task to train the context encoder and the scene boundary judgment module. The specific steps are as follows: Splice the feature vector sequence P = [p1, …, p n to X syn to supplement the position information of each shot, and finally feed it into the context relationship modeling module to obtain the feature vector sequence R syn = [r α , … r α+γ+1 , r β , …, r β+n-γ-1 : R syn = Context(concat(P, X syn )) where Context(·) is the context relationship modeling module, and concat(·) is the vector splicing operation; For video segment S syn , the starting shot s right of the pseudo-scene S β is regarded as a positive sample shot, and then a shot s syn is randomly selected from S i as a negative sample shot; finally, the feature vectors {r β , r i} corresponding to the positive and negative samples are input into the scene boundary judgment module, and the context encoder is pre-trained by minimizing the binary cross-entropy loss: Among them, h p (·) is the scene boundary judgment module, and the regularization term is used to combat overfitting, and λ represents the coefficient; S24: Use the real annotated scene boundary information to fine-tune the pre-trained VSM obtained in S25 to obtain the final VSM; S3: Detection. Pass the video segment to be detected through step S1 to obtain the shot sequence S’, then obtain the corresponding feature sequence X’ through step S21, input X’ into the final VSM, and output the confidence that the middle shot of the sequence is the scene boundary.
2. The self-supervised video scene boundary detection method based on scene montage according to claim 1, wherein: Each shot s in the S of S21 video is encoded as a feature vector x of dimension i The specific steps are as follows: i x i = concat(FGE(s i ), BGE(s i )) Among them, concat(·) represents the concatenation of feature vectors, FGE(·) and BGE(·) represent the foreground encoder and the background encoder respectively, and s i represents the i-th shot, and x i represents the feature of the i-th shot.
3. A self-supervised video scene boundary detection method based on scene montage according to claim 2, characterized in that: Fine-tune with the training dataset with scene boundary annotations in ImageNet. If the central shot s c in the dataset is a shot at the scene boundary, the label y c = 1; otherwise, y c = 0. The loss is calculated using the following formula: L sbd = -y c log(h sbd (r c )) + (1 - y c )log(1 - h sbd (r c )) Among which L sbd represents the loss in the fine-tuning stage, y c represents the central shot s in the shot sequence during fine-tuning, c the label of r c represents the context relationship feature of the central shot s c h sbd (·) represents the scene boundary judgment module.