Weakly-Supervised Video Clip Localization Method and System Based on Large-Scale Video Corpora
The semi-supervised video clip positioning method addresses the inefficiencies and inaccuracies of conventional methods by using self-supervised learning and multi-scale contrast learning to align text and video features, resulting in reduced labeling costs, improved accuracy, and enhanced efficiency.
Patent Information
- Application Number
- JP2024524993
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-05-11
- Filing Date
- 2023-09-05
- Publication Date
- 2025-06-26
- Estimated Expiration
- 2043-09-05
AI Technical Summary
Conventional video clip positioning methods in large-scale video corpora require labor-intensive labeling and suffer from low positioning accuracy and efficiency.
A semi-supervised positioning method and system that uses self-supervised learning to extract common semantic information between text and video, followed by multi-scale contrast learning to determine the spatial mapping relationship between video and text features, allowing for accurate and efficient video clip positioning with weak supervision.
The method significantly reduces the cost of dataset labeling, improves video positioning accuracy by aligning text and video feature representations, and enhances positioning efficiency by reducing computational complexity through hash-based search.
Smart Images

Figure 2025519301000001_ABST
Abstract
Description
Technical Field
[0001] (Cross - reference to related applications) This invention claims the priority of a Chinese patent application with the application number 202310523581.6 and the invention title "Weak - teacher - assisted method and system for positioning video clips based on large - scale video corpora", which was filed with the Chinese Patent Office on May 11, 2023, and all of its contents are incorporated into this invention by reference.
[0002] This invention relates to the technical field related to video data recognition, specifically, to a weak - teacher - assisted method and system for positioning video clips based on large - scale video corpora.
Background Art
[0003] The description of this part is only for providing information on the background art related to this invention and does not necessarily constitute the prior art.
[0004] The positioning of video clips based on large - scale video corpora refers to a technology that can determine the positions of video clips related to a search phrase in data having a large number of videos. Currently, there are many cases where video clip positioning technology is applied. For example, in the field of security measures, it is necessary to determine the position of a single video clip in a long video so as to quickly search for the clip to be searched. In this technology, it is necessary to artificially find the long video to be positioned and then use the search phrase for positioning. When a large number of long videos are included in a video database, it is very labor - intensive to artificially find this long video.
[0005] As the inventors have discovered through research, conventional video clip positioning methods often perform the positioning of video clips in a video corpus by means of a supervised method. In rare cases where a semi-supervised method is used, it is realized by metric learning, and the model is trained to learn an integrated feature space for one video and search, so as to measure the distance between the video and the search in the integrated feature space. Conventional video clip positioning methods have the problem that it is necessary to label the actual time level for the positioning of the dataset for the task of training video clip positioning, and the workload is very large. On the other hand, in the problem of video clip positioning in a large-scale video corpus, there is a problem that the positioning accuracy by the conventional method is not high and the positioning efficiency is low.
Summary of the Invention
Problems to be Solved by the Invention
[0006] In order to solve the above-mentioned problems, the present invention provides a semi-supervised positioning method and system for video clips based on a large-scale video corpus, which can directly determine the position of video clips accurately and quickly from a large-scale video database.
Means for Solving the Problems
[0007] In order to achieve the above-mentioned object, the present invention adopts the following technical means.
[0008] One or more embodiments include extracting common semantic information between text and video using self-supervised learning for the obtained training dataset, and obtaining fused semantic video features based on the semantic information; performing multi-scale contrast learning using a semi-supervised method for the fused semantic video features and corresponding text features, determining the spatial mapping relationship between the video features and the text features, mapping them to a metric space, and obtaining a trained metric space; A step of obtaining a search term, searching for text features similar to the search term in a trained metric space, and using the video clip corresponding to the text feature with the highest similarity as the video positioning result, and a method for positioning a video clip with weak supervision based on a large-scale video corpus is provided.
[0009] One or more embodiments are A common semantics detection module arranged to extract common semantic information between text and video using self-supervised learning for the obtained training dataset and to be used to obtain fused semantic video features based on the semantic information. A spatial mapping module for video features and text features arranged to perform multi-scale contrast learning using a weak supervision method on the fused semantic video features and corresponding text features, determine the spatial mapping relationship between the video features and the text features, map them to a metric space, and be used to obtain a trained metric space. A matching module arranged to obtain a search term, search for text features similar to the search term in a metric space, and search for a video clip with the highest similarity to the text features in the metric space trained based on the text features as the video positioning result, and a system for positioning a video clip with weak supervision based on a large-scale video corpus is provided.
[0010] An electronic device including a memory, a processor, and computer commands stored in the memory and executed by the processor, and when the computer commands are executed by the processor, the steps described in the above method are completed.
[0011] A computer-readable storage medium used to store computer commands that, when executed by a processor, complete the steps described in the above method.
Advantages of the Invention
[0012] Compared with the prior art, the beneficial effects of the present invention are as follows. (1) The present invention can use any video data including the title as the training data of the method by using the method with weak supervision regardless of the label of the dataset, and can greatly reduce the cost of labeling the dataset. (2) By using the self-supervised learning method for the training data to extract the common semantic information between the text and the video, better characterization information can be obtained, and the feature expressions of the same thing in the text mode and the video mode can be made more similar, thereby improving the accuracy of video positioning. (3) In the positioning process, instead of directly calculating the distance in the metric space between the video features using the new search term features, by preferentially searching for the text features similar to the new search term, the calculation amount can be reduced, and the video positioning efficiency can be greatly improved.
[0013] The video positioning method of the present invention can be embedded in any visual platform such as video entertainment, video monitoring, and autonomous driving, and can greatly improve the user experience.
[0014] The advantages of the present invention and the advantages of the additional forms will be described in detail in the following specific embodiments.
[0015] The drawings in the specification that are part of the present invention are for further understanding of the present invention, and the exemplary embodiments and their descriptions of the present invention are for interpreting the present invention and do not limit the present invention.
Brief Description of the Drawings
[0016]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Embodiments for Carrying Out the Invention
[0017] Hereinafter, the present invention will be further described with reference to the drawings and examples.
[0018] It should be pointed out that all the following detailed descriptions are illustrative and are intended to further describe the present invention. Unless otherwise specified, all technical and scientific terms used in this specification have the meanings commonly understood by those skilled in the art.
[0019] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. When used in this specification, the singular form is also intended to include the plural form unless otherwise clearly indicated in the context. Also, when terms such as "comprising" and / or "including" are used in this specification, it should also be understood that they indicate the presence of features, steps, operations, devices, components and / or combinations thereof. It is necessary to explain that the embodiments and features in the embodiments of the present invention can be combined with each other without conflict. Hereinafter, the embodiments will be described in detail with reference to the drawings.
Examples
[0020] In the technical means disclosed in one or more embodiments, as shown in FIGS. 1 to 5, For the obtained training dataset, use self-supervised learning to extract the common semantic information between text and video, and obtain fused semantic video features based on the semantic information in Step 1; For the fused semantic video features and corresponding text features, perform multi-scale contrast learning using the weak teacher method, determine the spatial mapping relationship between the video features and the text features, map them to the metric space, and obtain the trained metric space in Step 2; Obtain the search query, search for text features similar to the search query in the trained metric space, and use the video clip corresponding to the text feature with the highest similarity as the video positioning result in Step 3. This is a weak teacher-based video clip positioning method based on a large-scale video corpus.
[0021] In this embodiment, regardless of the labels of the dataset, the weak teacher method is used, and any video data including the title can be used as the training data of this method, which can greatly reduce the cost of labeling the dataset. By using the self-supervised learning method to extract the common semantic information between text and video for the training data, better characterization information can be obtained, making the feature representations of the same thing in the text mode and the video mode more similar, thereby improving the accuracy of video positioning. In the positioning process, instead of directly calculating the distance in the metric space between the new search query feature and the video features, preferentially searching for text features similar to the new search query can reduce the computational complexity and greatly improve the video positioning efficiency.
[0022] In Step 1, the training dataset includes past search queries and corresponding video data, and text features are obtained after performing feature extraction on the search queries.
[0023] Step 1 is for detecting the common semantic information between the two modes of search and video. Taking an example to explain, if the "desk" in the search and the "desk" in the video have the same semantic information, the feature expressions of the "desk" in the search and the "desk" in the video should be close to each other.
[0024] In some embodiments, the method of using self-supervised learning to extract the common semantic information between text and video and obtaining the fused semantic video features based on the semantic information includes the following steps 11 to 16 as shown in FIG. 1.
[0025] In step 11, feature extraction is performed on video data and search terms using a backbone network, and video features and text features can be obtained respectively.
[0026] In step 12, the obtained video features and text features (i.e., the features of the search terms) are fused.
[0027] Optionally, as a fusion method, addition and dot product operations are performed on the video features and text features and then concatenated.
[0028] In step 13, a convolution operation is performed on the fused features.
[0029] In this step, through the convolution operation, the feature data can be reduced in dimension and more detailed information can be obtained.
[0030] In step 14, the text features are predicted using the video features obtained after convolution, and self-supervised training is performed with the original text features as teacher information.
[0031] In step 15, the reconstruction loss is calculated based on the predicted text features and teacher information, and a larger weight value is assigned to the video clips with a reconstruction loss lower than the set value, and this weight value is the reconstruction reward.
[0032] In step 16, the reconstruction reward is weighted on the video features obtained after convolution to obtain fused semantic video features.
[0033] This embodiment uses self-supervised learning, and by predicting the text features again with the video features fused with text semantic information, more accurate text-video common semantic information can be obtained, and the feature expressions of the same thing in the text mode and the video mode become more similar. As a result, in the video positioning process, the corresponding video clip can be found more accurately by the features in the text mode, and the accuracy of video positioning is improved.
[0034] What is performed in step 2 is an integrated metric learning step. The purpose of integrated metric space learning is to learn a metric space, put all video clips and search texts into the metric space, and evaluate the corresponding video clips to be searched according to the similarity in the space.
[0035] Realize the spatial mapping of high-level video features and text features in the metric space. The schematic diagram of this part is as shown in Figure 2. First, by supplying the fused semantic video features obtained in step 1 to a GRU (long short-term memory network), the time-series information of the video is obtained. Map the training results of text features and high-level video features to the metric space every time. After several times of learning, the distance between the video clip with a high match degree to the text and the text becomes closer, and the distance to the video clip with a low match degree becomes farther. The distances in the space of different search pairs (that is, different text-video) become farther. Here, the text and the video clip with a high match degree to this text form a search pair.
[0036] In this embodiment, a more excellent mapping space can be obtained through multi-scale contrastive learning. The multi-scale contrastive learning includes clip-level learning and video-level learning. Clip-level learning is to reduce the distance of video clips similar to the search text and increase the distance of video clips not similar to the search text. Video-level learning is to increase the distance between the video corresponding to the search phrase and other videos.
[0037] In this embodiment, the learning for videos is divided into clip-level and video-level. Clip-level learning is the learning for detailed features in the video, and video-level learning is the learning for the complete video. For example, through video-level learning, it can be learned whether this video is a sports video or a car video, and clip-level feature learning is to learn whether what is in the video is high jump or weightlifting. These two-scale learning can be carried out independently.
[0038] In this embodiment, multi-scale contrastive learning is the key to realizing learning with a weak teacher. Through multi-scale contrastive learning, training for the metric space is realized.
[0039] The purpose of clip-level learning is to increase the matching score between text features and the corresponding video clips. The method of clip-level learning includes the following steps 21 to 25.
[0040] In step 21, the fused semantic video features are supplied to a GRU (long short-term memory network) to obtain the time-series information of the video. The obtained time-series information is added to the fused semantic video features to obtain high-level video features.
[0041] In step 22, the matching score calculated for the text feature and each high-level video clip feature is input into the locator. The video clip corresponding to the start and end times with a score higher than the set value is labeled as a positive sample, and the video that cannot be labeled is used as a negative sample.
[0042] Here, the locator can be a multi-layer perceptron integrated with a normalization exponential function, that is, MLP+softmax can be used. The normalization function outputs the confidence score for each text clip-search pair.
[0043] The calculation formula of the normalization function softmax is as follows:
Equation
[0044] The method for calculating the matching degree between the text feature and each high-level video clip feature is specifically as follows: The Tanh function can be used to calculate the matching degree between features. That is, the matching degree r t is calculated by the following formula:
Equation
[0045] In step 23, an adversarial generation network (GAN network) is set up to generate video clip features similar to the original video for clip-level learning, and these generated video clip features are used as negative samples.
[0046] In a specific example, for video A with n clips a1, a1... an is included, and clip a in video A is positioned by search term B i When training the GAN network, it is necessary to input video A and search B. Here, the GAN network is used to generate clip features similar to video A. Here, the "original video" refers to video A, that is, the GAN network generates video clip-level features similar to the training video ("original video").
[0047] Optionally, the number of negative samples generated using the adversarial generative network (GAN network) can be set. For example, setting the number N to 100 means that 100 pieces of video clip feature information can be generated separately.
[0048] In step 24, based on the recognized clip-level positive samples and negative samples, the similarity between samples is quantified by cosine similarity, and the similarity probability between two positive samples is maximized by minimizing the loss function to obtain the loss function.
[0049] The positive samples and negative samples of clip-level contrast learning are all clip-level samples, that is, one sample is one video clip.
[0050] The similarity between sample A and sample B is defined by the following formula.
Equation
[0051] Calculate the similarity probability between positive samples, and the probability calculation formula for one of the positive sample pairs (p i , p j ) is as follows
Equation
[0052] To maximize the similarity probability between two positive samples by minimizing the loss function, the following logarithmic loss formula can be used.
Number
[0053] When contrastive learning is performed by the above-described method, a metric space in which positive samples are closer to each other can be obtained.
[0054] In step 25, after obtaining the loss function, the clip-level loss function LOSS clip is optimized by the stochastic gradient descent algorithm. After the loss function converges, the optimization can be stopped. After about 1000 iterations and performing clip-level optimization, a metric space can be obtained.
[0055] The convergence of the loss means that as the model is trained, the value of the model's loss function gradually decreases until it approaches a stable state.
[0056] In another embodiment, the loss of clip-level contrastive learning may use the hinge loss.
[0057] Both clip-level learning and video-level learning are training processes. These two processes are performed in parallel, and the finally obtained result is a trained model.
[0058] Video-level contrast learning is almost exactly the same as clip-level contrast learning, except that the samples are different. The video-level contrast learning samples are one complete video, while the clip-level samples are one video clip.
[0059] Video-level learning aims to eliminate interference from other unrelated video features. The video-level learning method includes the following steps 2-1 to 2-4.
[0060] In step 2-1, the input video for video-level learning is used as the positive sample for video-level contrast learning, and a plurality of other video samples are randomly selected as negative samples.
[0061] Optionally, as negative samples, 100 other video samples can be randomly selected.
[0062] In step 2-2, based on the recognized video-level positive and negative samples, the similarity between samples is quantified by cosine similarity, and the similarity probability between two positive samples is maximized by minimizing the loss function to obtain the loss function.
[0063] In step 2-2, one of the video-level positive and negative samples is a complete video.
[0064] Quantifying the similarity between video-level samples by cosine similarity, that is, the similarity between video-level sample A and video-level sample B, is defined as follows.
Equation
[0065] Calculate the similarity probability between positive samples, and one of the positive sample pairs (p i , p j) The calculation formula for the probability is as follows:
Number
[0066] To maximize the similarity probability between two positive samples by minimizing the loss function, the following logarithmic loss formula can be used.
Number
[0067] When contrastive learning is performed by the method described above, a metric space where the distance between different videos becomes farther can be obtained.
[0068] In step 2-3, after obtaining the loss function, the video-level loss function LOSS video is optimized by the stochastic gradient descent algorithm. After the loss function converges, the optimization is stopped, and the loss can converge after about 1000 iterations.
[0069] In step 2-4, by optimizing the loss function, a metric space after video-level optimization is obtained.
[0070] In another embodiment, the loss of clip-level contrastive learning may use the hinge loss.
[0071] In this embodiment, a weakly supervised training process is realized by contrastive learning, and also, by creatively using the GAN network to generate similar video clip features, the negative sample data required for contrastive learning is increased.
[0072] In step 3, obtain the search term, search for text features similar to the search term in the trained metric space, and use the video clip corresponding to the text feature with the highest similarity as the video positioning result.
[0073] Here, searching for text features similar to the search term in the metric space is realized by the hash binary code. As specific steps, it includes the following steps 31, 32, and 33.
[0074] In step 31, convert the text features in the metric space into hash binary codes by a hash mapping function, and use the hash binary code as the search index of the text-video search pair.
[0075] In step 32, preferentially match the obtained search term with a predetermined number M of hash binary codes. Here, M may be set to 10.
[0076] In step 33, use the video clip corresponding to the hash binary code with the highest matching degree as the video clip positioning result. The start time and end time of the positioned video clip are the timestamp information to be output.
[0077] In this embodiment, instead of directly calculating the distance in the metric space between the new search term feature and the video feature using the text hash code as the index of the text-video pair, by preferentially searching for text features similar to the new search term, the search efficiency can be improved, the calculation amount can be reduced, about 10 times of acceleration is possible, and the time complexity is reduced by 90%.
[0078] In this embodiment, considering the problem of the positioning speed, by adopting the measure of using the text hash code as the index of the whole text-video, the calculation amount can be effectively saved, and the positioning speed can be significantly increased.
Example
[0079] Based on Example 1, in this example, For the obtained training dataset, a common semantics detection module is arranged to be used for extracting common semantic information between text and video using self-supervised learning and obtaining fused semantic video features based on the semantic information, For the fused semantic video features and corresponding text features, a spatial mapping module of video features and text features is arranged to be used for performing multi-scale contrastive learning using the weak-supervised method, determining the spatial mapping relationship between video features and text features, mapping them into a metric space, and obtaining a trained metric space, A matching module is provided, which is arranged to obtain a search phrase, search for text features similar to the search phrase in the metric space, and search for the video clip with the highest similarity to the text features in the metric space trained based on the text features to obtain a video positioning result, so as to provide a weak-supervised video clip positioning system based on a large-scale video corpus.
[0080] Here, each module in this example corresponds one-to-one to each step in Example 1, and its specific implementation process is the same. It should be noted that the repeated description here is omitted.
Example
[0081] This example provides an electronic device including a memory, a processor, and computer commands stored in the memory and executed by the processor. When the computer commands are executed by the processor, the steps described in the method of Example 1 are completed.
Example
[0082] This embodiment provides a computer-readable storage medium that is used to store computer commands, and when the computer commands are executed by a processor, the steps described in the method of Embodiment 1 are completed.
[0083] The above are only preferred embodiments of the present invention and do not limit the present invention. For those skilled in the art, various modifications and changes are possible to the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included within the protection scope of the present invention.
Claims
1. including the following steps 1, 2 and 3, In step 1, for the obtained training dataset, use self-supervised learning to extract the common semantic information between text and video, and obtain fused semantic video features based on the semantic information, Specific steps include the following steps 11 to 16, In step 11, use a backbone network to perform feature extraction on video data and search phrases, and obtain video features and text features respectively, In step 12, fuse the obtained video features and the text features which are the features of the search phrase, As the fusion method, perform addition and dot product operations on the video features and text features and then concatenate them, In step 13, perform a convolution operation on the fused features, Through the convolution operation, the dimensionality of the feature data can be reduced, and more detailed information can be obtained, In step 14, use the video features obtained after convolution to predict text features, and perform self-supervised training with the original text features as teacher information, In step 15, calculate the reconstruction loss based on the predicted text features and teacher information, and assign a larger weight value to the video clip with a reconstruction loss lower than the set value, and this weight value is the reconstruction reward, In step 16, weight the reconstruction reward to the video features obtained after convolution to obtain fused semantic video features, In step 2, perform multi-scale contrast learning on the fused semantic video features and the corresponding text features using the weak teacher method, determine the spatial mapping relationship between the video features and the text features, map them to a metric space, and obtain a trained metric space, Specifically, the multi-scale contrast learning includes clip-level learning and video-level learning. Clip-level learning is to make the distance of video clips similar to the search text closer and keep away from video clips not similar to the search text. Video-level learning is to make the distance between the video corresponding to the search phrase and other videos farther, Learning for videos is divided into clip level and video level. At the clip level, it is learning for detailed features in the video, and at the video level, it is learning for the complete video. These two scales of learning can be carried out independently. Multi-scale contrastive learning is crucial for realizing learning with a weak teacher. Through multi-scale contrastive learning, training for the metric space is realized. Clip-level learning aims to increase the matching score between text features and corresponding video clips. The method of clip-level learning includes the following steps 21 to 25. In step 21, the fused semantic video features are supplied to the GRU long short-term memory network to obtain the time-series information of the video. The obtained time-series information is added to the fused semantic video features to obtain high-level video features. In step 22, the matching scores calculated for the text features and each high-level video clip feature are put into the locator. The video clips corresponding to the start and end times with scores higher than the set value are labeled as positive samples, and the unlabeled videos are used as negative samples. Here, the locator is a combination of a multi-layer perceptron and a normalization exponential function, that is, MLP+softmax can be used to output the confidence scores of each text clip-search pair by the normalization function. The calculation formula of the normalization function softmax is as follows. 【Number 9】 Here, p is the confidence score, and x i is the output value of the multi-layer perceptron for the i-th clip, and k represents that there are a total of k clips. The method for calculating the matching degree between text features and each high-level video clip feature is specifically as follows. Calculate the matching degree between features using the TanH function, that is, the matching degree r t is calculated by the following formula 【Number 10】 Here, t represents a time step, and q t represents high-level video clip features, and k t represents text features, In step 23, a GAN network, which is an adversarial generation network, is set up to generate video clip features similar to the original video for clip-level learning, and these generated video clip features are used as negative samples. In step 24, based on the recognized positive and negative samples at the clip level, the similarity between samples is quantified by cosine similarity, and the loss function is minimized to maximize the similarity probability between two positive samples, thereby obtaining the loss function. The positive and negative samples of clip-level contrastive learning are all clip-level samples, that is, one sample is one video clip. The similarity between Sample A and Sample B is defined by the following formula, 【Number 11】 Calculate the similarity probability between positive samples, and the probability calculation formula for one of the positive sample pairs (p i , p j ) is as follows, 【Number 12】 Here, p i , p j is a positive sample, N is the number of all samples, N includes negative samples, a positive pair is a positive sample, a positive sample, and a negative pair is a positive sample, a negative sample, e is the base of the natural logarithm, a constant in mathematics, a non-repeating infinite decimal and a transcendental number, approximately 2.71828, To maximize the similarity probability between two positive samples by minimizing the loss function, the following logarithmic loss formula can be used. 【Number 13】 When contrastive learning is performed by the method described above, a metric space in which positive samples are closer to each other is obtained. In step 25, after obtaining the loss function, the clip-level loss function LOSS clip is optimized by the stochastic gradient descent algorithm, and after the loss function converges, the optimization can be stopped. After about 1000 iterations and performing the optimization of the clip level, a metric space is obtained. clip Video-level learning aims to eliminate interference from other unrelated video features. The video-level learning method includes the following steps 2-1 to 2-4. In step 2-1, the input video for video-level learning is used as a positive sample for video-level contrastive learning, and a plurality of other video samples are randomly selected as negative samples. As negative samples, 100 other video samples can be randomly selected. In step 2-2, based on the recognized video-level positive samples and negative samples, the similarity between samples is quantified by cosine similarity, and the similarity probability between two positive samples is maximized by minimizing the loss function to obtain a loss function. One of the video-level positive samples and negative samples is a complete video. Quantifying the similarity between video-level samples, i.e., the similarity between video-level sample A and video-level sample B, by cosine similarity is defined as follows. 【Number 14】 Calculate the similarity probability between positive samples, and the probability calculation formula for one of the positive sample pairs (p i , p j ) is as follows, 【Number 15】 Here, p i , p j are positive samples, N is the number of all samples, N includes negative samples, positive pairs are positive sample, positive sample, and negative pairs are positive sample, negative sample, To maximize the similarity probability between two positive samples by minimizing the loss function, the following logarithmic loss formula can be used. 【Number 16】 When contrastive learning is performed by the method described above, a metric space in which the distances between different videos are farther apart is obtained. In step 2-3, after obtaining the loss function, the video-level loss function LOSS video is optimized by the stochastic gradient descent algorithm. After the loss function converges, the optimization is stopped. After about 1000 iterations, the loss can converge. video In step 2-4, by optimizing the loss function, a metric space after video-level optimization is obtained. A method for positioning video clips with weak teachers based on a large-scale video corpus, characterized in that in step 3, a search query is obtained, text features similar to the search query are searched in the trained metric space, and the video clip corresponding to the text feature with the highest similarity is used as the video positioning result.
2. The obtained video features and text features are fused. As a fusion method, addition and dot product operations are performed on the video features and text features and then concatenated. The video clip positioning method with weak supervision based on the large-scale video corpus according to claim 1 is characterized by this.
3. Searching for text features similar to the search query in the metric space is realized by the hash binary code. As steps, Step 31 of converting the text features in the metric space into hash binary codes by a hash mapping function and using the hash binary code as the search index of the text-video search pair; Step 32 of preferentially matching the obtained search query with a predetermined number of hash binary codes; Step 33 of using the video clip corresponding to the hash binary code with the highest matching degree as the positioning result. The video clip positioning method with weak supervision based on the large-scale video corpus according to claim 1 is characterized by including these steps.
4. A video clip positioning system with weak supervision based on the large-scale video corpus, which is realized by using the video clip positioning method with weak supervision based on the large-scale video corpus according to claim 1, A common semantics detection module arranged to extract common semantic information between text and video using self-supervised learning for the obtained training dataset and used to obtain fused semantic video features based on the semantic information; A spatial mapping module of video features and text features arranged to perform multi-scale contrast learning using the weak-supervision method on the fused semantic video features and corresponding text features, determine the spatial mapping relationship between the video features and text features, map them to the metric space, and obtain the trained metric space. A matching module arranged to be used to obtain search terms, search for text features similar to the search terms in a metric space, and search for the video clip with the highest similarity to the text features in the metric space trained based on the text features to obtain a video positioning result, and a weakly-supervised video clip positioning system based on a large-scale video corpus, characterized by including the matching module.
5. An electronic device comprising a memory, a processor, and computer commands stored in the memory and executed by the processor, wherein when the computer commands are executed by the processor, the steps of the weakly-supervised video clip positioning method based on the large-scale video corpus according to any one of claims 1-3 are completed.
6. A computer-readable storage medium characterized by being used to store computer commands that, when executed by a processor, complete the steps of the weakly-supervised video clip positioning method based on the large-scale video corpus according to any one of claims 1-3.
Citation Information
Patent Citations
Multimedia data searching method and device, equipment and storage medium
CN113590850A