A method and system for weakly supervised location of video clips based on a large-scale video corpus
The weakly supervised method addresses the inefficiencies of conventional localization by using self-supervised learning and multi-scale contrastive learning to enhance video clip localization accuracy and speed in large-scale corpora.
Patent Information
- Application Number
- JP2024524993
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2023-05-11
- Filing Date
- 2023-09-05
- Publication Date
- 2026-02-04
- Estimated Expiration
- 2043-09-05
AI Technical Summary
Conventional video clip localization methods in large-scale video corpora require significant manual labeling and have low accuracy and efficiency, especially when using supervised methods.
A weakly supervised method employing self-supervised learning to extract common semantic information between text and video, followed by multi-scale contrastive learning to determine spatial mapping in a metric space, allowing for efficient and accurate video clip localization.
Reduces dataset labeling costs, improves feature representation accuracy, and significantly enhances localization efficiency by prioritizing text feature similarity searches in the metric space.
Smart Images

Figure 0007811038000017 
Figure 0007811038000018 
Figure 0007811038000019
Abstract
Description
[Technical Field]
[0001] (CROSS-REFERENCE TO RELATED APPLICATIONS) This invention claims priority to a Chinese patent application bearing application number 202310523581.6 and entitled "Method and system for weakly supervised positioning of video clips based on large-scale video corpus," filed with the China Patent Office on May 11, 2023, the entire contents of which are incorporated herein by reference.
[0002] The present invention relates to the technical field related to video data recognition, and in particular to a method and system for weakly supervised location of video clips based on a large-scale video corpus. [Background technology]
[0003] The description in this section is merely intended to provide background information related to the present invention and does not necessarily constitute prior art.
[0004] Video clip location based on a large-scale video corpus refers to a technology that can locate relevant video clips using a single search term in data containing a large number of videos. Currently, video clip location technology is often applied, for example, in the field of safety measures, where it is necessary to locate a single video clip in a long video so as to quickly search for the target clip. In this technology, the long video to be located must be manually found and then located using the search term. When a video database contains a large number of long videos, manually finding this long video is very time-consuming.
[0005] As the inventors have found through their research, conventional video clip localization methods often use supervised methods to locate video clips in a video corpus. In a few cases, weakly supervised methods are used, which are implemented through metric learning. A model is trained to learn a joint feature space between a video and a query, and the distance between the video and the query in the joint feature space is measured. Conventional video clip localization methods have the drawback of requiring a significant amount of work, since they require labeling the actual time of the localization in a dataset for the task of training the video clip localization. Meanwhile, when it comes to the problem of video clip localization in a large-scale video corpus, conventional methods have low localization accuracy and low localization efficiency. Summary of the Invention [Problem to be solved by the invention]
[0006] To solve the above-mentioned problems, the present invention provides a method and system for weakly supervised video clip location based on a large-scale video corpus, which can realize accurate and fast direct location of video clips from a large-scale video database. [Means for solving the problem]
[0007] In order to achieve the above object, the present invention employs the following technical means.
[0008] One or more embodiments may include: extracting common semantic information between the text and the video using self-supervised learning from the acquired training dataset, and obtaining fused semantic video features based on the semantic information; Performing multi-scale contrastive learning on the fused semantic video features and corresponding text features using a weakly supervised method to determine the spatial mapping relationship between the video features and the text features, and mapping them into a metric space to obtain a trained metric space; The present invention provides a method for weakly supervised localization of video clips based on a large-scale video corpus, including the steps of: obtaining a search term; searching for text features similar to the search term in a trained metric space; and determining the video clip corresponding to the text feature with the highest similarity as the video localization result.
[0009] One or more embodiments may include: a common semantics detection module configured to extract common semantic information between text and video using self-supervised learning from the acquired training dataset, and to obtain fused semantic video features based on the semantic information; a video feature and text feature spatial mapping module configured to perform multi-scale contrastive learning on the fused semantic video features and corresponding text features using a weakly supervised method, determine a spatial mapping relationship between the video features and the text features, map them into a metric space, and use the determined metric space to obtain a trained metric space; The present invention provides a weakly supervised video clip location system based on a large-scale video corpus, the system including: a matching module configured to receive a search term, search for text features similar to the search term in a metric space, and search for video clips that have the highest similarity to the text features in a metric space trained based on the text features to provide a video location result.
[0010] An electronic device comprising a memory, a processor, and computer instructions stored in the memory and executed by the processor, the computer instructions completing the steps of the method described above when executed by the processor.
[0011] A computer-readable storage medium used to store computer instructions that, when executed by a processor, complete the steps of the methods described above. [Effects of the Invention]
[0012] Compared with the prior art, the beneficial effects of the present invention are as follows: (1) The present invention uses a weakly supervised method regardless of the label of the dataset, and any video data, including titles, can be used as training data for the method, significantly reducing the cost of labeling the dataset. (2) By using a self-supervised learning method on the training data to extract common semantic information between text and video, better characterization information can be obtained, making the feature representation of the same object in text mode and video mode more similar, thereby improving the accuracy of video positioning. (3) During the positioning process, instead of directly using the new search term features to calculate the distance in metric space with the video features, we prioritize searching for text features similar to the new search term, which reduces the computational effort and significantly improves the efficiency of video positioning.
[0013] The video positioning method of the present invention can be embedded into any visual platform, such as video entertainment, video surveillance, unmanned driving, etc., and can greatly improve the user experience.
[0014] The advantages of the present invention and the advantages of the additional embodiments will be explained in detail in the following specific examples.
[0015] The drawings in the specification that form a part of this invention are intended to provide a further understanding of the invention, and the illustrative embodiments of the invention and the description thereof are intended to interpret the invention and are not intended to limit the invention. [Brief explanation of the drawings]
[0016] [Figure 1] 1 is a flowchart of a method for recognizing common semantic sensing information between modes according to a first embodiment of the present invention; [Figure 2] 1 is a schematic flowchart of a method for spatial mapping of video features and text features according to a first embodiment of the present invention; [Figure 3] 1 is a schematic flowchart of a method for clip-level learning in multi-scale contrastive learning according to a first embodiment of the present invention; [Figure 4] 1 is a schematic flowchart of a method for video level learning in multi-scale contrastive learning according to a first embodiment of the present invention; [Figure 5] 1 is an overall flowchart of a video clip positioning method according to a first embodiment of the present invention; DETAILED DESCRIPTION OF THE INVENTION
[0017] The present invention will now be further described with reference to the following drawings and examples.
[0018] It should be noted that the following detailed description is all exemplary and is intended to further explain the present invention. Unless otherwise specified, all technical and scientific terms used herein have the meanings commonly understood by those skilled in the art.
[0019] It should be noted that the terms used herein are merely for the purpose of describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, the singular is intended to include the plural unless the context clearly dictates otherwise, and it should also be understood that when terms such as "comprise" and / or "comprises" are used herein, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof. It should be noted that each embodiment and feature in each embodiment of the present invention can be combined with each other without conflict. Hereinafter, the embodiments will be described in detail with reference to the drawings. [Example]
[0020] In the technical means disclosed in one or more embodiments, as shown in FIGS. 1 to 5, Step 1: extracting common semantic information between text and video using self-supervised learning from the acquired training dataset, and obtaining fused semantic video features based on the semantic information; Step 2: Perform multi-scale contrastive learning on the fused semantic video features and corresponding text features using a weakly supervised method to determine the spatial mapping relationship between the video features and the text features, and map them into a metric space to obtain a trained metric space; and (3) obtaining a search term, searching for text features similar to the search term in the trained metric space, and determining the video clip corresponding to the text feature with the highest similarity as the video location result.
[0021] This embodiment uses a weakly supervised method regardless of the dataset label, allowing any video data, including titles, to be used as training data for this method, significantly reducing the cost of dataset labeling. By using a self-supervised learning method on the training data to extract common semantic information between text and video, better characterization information can be obtained, making the feature representations of the same object in text and video more similar, thereby improving the accuracy of video location. During the location process, rather than directly using the new search term features to calculate the distance between the video features in metric space, preferentially searching for text features similar to the new search term can reduce the amount of calculation and significantly improve the efficiency of video location.
[0022] In step 1, the training dataset contains video data corresponding to past search terms, and text features are obtained after feature extraction is performed on the search terms.
[0023] Step 1 is to detect common semantic information between the two modes, search and video. For example, if a desk in search and a desk in video have the same semantic information, the feature representation of the desk in search and the feature representation of the desk in video should be close to each other.
[0024] In some embodiments, a method for extracting common semantic information between text and video using self-supervised learning and obtaining fused semantic video features based on the semantic information includes the following steps 11 to 16, as shown in FIG. 1 :
[0025] In step 11, feature extraction is performed on the video data and the search terms using the backbone network to obtain video features and text features, respectively.
[0026] In step 12, the obtained video features and text features (i.e., search term features) are fused.
[0027] Optionally, the fusion method is to perform addition and dot product operations on the video and text features before concatenation.
[0028] In step 13, a convolution operation is performed on the fused features.
[0029] In this step, the dimension of the feature data is reduced by a convolution operation, and more detailed information can be obtained.
[0030] In step 14, the video features obtained after convolution are used to predict text features, and self-supervised training is performed using the original text features as training information.
[0031] In step 15, the reconstruction loss is calculated based on the predicted text features and the training information, and a larger weight value is assigned to the video clip whose reconstruction loss is lower than a set value, and the weight value is the reconstruction reward.
[0032] In step 16, the reconstruction reward is weighted to the video features obtained after convolution to obtain fused semantic video features.
[0033] This embodiment uses self-supervised learning to re-predict text features using video features that combine text semantic information, thereby obtaining more accurate text-video common semantic information, and making the feature representations of the same subject in text mode and video mode more similar. As a result, during the video location process, the features of the text mode can be used to more accurately find the corresponding video clip, thereby improving the accuracy of video location.
[0034] Step 2 is the integrated metric learning step, and the purpose of integrated metric space learning is to learn a single metric space, put all video clips and search texts into the metric space, and evaluate the corresponding video clips based on their similarity in the space.
[0035] The spatial mapping of high-level video features and text features is realized in metric space. The schematic diagram of this part is shown in Figure 2. First, the fused semantic video features obtained in step 1 are fed into a GRU (long short-term memory network) to obtain the video's time series information. The training results of each time for the text features and high-level video features are mapped to metric space. After several rounds of learning, the distance between the text and the video clips with high matching scores becomes closer, while the distance between the text and the video clips with low matching scores becomes farther. The distance in the space between different search pairs (i.e., different text-video) becomes greater. Here, the text and the video clips with high matching scores for this text constitute one search pair.
[0036] In this embodiment, a better mapping space can be obtained through multi-scale contrastive learning, and the multi-scale contrastive learning includes clip-level learning and video-level learning. The clip-level learning is to bring video clips similar to the search text closer and to move video clips that are not similar to the search text farther away. The video-level learning is to increase the distance between the video corresponding to the search term and other videos.
[0037] In this embodiment, video learning is divided into clip level and video level, where clip level is learning for detailed features in the video, and video level is learning for the entire video. For example, video level learning can learn whether the video is a physical education video or a car video, and clip level feature learning can learn whether the video is high jump or weightlifting. These two scales of learning can be performed separately.
[0038] In this embodiment, multi-scale contrastive learning is the key to realizing weakly supervised learning, and training on the metric space is realized by multi-scale contrastive learning.
[0039] The clip-level learning aims to improve the match score between text features and corresponding video clips, and the clip-level learning method includes the following steps 21 to 25.
[0040] In step 21, the fused semantic video features are fed to a GRU (long short-term memory network) to obtain time series information of the video, and the obtained time series information is added to the fused semantic video features to obtain high-level video features.
[0041] In step 22, the calculated match scores for the text features and each high-level video clip feature are entered into the locator, and video clips corresponding to start and end times with scores higher than a set value are labeled as positive samples, and videos that cannot be labeled are negative samples.
[0042] Here, the locator can be a combination of a multilayer perceptron and a normalized exponential function, i.e., MLP+softmax, which outputs a confidence score for each text clip-search pair according to the normalized function.
[0043] The formula for the normalization function softmax is as follows:
number
[0044] The method for calculating the match between the text features and each high-level video clip feature is as follows: The Tanh function can be used to calculate the match between features, i.e., the match r t is calculated by the following formula:
number
[0045] In step 23, a generative adversarial network (GAN network) is set up to generate video clip features similar to the original video through clip-level learning, and these generated video clip features are used as negative samples.
[0046] In a specific example, a video A contains n clips a1, a1...an Search for clip a in video A by search term B. i Therefore, when training the GAN network, video A and search B need to be input. Here, the GAN network is designed to generate clip features similar to video A, where "original video" refers to video A, i.e., the GAN network generates video clip-level features similar to the training video ("original video").
[0047] Optionally, the number of negative samples generated using the generative adversarial network (GAN network) can be set, for example, the number N can be set to 100, i.e., 100 pieces of video clip feature information can be generated separately.
[0048] In step 24, based on the recognized clip-level positive and negative samples, the similarity between the samples is quantified by cosine similarity, and the similarity probability between the two positive samples is maximized by minimizing the loss function, thereby obtaining the loss function.
[0049] The positive and negative samples in clip-level contrast training are all clip-level samples, that is, one sample corresponds to one video clip.
[0050] The similarity between sample A and sample B is defined by the following formula:
number
[0051] Calculate the similarity probability between positive samples and select one positive sample pair (p i , p j The formula for the probability of
number
[0052] To maximize the probability of similarity between two positive samples by minimizing a loss function, we can use the following equation, which is the logarithmic loss:
number
[0053] When contrastive learning is performed using the above method, a metric space in which positive samples are closer to each other is obtained.
[0054] In step 25, after obtaining the loss function, the clip-level loss function LOSS is calculated by the stochastic gradient descent algorithm. clip After the loss function converges, the optimization can be stopped and the metric space after clip-level optimization is obtained after about 1000 iterations.
[0055] Loss convergence means that as the model is trained, the value of the model's loss function becomes smaller and smaller until it approaches a stable state.
[0056] In another embodiment, the loss for clip-level contrast learning may use hinge loss.
[0057] Both clip-level learning and video-level learning are training processes, and these two processes are performed in parallel, and the final result is a trained model.
[0058] Video-level contrastive training is almost identical to clip-level contrastive training, except that the samples are different: a video-level contrastive training sample is a complete video, while a clip-level sample is a video clip.
[0059] The video level learning aims to eliminate interference from other irrelevant video features, and the video level learning method includes the following steps 2-1 to 2-4.
[0060] In step 2-1, the input video to be subjected to video-level learning is taken as a positive sample for video-level contrastive learning, and multiple other video samples are randomly selected as negative samples.
[0061] Optionally, one can randomly select another 100 video samples as negative samples.
[0062] In step 2-2, based on the recognized video-level positive and negative samples, the similarity between the samples is quantified by cosine similarity, and the similarity probability between the two positive samples is maximized by minimizing the loss function, thereby obtaining the loss function.
[0063] In step 2-2, one sample among the positive and negative samples at the video level is the complete video.
[0064] Quantifying the similarity between video level samples by cosine similarity, i.e., the similarity between video level sample A and video level sample B, is defined as follows:
number
[0065] Calculate the similarity probability between positive samples and select one positive sample pair (p i , p jThe formula for the probability of
number
[0066] To maximize the probability of similarity between two positive samples by minimizing a loss function, we can use the following equation, which is the logarithmic loss:
number
[0067] Contrastive learning using the above method results in a metric space in which different videos are more distant from each other.
[0068] In step 2-3, after obtaining the loss function, we use the stochastic gradient descent algorithm to calculate the video-level loss function LOSS. video After the loss function converges, the optimization is stopped and the loss can be converged after about 1000 iterations.
[0069] In steps 2-4, we obtain the metric space after video-level optimization by optimizing the loss function.
[0070] In another embodiment, the loss for clip-level contrast learning may use hinge loss.
[0071] In this embodiment, the weakly supervised training process is realized by contrastive learning, and the GAN network is creatively used to generate similar video clip features, thereby increasing the negative sample data required for contrastive learning.
[0072] In step 3, the search term is obtained, and text features similar to the search term are searched for in the trained metric space, and the video clip corresponding to the text feature with the highest similarity is taken as the video positioning result.
[0073] Here, searching for text features similar to the search term in the metric space is realized by hash binary code, and the specific steps include the following steps 31, 32 and 33.
[0074] In step 31, the text features in the metric space are converted into hashed binary codes by a hash mapping function, and the hashed binary codes are used as search indexes for text-video search pairs.
[0075] In step 32, the retrieved search phrase is preferentially matched with a predetermined number M of hashed binary codes, where M may be set to ten.
[0076] In step 33, the video clip corresponding to the hash binary code with the highest degree of match is determined as the video clip positioning result, and the start and end times of the video clip that is the positioning target are output as timestamp information.
[0077] In this embodiment, the text hash code is used as an index for the text-video pair, and instead of directly using the new search term feature to calculate the distance in metric space with the video feature, we prioritize searching for text features similar to the new search term, which improves search efficiency and reduces the amount of calculation, resulting in approximately 10 times faster search speed and a 90% reduction in time complexity.
[0078] In this embodiment, considering the problem of the speed of location, a measure is adopted in which the text hash code is used as an index of the entire text-video, which effectively reduces the amount of calculation and greatly increases the speed of location. [Example]
[0079] Based on Example 1, in this example, a common semantics detection module configured to extract common semantic information between text and video using self-supervised learning from the acquired training dataset, and to obtain fused semantic video features based on the semantic information; a video feature and text feature spatial mapping module configured to perform multi-scale contrastive learning on the fused semantic video features and corresponding text features using a weakly supervised method, determine a spatial mapping relationship between the video features and the text features, map them into a metric space, and use the determined metric space to obtain a trained metric space; The present invention provides a weakly supervised video clip location system based on a large-scale video corpus, the system including: a matching module configured to receive a search term, search for text features similar to the search term in a metric space, and search for video clips that have the highest similarity to the text features in a metric space trained based on the text features to provide a video location result.
[0080] Here, it should be noted that each module in this embodiment corresponds one-to-one to each step in the first embodiment, and the specific implementation steps are the same, so that the repeated explanation will be omitted here. [Example]
[0081] This embodiment provides an electronic device that includes a memory, a processor, and computer commands stored in the memory and executed by the processor, and that completes the steps of the method of embodiment 1 when the computer commands are executed by the processor. [Example]
[0082] This embodiment provides a computer-readable storage medium for storing computer commands that, when executed by a processor, complete the steps of the method of embodiment 1.
[0083] The above is only a preferred embodiment of the present invention, and is not intended to limit the present invention, and those skilled in the art can make various modifications and changes to the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention should be included within the protection scope of the present invention.
Claims
1. The method includes the following steps 1, 2, and 3: In step 1, for the acquired training dataset, common semantic information between text and video is extracted using self-supervised learning, and fused semantic video features are obtained based on the semantic information; Specific steps include the following steps 11 to 16: In step 11, feature extraction is performed on the video data and the search phrases using the backbone network to obtain video features and text features respectively; In step 12, the obtained video features are fused with text features, which are features of the search phrases; The fusion method is to perform addition and dot product operations on the video features and text features, then concatenate them. In step 13, a convolution operation is performed on the fused features; The convolution operation reduces the dimension of the feature data and allows you to obtain more detailed information. In step 14, the video features obtained after convolution are used to predict text features, and self-supervised training is performed using the original text features as training information. In step 15, a reconstruction loss is calculated according to the predicted text features and the teacher information, and a larger weight value is assigned to the video clip whose reconstruction loss is lower than a set value, and the weight value is the reconstruction reward; In step 16, weight the reconstruction reward to the video features obtained after convolution to obtain fused semantic video features; In step 2, multi-scale contrastive learning is performed on the fused semantic video features and corresponding text features using a weakly supervised method to determine the spatial mapping relationship between the video features and the text features, and then mapping them into a metric space to obtain a trained metric space; Specifically, the multi-scale contrastive learning includes clip-level learning and video-level learning. The clip-level learning is to bring video clips similar to the search text closer and to move video clips dissimilar to the search text farther away. The video-level learning is to increase the distance between the video corresponding to the search term and other videos. Video learning can be divided into clip-level and video-level learning. Clip-level learning is learning for detailed features in a video, while video-level learning is learning for the entire video. These two levels of learning can be performed independently. Multi-scale contrastive learning is the key to achieving weakly supervised learning. It allows training on a metric space. The clip-level learning aims to improve the matching score between text features and corresponding video clips. The clip-level learning method includes the following steps 21 to 25: In step 21, the fused semantic video features are fed into a GRU long short-term memory network to obtain time series information of the video, and the obtained time series information is added to the fused semantic video features to obtain high-level video features; In step 22, the calculated match scores for the text features and each high-level video clip feature are input into the locator, and the video clips corresponding to the start and end times whose scores are higher than a set value are labeled as positive samples, and the videos that cannot be labeled are negative samples; Here, the locator can be a combination of a multilayer perceptron and a normalized exponential function, i.e., MLP+softmax, and outputs a confidence score for each text clip-search pair using the normalized function. The formula for the normalization function softmax is as follows: [Equation 9] where p is the confidence score and x i is the output value of the multilayer perceptron for the i-th clip, and k indicates that there are k clips in total. The method for calculating the match between the text features and each high-level video clip feature is as follows: The Tanh function is used to calculate the match between features, i.e., the match r t is calculated by the following formula: [Equation 10] where t represents the time step and q t represents high-level video clip features, and k t represents text features, In step 23, a generative adversarial network (GAN) is set up to generate video clip features similar to the original video through clip-level learning, and these generated video clip features are used as negative samples; In step 24, based on the recognized clip-level positive samples and negative samples, quantify the similarity between the samples by cosine similarity, and maximize the similarity probability between the two positive samples by minimizing the loss function to obtain the loss function; The positive and negative samples in clip-level contrast training are all clip-level samples, i.e., one sample is one video clip. The similarity between sample A and sample B is defined by the following formula: [0011] Calculate the similarity probability between the positive samples, and select one of the positive sample pairs (p i , p j The formula for the probability of [0012] Here, p i , p j is a positive sample, N is the number of all samples, including negative samples, a positive pair is a positive sample, a positive sample, a negative pair is a positive sample, a negative sample, e is the base of the natural logarithm, a mathematical constant, a non-repeating infinite decimal and a transcendental number, approximately 2.71828, To maximize the probability of similarity between two positive samples by minimizing the loss function, we can use the following formula, which is the logarithmic loss: [0013] When contrastive learning was performed using the above method, a metric space was obtained in which positive samples were closer to each other. In step 25, after obtaining the loss function, the clip-level loss function LOSS is calculated by the stochastic gradient descent algorithm. clip After the loss function converges, the optimization can be stopped. After about 1000 iterations, the metric space after clip-level optimization is obtained. The video level learning aims to eliminate interference from other irrelevant video features, and the video level learning method includes the following steps 2-1 to 2-4: In step 2-1, the input video to be subjected to video-level learning is set as a positive sample for video-level contrast learning, and multiple other video samples are randomly selected as negative samples; We can randomly select the other 100 video samples as negative samples, In step 2-2, based on the recognized video-level positive samples and negative samples, the similarity between the samples is quantified by cosine similarity, and the similarity probability between the two positive samples is maximized by minimizing the loss function, thereby obtaining the loss function; One sample among the video-level positive and negative samples is a complete video, Quantifying the similarity between video level samples, i.e., the similarity between video level sample A and video level sample B, by cosine similarity is defined as follows: [0014] Calculate the similarity probability between the positive samples, and select one of the positive sample pairs (p i , p j The formula for the probability of [Equation 15] Here, p i , p j is a positive sample, N is the total number of samples, including negative samples, a positive pair is a positive sample, a positive sample, a negative pair is a positive sample, a negative sample, To maximize the probability of similarity between two positive samples by minimizing the loss function, we can use the following formula, which is the logarithmic loss: [0016] When contrastive learning was performed using the above method, a metric space was obtained in which the distance between different videos was greater. In step 2-3, after obtaining the loss function, we use the stochastic gradient descent algorithm to calculate the video-level loss function LOSS video After the loss function converges, the optimization is stopped. After about 1000 iterations, the loss can be converged. In step 2-4, we optimize the loss function to obtain the metric space after video-level optimization. In step 3, a search term is obtained, text features similar to the search term are searched for in the trained metric space, and the video clip corresponding to the text feature with the highest similarity is determined as the video location result.
2. The method for weakly supervised localization of video clips based on a large-scale video corpus according to claim 1, characterized in that the obtained video features and text features are fused, and the fusion method is to perform addition and dot product operations on the video features and text features before concatenation.
3. Searching for text features similar to the search term in the metric space is achieved by hashing binary codes, and the steps are: Step 31: converting the text features in the metric space into hashed binary codes by a hash mapping function, and using the hashed binary codes as search indexes for text-video search pairs; a step 32 of preferentially matching the obtained search phrase with a predetermined number of hashed binary codes; and determining the video clip corresponding to the hashed binary code with the highest degree of match as the location result (33).
4. A system for weakly supervised localization of video clips based on a large-scale video corpus, which is implemented using the method for weakly supervised localization of video clips based on a large-scale video corpus according to claim 1, comprising: a common semantics detection module configured to extract common semantic information between text and video using self-supervised learning from the acquired training dataset, and to obtain fused semantic video features based on the semantic information; a video feature and text feature spatial mapping module configured to perform multi-scale contrastive learning on the fused semantic video features and corresponding text features using a weakly supervised method, determine a spatial mapping relationship between the video features and the text features, map them into a metric space, and use the determined metric space to obtain a trained metric space; a matching module configured to obtain a search term, search for text features similar to the search term in a metric space, and search for video clips that have the highest similarity to the text features in a metric space trained based on the text features to provide a video location result;
5. 10. An electronic device comprising: a memory; a processor; and computer commands stored in the memory and executed by the processor, the computer commands completing the steps of the method for weakly supervised localization of video clips based on a large-scale video corpus according to any one of claims 1 to 3 when executed by the processor.
6. A computer-readable storage medium having stored thereon computer instructions that, when executed by a processor, complete the steps of the method for weakly supervised localization of video clips based on a large video corpus according to any one of claims 1-3.
Citation Information
Patent Citations
Multimedia data searching method and device, equipment and storage medium
CN113590850A