Video Clip Localization Method, Device, Equipment and Storage Medium
By jointly training the identification model and the positioning model, the prediction results of sample data are used to improve the accuracy of video clip positioning, and the problems of complex video retrieval and inaccurate positioning in the prior art are solved.
Patent Information
- Application Number
- CN202111155107.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-29
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2041-09-29
AI Technical Summary
When performing video retrieval based on text, the prior art needs to train the recognition model and the positioning model separately, and the search process is complex, which affects the accuracy of video clip positioning.
A video clip positioning method is provided. By determining the sample results, predicted recognition results and predicted positioning results of sample data, the identification model and the positioning model are trained respectively to improve the accuracy of the model.
The accuracy of the identification model and positioning model is improved, thereby improving the accuracy of video clip positioning and simplifying the video retrieval process.
Smart Images

Figure CN113918767B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of video processing, and particularly to a method, apparatus, device, and storage medium for video segment localization. Background Art
[0002] With the popularization of video applications, more and more videos appear on the network. To facilitate users in finding videos of interest, videos can be retrieved based on text input by users that can describe a certain video segment in the video, so as to find video segments matching the text.
[0003] Currently, before retrieving based on text, an identification model and a localization model need to be trained separately; when retrieving based on text, the identification model is used to screen out videos matching the text from a large number of videos; and then the localization model is used to locate video segments matching the text from the screened-out videos. Summary of the Invention
[0004] Embodiments of this application provide a method, apparatus, device, and storage medium for video segment localization, which can improve the accuracy of the identification model and the localization model and the accuracy of video segment localization. The technical solution is as follows:
[0005] On the one hand, a method for video segment localization is provided, and the method includes:
[0006] Determine the sample result, predicted identification result, and predicted localization result of the sample data, where the sample data includes sample text and a sample video, the predicted identification result is obtained by processing the sample data through an identification model to be trained, and is used to represent whether the sample video and the sample text match, and the predicted localization result is obtained by processing the sample data through a localization model to be trained, and is used to represent the video segment in the sample video that matches the sample text;
[0007] Based on the sample result, the predicted identification result, and the predicted localization result, train the identification model and the localization model respectively to obtain a trained identification model and a trained localization model;
[0008] When obtaining text for retrieving a video segment, determine a target video in candidate videos that matches the text through the trained identification model, and determine a target video segment in the target video that matches the text through the trained localization model.
[0009] In a possible implementation manner, the identification model includes a static feature extraction layer, a dynamic feature extraction layer, and a prediction layer. The step of determining a target video in candidate videos that matches the text through the trained identification model includes:
[0010] Through the static feature extraction layer, static features of multiple words in the text are extracted to obtain word vectors of each word, and static features of multiple video frames in the candidate video are extracted to obtain video frame features of each video frame;
[0011] Through the dynamic feature extraction layer, dynamic features of the word vectors of each word are extracted to obtain text features of the text, and dynamic features of the video frame features of each video frame are extracted to obtain video features of the candidate video;
[0012] Through the prediction layer, based on the correlation between the text features and the video features, a predicted recognition result of the candidate video is determined, and the predicted recognition result indicates whether the candidate video is a target video matching the text.
[0013] In a possible implementation manner, the dynamic feature extraction layer is a bidirectional gating layer. The step of extracting dynamic features of the word vectors of each word through the dynamic feature extraction layer to obtain text features of the text and extracting dynamic features of the video frame features of each video frame to obtain video features of the candidate video includes:
[0014] Through the bidirectional gating layer, features of the word vectors of each word are extracted to obtain the text features;
[0015] Through the bidirectional gating layer, features of the video frame features of each video frame are extracted to obtain the video features.
[0016] In a possible implementation manner, the step of determining a target video segment in the target video that matches the text through the trained localization model includes:
[0017] Through the localization model, the text features of the text are respectively processed with the video frame features of each video frame in the target video to obtain a Hadamard matrix, and each column vector of the Hadamard matrix is the Hadamard product of the text feature and a video frame feature;
[0018] Through the localization model, the Hadamard matrix is predicted to obtain the matching degree between each video frame and the text;
[0019] Through the localization model, based on the matching degree between each video frame and the text, a target video segment in the target video that matches the text is determined.
[0020] On the one hand, a video segment localization device is provided, and the device includes:
[0021] A determination module, configured to determine a sample result, a prediction recognition result, and a prediction localization result of sample data, where the sample data includes sample text and a sample video, the prediction recognition result is obtained by processing the sample data through a recognition model to be trained, and is used to indicate whether the sample video and the sample text match, and the prediction localization result is obtained by processing the sample data through a localization model to be trained, and is used to indicate a video segment in the sample video that matches the sample text;
[0022] A training module, configured to train the recognition model and the localization model respectively based on the sample result, the prediction recognition result, and the prediction localization result, to obtain a trained recognition model and a trained localization model;
[0023] A retrieval module, configured to, when obtaining text for retrieving a video segment, determine a target video in candidate videos that matches the text through the trained recognition model, and determine a target video segment in the target video that matches the text through the trained localization model.
[0024] In a possible implementation manner, the sample result includes a sample localization result, and the training module includes:
[0025] A first determination unit, configured to determine a first loss value of the recognition model based on the prediction recognition result;
[0026] A second determination unit, configured to determine a second loss value of the localization model based on the sample localization result and the prediction localization result;
[0027] A training unit, configured to train the recognition model and the localization model respectively based on the first loss value and the second loss value, to obtain the trained recognition model and the trained localization model.
[0028] In a possible implementation manner, the first determination unit is configured to process the prediction recognition result through a first loss function of the recognition model to obtain the first loss value;
[0029] The second determination unit is configured to process the sample localization result and the prediction localization result through a second loss function of the localization model to obtain the second loss value.
[0030] In a possible implementation manner, the training unit is configured to perform weighted processing on the first loss value and the second loss value to obtain a third loss value; and train the recognition model and the localization model based on the third loss value, to obtain the trained recognition model and the trained localization model.
[0031] In a possible implementation manner, the training module is configured to construct alignment features of the sample data based on the sample text and the sample video, where the alignment features represent features similar to the sample text and the sample video; and train the recognition model and the localization model respectively based on the sample result, the predicted recognition result, the predicted localization result, the alignment features, the text features of the sample text, and the video features of the sample video, to obtain the trained recognition model and the trained localization model.
[0032] In a possible implementation manner, the training module includes:
[0033] The first determination unit is configured to determine a first loss value of the recognition model based on the predicted recognition result;
[0034] The second determination unit is configured to determine a second loss value of the localization model based on the sample localization result and the predicted localization result;
[0035] The third determination unit is configured to determine a fourth loss value of the recognition model based on the alignment features, the text features of the sample text, and the video features of the sample video;
[0036] The training unit is configured to train the recognition model and the localization model respectively based on the first loss value, the second loss value, and the fourth loss value, to obtain the trained recognition model and the trained localization model.
[0037] In a possible implementation manner, the third determination unit is configured to process the alignment features, the text features of the sample text, and the video features of the sample video based on a third loss function, to obtain the fourth loss value, where the third loss function is used to use the alignment features as an anchor point to reduce the distance between the text features of the sample text and the video features of the matching sample video, and increase the distance between the text features of the sample text and the video features of the non-matching sample video.
[0038] In a possible implementation, the sample data is multiple pieces, and the multiple pieces of sample data include positive sample data and negative sample data. The third determination unit is configured to obtain a first similarity between the text feature of the positive sample data and the first alignment feature of the positive sample data; obtain a second similarity between the text feature of the positive sample data and the second alignment feature of the negative sample data; obtain a third similarity between the video feature of the positive sample data and the first alignment feature of the positive sample data; obtain a fourth similarity between the video feature of the positive sample data and the second alignment feature of the negative sample data; and determine the fourth loss value based on the difference between the first similarity and the second similarity, and the difference between the third similarity and the fourth similarity.
[0039] In a possible implementation, the training module is configured to obtain the text feature of the sample text based on the sample text; obtain the video frame features of the multiple video frames based on the multiple video frames of the sample video; and fuse the obtained text feature and the multiple video frame features to obtain the alignment feature.
[0040] In a possible implementation, the training module is configured to obtain the Hadamard matrix of the sample data based on the text feature and each video frame feature, where each column vector of the Hadamard matrix is the Hadamard product of the text feature and a video frame feature; and perform weighted averaging on the column vectors in the Hadamard matrix based on the correlation degree between each video frame and the sample text to obtain the alignment feature.
[0041] On the one hand, a computer device is provided, which includes one or more processors and one or more memories. At least one program code is stored in the one or more memories, and the at least one program code is loaded and executed by the one or more processors to implement the operations performed by the video segment localization method in any of the above possible implementations.
[0042] On the one hand, a computer-readable storage medium is provided, in which at least one program code is stored, and the at least one program code is loaded and executed by a processor to implement the operations performed by the video segment localization method in any of the above possible implementations.
[0043] On the one hand, a computer program or a computer program product is provided, where the computer program or the computer program product includes: computer program code, and when the computer program code is executed by a computer, the computer is caused to implement the operations performed by the video segment localization method in any of the above possible implementations.
[0044] The beneficial effects brought by the technical solution provided in the embodiments of the present application at least include:
[0045] In the video clip localization method, device, equipment, and storage medium provided by the embodiments of the present application, sample results, predicted recognition results, and predicted localization results are all used when training the recognition model and the localization model, enabling the recognition model and the localization model to learn not only information about video recognition but also information about video clip localization. Since the recognition model and the localization model learn more information, the accuracy of the recognition model and the localization model is improved, and thus the accuracy of video clip localization is enhanced. Description of the Drawings
[0046] To more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for description in the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0047] Figure 1 It is a schematic diagram of an implementation environment provided by the embodiments of the present application;
[0048] Figure 2 It is a flowchart of a video clip localization method provided by the embodiments of the present application;
[0049] Figure 3 It is a flowchart of a video clip localization method provided by the embodiments of the present application;
[0050] Figure 4 It is a processing flowchart of a recognition model and a localization model provided by the embodiments of the present application;
[0051] Figure 5 It is a flowchart of a video clip localization method provided by the embodiments of the present application;
[0052] Figure 6 It is a schematic structural diagram of a video clip localization device provided by the embodiments of the present application;
[0053] Figure 7 It is a schematic structural diagram of another video clip localization device provided by the embodiments of the present application;
[0054] Figure 8 It is a schematic structural diagram of a terminal provided by the embodiments of the present application;
[0055] Figure 9 It is a schematic structural diagram of a server provided by the embodiments of the present application. Detailed Embodiments
[0056] To make the objectives, technical solutions, and advantages of this application clearer, the following will further describe the embodiments of this application in detail with reference to the accompanying drawings.
[0057] It can be understood that the terms "first", "second", etc. used in this application may be used herein to describe various concepts, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of this application, the first sample data may be referred to as the second sample data, and similarly, the second sample data may be referred to as the first sample data.
[0058] The terms "at least one", "multiple", "each", and "any one" used in this application, at least one includes one, two, or more than two, multiple includes two or more than two, and each refers to each of the corresponding multiple, and any one refers to any one of the multiple. For example, multiple sample data includes 3 sample data, and each refers to each of these 3 sample data, and any one refers to any one of these 3 sample data, which can be the first, the second, or the third.
[0059] The video clip localization method provided by the embodiments of this application is executed by a computer device. In one possible implementation, the computer device is a terminal. Optionally, the terminal is any type of terminal such as a desktop computer, a tablet computer, or a mobile phone. In another possible implementation, the computer device is a server. Optionally, the server is a single server, or a server cluster composed of several servers, or a cloud computing service center. In another possible implementation, the computer device includes a terminal and a server.
[0060] Figure 1 is a schematic diagram of an implementation environment provided by the embodiments of this application, as Figure 1 shown, this implementation environment includes a terminal 101 and a server 102. The terminal 101 and the server 102 are connected through a wireless or wired network.
[0061] The target application provided by the server 102 is installed on the terminal 101, and the terminal 101 can implement functions such as data transmission and message interaction through this target application. Optionally, the terminal 101 is a computer, a mobile phone, a tablet computer, or other terminal. Optionally, the target application is a target application in the operating system of the terminal 101, or a target application provided by a third party. For example, the target application is a video processing application, and this video processing application has the function of processing videos. Of course, this video processing application can also have other functions, such as a sharing function, a video playback function, a game function, etc. Optionally, the server 102 is the background server of this target application or a cloud server that provides services such as cloud computing and cloud storage.
[0062] Optionally, the terminal 101 obtains the text input by the user, sends the text to the server 102, the server 102 filters out the target video matching the text from the video library based on the text, then determines the target video segment corresponding to the text from the filtered target videos, and returns the filtered target video and the target video segment to the terminal 101, and the terminal displays the target video and the target video segment.
[0063] The video segment positioning method provided by the embodiments of the present application can be applied to scenarios for processing videos.
[0064] For example, it is applied to the video retrieval scenario.
[0065] The user inputs text in the client. If the video segment positioning method provided by the embodiments of the present application is adopted, the video matching the text can be accurately retrieved, and the video segment matching the text can be located in the video, and the retrieved video and the located video segment are displayed to the user, so that the user can find the video they want to watch according to the located video segment.
[0066] It should be noted that the embodiments of the present application only take the video retrieval scenario as an example to exemplarily illustrate the scenarios for video processing, and do not limit the scenarios for video processing.
[0067] Figure 2 It is a flowchart of a video segment positioning method provided by the embodiments of the present application. The embodiments of the present application are exemplarily described by taking the execution entity as a computer device. The embodiment includes:
[0068] 201. The computer device determines the sample result, the prediction recognition result, and the prediction positioning result of the sample data. The sample data includes the sample text and the sample video. The prediction recognition result is obtained by processing the sample data by the recognition model to be trained, and is used to indicate whether the sample video and the sample text match. The prediction positioning result is obtained by processing the sample data by the positioning model to be trained, and is used to indicate the video segment in the sample video that matches the sample text.
[0069] Among them, the sample text and the sample video in the sample data may or may not match. The embodiments of the present application do not limit this. For example, if the sample data is positive sample data, the sample text and the sample video match; if the sample text is negative sample data, the sample text and the sample video do not match.
[0070] The sample result is obtained by annotating the sample video based on the sample text. Optionally, the sample result includes a sample recognition result. The sample recognition result is used to indicate whether the sample video is a video that matches the sample text. Among them, the sample recognition result can be in any form. Optionally, the sample recognition result is 0 or 1, where 0 indicates that the sample video is not a video that matches the sample text; 1 indicates that the sample video is a video that matches the sample text. Optionally, the sample recognition result is the matching degree between the sample video and the sample text. If the matching degree is greater than the first threshold, it indicates that the sample video is a video that matches the sample text; if the matching degree is not greater than the first threshold, it indicates that the sample video is not a video that matches the sample text.
[0071] Optionally, the sample result includes a sample localization result. The sample localization result is obtained by annotating the video segment in the sample video that matches the sample text. If the sample video matches the sample text, then the sample localization result is used to indicate the video segment in the sample video that matches the sample text; if the sample video does not match the sample text, then the sample localization result is used to indicate that the sample video does not include a video segment that matches the sample text. The sample localization result can be in any form. Optionally, the sample localization result can include the start time and end time of the video segment in the sample video that matches the sample text. Optionally, if the sample video does not include a video segment that matches the sample text, both the start time and end time in the sample localization result are 0.
[0072] The recognition model is a model used to identify whether a video and text match. In the embodiments of the present application, the computer device processes the sample text and the sample video through the recognition model to obtain a predicted recognition result, and the predicted recognition result is used to indicate whether the sample video is a video that matches the sample text.
[0073] The localization model is a model used to locate the video segment that matches the text from the video. In the embodiments of the present application, the computer device processes the sample text and the sample video through the localization model to obtain a predicted localization result, and the predicted localization result is used to indicate the video segment in the sample video that matches the sample text.
[0074] The recognition model to be trained and the localization model to be trained can be models that have not been trained, models that have been trained but not yet completed, or models that have been trained and completed and are trained again to improve the accuracy of the models. For example, in the embodiments of the present application, the recognition model to be trained and the localization model to be trained are models that have been trained and put into use. The recognition model and the localization model that have been put into use are used as the models to be trained again, and the method provided in the embodiments of the present application is used to train the models to be trained.
[0075] 202. The computer device trains the recognition model and the localization model respectively based on the sample result, the predicted recognition result, and the predicted localization result, to obtain the trained recognition model and the trained localization model.
[0076] Among them, the computer device trains the recognition model and the localization model respectively based on the sample result, the predicted recognition result, and the predicted localization result, to obtain the trained recognition model and the trained localization model means that: the computer device trains the recognition model based on the sample result, the predicted recognition result, and the predicted localization result, to obtain the trained recognition model; the computer device trains the localization model based on the sample result, the predicted recognition result, and the predicted localization result, to obtain the trained localization model.
[0077] During the process of retrieving based on a file, the recognition model completes the retrieval task, and the localization model completes the localization task. Among them, the recognition model is used to retrieve videos matching the text from the video library, and the localization model is used to locate video segments matching the text from the retrieved videos. In the embodiments of the present application, based on the sample result, the predicted recognition result, and the predicted localization result, the recognition model and the localization model are trained respectively, so that the recognition model can also learn information about the localization task, and the localization model can also learn information about the retrieval task. Since the recognition model and the localization model learn more information, the accuracy of the recognition model and the localization model is improved.
[0078] 203. When obtaining the text for retrieving a video segment, the computer device determines the target video matching the text in the candidate videos through the trained recognition model, and determines the target video segment matching the text in the target video through the trained localization model.
[0079] After training the recognition model and the localization model, based on the input text, the target video matching the text and the target segment matching the text in the target video can be determined through the trained recognition model and the localization model.
[0080] In the video segment localization method provided by the embodiments of the present application, when training the recognition model and the localization model, the sample result, the predicted recognition result, and the predicted localization result are all used, so that the recognition model and the localization model can not only learn information about video recognition, but also learn information about video segment localization. Since the recognition model and the localization model learn more information, the accuracy of the recognition model and the localization model is improved, and thus the accuracy of video segment localization is improved.
[0081] Figure 3 It is a flowchart of a video segment localization method provided by the embodiments of the present application. The embodiments of the present application are exemplarily described by taking the execution entity as a terminal. This embodiment includes:
[0082] 301. The computer device obtains sample data and the sample results of the sample data. The sample data includes sample text and sample video, and the sample results are obtained by annotating the video segments in the sample video that match the sample text.
[0083] The sample data in step 301 is the same as the sample data in step 201, and the sample results in step 301 are the same as the sample results in step 201, which will not be elaborated here one by one.
[0084] In a possible implementation, there is one piece of sample data. The computer device obtains one piece of sample data each time, and processes the sample data through the recognition model to be trained and the localization model to be trained; trains the recognition model and the localization model based on the processing results and the sample results of the sample data. By obtaining sample data and training multiple times, the accuracy of the recognition model and the localization model is continuously improved.
[0085] In another possible implementation, there are multiple pieces of sample data, and the multiple pieces of sample data include positive sample data and negative sample data. Among them, the positive sample data includes sample text and sample video that match each other, and the negative sample data includes sample text and sample video that do not match each other. The computer device obtains multiple pieces of sample data each time, and processes the multiple pieces of sample data through the recognition model to be trained and the localization model to be trained in sequence to obtain the processing results of each piece of sample data, and trains the recognition model and the localization model based on the processing results of the positive sample data and the negative sample data among the multiple pieces of sample data. By obtaining multiple pieces of sample data and training multiple times, the accuracy of the recognition model and the localization model is continuously improved.
[0086] 302. The computer device processes the sample data through the recognition model to be trained to obtain a predicted recognition result, and the predicted recognition result is used to indicate whether the sample video matches the sample text.
[0087] If the sample video matches the sample text, it means that the content described by the sample video is similar to the content described by the sample text, or the content described by the sample video includes the content described by the sample text. Therefore, it is possible to determine whether the sample video matches the sample text through the video features of the sample video and the text features of the sample text. In a possible implementation, the computer device processes the sample data through the recognition model to be trained to obtain a predicted recognition result, including: extracting the text features of the sample text through the recognition model; extracting the video features of the sample video through the recognition model; processing the text features and the video features through the recognition model to obtain a predicted recognition result.
[0088] Among them, after obtaining the text features and video features, the recognition model can determine the predicted recognition result based on the similarity between the text features and the video features. For example, directly determine the similarity between the text features and the video features as the predicted recognition result. Optionally, in order to more accurately determine the similarity between the text features and the video features, the text features and the video features can be mapped into the same feature space, and then the similarity between the text features and the video features is determined, as Figure 4 shown.
[0089] In the embodiments of the present application, the video can describe a dynamic information. For example, if the sample video is a video of a child sliding down a slide, then the sample video is used to describe the dynamic information of the child sliding down the slide. The text is used to describe the content in the video that the user wants to watch, usually describing an event. Therefore, the text is also used to describe dynamic information. In order to more accurately obtain the text features and video features, the embodiments of the present application adopt a dynamic feature extraction method when extracting features from the sample text and the sample video. In one possible implementation, the computer device processes the sample data through the recognition model to be trained to obtain the predicted recognition result, including: performing dynamic feature extraction on the sample text through the recognition model to obtain the text features of the sample text; performing dynamic feature extraction on the sample video through the recognition model to obtain the video features of the sample video; and processing the text features and the video features through the recognition model to obtain the predicted recognition result.
[0090] Among them, the sample text includes multiple words, and the multiple words constitute the dynamic information of the sample text. Therefore, when performing dynamic feature extraction on the sample text, the sample text can be first divided into multiple words, the word vectors of each word are obtained in sequence, and then the dynamic features of the sample text are extracted according to the word vectors of adjacent words. The sample video includes multiple video frames, each video frame contains static image information, and consecutive video frames include motion information. Therefore, feature extraction can be performed on the video frames of the sample video, and then the motion information of the sample video is extracted based on the video frame features of adjacent video frames.
[0091] In a possible implementation, the recognition model includes a static feature extraction layer, a dynamic feature extraction layer, and a prediction layer; the computer device processes the sample data through the recognition model to be trained to obtain a predicted recognition result, including: through the static feature extraction layer, performing static feature extraction on multiple words in the sample text to obtain a word vector for each word, and performing static feature extraction on multiple video frames in the sample video to obtain a video frame feature for each video frame; through the dynamic feature extraction layer, performing dynamic feature extraction on the word vector of each word to obtain a text feature of the sample text, and performing dynamic feature extraction on the video frame feature of each video frame to obtain a video feature of the sample video; through the prediction layer, based on the correlation between the text feature and the video feature, obtaining a predicted recognition result, where the predicted recognition result indicates whether the sample video is a video matching the sample text.
[0092] Among them, the multiple video frames in the sample video can be all the video frames of the sample video or some video frames. Optionally, since the difference between two adjacent frames in a video is small, to reduce the amount of calculation, the computer device extracts key frames from the sample video, performs static feature extraction on the multiple key frames through the static feature extraction layer to obtain a video frame feature for each key frame; through the dynamic feature extraction layer, performs dynamic feature extraction on the video frame feature of each key frame to obtain a video feature of the sample video. Among them, extracting key frames from the sample video may include: extracting a video frame every first time period in the sample video and using the extracted video frame as a key frame; extracting a video frame every first number of video frames in the sample video and using the extracted video frame as a key frame. Among them, the first time period can be any time period, for example, 1 second, 0.5 second, etc. The first number can be any number, for example, 30, 50, etc.
[0093] Optionally, the dynamic feature extraction layer is a memory network. In this way, when processing the input information, it can refer to the previously input information to obtain dynamic features. Optionally, the dynamic feature extraction layer is a gating layer. The computer device performs dynamic feature extraction on the word vector of each word through the dynamic feature extraction layer to obtain the text feature of the sample text, including: inputting the first word vector of the sample text into the gating layer, processing the first word vector through the gating layer, and outputting a first feature; inputting the second word vector into the gating layer, processing the second word vector based on the first feature, and outputting a second feature; inputting the third word vector into the gating layer, processing the third word vector based on the second feature to obtain a third feature; sequentially inputting the word vectors of the sample text until the last word vector of the sample text is input into the gating layer, and processing the last word vector through the gating layer based on the previously output feature to output the text feature of the sample text.
[0094] It should be noted that when dynamically extracting the feature vectors of the sample text, referring to the feature vectors before the feature vector of interest can improve the accuracy of the dynamic feature extraction of this feature vector. However, referring to the feature vectors before and after the feature vector of interest can further improve the accuracy of the dynamic feature extraction of this feature vector. Optionally, the dynamic feature extraction layer is a bidirectional gating layer. The computer device dynamically extracts features from the feature vectors of each word through the dynamic feature extraction layer to obtain the text features of the sample text, and dynamically extracts features from the video frame features of each video frame to obtain the video features of the sample video, including: extracting features from the feature vectors of each word through the bidirectional gating layer to obtain text features; extracting features from the video frame features of each video frame through the bidirectional gating layer to obtain video features.
[0095] It should be noted that the embodiments of this application only use the dynamic features of the sample text as an example for illustrative purposes of the text features of the sample text. In another embodiment, in order to fully explore tiny short-time feature patterns and make the expression ability of the features stronger, we continue to extract features from the extracted dynamic features to obtain sample features and video features. In one possible implementation, the sample data is processed through a recognition model to be trained to obtain a predicted recognition result, including: statically extracting features from multiple words in the sample text through a static feature extraction layer to obtain the feature vectors of each word, and statically extracting features from multiple video frames in the sample video to obtain the video frame features of each video frame; dynamically extracting features from the feature vectors of each word through a dynamic feature extraction layer to obtain the dynamic features of these multiple words, statically extracting features from the dynamic features of these multiple words to obtain the text features of the sample text, dynamically extracting features from each video frame to obtain the dynamic features of these multiple video frames, and statically extracting features from the dynamic features of these multiple video frames to obtain the video features of the sample video.
[0096] In another embodiment, to enrich the information of text features and video features, the static features and dynamic features of the sample text can be fused to obtain the text features of the sample text; the static features and dynamic features of the sample video can be fused to obtain the video features of the sample video. In a possible implementation, the sample data is processed by a recognition model to be trained to obtain a predicted recognition result, including: through a static feature extraction layer, static features of multiple words in the sample text are extracted to obtain the static features of each word, and static features of multiple video frames in the sample video are extracted to obtain the static features of each video frame; through a dynamic feature extraction layer, dynamic features of the static features of each word are extracted to obtain the dynamic features of each word, and dynamic features of the static features of each video frame are extracted to obtain the dynamic features of each video frame; based on the static features and dynamic features of each word, the text features of the sample text are obtained; based on the static features and dynamic features of each video frame, the video features of the sample video are obtained.
[0097] Among them, obtaining the text features of the sample text based on the static features and dynamic features of each word can be to splice the static features and dynamic features of each word, or to fuse the static features and dynamic features of each word. The embodiments of the present application do not limit this. In the embodiments of the present application, the process of obtaining the video features of the sample video based on the static features and dynamic features of each video frame is the same as the process of obtaining the text features of the sample text based on the static features and dynamic features of each word, and will not be elaborated here one by one.
[0098] It should be noted that the embodiments of the present application only take the recognition model including a static feature extraction layer and a dynamic feature extraction layer as an example to exemplarily illustrate the process of obtaining sample features and video features. Among them, the process of obtaining sample features can be the same as or different from the process of obtaining video features. The embodiments of the present application do not limit this.
[0099] In a possible implementation, the static feature extraction layer includes a text processing sub-layer and a video processing sub-layer, and the dynamic feature extraction layer includes a text processing sub-layer and a video processing sub-layer. Among them, the text processing sub-layer of the static feature extraction layer is connected to the text processing sub-layer of the dynamic feature extraction layer; the video processing sub-layer of the static feature extraction layer is connected to the video processing sub-layer of the dynamic feature extraction layer. The processing processes of the sample text and the sample video do not interfere with each other.
[0100] For example, a computer device decomposes a sample video into a sequence of key frames, and extracts first key frame features (static features) of each key frame in the key frame sequence through a static feature extraction layer (such as a convolutional neural network, etc.). The obtained key frame features are input into a bidirectional gating layer (dynamic feature extraction layer) to capture the temporal information between consecutive key frames, and second key frame features (dynamic features) of each key frame are obtained. 1D convolutions with different receptive fields (convolution kernel sizes are {3, 5, 7} respectively) are used to perform convolution processing on the output result of the bidirectional gating layer, and third key frame features of multiple key frames are obtained. The third key frame features of each key frame are concatenated together, denoted as {v1, v2, … v n}, where v i represents the third key frame feature of the i-th key frame. The third key frame features of each key frame are subjected to weighted average pooling to obtain the video feature of the sample video:
[0101]
[0102] where n is the number of key frames, w i is the weight of the i-th key frame, and v i is the third key frame feature of the i-th key frame. Among them, w i is a parameter to be learned by the recognition model, and w i = sigmoid(w T v i ). Among them, sigmoid is an activation function.
[0103] Another example is that a computer device divides a sample text into multiple words, and in the way of word2vec, converts the multiple words into multiple word vectors. The obtained word vectors are processed through a bidirectional gating layer to obtain the first text feature f t (1) = pooling(h1,..., h l ). Among them, l is the number of multiple words, and h i is the output of the i-th step of the bidirectional gating layer. In addition, the computer device further performs convolution processing on the output result of each step of the bidirectional gating layer by using a 1D convolutional neural network with convolution kernel sizes of {1, 3, 5} respectively, and fuses the features obtained after convolution through max pooling to obtain the second text feature f t (2) . The first text feature and the second text feature of the sample text are concatenated to obtain the text feature t of the sample text = (f t (1) , f t (2) ).
[0104] 303. The computer device processes the sample data through the positioning model to be trained, and obtains a predicted positioning result, which is used to represent the video segment in the sample video that matches the sample text.
[0105] In the process of the computer device processing the sample data through the recognition model to be trained, the video frame features of multiple video frames in the sample video and the text features of the sample text are obtained. Therefore, the positioning model can directly use the multiple video frame features and text features obtained by the recognition model, or reprocess the sample video and the sample text to obtain multiple video frame features and text features. The embodiments of the present application do not make any limitations in this regard.
[0106] In the embodiments of the present application, for each video frame, the positioning model determines the video segment in the sample video that matches the sample text by determining whether the content in the video frame matches the sample text. In a possible implementation manner, the computer device processes the sample data through the positioning model to be trained, and obtains a predicted positioning result, including: through the positioning model, fusing each video frame feature with the text feature respectively, and using the attention mechanism to determine the response value of the sample text and each video frame; based on the response value of the sample text and each video frame, determining the predicted positioning result. Among them, the attention mechanism uses the start and end time of the action as the supervision information. The start and end time of the action is the sample positioning result of the sample data.
[0107] Optionally, if the predicted positioning result is the start time and end time of the video segment in the sample video that matches the sample text, then the video frame time of the first video frame corresponding to the high response value is used as the start time, and the video frame time of the last video frame corresponding to the high response value is used as the end time. Optionally, the computer device extracts the video frames with high response values to generate a video segment, and this video segment is the video segment in the sample video that matches the sample text.
[0108] In a possible implementation manner, when the positioning model processes the video frame features and text features, it can extract the similar features between the two from the video frame features and text features, and determine whether the content in the video frame matches the sample text according to the extracted similar features. Optionally, the computer device, through the positioning model, fuses each video frame feature with the text feature respectively, and uses the attention mechanism to determine the response value of the sample text and each video frame; based on the response value of the sample text and each video frame, determining the predicted positioning result, including: fusing each video frame feature with the text feature respectively to obtain an alignment feature, which represents the similar features of the sample text and the sample video, and the alignment feature includes the similar features of each video frame in the sample video and the sample text; determining the response value of each video frame based on the alignment feature; and determining the predicted positioning result based on the response value of each video frame.
[0109] Among them, the alignment feature is A = [a1,..., a n , where a1 represents the similarity feature between the first video frame and the sample text, and n represents the number of video frames of the sample text. The response value of each video frame is s = (s1,..., s n ). Among them, s i is the response value corresponding to the i-th video frame. Among them, s i = sigmoid(u T a i ). sigmoid is the activation function; u is a parameter to be learned.
[0110] Optionally, the computer device can obtain the alignment feature by obtaining the Hadamard product of the text feature and the key frame feature. a i = t ⊙ v i ; where t is the text feature of the sample text, and v i is the video frame feature of the i-th video frame; a i represents the Hadamard product of the i-th video frame and the sample text.
[0111] In a possible implementation, the computer device processes the sample data through a positioning model to be trained to obtain a predicted positioning result, including: through the positioning model, processing the text feature of the sample text and the video frame feature of each video frame in the sample video respectively to obtain a Hadamard matrix, and each column vector of the Hadamard matrix is the Hadamard product of the text feature and a video frame feature; through the positioning model, predicting the Hadamard matrix to obtain the matching degree of each video frame and the sample text; through the positioning model, based on the matching degree of each video frame and the sample text, determining the predicted positioning result, and the predicted positioning result is used to indicate the video segment in the sample video that matches the sample text.
[0112] Among them, the positioning model determines the predicted positioning result based on the matching degree of each video frame and the sample text, which can be: the positioning model determines the video frame corresponding to the matching degree greater than the first threshold as the video frame of the video segment that matches the sample text. Among them, the first threshold can be any value, for example, 0.4, 0.5, 0.6, etc. Among them, the first threshold can be set by those skilled in the art or can be usually calculated, and the embodiments of the present application do not limit the first threshold.
[0113] It should be noted that the matching degree sequence between the video frame and the sample text usually shows a unimodal shape. Therefore, the first threshold can be determined by the maximum matching degree in the matching degree sequence. The maximum matching degree in the matching degree sequence is V max , and the first threshold is set to δ = γV max , where γ is an empirical value.
[0114] 304. The computer device determines a first loss value of the recognition model based on the predicted recognition result.
[0115] The first loss value is the loss value of the recognition model and can be obtained from the predicted recognition result determined by the recognition model. Optionally, the first loss value of the recognition model is determined based on the predicted recognition result of the recognition model and the sample recognition result of the sample data, and this first loss value represents the difference between the predicted recognition result and the sample recognition result. Optionally, the recognition model is trained with positive sample data and negative sample data, and the recognition model is trained with the predicted recognition results of the positive sample data and the predicted recognition results of the negative sample data.
[0116] In a possible implementation manner, the first loss value of the recognition model is determined by the predicted recognition result and the sample recognition result. The computer device determines the first loss value of the recognition model based on the predicted recognition result, including: the sample data includes the sample recognition result, and the first loss value of the recognition model is determined based on the sample recognition result and the predicted recognition result.
[0117] In another possible implementation manner, the first loss value of the recognition model is determined by the predicted recognition results of the positive sample data and the negative sample data. The computer device determines the first loss value of the recognition model based on the predicted recognition result, including: there are multiple pieces of sample data, and the multiple pieces of sample data include positive sample data and negative sample data, and the first loss value of the recognition model is determined based on the predicted recognition results of the positive sample data and the predicted recognition results of the negative sample data.
[0118] It should be noted that when the recognition model is trained alone, the first loss function is used for training. When the recognition model is trained together with the positioning model, the first loss function can also be used to process the predicted recognition result to obtain the first loss value of the recognition model. In a possible implementation manner, determining the first loss value of the recognition model based on the predicted recognition result includes: processing the predicted recognition result through the first loss function of the recognition model to obtain the first loss value.
[0119] It should be noted that in the embodiments of the present application, the first loss function can be any loss function, and the embodiments of the present application do not limit the first loss function. Taking the first loss function as the triplet loss function as an example, the embodiments of the present application exemplarily illustrate the process of obtaining the first loss value.
[0120] For example, the sample data includes positive sample data, first negative sample data, and second negative sample data; the positive sample data includes positive sample text (a group of puppies sliding down the slide with their owner) and positive sample video (a video of a group of puppies sliding down the slide with their owner); the first negative sample data includes positive sample text and negative sample video; the second negative sample data includes negative sample text and positive sample video, and the information described by the negative sample text is different from the information described by the positive sample text. After processing one positive sample data, multiple first negative sample data, and multiple second negative sample data through the recognition model, the similarity between the sample text and the sample video in each sample data is determined, and the difficult negative sample data is determined. Based on the triplet loss function and the difficult negative sample data, the recognition model is trained. Among them, the difficult negative sample data refers to the situation where the sample text and the video text in the sample data do not match, but have a relatively high similarity.
[0121] The triplet loss function is
[0122] where L tri is the first loss value of the recognition model, α is an empirical value used to widen the similarity between the sample text and the sample video in the positive sample data and the similarity between the sample text and the sample video in the negative sample data. S(v, t - ) represents the maximum similarity between the negative sample text and the positive sample video in the second negative sample data, S(v, t) represents the similarity between the positive sample text and the positive sample video in the positive sample data, and S(v - , t) represents the maximum similarity between the positive sample text and the negative sample video in the first negative sample data.
[0123] 305. The computer device determines the second loss value of the positioning model based on the sample positioning result and the predicted positioning result.
[0124] The second loss value is the loss value of the positioning model. In the embodiments of the present application, any loss function can be used to process the sample positioning result and the predicted positioning result to obtain the second loss value of the positioning model.
[0125] It should be noted that when the positioning model is trained alone, the second loss function is used for training. When the positioning model and the recognition model are trained together, the second loss function can also be used to process the predicted positioning result to obtain the second loss value of the positioning model. In a possible implementation manner, determining the second loss value of the positioning model based on the sample positioning result and the predicted positioning result includes: processing the sample positioning result and the predicted recognition positioning result through the second loss function of the positioning model to obtain the second loss value.
[0126] It should be noted that the sample positioning result and the predicted positioning result can be the start and end times of the video segment in the sample video that matches the sample text. When training the positioning model based on the sample positioning result and the predicted positioning result, the sample positioning result can be converted into a sequence of matching degrees corresponding to the video frames, and based on the sample matching degree sequence and the predicted matching degree sequence, the second loss value of the positioning model can be determined.
[0127] For example, represent the sample positioning result S as a sequence of matching degrees where n represents n video frames. If the video frame time of the i-th video frame is within the start and end times of the sample positioning result, then is 1; if the video frame time of the i-th video frame is not within the start and end times of the sample positioning result, then is 0. Based on the second loss function, and the sequence of matching degrees generated by the positioning model during processing are processed to obtain the second loss value.
[0128] This second loss value
[0129] 306. The computer device trains the recognition model and the positioning model respectively based on the first loss value and the second loss value to obtain the trained recognition model and the trained positioning model.
[0130] Among them, the computer device trains the recognition model and the positioning model respectively based on the first loss value and the second loss value. Through multiple trainings, the first loss value of the recognition model and the second loss value of the positioning model converge.
[0131] In a possible implementation manner, the computer device trains the recognition model and the positioning model respectively based on the sum of the first loss value and the second loss value to obtain the trained recognition model and the trained positioning model. In a possible implementation manner, the computer device trains the recognition model and the positioning model respectively based on the first loss value and the second loss value to obtain the trained recognition model and the trained positioning model, including: obtaining the sum value of the first loss value and the second loss value, and training the recognition model and the positioning model respectively based on this sum value to obtain the trained recognition model and the trained positioning model.
[0132] It should be noted that the first loss value and the second loss value have different impacts on the model accuracy. Therefore, weights can also be assigned to the first loss value and the second loss value, and the weight represents the degree of influence of the loss value on the model accuracy. In one possible implementation, based on the first loss value and the second loss value, the recognition model and the localization model are respectively trained to obtain the trained recognition model and the trained localization model, including: performing weighted processing on the first loss value and the second loss value to obtain a third loss value; based on the third loss value, training the recognition model and the localization model to obtain the trained recognition model and the trained localization model.
[0133] Among them, the weight of the first loss value and the weight of the second loss value can be set by those skilled in the art or determined by calculation, and the embodiments of the present application do not limit this.
[0134] It should be noted that in the embodiments of the present application, taking the example of using the same loss value to train the recognition model and the localization model for exemplary illustration. In another embodiment, different loss values are used to train the recognition model and the localization model. In one possible implementation, based on the first loss value and the second loss value, the recognition model and the localization model are respectively trained to obtain the trained recognition model and the trained localization model, including: obtaining the first weight of the first loss value and the second weight of the second loss value, and performing weighted processing on the first loss value and the second loss value based on the first weight and the second weight to obtain a fifth loss value; obtaining the third weight of the first loss value and the fourth weight of the second loss value, and performing weighted processing on the first loss value and the second loss value based on the third weight and the fourth weight to obtain a sixth loss value; based on the fifth loss value, training the recognition model to obtain the trained recognition model; based on the sixth loss value, training the localization model to obtain the trained localization model.
[0135] For example, when training the recognition model, the weight of the first loss value is greater than the weight of the second loss value. In this way, when training the recognition model based on the weighted loss value, the recognition model mainly learns information about video recognition and secondarily learns information about video segment localization. When training the localization model, the weight of the first loss value is less than the weight of the second loss value. In this way, when training the recognition model based on the weighted loss value, the recognition model mainly learns information about video segment localization and secondarily learns information about video recognition.
[0136] 307. In the case of obtaining the text for retrieving the video segment, the computer device determines the target video in the candidate videos that matches the text through the trained recognition model, and determines the target video segment in the target video that matches the text through the trained localization model.
[0137] Step 307 above is a process of using the recognition model and the localization model to retrieve videos and video clips after the training of the recognition model and the localization model is completed.
[0138] Among them, when the text for retrieving video clips is obtained, the recognition model is first used to perform a preliminary screening on the video, and then the results after the preliminary screening are further screened by the localization model.
[0139] In a possible implementation manner, the recognition model includes a static feature extraction layer, a dynamic feature extraction layer, and a prediction layer. Determining a target video that matches the text in the candidate videos through the trained recognition model includes: through the static feature extraction layer, performing static feature extraction on multiple words in the text to obtain word vectors for each word, and performing feature extraction on multiple video frames in the candidate video to obtain video frame features for each video frame; through the dynamic feature extraction layer, performing dynamic feature extraction on the word vectors of each word to obtain text features of the text, and performing dynamic feature extraction on the video frame features of each video frame to obtain video features of the candidate video; through the prediction layer, based on the correlation degree between the text features and the video features, determining a predicted recognition result of the candidate video, and this predicted recognition result indicates whether the candidate video is the target video that matches the text.
[0140] Optionally, the dynamic feature extraction layer is a bidirectional gating layer; the computer device performs dynamic feature extraction on the word vectors of each word through the dynamic feature extraction layer to obtain text features of the text, and performs dynamic feature extraction on the video frame features of each video frame to obtain video features of the candidate video, including: through the bidirectional gating layer, performing feature extraction on the word vectors of each word to obtain text features; through the bidirectional gating layer, performing feature extraction on the video frame features of each video frame to obtain video features.
[0141] In a possible implementation manner, the computer device determines a target video clip that matches the text in the target video through the trained localization model, including: through the localization model, processing the text features of the text with the video frame features of each video frame in the target video respectively to obtain a Hadamard matrix, and each column vector of this Hadamard matrix is the Hadamard product of the text feature and a video frame feature; through the localization model, predicting the Hadamard matrix to obtain the matching degree between each video frame and the text; through the localization model, based on the matching degree between each video frame and the text, determining the target video clip that matches the text in the target video.
[0142] It should be noted that in the embodiments of the present application, the processing process of determining the target video and the target video clip in the target video through the trained recognition model and localization model is the same as the processing process of the sample data through the recognition model and localization model to be trained, and will not be elaborated here one by one.
[0143] In the video segment localization method provided by the embodiments of the present application, when training the recognition model and the localization model, sample results, predicted recognition results, and predicted localization results are all used, enabling the recognition model and the localization model to not only learn information about video recognition but also learn information about video segment localization. Since the recognition model and the localization model learn more information, the accuracy of the recognition model and the localization model is improved, and thus the accuracy of video segment localization is enhanced.
[0144] Figure 5 FIG. 4 is a flowchart of a method for video segment localization provided by the embodiments of the present application. Taking the execution entity as a terminal as an example for illustrative purposes, this embodiment includes:
[0145] 501. The computer device obtains sample data and the sample results of the sample data. The sample data includes sample text and a sample video, and the sample results are obtained by annotating the video segments in the sample video that match the sample text.
[0146] 502. The computer device processes the sample data through a recognition model to be trained, and obtains a predicted recognition result, which is used to indicate whether the sample video matches the sample text.
[0147] 503. The computer device processes the sample data through a localization model to be trained, and obtains a predicted localization result, which is used to indicate the video segments in the sample video that match the sample text.
[0148] It should be noted that steps 501 to 503 are the same as steps 301 to 303 in the embodiment shown in Figure 3 and will not be elaborated here one by one.
[0149] 504. The computer device obtains the alignment feature of the sample data, and this alignment feature represents the feature where the sample text and the sample video are similar.
[0150] It should be noted that in a possible implementation manner, during the process of the localization model processing the sample data, the alignment feature of the sample data will be calculated. Therefore, in step 504 above, obtaining the alignment feature of the sample data can be obtaining the alignment feature of the sample data from the recognition model. In another possible implementation manner, the computer device processes the sample data to obtain the alignment feature of the sample data. The computer device obtaining the alignment feature of the sample data includes: the computer device constructs the alignment feature of the sample data based on the sample text and the sample video.
[0151] Another point to note is that the process by which the computer device obtains the alignment features of the sample data may be the same as or different from the process by which the positioning model processes the sample data. In one possible implementation, the process by which the computer device obtains the alignment features of the sample data is the same as the process by which the positioning model processes the sample data. Optionally, the computer device constructs the alignment features of the sample data based on the sample text and the sample video, including: the computer device obtains the text features of the sample text based on the sample text; obtains the video frame features of multiple video frames based on the multiple video frames of the sample video; and fuses the obtained text features and the multiple video frame features to obtain the alignment features.
[0152] Optionally, the computer device fuses the obtained text features and the multiple video frame features to obtain the alignment features, including: obtaining the Hadamard matrix of the sample data based on the text features and each video frame feature, where each column vector of the Hadamard matrix is the Hadamard product of the text feature and a video frame feature; and performing weighted averaging on the column vectors in the Hadamard matrix based on the correlation degree between each video frame and the sample text to obtain the alignment features.
[0153] It should be noted that the process by which the computer device constructs the alignment features of the sample data based on the sample text and the sample video can refer to Figure 3 the embodiments shown, and will not be elaborated here one by one.
[0154] 505. The computer device trains the recognition model and the positioning model respectively based on the sample result, the predicted recognition result, the predicted positioning result, the alignment features, the text features of the sample text, and the video features of the sample video to obtain the trained recognition model and the trained positioning model.
[0155] Since the alignment features represent the features similar to the sample text and the sample video, if the sample text and the sample video are matched, then the text features of the sample text and the video features of the sample video will be relatively similar, and both the text features of the sample text and the video features of the sample video will be relatively similar to the alignment features. If the sample text and the sample video are not matched, then the similarity between the text features of the sample text and the video features of the sample video will be small, and the alignment features that can represent the similar features of the two will be quite different from the text features of the sample text and the video features of the sample video. Therefore, the recognition model and the positioning model can be trained based on the alignment features, the text features of the sample text, and the video features of the sample video, so that the absolute distance between the recognition model and the positioning model regarding the text features and the video features, and then accurately identify similar samples. Among them, similar samples refer to: the sample text and the sample video do not match, but the similarity between the text features of the sample text and the video features of the sample video is relatively high.
[0156] Among them, the computer device trains the recognition model and the localization model respectively based on the sample result, the predicted recognition result, the predicted localization result, the alignment feature, the text feature of the sample text, and the video feature of the sample video to obtain the trained recognition model and the trained localization model, which means: the computer device trains the recognition model based on the sample result, the predicted recognition result, the predicted localization result, the alignment feature, the text feature of the sample text, and the video feature of the sample video to obtain the trained recognition model; the computer device trains the localization model based on the sample result, the predicted recognition result, the predicted localization result, the alignment feature, the text feature of the sample text, and the video feature of the sample video to obtain the trained localization model.
[0157] In a possible implementation manner, the computer device trains the recognition model and the localization model respectively based on the sample result, the predicted recognition result, the predicted localization result, the alignment feature, the text feature of the sample text, and the video feature of the sample video to obtain the trained recognition model and the trained localization model, including: determining a first loss value of the recognition model based on the predicted recognition result; determining a second loss value of the localization model based on the sample localization result and the predicted localization result; determining a fourth loss value of the recognition model based on the alignment feature, the text feature of the sample text, and the video feature of the sample video; and training the recognition model and the localization model respectively based on the first loss value, the second loss value, and the fourth loss value to obtain the trained recognition model and the trained localization model.
[0158] Among them, the process of determining the first loss value and the second loss value can refer to Figure 3 the embodiments shown, which will not be elaborated here one by one. In a possible implementation manner, the fourth loss value can represent the difference between the text feature of the sample text and the alignment feature and the difference between the video feature of the sample video and the alignment feature. If the sample text and the sample video are matched, but the difference between the text feature of the sample text and the alignment feature is large, or the difference between the video feature of the sample video and the alignment feature is large, then the recognition model and the localization model need to be trained to reduce the difference between the text feature of the sample text and the alignment feature and the difference between the video feature of the sample video and the alignment feature. If the sample text and the sample video are not matched, but the difference between the text feature of the sample text and the alignment feature is small, or the difference between the video feature of the sample video and the alignment feature is small, then the recognition model and the localization model need to be trained to increase the difference between the text feature of the sample text and the alignment feature and enhance the difference between the video feature of the sample video and the alignment feature.
[0159] Optionally, the computer device determines a fourth loss value of the recognition model based on the alignment feature, the text feature of the sample text, and the video feature of the sample video, including: obtaining a fifth similarity between the text feature and the alignment feature; obtaining a sixth similarity between the video feature and the alignment feature; and determining the fourth loss value according to the predicted recognition result, the fifth similarity, and the sixth similarity of the sample data.
[0160] In another possible implementation, there are multiple pieces of sample data, and the multiple pieces of sample data include positive sample data and negative sample data. Among them, the positive sample data includes sample text and sample video that match each other, and the negative sample data includes sample text and sample video that do not match. The differences between the text feature of the sample text and the video feature of the sample video in the negative sample data and the alignment feature should be relatively large, while the differences between the text feature of the sample text and the video feature of the sample video in the positive sample data and the alignment feature should be relatively small. Therefore, the loss value can be determined by the differences between the text feature of the sample text and the video feature of the sample video in the positive and negative sample data and the alignment feature. Optionally, the computer device determines a fourth loss value of the recognition model based on the alignment feature, the text feature of the sample text, and the video feature of the sample video, including: obtaining a first similarity between the text feature of the positive sample data and the first alignment feature of the positive sample data; obtaining a second similarity between the text feature of the positive sample data and the second alignment feature of the negative sample data; obtaining a third similarity between the video feature of the positive sample data and the first alignment feature of the positive sample data; obtaining a fourth similarity between the video feature of the positive sample data and the second alignment feature of the negative sample data; and determining the fourth loss value based on the difference between the first similarity and the second similarity and the difference between the third similarity and the fourth similarity. In the embodiments of the present application, the second alignment feature is the alignment feature of the negative sample data. Among them, the negative sample data corresponding to the second alignment feature in the process of obtaining the second similarity and the negative sample data corresponding to the second alignment feature in the process of obtaining the fourth similarity can be the same piece of sample data or different pieces of sample data. Therefore, the second alignment features in the two processes can be the same or different. Among them, the negative sample data corresponding to the second alignment feature in the process of obtaining the second similarity may include positive sample text and negative sample video; the negative sample data corresponding to the second alignment feature in the process of obtaining the fourth similarity may include negative sample text and positive sample video.
[0161] For example, the sample data includes first positive sample data, first negative sample data, and second negative sample data. Among them, the first positive sample data includes sample text A and sample video B, and the sample text A matches the sample video B. The first negative sample data includes sample text A and sample video C, and the sample text A does not match the sample video C. The second negative sample data includes sample text D and sample video B, and the sample text D does not match the sample video B. Among them, the sample text A is a positive sample text, the sample text D is a negative sample text, the sample video B is a positive sample video, and the sample video C is a negative sample video. The computer device obtains similarity 1 between the text feature of the positive sample text and the alignment feature of the positive sample data, and obtains similarity 2 between the text feature of the positive sample text and the alignment feature of the first negative sample data; obtains similarity 3 between the video feature of the positive sample video and the alignment feature of the positive sample data; obtains similarity 4 between the video feature of the positive sample video and the alignment feature of the second negative sample data; determines a fourth loss value based on the difference between similarity 1 and similarity 2 and the difference between similarity 3 and similarity 4.
[0162] It should be noted that in a possible implementation manner, the first loss value is obtained after processing multiple sample data. Therefore, the fourth loss value is also obtained after processing multiple sample data. Therefore, multiple sample data can be divided into multiple groups, and the fourth loss value is determined for each group of sample data respectively. Based on the first loss value, the second loss value, and the fourth loss value of each group of sample data, the recognition model and the positioning model are trained respectively.
[0163] For example, the fourth loss value is:
[0164]
[0165] Among them, L CR is the fourth loss value, α is empirical data, S represents similarity, t represents the text feature of the positive sample text, v represents the video feature of the positive sample video, t - represents the text feature of the negative sample text, v - represents the video feature of the negative sample video, E(v, t) represents the alignment feature of the positive sample data, E(v - , t) represents the alignment feature of the negative sample video and the positive sample text, E(v, t - ) represents the alignment feature of the positive sample video and the negative sample text.
[0166] It should be noted that in the embodiments of the present application, any loss function can be used to determine the fourth loss value. In a possible implementation manner, based on the alignment feature, the text feature of the sample text, and the video feature of the sample video, determining the fourth loss value of the recognition model includes: processing the alignment feature, the text feature of the sample text, and the video feature of the sample video based on the third loss function to obtain the fourth loss value. The third loss function is used to use the alignment feature as an anchor point to reduce the distance between the text feature of the sample text and the video feature of the matching sample video, and expand the distance between the text feature of the sample text and the video feature of the non-matching sample video.
[0167] In a possible implementation manner, based on the first loss value, the second loss value, and the fourth loss value, training the recognition model and the localization model respectively to obtain the trained recognition model and the trained localization model includes: performing weighted processing on the first loss value, the second loss value, and the fourth loss value to obtain the sixth loss value; based on the sixth loss value, training the recognition model and the localization model respectively to obtain the trained recognition model and the trained localization model.
[0168] For example, the sixth loss value is L. Wherein, L = λ1L tri +λ2L CR +λ3L reg . Wherein, λ1, λ2, and λ3 can be any values. For example, λ1 = 0.3, λ2 = 0.3, and λ3 = 0.4.
[0169] 506. When the text for retrieving the video segment is obtained, the computer device determines the target video in the candidate videos that matches the text through the trained recognition model, and determines the target video segment in the target video that matches the text through the trained localization model.
[0170] This step 506 is the same as step 307 and will not be elaborated here one by one.
[0171] In the video segment localization method provided by the embodiments of the present application, when training the recognition model and the localization model, the sample result, the predicted recognition result, and the predicted localization result are all used, so that the recognition model and the localization model can not only learn the information about video recognition, but also learn the information about video segment localization. Since the recognition model and the localization model learn more information, the accuracy of the recognition model and the localization model is improved, and thus the accuracy of video segment localization is improved.
[0172] Figure 6 It is a schematic structural diagram of a video segment localization device provided by the embodiments of the present application. Refer to Figure 6 , and the device includes:
[0173] A determination module 601, configured to determine a sample result, a prediction recognition result, and a prediction localization result of sample data, where the sample data includes sample text and a sample video, the prediction recognition result is obtained by processing the sample data through a recognition model to be trained, and is used to indicate whether the sample video and the sample text match, and the prediction localization result is obtained by processing the sample data through a localization model to be trained, and is used to indicate a video segment in the sample video that matches the sample text;
[0174] A training module 602, configured to train the recognition model and the localization model respectively based on the sample result, the prediction recognition result, and the prediction localization result, so as to obtain a trained recognition model and a trained localization model;
[0175] A retrieval module 603, configured to, when obtaining text for retrieving a video segment, determine a target video that matches the text in candidate videos through the trained recognition model, and determine a target video segment that matches the text in the target video through the trained localization model.
[0176] As Figure 7 shown, in a possible implementation manner, the sample result includes a sample localization result, and the training module 602 includes:
[0177] A first determination unit 6021, configured to determine a first loss value of the recognition model based on the prediction recognition result;
[0178] A second determination unit 6022, configured to determine a second loss value of the localization model based on the sample localization result and the prediction localization result;
[0179] A training unit 6023, configured to train the recognition model and the localization model respectively based on the first loss value and the second loss value, so as to obtain the trained recognition model and the trained localization model.
[0180] In a possible implementation manner, the first determination unit 6021 is configured to process the prediction recognition result through a first loss function of the recognition model to obtain the first loss value;
[0181] The second determination unit 6022 is configured to process the sample localization result and the prediction localization result through a second loss function of the localization model to obtain the second loss value.
[0182] In a possible implementation, the training unit 6023 is configured to perform a weighted processing on the first loss value and the second loss value to obtain a third loss value; and based on the third loss value, train the recognition model and the localization model to obtain the trained recognition model and the trained localization model.
[0183] In a possible implementation, the training module 602 is configured to construct an alignment feature of the sample data based on the sample text and the sample video, where the alignment feature represents a feature similar to the sample text and the sample video; and based on the sample result, the predicted recognition result, the predicted localization result, the alignment feature, the text feature of the sample text, and the video feature of the sample video, train the recognition model and the localization model respectively to obtain the trained recognition model and the trained localization model.
[0184] In a possible implementation, the training module 602 includes:
[0185] The first determination unit 6021 is configured to determine a first loss value of the recognition model based on the predicted recognition result;
[0186] The second determination unit 6022 is configured to determine a second loss value of the localization model based on the sample localization result and the predicted localization result;
[0187] The third determination unit 6024 is configured to determine a fourth loss value of the recognition model based on the alignment feature, the text feature of the sample text, and the video feature of the sample video;
[0188] The training unit 6023 is configured to train the recognition model and the localization model respectively based on the first loss value, the second loss value, and the fourth loss value to obtain the trained recognition model and the trained localization model.
[0189] In a possible implementation, the third determination unit 6024 is configured to process the alignment feature, the text feature of the sample text, and the video feature of the sample video based on a third loss function, where the third loss function is used to use the alignment feature as an anchor point to reduce the distance between the text feature of the sample text and the video feature of the matching sample video, and increase the distance between the text feature of the sample text and the video feature of the non-matching sample video.
[0190] In a possible implementation, there are multiple pieces of sample data, and the multiple pieces of sample data include positive sample data and negative sample data. The third determination unit 6024 is configured to obtain a first similarity between the text feature of the positive sample data and the first alignment feature of the positive sample data; obtain a second similarity between the text feature of the positive sample data and the second alignment feature of the negative sample data; obtain a third similarity between the video feature of the positive sample data and the first alignment feature of the positive sample data; obtain a fourth similarity between the video feature of the positive sample data and the second alignment feature of the negative sample data; and determine the fourth loss value based on the difference between the first similarity and the second similarity and the difference between the third similarity and the fourth similarity.
[0191] In a possible implementation, the training module 602 is configured to obtain the text feature of the sample text based on the sample text; obtain the video frame features of the multiple video frames based on the multiple video frames of the sample video; and fuse the obtained text feature and the multiple video frame features to obtain the alignment feature.
[0192] In a possible implementation, the training module 602 is configured to obtain the Hadamard matrix of the sample data based on the text feature and each video frame feature, where each column vector of the Hadamard matrix is the Hadamard product of the text feature and a video frame feature; and perform weighted averaging on the column vectors in the Hadamard matrix based on the correlation between each video frame and the sample text to obtain the alignment feature.
[0193] In an exemplary embodiment, a computer device is provided, which includes one or more processors and one or more memories. At least one program code is stored in the one or more memories, and the at least one program code is loaded and executed by the one or more processors to implement the video segment location method in the above embodiment.
[0194] Optionally, the computer device is provided as a terminal. Figure 8 The structural block diagram of a terminal 800 provided in an exemplary embodiment of the present application is shown. The terminal 800 may be: a smart phone, a tablet computer, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 (Moving Picture Experts Group Audio Layer IV) player, a laptop computer, or a desktop computer. The terminal 800 may also be referred to by other names such as a user equipment, a portable terminal, a laptop terminal, a desktop terminal, etc.
[0195] The terminal 800 includes a processor 801 and a memory 802.
[0196] The processor 801 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. The processor 801 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), or PLA (Programmable Logic Array). The processor 801 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 801 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 801 may further include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.
[0197] The memory 802 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 802 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 802 is used to store at least one program code, and the at least one program code is used to be executed by the processor 801 to implement the video segment localization method provided in the method embodiments of the present application.
[0198] In some embodiments, the terminal 800 may further optionally include a peripheral device interface 803 and at least one peripheral device. The processor 801, the memory 802, and the peripheral device interface 803 may be connected through a bus or signal lines. Each peripheral device may be connected to the peripheral device interface 803 through a bus, signal lines, or a circuit board. Specifically, the peripheral device includes at least one of a radio frequency circuit 804, a display screen 805, a camera 806, an audio circuit 807, a positioning component 808, and a power supply 809.
[0199] The peripheral device interface 803 can be used to connect at least one I / O (Input / Output) related peripheral device to the processor 801 and the memory 802. In some embodiments, the processor 801, the memory 802, and the peripheral device interface 803 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 801, the memory 802, and the peripheral device interface 803 can be implemented on a separate chip or circuit board, and this embodiment does not limit this.
[0200] The radio frequency circuit 804 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 804 communicates with the communication network and other communication devices through electromagnetic signals. The radio frequency circuit 804 converts an electrical signal into an electromagnetic signal for transmission, or converts the received electromagnetic signal into an electrical signal. Optionally, the radio frequency circuit 804 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a subscriber identity module card, and so on. The radio frequency circuit 804 can communicate with other terminals through at least one wireless communication protocol. The wireless communication protocol includes but is not limited to: metropolitan area network, generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area network, and / or WiFi (Wireless Fidelity) network. In some embodiments, the radio frequency circuit 804 may further include a circuit related to NFC (Near Field Communication), and this application does not limit this.
[0201] The display screen 805 is used to display the UI (User Interface). The UI may include graphics, text, icons, videos, and any combination thereof. When the display screen 805 is a touch display screen, the display screen 805 also has the ability to collect touch signals on or above the surface of the display screen 805. The touch signals can be input as control signals to the processor 801 for processing. At this time, the display screen 805 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there can be one display screen 805, which is provided on the front panel of the terminal 800; in other embodiments, there can be at least two display screens 805, which are respectively provided on different surfaces of the terminal 800 or are in a foldable design; in still other embodiments, the display screen 805 can be a flexible display screen, which is provided on the curved surface or the folding surface of the terminal 800. Even further, the display screen 805 can also be set to an irregular non-rectangular shape, that is, a special-shaped screen. The display screen 805 can be prepared using materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).
[0202] The camera module 806 is used to collect images or videos. Optionally, the camera module 806 includes a front camera and a rear camera. The front camera is provided on the front panel of the terminal, and the rear camera is provided on the back of the terminal. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth camera, a wide-angle camera, and a telephoto camera respectively, to achieve functions such as background blurring by fusing the main camera and the depth camera, panoramic shooting by fusing the main camera and the wide-angle camera, and VR (Virtual Reality) shooting function or other fused shooting functions. In some embodiments, the camera module 806 can also include a flash. The flash can be a single-color temperature flash or a two-color temperature flash. The two-color temperature flash refers to the combination of a warm light flash and a cold light flash, which can be used for light compensation under different color temperatures.
[0203] The audio circuit 807 may include a microphone and a speaker. The microphone is used to collect sound waves of the user and the environment, and convert the sound waves into electrical signals for input to the processor 801 for processing, or input to the radio frequency circuit 804 to achieve voice communication. For the purpose of stereo collection or noise reduction, there may be multiple microphones, which are respectively arranged at different parts of the terminal 800. The microphone may also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert the electrical signal from the processor 801 or the radio frequency circuit 804 into sound waves. The speaker may be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert the electrical signal into sound waves audible to humans, but also convert the electrical signal into sound waves inaudible to humans for uses such as ranging. In some embodiments, the audio circuit 807 may further include a headphone jack.
[0204] The positioning component 808 is used to locate the current geographical location of the terminal 800 to achieve navigation or LBS (Location Based Service). The positioning component 808 may be a positioning component based on the GPS (Global Positioning System) of the United States, the Beidou system of China, the GLONASS system of Russia, or the Galileo system of the European Union.
[0205] The power supply 809 is used to supply power to each component in the terminal 800. The power supply 809 may be alternating current, direct current, a disposable battery, or a rechargeable battery. When the power supply 809 includes a rechargeable battery, the rechargeable battery may support wired charging or wireless charging. The rechargeable battery may also be used to support fast charging technology.
[0206] In some embodiments, the terminal 800 further includes one or more sensors 810. The one or more sensors 810 include but are not limited to: an acceleration sensor 811, a gyroscope sensor 812, a pressure sensor 813, a fingerprint sensor 814, an optical sensor 815, and a proximity sensor 816.
[0207] The acceleration sensor 811 can detect the magnitudes of accelerations on the three coordinate axes of the coordinate system established with the terminal 800. For example, the acceleration sensor 811 can be used to detect the components of the gravitational acceleration on the three coordinate axes. The processor 801 can control the display screen 805 to display the user interface in a landscape view or a portrait view according to the gravitational acceleration signal collected by the acceleration sensor 811. The acceleration sensor 811 can also be used for collecting game or user movement data.
[0208] The gyroscope sensor 812 can detect the body direction and rotation angle of the terminal 800. The gyroscope sensor 812 can cooperate with the acceleration sensor 811 to collect the 3D actions of the user on the terminal 800. Based on the data collected by the gyroscope sensor 812, the processor 801 can implement the following functions: motion sensing (such as changing the UI according to the user's tilt operation), image stabilization during shooting, game control, and inertial navigation.
[0209] The pressure sensor 813 can be disposed on the side frame of the terminal 800 and / or the lower layer of the display screen 805. When the pressure sensor 813 is disposed on the side frame of the terminal 800, it can detect the holding signal of the user on the terminal 800, and the processor 801 can perform left / right hand recognition or shortcut operations based on the holding signal collected by the pressure sensor 813. When the pressure sensor 813 is disposed on the lower layer of the display screen 805, the processor 801 can control the operable controls on the UI interface according to the pressure operation of the user on the display screen 805. The operable controls include at least one of button controls, scroll bar controls, icon controls, and menu controls.
[0210] The fingerprint sensor 814 is used to collect the fingerprints of the user. The processor 801 can identify the user's identity based on the fingerprints collected by the fingerprint sensor 814, or the fingerprint sensor 814 can identify the user's identity based on the collected fingerprints. When the identity of the user is identified as a trusted identity, the processor 801 authorizes the user to perform relevant sensitive operations, and the sensitive operations include unlocking the screen, viewing encrypted information, downloading software, making payments, and changing settings, etc. The fingerprint sensor 814 can be disposed on the front, back, or side of the terminal 800. When there are physical buttons or manufacturer logos on the terminal 800, the fingerprint sensor 814 can be integrated with the physical buttons or manufacturer logos.
[0211] The optical sensor 815 is used to collect the ambient light intensity. In one embodiment, the processor 801 can control the display brightness of the display screen 805 according to the ambient light intensity collected by the optical sensor 815. Specifically, when the ambient light intensity is high, the display brightness of the display screen 805 is increased; when the ambient light intensity is low, the display brightness of the display screen 805 is decreased. In another embodiment, the processor 801 can also dynamically adjust the shooting parameters of the camera module 806 according to the ambient light intensity collected by the optical sensor 815.
[0212] A proximity sensor 816, also known as a distance sensor, is disposed on the front panel of the terminal 800. The proximity sensor 816 is used to collect the distance between the user and the front of the terminal 800. In one embodiment, when the proximity sensor 816 detects that the distance between the user and the front of the terminal 800 is gradually decreasing, the processor 801 controls the display screen 805 to switch from the lit state to the off state; when the proximity sensor 816 detects that the distance between the user and the front of the terminal 800 is gradually increasing, the processor 801 controls the display screen 805 to switch from the off state to the lit state.
[0213] Those skilled in the art can understand that Figure 8 the structure shown in does not constitute a limitation on the terminal 800, and it may include more or fewer components than shown in the figure, or combine certain components, or adopt a different component arrangement.
[0214] Optionally, the computer device is provided as a server. Figure 9 is a schematic structural diagram of a server provided by an embodiment of the present application. The server 900 may vary greatly due to different configurations or performances, and may include one or more processors (Central Processing Units, CPUs) 901 and one or more memories 902. Among them, at least one program code is stored in the memory 902, and the at least one program code is loaded and executed by the processor 901 to implement the methods provided by the above various method embodiments. Of course, the server may also have components such as a wired or wireless network interface, a keyboard, and an input / output interface for input / output. The server may also include other components for implementing the functions of the device, which will not be elaborated here.
[0215] The server 900 is used to execute the steps performed by the server in the above method embodiments.
[0216] In an exemplary embodiment, a computer-readable storage medium is also provided, such as a memory including program code. The above program code can be executed by a processor in a computer device to complete the video segment positioning method in the above embodiments. For example, the computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0217] In an exemplary embodiment, a computer program or a computer program product is also provided. The computer program or the computer program product includes computer program code. When the computer program code is executed by a computer, the computer implements the video segment positioning method in the above embodiments.
[0218] Those of ordinary skill in the art can understand that all or part of the steps to implement the above embodiments can be completed by hardware, or can be completed by instructing relevant hardware through a program. This program can be stored in a computer-readable storage medium, and the above-mentioned storage medium can be a read-only memory, a magnetic disk or an optical disc, etc.
[0219] The above are only alternative embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.
Claims
1. A method for video segment location, characterized in that The method includes: determining a sample result, a predicted recognition result, and a predicted localization result of sample data, where the sample data includes sample text and a sample video, the predicted recognition result is obtained by processing the sample data through a recognition model to be trained, and is used to indicate whether the sample video and the sample text match, and the predicted localization result is obtained by processing the sample data through a localization model to be trained, and is used to indicate a video segment in the sample video that matches the sample text; training the recognition model and the localization model respectively based on the sample result, the predicted recognition result, and the predicted localization result to obtain a trained recognition model and a trained localization model; when obtaining text for retrieving a video segment, determining a target video that matches the text in a candidate video through the trained recognition model, and determining a target video segment that matches the text in the target video through the trained localization model; The training the recognition model and the localization model respectively based on the sample result, the predicted recognition result, and the predicted localization result to obtain a trained recognition model and a trained localization model includes: constructing an alignment feature of the sample data based on the sample text and the sample video, where the alignment feature represents a feature in which the sample text and the sample video are similar; training the recognition model and the localization model respectively based on the sample result, the predicted recognition result, the predicted localization result, the alignment feature, a text feature of the sample text, and a video feature of the sample video to obtain the trained recognition model and the trained localization model; The training the recognition model and the localization model respectively based on the sample result, the predicted recognition result, the predicted localization result, the alignment feature, a text feature of the sample text, and a video feature of the sample video to obtain the trained recognition model and the trained localization model includes: determining a first loss value of the recognition model based on the predicted recognition result; determining a second loss value of the localization model based on the sample localization result and the predicted localization result; determining a fourth loss value of the recognition model based on the alignment feature, the text feature of the sample text, and the video feature of the sample video; training the recognition model and the localization model respectively based on the first loss value, the second loss value, and the fourth loss value to obtain the trained recognition model and the trained localization model.
2. The method according to claim 1, characterized in that The determining a first loss value of the recognition model based on the predicted recognition result includes: processing the predicted recognition result through a first loss function of the recognition model to obtain the first loss value; the determining a second loss value of the localization model based on the sample localization result and the predicted localization result includes: processing the sample localization result and the predicted localization result through a second loss function of the localization model to obtain the second loss value.
3. The method according to claim 1, characterized in that, Determining a fourth loss value of the recognition model based on the alignment feature, the text feature of the sample text, and the video feature of the sample video includes: processing the alignment feature, the text feature of the sample text, and the video feature of the sample video based on a third loss function to obtain the fourth loss value. The third loss function is used to use the alignment feature as an anchor point to reduce the distance between the text feature of the sample text and the video feature of the matching sample video, and to increase the distance between the text feature of the sample text and the video feature of the non-matching sample video.
4. The method according to claim 1, wherein There are multiple pieces of sample data, and the multiple pieces of sample data include positive sample data and negative sample data. Determining a fourth loss value of the recognition model based on the alignment feature, the text feature of the sample text, and the video feature of the sample video includes: obtaining a first similarity between the text feature of the positive sample data and the first alignment feature of the positive sample data; obtaining a second similarity between the text feature of the positive sample data and the second alignment feature of the negative sample data; obtaining a third similarity between the video feature of the positive sample data and the first alignment feature of the positive sample data; obtaining a fourth similarity between the video feature of the positive sample data and the second alignment feature of the negative sample data; determining the fourth loss value based on the difference between the first similarity and the second similarity, and the difference between the third similarity and the fourth similarity.
5. The method according to claim 1, wherein Constructing the alignment feature of the sample data based on the sample text and the sample video includes: obtaining the text feature of the sample text based on the sample text; obtaining the video frame features of multiple video frames of the sample video; and fusing the obtained text feature and the multiple video frame features to obtain the alignment feature.
6. The method according to claim 5, characterized in that Fusing the obtained text feature and the multiple video frame features to obtain the alignment feature includes: obtaining a Hadamard matrix of the sample data based on the text feature and each video frame feature, where each column vector of the Hadamard matrix is the Hadamard product of the text feature and a video frame feature; and performing weighted averaging on the column vectors in the Hadamard matrix based on the correlation degree of each video frame with the sample text to obtain the alignment feature.
7. A video clip positioning device, characterized in that, The device includes: a determination module, configured to determine a sample result, a prediction recognition result, and a prediction localization result of sample data, where the sample data includes sample text and a sample video, the prediction recognition result is obtained by processing the sample data through a recognition model to be trained, and is used to indicate whether the sample video and the sample text match, and the prediction localization result is obtained by processing the sample data through a localization model to be trained, and is used to indicate a video segment in the sample video that matches the sample text; a training module, configured to train the recognition model and the localization model respectively based on the sample result, the prediction recognition result, and the prediction localization result, to obtain a trained recognition model and a trained localization model. The training of the recognition model and the localization model respectively based on the sample result, the prediction recognition result, and the prediction localization result to obtain a trained recognition model and a trained localization model includes: constructing an alignment feature of the sample data based on the sample text and the sample video, where the alignment feature represents a feature in which the sample text and the sample video are similar; training the recognition model and the localization model respectively based on the sample result, the prediction recognition result, the prediction localization result, the alignment feature, the text feature of the sample text, and the video feature of the sample video, to obtain the trained recognition model and the trained localization model. The training of the recognition model and the localization model respectively based on the sample result, the prediction recognition result, the prediction localization result, the alignment feature, the text feature of the sample text, and the video feature of the sample video, to obtain the trained recognition model and the trained localization model includes: determining a first loss value of the recognition model based on the prediction recognition result; determining a second loss value of the localization model based on the sample localization result and the prediction localization result; determining a fourth loss value of the recognition model based on the alignment feature, the text feature of the sample text, and the video feature of the sample video; training the recognition model and the localization model respectively based on the first loss value, the second loss value, and the fourth loss value, to obtain the trained recognition model and the trained localization model; a retrieval module, configured to, when obtaining text for retrieving a video segment, determine a target video in candidate videos that matches the text through the trained recognition model, and determine a target video segment in the target video that matches the text through the trained localization model.
8. A computer device, characterized in that, The computer device includes one or more processors and one or more memories, and at least one program code is stored in the one or more memories, and the at least one program code is loaded and executed by the one or more processors to implement the operations performed by the video segment localization method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, At least one program code is stored in the storage medium, and the at least one program code is loaded and executed by a processor to implement the operations performed by the video clip positioning method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Video clip positioning method and device, computer equipment and storage medium
CN110121118A
Rhythm phrase recognition method and device and electronic equipment
CN111640418A
Vehicle damage detection model training method and device, vehicle damage detection method and device, equipment and medium
CN112668462A