Video segment positioning method, system, control device and readable storage medium

By acquiring multimodal features in video clip positioning, constructing effective candidate video clips, and performing fine-grained coding and deep fusion, the problems of low positioning efficiency and low accuracy in the prior art are solved, and more efficient and accurate video clip positioning is achieved.

CN114896451BActive Publication Date: 2025-05-16GUANGZHOU YUNCONG INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210583620.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-25
Publication Date
2025-05-16
Estimated Expiration
2042-05-25

AI Technical Summary

Technical Problem

The prior art has problems such as low efficiency, large calculation volume and inaccurate positioning of video clips, especially when integrating video and language features, it is difficult to effectively solve.

Method used

By obtaining the multimodal features of the video to be queried and the query statements, multiple valid candidate video clips are constructed, and the relationship perception characteristics of the video clips are obtained through fine-grained encoding and deep fusion, and the precise positioning of the video clips is finally achieved.

Benefits of technology

It improves the efficiency and accuracy of video clip positioning, can better integrate video and language features, enhances the positioning ability of the model, and reduces the calculation amount.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114896451B_ABST
    Figure CN114896451B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of cross-modal perception technology, and specifically provides a video clip positioning method, system, control device and readable storage medium, aiming to solve the problem of how to efficiently, quickly and accurately locate video clips. To this end, the present invention compares the video clip positioning task to the human reading comprehension task, and draws on the reading strategy of the reading comprehension task of first rough reading and then detailed reading to process the video positioning task, so that multimodal features can be integrated in the video positioning process, and the semantic information within and between the language modality and the visual modality can be deeply excavated, which can be more in line with the strategy of human reading comprehension tasks and obtain better positioning effects. At the same time, since effective candidate video clips are constructed, it can help further distinguish visually similar video clips, while ensuring the accuracy of video clip positioning, it can also improve the efficiency of video clip positioning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of cross-modal perception technology, and specifically provides a video segment positioning method, system, control device and readable storage medium. Background Art

[0002] With the popularity of high-definition cameras and the rapid development of short videos, related technologies in the video field, such as action recognition, temporal action detection, video retrieval, and video subtitles, have received widespread attention. Related applications, including video classification, intelligent subtitles, intelligent covers, text video retrieval, and video highlights extraction, have gradually become an important part of people's lives. Among them, the task of video segment localization based on language query is a relatively new research topic in recent years. The purpose of this task is to give an uncut video and a language description, which requires the perception and interaction of information in two modalities of visual language, and then locate the video segment where the action described by this language occurs in the video. This task not only needs to pay attention to the characteristics of the video content, but also needs to integrate the characteristics of the language. It is a multimodal task with certain challenges. It has attracted the attention of both the computer vision field and the natural language processing field in academia, and has certain application prospects in the industry. For example, a certain behavior in a long video can be located through a language description, which effectively reduces manpower time; for example, in online entertainment, we can retrieve the movie clips of interest through a language description, which is convenient for editing and watching.

[0003] In the prior art, the video segment location methods based on language query are mainly divided into three categories:

[0004] (1) One-stage method:

[0005] This method usually predicts each frame in the video to determine whether it is the start frame or the end frame, or regresses the distance from each frame to the boundary. The one-stage method is very efficient because it only needs to predict or regress each video frame, but it ignores the context information of the video frame and cannot obtain the global features of the video well, so the positioning effect is not very good.

[0006] (2) Two-stage approach:

[0007] This method usually uses technologies such as sliding windows to predefine a series of candidate video clip proposals of different lengths, then calculates the similarity between these video clips and the description sentences in the same space, sorts these candidate video clips according to the similarity, and selects the video clip that best matches the description sentence. This method can usually obtain good positioning results because it perceives the global characteristics of the video, but there are still two defects: ① Since the sliding window is predefined, the boundaries of the candidate video clips are not flexible enough, and the final positioning results largely depend on the quality of these pre-generated candidate video clips. ② Since some candidate video clips of different sizes need to be predefined at each position in the video, a large number of candidate video clips need to be densely sampled for the entire video, and the huge amount of calculation will also affect the implementation of the model.

[0008] (3) Reinforcement Learning:

[0009] This method regards the video clip location task as a sequential decision task and uses reinforcement learning to handle the task. Given an initial window, each iteration decides whether to move the window left or right, how many steps to move, and whether to expand or shrink the window based on the feedback value. This method can be trained using reinforcement learning, but there are two disadvantages: ① The training method of reinforcement learning is not stable enough, difficult to train, and not easy to find the best matching video clip; ② Due to the limited decision space, the movement and scaling of the window are limited to predefined strategies, and the quality of the retrieved video clip may not be high, affecting the model performance.

[0010] Accordingly, the art needs a new video segment positioning solution to solve the above problem. Summary of the invention

[0011] In order to overcome the above-mentioned defects, the present invention is proposed to provide a solution or at least partially solve the problem of how to locate the video segment efficiently, quickly and accurately.

[0012] In a first aspect, the present invention provides a method for locating a video segment, the method comprising:

[0013] According to the video to be queried and the query sentence, a query-aware video representation and a video-aware language representation are obtained;

[0014] Constructing a plurality of valid candidate video segments of the video to be queried according to the video to be queried; and acquiring content features and boundary features of each valid candidate video segment according to the video representation perceived by the query;

[0015] Fine-grained encoding is performed on the query-aware video representation and the video-aware language representation respectively to obtain fine-grained video encoding features and fine-grained language encoding features; and the fine-grained video encoding features and the fine-grained language encoding features are deeply fused to obtain fine-grained fusion features;

[0016] Acquire a relation-aware feature of each valid candidate video segment according to the fine-grained fusion feature, the content feature, and the boundary feature;

[0017] According to the relationship perception feature, a final positioning result of the video segment is obtained.

[0018] In a technical solution of the above-mentioned video segment positioning method, the step of "constructing multiple valid candidate video segments of the to-be-queried video according to the to-be-queried video" includes:

[0019] Construct a two-dimensional time network graph of T×T grids; wherein T is the characteristic length of the video representation perceived by the query, the ordinate of the two-dimensional time network graph represents the start time of the candidate video segment in the video to be queried, and the abscissa represents the end time of the candidate video segment in the video to be queried, and the network in the two-dimensional time network graph whose start time is less than the end time is a valid grid;

[0020] According to the time interval between the candidate video segments corresponding to each valid grid and the candidate video segments corresponding to other valid grids, the valid grids are sparsely sampled to obtain multiple sampled valid grids, and the candidate video segments corresponding to the sampled valid grids are used as valid candidate video segments.

[0021] In a technical solution of the above-mentioned video segment positioning method, the step of "obtaining content features and boundary features of each valid candidate video segment according to the query-perceived video representation" includes:

[0022] Obtain the content features of the nth valid candidate video segment according to the following formula and boundary features

[0023]

[0024]

[0025] in, is the query-aware video representation of the start time of the nth valid candidate video segment, It is the query-aware video representation of the end time of the nth valid candidate video segment, MaxPooling is the maximum pooling operation, and Addition is the addition operation.

[0026] In a technical solution of the above-mentioned video segment positioning method, the step of "respectively fine-grained encoding of the query-aware video representation and the video-aware language representation to obtain fine-grained video encoding features and fine-grained language encoding features" includes:

[0027] The fine-grained video coding features are obtained according to the following formula:

[0028]

[0029] in, is the fine-grained video coding feature, is the query-aware video representation, Linear is a linear fully connected layer operation, and ReLU is a linear rectification function.

[0030] A one-dimensional convolutional network is used to encode the language representation of video perception to obtain unigram language features, bigram language features, and trigram language features.

[0031] According to the one-gram language features, two-gram language features and three-gram language features, the following formula is applied to obtain the fine-grained language encoding features:

[0032]

[0033] in, encoding features for the fine-grained language, Let be the uni-gram feature, the bigram feature and the tri-gram feature, and Concat be the feature fusion operation.

[0034] In a technical solution of the above-mentioned video segment positioning method, the step of "deeply fusing the fine-grained video coding features and the fine-grained language coding features to obtain fine-grained fusion features" includes:

[0035] The fine-grained fusion feature is obtained according to the following formula:

[0036]

[0037] in, Hook the fine-grained fusion feature, Check the perceived video segment features of the query, is the video segment feature perceived by the video; the query perceived video segment feature is obtained according to the following formula:

[0038]

[0039] A C is the set of content features of valid candidate video clips, G Qis a gated language feature, which is obtained by the following formula:

[0040]

[0041] σ is the gate function, A N is the set of boundary features of valid candidate video clips, To transfer language features, the transfer language features are obtained by the following formula:

[0042]

[0043] Linear is a linear fully connected layer operation, and MaxPooling is a maximum pooling layer operation;

[0044] The video segment features described by video perception are as follows:

[0045]

[0046] Avgpooling is the average pooling layer operation.

[0047] In a technical solution of the above-mentioned video segment positioning method, the step of "obtaining the relationship perception feature of each valid candidate video segment according to the fine-grained fusion feature, the content feature and the boundary feature" includes:

[0048] The fine-grained fusion feature, the content feature and the boundary feature are fused, and the enhanced fusion feature is obtained according to the following formula:

[0049]

[0050] The enhanced fusion features are input into a stacked multi-layer grouped convolutional network, and the relation-aware features of each valid candidate video segment are obtained according to the following formula:

[0051]

[0052] A collection of relation-aware features that identify valid candidate video clips.

[0053] In a technical solution of the above-mentioned video segment positioning method, the step of "obtaining the final video segment positioning result according to the relationship perception feature" includes:

[0054] Scoring each valid candidate video segment according to the relation-perception feature of each valid candidate video segment to obtain a score for each valid candidate video segment;

[0055] The final positioning result of the video segment is determined according to the score of the valid candidate video segment.

[0056] In a technical solution of the above-mentioned video segment positioning method, the step of "scoring each valid candidate video segment according to the relational perception feature of each valid candidate video segment to obtain a score of each valid candidate video segment" includes:

[0057] The score of each valid candidate video segment is obtained according to the following formula:

[0058]

[0059] Among them, P A is a set of scores of valid candidate video segments.

[0060] In a technical solution of the above-mentioned video segment positioning method, the step of "determining the final video segment positioning result according to the score of the valid candidate video segment" includes:

[0061] Sorting the scores of the valid candidate video clips in descending order;

[0062] According to the sorting results and preset requirements, select the valid candidate video segment with the highest score as the final video segment positioning result; or,

[0063] According to the sorting results and preset requirements, the first k valid candidate video segments are used as the final video segment positioning results.

[0064] In a technical solution of the above-mentioned video segment location method, the step of "obtaining a query-aware video representation and a video-aware language representation according to the video to be queried and the query statement" includes:

[0065] Dividing the video to be queried into multiple video segments;

[0066] Use a preset video feature extraction model to extract features from each video clip to obtain video features of each video clip;

[0067] Coarse-grained encoding is performed on the video features of all video clips to obtain coarse-grained video features;

[0068] Using a preset language feature extraction model to extract features from the query statement to obtain language features of the query statement;

[0069] Coarse-grained encoding is performed on the language features of the query statement to obtain coarse-grained language features;

[0070] The coarse-grained video features and the coarse-grained language features are modally interacted to obtain query-aware video representations and video-aware language representations.

[0071] In a technical solution of the above-mentioned video segment positioning method, the step of "coarse-grained encoding of video features of all video segments to obtain coarse-grained video features" includes:

[0072] A one-dimensional convolutional network is used to encode the video features of the video clips, and the encoded video features are reduced in dimension through an average pooling layer network to obtain reduced-dimensional video features;

[0073] Applying a Bi-GRU network to encode the reduced-dimensional video features to obtain the coarse-grained video features; and / or,

[0074] The steps of “coarse-grained encoding of the language features of the query statement to obtain the coarse-grained language features” include:

[0075] A Bi-GRU network is applied to encode the language features of the query sentence to obtain the coarse-grained language features.

[0076] In a technical solution of the above-mentioned video segment localization method, the step of "modally interacting the coarse-grained video features and the coarse-grained language features to obtain the query-aware video representation and the video-aware language representation" includes obtaining the query-aware video representation and the video-aware language representation according to the following formula:

[0077]

[0078]

[0079] in, a video representation perceived by the query, Check the coarse-grained video features, The language representation of video perception, v atten is the weighted sum of the coarse-grained video features, The coarse-grained language feature, T is the length of the coarse-grained video feature, C is the dimension of the coarse-grained video feature and the coarse-grained language feature, L is the length of the coarse-grained language feature, q atten is the weighted sum of the coarse-grained language features, and q is obtained according to the following formula atten :

[0080]

[0081] Check the jth element of the coarse-grained language feature Average attention weight, the attention weight is obtained according to the following formula:

[0082]

[0083] a Q is the attention weight matrix of the coarse-grained language feature, The transpose of the coarse-grained language feature matrix; Q is the first learnable parameter matrix, b Q is the second learnable parameter matrix;

[0084] v atten is the weighted sum of the coarse-grained video features, and v is obtained according to the following formula atten :

[0085]

[0086] is the jth element of the coarse-grained video feature The attention weight is obtained according to the following formula:

[0087]

[0088] a V is the attention weight matrix of the coarse-grained video feature, Hook the transpose of the coarse-grained video feature matrix.

[0089] In a second aspect, the present invention provides a video segment positioning system, the system comprising:

[0090] A coarse-grained encoding and modality interaction module, which is configured to obtain a query-aware video representation and a video-aware language representation according to a video to be queried and a query statement;

[0091] A candidate video frequency band construction module is configured to construct a plurality of candidate video segments of the video to be queried according to the video to be queried; and obtain content features and boundary features of each candidate video segment according to the video representation perceived by the query;

[0092] A fine-grained fusion feature acquisition module is configured to perform fine-grained encoding on the query-aware video representation and the video-aware language representation, respectively, to obtain fine-grained video encoding features and fine-grained language encoding features; and deeply fuse the fine-grained video encoding features and the fine-grained language encoding features to obtain fine-grained fusion features;

[0093] a relationship-aware feature acquisition module, configured to acquire a relationship-aware feature of each candidate video segment according to the fine-grained fusion feature, the content feature and the boundary feature;

[0094] The video positioning result acquisition module is configured to obtain the final positioning result of the video segment according to the relationship perception feature.

[0095] In a third aspect, a control device is provided, which includes a processor and a storage device, wherein the storage device is suitable for storing multiple program codes, and the program codes are suitable for being loaded and run by the processor to execute the video segment positioning method described in any one of the technical solutions of the above-mentioned video segment positioning method.

[0096] In a fourth aspect, a computer-readable storage medium is provided, wherein a plurality of program codes are stored therein, wherein the program codes are suitable for being loaded and run by a processor to execute the video segment positioning method described in any one of the technical solutions of the above-mentioned video segment positioning method.

[0097] The above one or more technical solutions of the present invention have at least one or more of the following beneficial effects:

[0098] In the technical solution of the present invention, the present invention can obtain the query-perceived video representation that integrates the query statement information and the video-perceived language representation that integrates the video features according to the video to be queried and the query statement, respectively perform fine-grained encoding on the query-perceived video representation and the video-perceived language representation, and fuse the encoded features to obtain fine-grained fusion features. According to the video to be queried, multiple valid candidate video segments are constructed, and the relationship-perceived features of the valid candidate video segments are obtained according to the fine-grained fusion features, as well as the content features and boundary features of the valid candidate video segments, and the final video segment positioning results are obtained according to the relationship-perceived features. Through the above configuration, the present invention compares the video segment positioning task to the human reading comprehension task, and draws on the reading strategy of the reading comprehension task of first rough reading and then detailed reading to process the video positioning task, so that multimodal features can be integrated in the video positioning process, and at the same time, the semantic information within and between the language modality and the visual modality can be deeply excavated, so that the video segment positioning method can be more in line with the strategy of human reading comprehension tasks, and better positioning effects can be obtained. At the same time, since multiple valid candidate video clips are constructed, the relationship-aware features of the valid candidate video clips contain the relationships with other valid candidate video clips. The relationship-aware features can help further distinguish visually similar video clips, while ensuring the accuracy of video clip positioning, it can also improve the efficiency of video clip positioning. BRIEF DESCRIPTION OF THE DRAWINGS

[0099] The disclosure of the present invention will become more easily understood with reference to the accompanying drawings. It is easy for those skilled in the art to understand that these drawings are only for illustrative purposes and are not intended to limit the scope of protection of the present invention. Among them:

[0100] Figure 1 is a schematic flow chart of main steps of a method for locating a video clip according to an embodiment of the present invention;

[0101] Figure 2 It is a flowchart of the main steps of analogizing the video segment localization task to the human reading comprehension task;

[0102] Figure 3 is a schematic diagram of a two-dimensional time network diagram according to an implementation of an embodiment of the present invention;

[0103] Figure 4 is a schematic diagram of a method for obtaining content features and boundary features of valid candidate video segments according to an implementation of an embodiment of the present invention;

[0104] Figure 5 is a schematic diagram of positioning results of a video clip positioning method according to an example of an embodiment of the present invention;

[0105] Figure 6 is a schematic diagram of positioning results of a video clip positioning method according to another example of an embodiment of the present invention;

[0106] Figure 7 is a main structural block diagram of a video clip positioning system according to an embodiment of the present invention;

[0107] Figure 8 It is a main structural block diagram of a video segment positioning system according to an implementation of an embodiment of the present invention. DETAILED DESCRIPTION

[0108] Some embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood by those skilled in the art that these embodiments are only used to explain the technical principles of the present invention and are not intended to limit the protection scope of the present invention.

[0109] In the description of the present invention, "module" and "processor" may include hardware, software or a combination of the two. A module may include hardware circuits, various suitable sensors, communication ports, and memories, and may also include software parts, such as program codes, or a combination of software and hardware. The processor may be a central processing unit, a microprocessor, an image processor, a digital signal processor, or any other suitable processor. The processor has data and / or signal processing functions. The processor may be implemented in software, hardware, or a combination of the two. Non-temporary computer-readable storage media include any suitable medium that can store program codes, such as a magnetic disk, a hard disk, an optical disk, a flash memory, a read-only memory, a random access memory, and the like. The term "A and / or B" means all possible combinations of A and B, such as only A, only B, or A and B. The term "at least one A or B" or "at least one of A and B" has a similar meaning to "A and / or B", and may include only A, only B, or A and B. The singular terms "one" and "the" may also include plural forms.

[0110] In view of the problems existing in the prior art, the present invention proposes a new video segment localization method, which draws on the method of human beings to handle reading comprehension tasks. Figure 2 , Figure 2 This is a flowchart of the main steps of analogizing the video segment localization task to the human reading comprehension task. Figure 2 As shown in the figure, the query-based video segment localization task is analogous to the multiple-choice reading task in the natural language reading comprehension task. The input of the video segment localization task includes video, query, and predefined candidate video segments, where the video corresponds to the article in the reading comprehension, the query corresponds to the question in the reading comprehension, and the candidate video segments correspond to the candidate answers in the reading comprehension. The final output is to select the K candidate video segments with the highest matching probability.

[0111] See attached Figure 1 , Figure 1 FIG. 1 is a schematic flow chart of the main steps of a method for locating a video segment according to an embodiment of the present invention. Figure 1 As shown, the video segment positioning method in the embodiment of the present invention mainly includes the following steps S101 to S105.

[0112] Step S101: obtaining a query-aware video representation and a video-aware language representation according to the video to be queried and the query statement.

[0113] In this embodiment, query-aware video representation and video-aware language representation can be obtained according to the video to be queried and the query statement, wherein the query-aware video representation refers to video features integrated with language features, and the video-aware language representation refers to language features integrated with video features.

[0114] In one implementation, feature extraction can be performed on the query video and query statement respectively, and then the extracted features can be coarse-grained encoded, and the encoded features can be subjected to modal interaction to obtain query-perceived video representation and video-perceived language representation. Coarse-grained encoding refers to a coding process in which the extracted features are simply mapped and temporally dependent. Temporal dependency refers to the temporal relationship between features. Modal interaction refers to the mutual fusion of features of different modalities. Modality refers to senses, and different modalities refer to information obtained through different senses, such as video modality is video sense, language modality is text sense, etc.

[0115] Step S102: construct multiple valid candidate video segments of the video to be queried according to the video to be queried; and obtain content features and boundary features of each valid candidate video segment according to the query-perceived video representation.

[0116] In this embodiment, multiple valid candidate video segments of the video to be queried can be constructed based on the video to be queried, that is, multiple valid candidate video segments are generated based on the video to be queried. At the same time, in order to obtain a more comprehensive representation of the valid candidate video segments, the content features and boundary features of the valid video segments can be obtained based on the video representation perceived by the query. Among them, the content features refer to the features obtained based on the video representation of all moments in the valid candidate video segments; the boundary features refer to the features obtained based on the video representation of the start time and the end time of the valid candidate video segments.

[0117] Step S103: Fine-grained encoding is performed on the query-perceived video representation and the video-perceived language representation respectively to obtain fine-grained video coding features and fine-grained language coding features; and the fine-grained video coding features and the fine-grained language coding features are deeply fused to obtain fine-grained fusion features.

[0118] In this embodiment, the query-perceived video representation and the video-perceived language representation can be fine-grained encoded respectively to obtain fine-grained video coding features and fine-grained language coding features, and the obtained fine-grained video coding features and fine-grained language coding features are deeply fused to obtain fine-grained fusion features. Among them, fine-grained coding is relative to coarse-grained coding, that is, to further refine the coding of video representation and language representation to obtain deeper features of video modality and language modality. And the fine-grained video coding features and fine-grained language coding features are deeply fused to obtain fine-grained fusion features. Among them, fine-grained fusion features are deeper features obtained after further interaction (deep fusion) between video modality and language modality.

[0119] In one implementation, a Concat function may be used to achieve deep fusion of fine-grained video coding features and fine-grained language coding features.

[0120] Step S104: Obtain the relationship-aware features of each valid candidate video segment based on the fine-grained fusion features, content features, and boundary features.

[0121] In this embodiment, the relationship perception feature of each valid candidate video segment can be obtained based on the content features and boundary features of the valid candidate video segments and combined with the fine-grained fusion features obtained in step S103. The relationship perception feature includes the features of the relationship between different candidate video segments. That is, the context information of different valid candidate video segments can be learned based on the fine-grained fusion features, content features and boundary features, thereby obtaining the relationship between different candidate video segments.

[0122] Step S105: obtaining the final positioning result of the video segment according to the relationship perception feature.

[0123] In this embodiment, each valid candidate video segment may be evaluated according to the relational perception features of each valid candidate video segment, and the final positioning result of the video segment may be obtained according to the evaluation result.

[0124] In one implementation, the final video segment positioning result can be determined according to a preset requirement. That is, when the preset requirement is to obtain only one valid candidate video segment, the valid candidate video segment that best matches the relationship perception feature can be used as the final video segment positioning result; when the preset requirement is to obtain the first K valid candidate video segments, the first K valid candidate video segments that best match the relationship perception feature can be used as the final video segment positioning result.

[0125] Based on the above steps S101 to S105, the embodiment of the present invention can obtain the query-perceived video representation that integrates the query statement information and the video-perceived language representation that integrates the video features according to the video to be queried and the query statement, respectively perform fine-grained encoding on the query-perceived video representation and the video-perceived language representation, and fuse the encoded features to obtain fine-grained fusion features. According to the video to be queried, multiple valid candidate video segments are constructed, and the relationship-perceived features of the valid candidate video segments are obtained according to the fine-grained fusion features, as well as the content features and boundary features of the valid candidate video segments, and the final video segment positioning results are obtained according to the relationship-perceived features. Through the above configuration, the embodiment of the present invention compares the video segment positioning task to the human reading comprehension task, and draws on the reading strategy of the reading comprehension task of first rough reading and then detailed reading to process the video positioning task, so that multimodal features can be integrated in the video positioning process, and the semantic information within and between the language modality and the visual modality can be deeply excavated, so that the video segment positioning method can be more in line with the strategy of human reading comprehension tasks, and can obtain better positioning effects. At the same time, since multiple valid candidate video clips are constructed, the relationship-aware features of the valid candidate video clips contain the relationships with other valid candidate video clips. The relationship-aware features can help further distinguish visually similar video clips, while ensuring the accuracy of video clip positioning, it can also improve the efficiency of video clip positioning.

[0126] Steps S101 to S105 are further described below.

[0127] In one implementation of the embodiment of the present invention, step S101 may further include the following steps S1011 to S1016:

[0128] Step S1011: Divide the video to be queried into multiple video segments.

[0129] In this embodiment, the query video can be firstly cut into frames to obtain a sequence of pictures after cutting. The continuous Tc pictures can be regarded as a video clip. c It can be expressed as Among them, v i is the i-th video clip in the query video, n c is the number of video clips in the video to be queried.

[0130] In one implementation, when the total number of pictures Tv in the picture sequence after frame cutting is not an integer multiple of Tc, the last remaining picture sequence with less than Tc pictures is discarded.

[0131] Step S1012: Use a preset video feature extraction model to perform feature extraction on each video clip to obtain video features of each video clip.

[0132] In this embodiment, a video feature extraction model can be used to extract features from video clips to obtain video features of each video clip. The extracted video features can be recorded as Among them, C V The feature dimensions of each video segment.

[0133] In one embodiment, the video feature extraction model includes but is not limited to a VGG (Visual Geometry Group) model, a C3D (Convolution 3D) model, or an I3D (Inflated 3DConvNets) model.

[0134] Step S1013: coarse-grained encoding is performed on the video features of all video clips to obtain coarse-grained video features.

[0135] In this implementation, coarse-grained encoding may be performed on multiple pairs of video features of the video clips to obtain coarse-grained video features of the video clips.

[0136] In one implementation, step S1013 may further include the following steps S10131 and S10132:

[0137] Step S10131: Apply a one-dimensional convolutional network to encode the video features of the video clip, and reduce the dimension of the encoded video features through an average pooling layer network to obtain video features after dimension reduction.

[0138] In this embodiment, a one-dimensional convolutional network can be used to train the video features. Encode and map the encoded video features to R through the average pooling layer T×C In the space, the dimension of the video feature is reduced to obtain the video feature after dimension reduction. Among them, T is the feature length of the video feature after dimension reduction, and C is the dimension of the video feature after dimension reduction. The average pooling layer refers to taking the average value of the pooling area. The average pooling layer operation can reduce the dimension of the feature.

[0139] Step S10132: Apply a Bi-GRU (Bi-Gated Recurrent Unit) network to encode the reduced-dimensional video features to obtain coarse-grained video features.

[0140] In this embodiment, considering the temporal characteristics between video clips, the Bi-GRU network can be used to encode the reduced-dimensional video features to obtain the temporal dependencies between video clips. After encoding, the coarse-grained video features can be obtained.

[0141] Step S1014: extracting features from the query statement using a preset language feature extraction model to obtain language features of the query statement.

[0142] In this embodiment, a language feature extraction model can be used to extract the language features of the query statement. The query statement can be expressed as Among them, q i is the i-th element of the query statement, n q The number of elements in the query statement.

[0143] In one implementation, the language feature extraction model includes but is not limited to a GloVe (Global Vectors) model, a BERT (Bidirectional Transformers) model, and the like.

[0144] Step S1015: coarse-grained encoding is performed on the language features of the query statement to obtain coarse-grained language features.

[0145] In this embodiment, the Bi-GRU network can also be used to encode the language features of the query sentence to obtain coarse-grained language features.

[0146] Step S1016: Modally interact the coarse-grained video features and the coarse-grained language features to obtain query-aware video representations and video-aware language representations.

[0147] In this embodiment, after obtaining coarse-grained video features and coarse-grained language features, the features of these two modalities can be first modally interacted so that the video features are fused with language information, and the language features are fused with video information, that is, query-aware video representations and video-aware language representations are obtained.

[0148] In one implementation, the query-aware video representation may be obtained according to the following formulas (1) to (3):

[0149]

[0150]

[0151]

[0152] in, is a query-aware video representation, a Q is the attention weight matrix of the coarse-grained language features, is the transpose of the coarse-grained language feature matrix; W Q is the first learnable parameter matrix, b Q is the second learnable parameter matrix; is the jth element of the coarse-grained language feature The attention weight of atten is the weighted sum of coarse-grained language features.

[0153] That is to say, the coarse-grained language features can be input into the linear layer network for linear transformation, and the softmax function can be used to obtain the attention weight of the coarse-grained language features, where the attention weight represents the importance of each element in the coarse-grained language features. Then, the attention weight and the language features in the coarse-grained language features are multiplied and accumulated, and then dot-multiplied with the coarse-grained video features to obtain the query-aware video representation. The query-aware video representation can be normalized using L2normalization.

[0154] Since modal interaction is symmetrical, the same method can be used to obtain the language representation of video perception, specifically formula (4)-formula (6):

[0155]

[0156]

[0157]

[0158] Among them, a V is the attention weight matrix of coarse-grained video features, The transpose of the coarse-grained video feature matrix, v atten is the weighted sum of coarse-grained video features, is the jth element of the coarse-grained video feature The attention weight, is the language representation of video perception, Coarse-grained language features.

[0159] In one implementation of the embodiment of the present invention, step S102 may further include the following steps S1021 to S1022:

[0160] Step S1021: construct a two-dimensional time network graph of T×T grids; wherein T is the characteristic length of the video representation perceived by the query, the ordinate of the two-dimensional time network graph represents the start time of the candidate video segment in the video to be queried, and the abscissa represents the end time of the candidate video segment in the video to be queried, and the network in the two-dimensional time network graph whose start time is less than the end time is a valid grid.

[0161] In this embodiment, please refer to the attached Figure 3 , Figure 3 is a schematic diagram of a two-dimensional time network diagram according to an implementation of an embodiment of the present invention, wherein: Figure 3 The horizontal axis is the end time of the candidate video segment, and the vertical axis is the start time of the candidate video segment. Figure 3 As shown, a two-dimensional time network diagram of T×T grids can be constructed. Since the start time must be less than the end time to be meaningful, the grids in the two-dimensional time network diagram whose start time is less than the end time are valid grids.

[0162] Step S1022: Sparsely sample the valid grids according to the time interval between the candidate video segments corresponding to each valid grid and the candidate video segments corresponding to other valid grids to obtain multiple sampled valid grids, and use the candidate video segments corresponding to the sampled valid grids as valid candidate video segments.

[0163] In this embodiment, since the number of valid grids is large, the valid grid can be sparsely sampled according to the time interval between the candidate video segments corresponding to each valid grid and the candidate video segments corresponding to other valid grids. That is, when the time interval between the candidate video segments changes from short to long, the sampling of the valid grids is also adjusted from dense to sparse.

[0164] In one implementation of the embodiment of the present invention, step S102 may further include:

[0165] The content features of the nth valid candidate video segment are obtained according to the following formulas (7) and (8): and boundary features

[0166]

[0167]

[0168] in, Query the perceptual video representation for the start time of the nth valid candidate video segment, It is the video representation perceived by querying the end time of the nth valid candidate video segment, MaxPoolin9 is the maximum pooling operation, and Addition is the addition operation.

[0169] In this embodiment, see the attached Figure 4 , Figure 4 is a schematic diagram of a method for obtaining content features and boundary features of valid candidate video segments according to an implementation of an embodiment of the present invention, wherein: Figure 4 The horizontal axis is the end time of the candidate video segment, and the vertical axis is the start time of the candidate video segment. Figure 4As shown, the query-perceived video representations at all times between the start time and the end time of the valid candidate video segment can be subjected to a maximum pooling operation to obtain the content features of the valid candidate video segment; the query-perceived video representations within the start time and the end time of the valid candidate video segment are added together to obtain the boundary features of the valid candidate video segment. The content features and boundary features of all valid candidate video segments obtained in step S1022 can be integrated to obtain a set A of content features of the valid candidate video segment. C and the set A of boundary features B , where: A C and A B It can be expressed by the following formula (9) and formula (10):

[0170]

[0171]

[0172] In one implementation of the embodiment of the present invention, step S103 may further include the following steps S1031 to S1033:

[0173] Step S1031: Obtain fine-grained video coding features according to the following formula (11):

[0174]

[0175] in, Fine-grained video coding features, It is a query-aware video representation. Linear is a linear fully connected layer operation, and ReLU is a linear rectification function.

[0176] In this embodiment, in order to further perceive the information within the modality, the query-perceived video representation can be fine-grainedly encoded by imitating the human habit of processing reading comprehension tasks. Specifically, the query-perceived video representation can be encoded using a feedforward neural network and added to the query-perceived video representation to obtain fine-grained video encoding features. The feedforward neural network includes a linear fully connected layer, whose activation function is a linear rectifier function, namely, a ReLU function.

[0177] Step S1032: Apply a one-dimensional convolutional network to encode the language representation of video perception to obtain unigram language features, bigram language features and trigram language features respectively.

[0178] In this embodiment, in order to further mine the features at the word level and the phrase level, thereby obtaining more accurate fine-grained features, a one-dimensional convolutional network can be applied to encode the language representation of video perception. Convolution kernels of different sizes can be used to implement different encoding processes, that is, a convolution operation is performed using a convolution kernel of size 1 to obtain a unigram language feature; a convolution operation is performed using a convolution kernel of size 3 to obtain a bigram language feature; and a convolution operation is performed using a convolution kernel of size 5 to obtain a trigram language feature.

[0179] Step S1033: According to the uni-gram language features, the bigram language features and the tri-gram language features, the following formula (12) is applied to obtain fine-grained language coding features:

[0180]

[0181] in, Encoding features for fine-grained languages, They are unary language features, bigram language features and tri-ary language features respectively, and Concat is a feature fusion operation.

[0182] In this embodiment, the Concat function can be used to concatenate uni-gram language features, bi-gram language features, and tri-gram language features and pass them into a linear fully connected layer network to obtain fine-grained language encoding features.

[0183] In one implementation of the embodiment of the present invention, step S103 may include step S1034 in addition to step S1031 to step S1033:

[0184] Step S1034: Obtain fine-grained fusion features according to the following formulas (13) to (17):

[0185]

[0186]

[0187]

[0188]

[0189]

[0190] in, Check the fine-grained fusion features, To query perceived video segment features, is the video segment feature of video perception, G Q is the gated language feature, σ is the gate function, To transfer language features, Avgpooling is an average pooling layer operation.

[0191] In this embodiment, after fine-grained perception of the internal information of the modality, the information between the interactive modalities can be further obtained, and the relationship between the two modalities can be mined using a gate function. In this embodiment, the gate function is a sigmoid function. The query-perceived video segment features can be obtained by formulas (13) to (15), the video-perceived video segment features can be obtained by formula (16), and the query-perceived video segment features and the video-perceived video segment features can be feature fused according to formula (17) to obtain fine-grained fused features.

[0192] In one implementation of the embodiment of the present invention, step S104 may include the following steps S1041 and S1042:

[0193] Step S1041: Fuse the fine-grained fusion features, content features and boundary features, and obtain enhanced fusion features according to the following formula (18):

[0194]

[0195] In this embodiment, after obtaining the fine-grained fusion features, the features of different valid candidate video clips can be compared, and context information can be learned to accurately distinguish similar valid candidate video clips, in accordance with the human habit of comparing different options and then drawing conclusions when processing reading comprehension tasks. Specifically, the fusion features obtained by feature fusion of fine-grained fusion features, content features, and boundary features can be integrated through a two-dimensional convolutional network (Conv2d) to obtain enhanced fusion features.

[0196] Step S1042: Input the enhanced fusion features into the stacked multi-layer group convolutional network, and obtain the relation-aware features of each valid candidate video segment according to the following formula (19):

[0197]

[0198] A collection of relation-aware features that identify valid candidate video clips.

[0199] In this embodiment, more context information can be perceived from adjacent valid candidate video segments by stacking multiple layers of grouped convolutional networks to obtain the relational perception features of each valid candidate video segment. Since the number of parameters of the grouped convolutional network is much smaller than that of a general convolutional network and the size of the hidden layer of the grouped convolutional network in the present invention is only half of that of a conventional solution, the positioning capability of the video segments can be further improved based on the relational perception features, thereby improving the efficiency of the operation process.

[0200] In one implementation of the embodiment of the present invention, step S105 may further include the following steps S1051 and S1052:

[0201] Step S1051: scoring each valid candidate video segment according to the relational perception features of each valid candidate video segment to obtain a score for each valid candidate video segment.

[0202] In this embodiment, the score of each valid candidate video segment can be obtained according to the following formula (20):

[0203]

[0204] Among them, PA is the set of scores of valid candidate video clips, and σ can be the sigmoid activation function.

[0205] Step S1052: Determine the final positioning result of the video segment according to the scores of the valid candidate video segments.

[0206] In one implementation, step S1052 may include the following steps S10521 and S10522:

[0207] Step S10521: sorting the scores of the valid candidate video segments in descending order;

[0208] Step S10522: According to the sorting result and preset requirements, select the valid candidate video segment with the highest score as the final video segment positioning result.

[0209] In this embodiment, when the preset requirement is to obtain a valid candidate video segment, the valid candidate video segment with the highest score is used as the final video segment positioning result.

[0210] In one implementation, step S1052 may include the following steps S10521 and S10523:

[0211] Step S10521: sorting the scores of the valid candidate video segments in descending order;

[0212] Step S10523: According to the sorting result and preset requirements, the first k valid candidate video segments are used as the final video segment positioning results.

[0213] In this embodiment, when the preset requirement is to obtain k valid candidate video segments, the first k valid candidate video segments are obtained according to the sorting result as the final video segment positioning result.

[0214] In one implementation, a training video and a corresponding annotation file may be prepared as a training set to train a model for implementing the video segment positioning method of an embodiment of the present invention. The training video contains multiple human behaviors, and the annotation file annotates the start time and end time of each human behavior, as well as the sentence describing the human behavior. The video segments corresponding to each human behavior may overlap and have different durations, and each description sentence corresponds to only one video segment. The behavior start time and end time in each annotation file may be normalized so that the normalized timestamp is between [0, 1]. For each video sample in the training set, the overlap percentage between each video segment in the video and the behavior time in the annotation file corresponding to the video may be calculated. If the overlap percentage is greater than or equal to the first threshold, the label of the video segment sample is set to 1; if the overlap percentage is less than or equal to the second threshold, the label of the video segment sample is set to 0; if the overlap percentage is between the second threshold and the first threshold, the label value may be normalized so that the label value is between 0 and 1. The loss function used in the training process is a binary cross entropy loss function, as shown in formulas (21) and (22):

[0215]

[0216]

[0217] Among them, g i is the label of the sample, L is the binary cross entropy loss function, p i The score of the i-th valid candidate video segment, N is the number of valid candidate video segments, θ max is the first threshold, θ min is the second threshold.

[0218] During the model training, the learning rate is set to 1×10 -3 , using the Adam optimizer, and iterating training for a total of 15 times.

[0219] See attached Figure 5 and attached Figure 6 , Figure 5 is a schematic diagram of positioning results of a video clip positioning method according to an example of an embodiment of the present invention; Figure 6 FIG. 1 is a schematic diagram of positioning results of another example of a video segment positioning method according to an embodiment of the present invention. Figure 5 and Figure 6 The result corresponding to model 3 is the video segment positioning result after removing the fine-grained encoding step, and the result corresponding to model 4 is the video segment positioning result after removing the step of obtaining fine-grained fusion features. Figure 5 and Figure 6As shown, when a query sentence is given, the positioning result obtained by the video positioning method of the embodiment of the present invention is closer to the data marked in the training process.

[0220] It should be pointed out that although the various steps in the above embodiments are described in a specific order, those skilled in the art can understand that in order to achieve the effects of the present invention, different steps do not have to be performed in such an order. They can be performed simultaneously (in parallel) or in other orders. These changes are within the scope of protection of the present invention.

[0221] Furthermore, the present invention also provides a video clip positioning system.

[0222] See attached Figure 7 , Figure 7 FIG. 1 is a main structural block diagram of a video segment positioning system according to an embodiment of the present invention. Figure 7 As shown, the video segment positioning system in the embodiment of the present invention may include a coarse-grained coding and modal interaction module, a candidate video frequency band construction module, a fine-grained fusion feature acquisition module, a relationship-aware feature acquisition module, and a video positioning result acquisition module. In this embodiment, the coarse-grained coding and modal interaction module may be configured to acquire a query-aware video representation and a video-aware language representation according to the video to be queried and the query statement. The candidate video frequency band construction module may be configured to construct a plurality of candidate video segments of the video to be queried according to the video to be queried; and acquire content features and boundary features of each candidate video segment according to the query-aware video representation. The fine-grained fusion feature acquisition module may be configured to perform fine-grained coding on the query-aware video representation and the video-aware language representation respectively to acquire fine-grained video coding features and fine-grained language coding features; and perform deep fusion of the fine-grained video coding features and the fine-grained language coding features to acquire fine-grained fusion features. The relationship-aware feature acquisition module may be configured to acquire relationship-aware features of each candidate video segment according to the fine-grained fusion features, content features, and boundary features. The video positioning result acquisition module may be configured to acquire the final positioning result of the video segment according to the relationship-aware features.

[0223] In one embodiment, see the attached Figure 8 , Figure 8 FIG. 1 is a main structural block diagram of a video segment positioning system according to an implementation of an embodiment of the present invention. Figure 8As shown, the video segment positioning system may include a video feature extraction module, a coarse-grained video feature acquisition module, a language feature extraction module, a coarse-grained language feature acquisition module, a modal interaction module, a fine-grained video coding feature acquisition module, a fine-grained language coding feature acquisition module, an effective candidate video segment generation module, a fine-grained fusion feature acquisition module, a relationship perception feature acquisition module and a video segment positioning module. The video to be queried is input into the video feature extraction module, the query statement is input into the language feature extraction module, and the video segment positioning module can output the positioning result of the video segment.

[0224] The video clip positioning system is used to perform Figure 1 The video clip positioning method embodiment shown in the figure has similar technical principles, technical problems solved and technical effects produced. Technical personnel in this technical field can clearly understand that for the convenience and conciseness of description, the specific working process and related instructions of the video clip positioning system can refer to the contents described in the embodiment of the video clip positioning method, which will not be repeated here.

[0225] It is understood by those skilled in the art that the present invention implements all or part of the processes in the method of the above embodiment, and can also be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable storage medium may include: any entity or device, medium, U disk, mobile hard disk, disk, optical disk, computer memory, read-only memory, random access memory, electric carrier signal, telecommunication signal and software distribution medium that can carry the computer program code. It should be noted that the content contained in the computer-readable storage medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable storage media do not include electric carrier signals and telecommunication signals.

[0226] Furthermore, the present invention also provides a control device. In an embodiment of a control device according to the present invention, the control device includes a processor and a storage device. The storage device can be configured to store a program for executing the video segment positioning method of the above method embodiment, and the processor can be configured to execute the program in the storage device, which includes but is not limited to the program for executing the video segment positioning method of the above method embodiment. For ease of explanation, only the parts related to the embodiment of the present invention are shown. For specific technical details not disclosed, please refer to the method part of the embodiment of the present invention. The control device can be a control device device formed by various electronic devices.

[0227] Furthermore, the present invention also provides a computer-readable storage medium. In a computer-readable storage medium embodiment according to the present invention, the computer-readable storage medium can be configured to store a program for executing the video segment positioning method of the above method embodiment, and the program can be loaded and run by a processor to implement the above video segment positioning method. For ease of explanation, only the parts related to the embodiment of the present invention are shown. For specific technical details not disclosed, please refer to the method part of the embodiment of the present invention. The computer-readable storage medium can be a storage device formed by various electronic devices. Optionally, the computer-readable storage medium in the embodiment of the present invention is a non-temporary computer-readable storage medium.

[0228] Further, it should be understood that since the setting of each module is only for illustrating the functional units of the device of the present invention, the physical devices corresponding to these modules may be the processor itself, or a part of the software in the processor, a part of the hardware, or a part of the combination of software and hardware. Therefore, the number of each module in the figure is only schematic.

[0229] Those skilled in the art will appreciate that the modules in the device can be adaptively split or merged. Such splitting or merging of specific modules will not cause the technical solution to deviate from the principle of the present invention, and therefore, the technical solutions after splitting or merging will fall within the protection scope of the present invention.

[0230] So far, the technical solutions of the present invention have been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will fall within the protection scope of the present invention.

Claims

1. A video segment positioning method, characterized in that: The method comprises: According to the video to be queried and the query sentence, a query-aware video representation and a video-aware language representation are obtained; Constructing a plurality of valid candidate video segments of the video to be queried according to the video to be queried; and acquiring content features and boundary features of each valid candidate video segment according to the video representation perceived by the query; Fine-grained encoding is performed on the query-aware video representation and the video-aware language representation respectively to obtain fine-grained video encoding features and fine-grained language encoding features; and the fine-grained video encoding features and the fine-grained language encoding features are deeply fused to obtain fine-grained fusion features; Acquire a relation-aware feature of each valid candidate video segment according to the fine-grained fusion feature, the content feature, and the boundary feature; Obtaining a final positioning result of the video segment according to the relationship perception feature; The step of "obtaining a query-aware video representation and a video-aware language representation according to the video to be queried and the query statement" includes: Dividing the video to be queried into multiple video segments; Use a preset video feature extraction model to extract features from each video clip to obtain video features of each video clip; Coarse-grained encoding is performed on the video features of all video clips to obtain coarse-grained video features; Using a preset language feature extraction model to extract features from the query statement to obtain language features of the query statement; Coarse-grained encoding is performed on the language features of the query statement to obtain coarse-grained language features; The coarse-grained video features and the coarse-grained language features are modally interacted to obtain query-aware video representations and video-aware language representations.

2. The video segment positioning method according to claim 1, characterized in that: The step of "constructing a plurality of valid candidate video segments of the video to be queried according to the video to be queried" includes: Construct a two-dimensional time network graph of T×T grids; wherein T is the characteristic length of the video representation perceived by the query, the ordinate of the two-dimensional time network graph represents the start time of the candidate video segment in the video to be queried, and the abscissa represents the end time of the candidate video segment in the video to be queried, and the network in the two-dimensional time network graph whose start time is less than the end time is a valid grid; According to the time interval between the candidate video segments corresponding to each valid grid and the candidate video segments corresponding to other valid grids, the valid grids are sparsely sampled to obtain multiple sampled valid grids, and the candidate video segments corresponding to the sampled valid grids are used as valid candidate video segments.

3. The video segment positioning method according to claim 2, characterized in that: The step of "obtaining content features and boundary features of each valid candidate video segment according to the query-perceived video representation" includes: Obtain the content features of the nth valid candidate video segment according to the following formula and boundary features in, is the query-aware video representation of the start time of the nth valid candidate video segment, It is the query-aware video representation of the end time of the nth valid candidate video segment, MaxPooling is the maximum pooling operation, and Addition is the addition operation.

4. The video segment positioning method according to claim 1, characterized in that: The step of "respectively fine-grained encoding of the query-aware video representation and the video-aware language representation to obtain fine-grained video encoding features and fine-grained language encoding features" includes: The fine-grained video coding features are obtained according to the following formula: in, is the fine-grained video coding feature, is the query-aware video representation, Linear is a linear fully connected layer operation, and ReLU is a linear rectification function; A one-dimensional convolutional network is used to encode the language representation of video perception to obtain unigram language features, bigram language features, and trigram language features. According to the one-gram language features, two-gram language features and three-gram language features, the following formula is applied to obtain the fine-grained language encoding features: in, encoding features for the fine-grained language, They are the uni-gram language feature, the bi-gram language feature and the tri-gram language feature respectively, and Concat is a feature fusion operation.

5. The video segment positioning method according to claim 4, characterized in that: The step of "deeply fusing the fine-grained video coding features and the fine-grained language coding features to obtain fine-grained fusion features" includes: The fine-grained fusion feature is obtained according to the following formula: in, is the fine-grained fusion feature, To query perceived video segment features, is the video segment feature perceived by the video; the query perceived video segment feature is obtained according to the following formula: A C is the set of content features of valid candidate video clips, G Q is a gated language feature, which is obtained by the following formula: σ is the gate function, A B is the set of boundary features of valid candidate video clips, To transfer language features, the transfer language features are obtained by the following formula: Linear is a linear fully connected layer operation, and MaxPooling is a maximum pooling layer operation; The video segment features described by video perception are as follows: Avgpooling is the average pooling layer operation; C is the dimension of the feature.

6. The video segment positioning method according to claim 5, characterized in that: The step of “obtaining the relation-aware feature of each valid candidate video segment according to the fine-grained fusion feature, the content feature and the boundary feature” comprises: The fine-grained fusion feature, the content feature and the boundary feature are fused, and the enhanced fusion feature is obtained according to the following formula: The enhanced fusion features are input into a stacked multi-layer grouped convolutional network, and the relation-aware features of each valid candidate video segment are obtained according to the following formula: is a set of relation-aware features of valid candidate video clips.

7. The video segment positioning method according to claim 6, characterized in that: The step of “obtaining the final positioning result of the video segment according to the relationship perception feature” includes: Scoring each valid candidate video segment according to the relation-perception feature of each valid candidate video segment to obtain a score for each valid candidate video segment; The final positioning result of the video segment is determined according to the score of the valid candidate video segment.

8. The video segment positioning method according to claim 7, characterized in that: The step of "scoring each valid candidate video segment according to the relation-aware features of each valid candidate video segment to obtain a score of each valid candidate video segment" includes: The score of each valid candidate video segment is obtained according to the following formula: Among them, P A is a set of scores of valid candidate video segments.

9. The video segment positioning method according to claim 7, characterized in that: The step of "determining the final video segment positioning result according to the score of the valid candidate video segment" includes: Sorting the scores of the valid candidate video clips in descending order; According to the sorting results and preset requirements, select the valid candidate video segment with the highest score as the final video segment positioning result; or, According to the sorting results and preset requirements, the first k valid candidate video segments are used as the final video segment positioning results.

10. The video segment positioning method according to claim 1, characterized in that: The step of "coarse-grained encoding of video features of all video clips to obtain coarse-grained video features" includes: A one-dimensional convolutional network is used to encode the video features of the video clips, and the encoded video features are reduced in dimension through an average pooling layer network to obtain reduced-dimensional video features; Applying a Bi-GRU network to encode the reduced-dimensional video features to obtain the coarse-grained video features; and / or, The step of "coarse-grained encoding of the language features of the query statement to obtain the coarse-grained language features" includes: A Bi-GRU network is applied to encode the language features of the query sentence to obtain the coarse-grained language features.

11. The video segment positioning method according to claim 1, characterized in that: The step of "modally interacting the coarse-grained video features and the coarse-grained language features to obtain a query-aware video representation and a video-aware language representation" includes obtaining the query-aware video representation and the video-aware language representation according to the following formula: in, a video representation perceived by the query, is the coarse-grained video feature, is the language representation of the video perception, v atten is the weighted sum of the coarse-grained video features, The coarse-grained language feature, T is the length of the coarse-grained video feature, C is the dimension of the coarse-grained video feature and the coarse-grained language feature, L is the length of the coarse-grained language feature, q atten is the weighted sum of the coarse-grained language features, and q is obtained according to the following formula atten : is the jth element of the coarse-grained language feature The attention weight is obtained according to the following formula: a Q is the attention weight matrix of the coarse-grained language feature, is the transpose of the coarse-grained language feature matrix; W Q is the first learnable parameter matrix, b Q is the second learnable parameter matrix; v atten is the weighted sum of the coarse-grained video features, and v is obtained according to the following formula atten : is the jth element of the coarse-grained video feature The attention weight is obtained according to the following formula: a V is the attention weight matrix of the coarse-grained video feature, is the transpose of the coarse-grained video feature matrix.

12. A video clip positioning system, characterized in that: The system comprises: A coarse-grained encoding and modality interaction module, which is configured to obtain a query-aware video representation and a video-aware language representation according to a video to be queried and a query statement; A candidate video frequency band construction module is configured to construct a plurality of candidate video segments of the video to be queried according to the video to be queried; and obtain content features and boundary features of each candidate video segment according to the video representation perceived by the query; A fine-grained fusion feature acquisition module is configured to perform fine-grained encoding on the query-aware video representation and the video-aware language representation, respectively, to obtain fine-grained video encoding features and fine-grained language encoding features; and deeply fuse the fine-grained video encoding features and the fine-grained language encoding features to obtain fine-grained fusion features; a relationship-aware feature acquisition module, configured to acquire a relationship-aware feature of each candidate video segment according to the fine-grained fusion feature, the content feature and the boundary feature; A video positioning result acquisition module, which is configured to acquire a positioning result of a final video segment according to the relationship perception feature; The coarse-grained coding and modal interaction module is further configured as follows: Dividing the video to be queried into multiple video segments; Use a preset video feature extraction model to extract features from each video clip to obtain video features of each video clip; Coarse-grained encoding is performed on the video features of all video clips to obtain coarse-grained video features; Using a preset language feature extraction model to extract features from the query statement to obtain language features of the query statement; Coarse-grained encoding is performed on the language features of the query statement to obtain coarse-grained language features; The coarse-grained video features and the coarse-grained language features are modally interacted to obtain query-aware video representations and video-aware language representations.

13. A control device, comprising a processor and a storage device, wherein the storage device is suitable for storing a plurality of program codes, characterized in that: The program code is suitable for being loaded and run by the processor to execute the video segment positioning method according to any one of claims 1 to 11.

14. A computer-readable storage medium storing a plurality of program codes, characterized in that: The program code is suitable for being loaded and run by a processor to execute the video segment positioning method according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • Method for solving video question and answer task based on multi-mode progressive attention model

    CN113688296A

  • Video clip positioning method and device and computer readable storage medium

    CN113806589A