Weakly Supervised Temporal Boundary Localization Method, Device, Electronic Device and Storage Medium
By constructing a mask reconstruction module, using the complementary reconstruction of visual and text features, the problem of timing boundary positioning under weak supervision is solved, the alignment of video and text features is achieved, and the accuracy and efficiency of timing boundary positioning is improved.
Patent Information
- Application Number
- CN202310118829.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-31
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2043-01-31
AI Technical Summary
The prior art cannot realize timing language positioning and timing action positioning at the same time under weak supervision conditions, resulting in the inability to align language and video features and the inaccurate timing boundary positioning.
By constructing a mask reconstruction module, using the complementary reconstruction of visual and text features, the video frames and text descriptions are extracted and aligned respectively, and positive and negative correlation features are obtained, and the timing boundaries in the video are positioned.
Effective alignment of language and video features under weak supervision conditions is achieved, the accuracy and efficiency of timing boundary positioning is improved, and the cost of manual boundary marking is reduced.
Smart Images

Figure CN116129319B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of video analysis, and particularly relates to a weakly supervised temporal boundary localization method, device, electronic device, and storage medium. Background Art
[0002] Temporal boundary localization, as a research hotspot in the field of video analysis, is crucial for uncropped videos and has great application potential in various scenarios. Temporal boundary localization not only requires annotating the segment intervals where actions occur but also identifying the categories of actions. For example, localizing a video of an athlete sprinting requires determining the start and end intervals of the running segment and simultaneously identifying the action category within that segment as running. Since manually marking the boundaries of videos is time-consuming and laborious, recent research has mainly focused on weakly supervised settings, that is, during the training process, no explicit supervision information about temporal boundaries is provided, but only text descriptions or video-level action labels.
[0003] However, in related technologies, it is difficult to align language and video features, that is, for now, only one of temporal language localization or temporal action localization can be achieved, and it is impossible to complete both localizations simultaneously, and thus it is impossible to align language and video features, and temporal boundary localization cannot be achieved.
[0004] Therefore, there is an urgent need for a weakly supervised solution that can align language and video features to achieve temporal boundary localization. Summary of the Invention
[0005] Embodiments of this application provide a weakly supervised temporal boundary localization method, device, electronic device, and storage medium to solve the problem in related technologies that language and video features cannot be aligned and temporal boundary localization cannot be achieved.
[0006] To solve the above technical problems, the technical solutions adopted in this application are as follows:
[0007] According to one aspect of the present application, a weakly supervised temporal boundary localization method, the method comprising: obtaining a video, and respectively extracting features from each video frame in the video and the text description corresponding to the video to obtain the original features of each video frame and the text features of the text description; the text description is used to describe the action label corresponding to the video; obtaining the positive correlation features and negative correlation features of each video frame according to the correlation between the original features of each video frame and the text features of the text description; using mask reconstruction to align the video features of the video with the text features of the text description to respectively obtain the reconstructed text features and reconstructed video features of each video frame; the video features of the video include the original features, positive correlation features and negative correlation features of each video frame; and localizing the temporal boundary of the action in the video according to the reconstructed text features and reconstructed video features of each video frame to obtain a boundary localization result.
[0008] According to one aspect of the present application, a weakly supervised temporal boundary localization device, the device comprising: an original feature extraction module, configured to obtain a video, and respectively extract features from each video frame in the video and the text description corresponding to the video to obtain the original features of each video frame and the text features of the text description; the text description is used to describe the action label corresponding to the video; a correlation feature extraction module, configured to obtain the positive correlation features and negative correlation features of each video frame according to the correlation between the original features of each video frame and the text features of the text description; a mask reconstruction module, configured to use mask reconstruction to align the video features of the video with the text features of the text description to respectively obtain the reconstructed text features and reconstructed video features of each video frame; the video features of the video include the original features, positive correlation features and negative correlation features of each video frame; and a boundary localization module, configured to localize the temporal boundary of the action in the video according to the reconstructed text features and reconstructed video features of each video frame to obtain a boundary localization result.
[0009] In an exemplary embodiment, the original feature extraction module includes: an original feature extraction unit, configured to extract features from each video frame in the video to obtain the original features of each video frame; a text conversion unit, configured to perform action category prediction on the original features of each video frame to obtain the action label corresponding to the video, and convert the action label corresponding to the video into the text description corresponding to the video; and a text feature extraction unit, configured to extract features from the text description corresponding to the video to obtain the text features of the text description.
[0010] In an exemplary embodiment, the correlation feature extraction module includes: a score calculation unit, configured to calculate a positive correlation score and a negative correlation score between each video frame and the text description respectively according to the correlation between the original features of each video frame and the text features of the text description; a first feature calculation unit, configured to calculate the positive correlation features of each video frame according to the calculated positive correlation score and the original features of each video frame; and a second feature calculation unit, configured to calculate the negative correlation features of each video frame according to the calculated negative correlation score and the original features of each video frame.
[0011] In an exemplary embodiment, the mask reconstruction module includes: a text reconstruction unit, configured to guide the reconstruction of the masked text features from the video features of the video based on mask reconstruction to obtain the reconstructed text features of each video frame; and a video reconstruction unit, configured to guide the reconstruction of the masked video features from the text features of the text description to obtain the reconstructed video features of each video frame.
[0012] In an exemplary embodiment, the text reconstruction unit includes a text feature reconstruction subunit, configured to perform mask reconstruction on the masked text features respectively according to the original features, positive correlation features and negative correlation features of each video frame to obtain multiple reconstructed text features of each video frame; and the video reconstruction unit includes a video feature reconstruction subunit, configured to perform mask reconstruction on the original features, positive correlation features and negative correlation features of the masked video frame respectively according to the text features of the text description to obtain multiple reconstructed video features of each video frame.
[0013] In an exemplary embodiment, the temporal boundary localization is implemented by invoking a temporal boundary localization network, which is a machine learning model that has been trained and has the ability to locate the temporal boundaries of actions in the video. Among them, the temporal boundary localization network includes a proposal generation module for feature extraction, a text mask reconstruction module for text feature reconstruction, and a video mask reconstruction module for video feature reconstruction.
[0014] In an exemplary embodiment, the device includes a training module, and the training module includes: a text loss calculation unit, configured to calculate a text ranking loss and a text reconstruction loss according to the reconstructed text features of the video frame in the text mask reconstruction module, and use the text ranking loss to constrain the text reconstruction loss of the video frame with respect to the reconstructed text features; a video loss calculation unit, configured to calculate a video ranking loss and a video reconstruction loss according to the reconstructed video features of the video frame in the video mask reconstruction module, and use the video ranking loss to constrain the video reconstruction loss of the video frame with respect to the reconstructed video features; a training unit, configured to complete the training and obtain the temporal boundary localization network if the text reconstruction loss indicates that the reconstruction effect of the reconstructed text features reconstructed from the positive correlation features of the video frame is the best, and the video reconstruction loss indicates that the reconstruction effect of the reconstructed video features related to the positive correlation features is the best.
[0015] According to one aspect of the present application, an electronic device includes a processor and a memory, and a computer-readable instruction is stored on the memory. When the computer-readable instruction is executed by the processor, the weakly supervised temporal boundary localization method as described above is implemented.
[0016] According to one aspect of the present application, a storage medium stores a computer program, and when the computer program is executed by a processor, the weakly supervised temporal boundary localization method as described above is implemented.
[0017] According to one aspect of the present application, a computer program product includes a computer program. The computer program is stored in a storage medium, and a processor of a computer device reads the computer program from the storage medium and executes the computer program, so that the computer device implements the weakly supervised temporal boundary localization method as described above when executed.
[0018] In the above technical solution, the present application realizes a weakly supervised temporal boundary localization method that can align language and video features to achieve temporal boundary localization.
[0019] Specifically, the present application first constructs a proposal generation module and a mask reconstruction module. The mask reconstruction module includes two branches, namely a visual mask reconstruction branch and a text mask reconstruction branch, corresponding to the text mask reconstruction module and the video mask reconstruction module respectively. The text mask reconstruction module for visual perception uses video features as visual auxiliary information to help reconstruct the text features with some information masked. The video mask reconstruction module for text perception uses text features as language auxiliary information to help reconstruct the video features with some information masked. The two modules can capture highly matching video features and text features, thereby achieving accurate temporal boundary localization of actions in the video. In other words, the present application complements the visual mask reconstruction branch with the text mask reconstruction branch, unifies the temporal language localization and the temporal action localization, and designs a temporal boundary localization network to complete the temporal boundary localization of actions in the video. By using the way of mutual complementation and mutual guidance between text and visual reconstruction, the alignment effect of text features and video features is improved, so as to better achieve the temporal boundary localization effect.
[0020] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.
[0022] Figure 1 is a schematic diagram of the implementation environment related to the present application.
[0023] Figure 2 is a flowchart of a weak supervision temporal boundary localization method shown according to an exemplary embodiment;
[0024] Figure 3 is Figure 2 a schematic diagram of steps 110 and 130 in the corresponding embodiment in one embodiment;
[0025] Figure 4 is Figure 2 a schematic diagram of step 150 in the corresponding embodiment in one embodiment;
[0026] Figure 5 is Figure 2 a schematic diagram of step 150 in the corresponding embodiment in one embodiment;
[0027] Figure 6 is a comparison diagram of the effects with various prior arts in the application scenario according to an exemplary embodiment;
[0028] Figure 7It is a block diagram of a weakly supervised temporal boundary localization device shown according to an exemplary embodiment;
[0029] Figure 8 It is a hardware structure diagram of an electronic device shown according to an exemplary embodiment;
[0030] Figure 9 It is a block diagram of an electronic device shown according to an exemplary embodiment. Detailed implementation manners
[0031] Embodiments of the present application will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are only used to explain the present application, and should not be construed as a limitation to the present application.
[0032] Those skilled in the art of the present technology can understand that, unless specifically stated otherwise, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the term "including" used in the specification of the present application means the presence of the described features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or their groups. It should be understood that when we say an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any unit and all combinations of one or more related listed items.
[0033] The following is an introduction and explanation of several terms related to the present application:
[0034] Mask reconstruction: Mask reconstruction refers to the process of covering a certain proportion of data in a given data, such as video, text, etc., and inputting the masked data into a certain model to try to reconstruct the original data.
[0035] Weakly supervised temporal language localization: Weakly supervised temporal language localization refers to the ability to detect the segment interval in a video that is semantically consistent with a given description given an uncropped video and a description, without providing supervision information for the segment interval during the entire training process.
[0036] Weakly-Supervised Temporal Action Localization: Weakly-supervised temporal action localization refers to the ability to detect the start and end intervals of an action in an uncropped long video containing the action, and to determine the category of the action within the segment interval from the start interval to the end interval, while only providing video-level action category supervision during the training phase of the entire process, without providing supervision information for the segment interval.
[0037] Currently, temporal boundary localization is crucial for uncropped videos, especially long uncropped videos containing actions. Among them, temporal boundary localization includes temporal language localization and temporal action localization, etc. Since manually marking the boundaries of videos is time-consuming and laborious, existing temporal boundary localization mainly focuses on weakly-supervised settings, that is, during the training phase, no supervision information indicating the temporal boundaries is provided, but only video-level action labels or descriptions related to the actions in the video.
[0038] However, in most weakly-supervised schemes, it is difficult to align language and video features. In other words, in most weakly-supervised schemes, only one of temporal language localization or temporal action localization can be achieved temporarily, and the two localizations cannot be completed simultaneously, thus unable to align language and video features, and thus unable to achieve relatively accurate temporal boundary localization.
[0039] As can be seen from the above, there are still limitations in the related technologies that language and video features cannot be aligned and temporal boundary localization cannot be achieved.
[0040] Therefore, the weakly-supervised temporal boundary localization method provided in this application first extracts features from each video frame in the video and the corresponding text description of the video to obtain the original features of each video frame and the text features of the text description; then, according to the correlation between the original features of each video frame and the text features of the text description, the positive correlation features and negative correlation features of each video frame are obtained; furthermore, mask reconstruction is used to align the video features of the video with the text features of the text description to obtain the reconstructed text features and reconstructed video features of each video frame respectively, where the video features of the video include the original features, positive correlation features, and negative correlation features of each video frame; then, according to the reconstructed text features and reconstructed video features of each video frame, the temporal boundaries of the actions in the video are located to obtain the boundary localization result, which is used to indicate the start video frame and end video frame visual perception video features of the action in the video, thus realizing a relatively accurate weakly-supervised temporal boundary localization method, which can effectively enhance the alignment degree of the video and the text and improve the temporal boundary localization ability. This weakly-supervised temporal boundary localization method is applicable to a weakly-supervised temporal boundary localization device, which can be deployed on an electronic device, and the electronic device can be a computer device configured with a von Neumann architecture. For example, the computer device can be a desktop computer, a laptop computer, a server, etc.
[0041] Figure 1 It is a schematic diagram of the implementation environment of a weakly supervised temporal boundary localization method. This implementation environment includes a user terminal 110, a collection end 130, a gateway 150, a server end 170, and a router 190.
[0042] Specifically, the user terminal 110, which can also be considered as the user side or the terminal, can receive and / or view the temporal boundary localization result. This user terminal 110 can be an electronic device such as a smart phone, a tablet computer, a notebook computer, a desktop computer, etc., and is not limited here.
[0043] The collection end 130 is used to collect videos. This collection end can be an electronic device with an image collection function such as a camera, a video recorder, etc., and is not limited here either.
[0044] The interaction between the server end 170 and the collection end 130 and the user terminal 110 can be realized through a local area network or a wide area network. In an application scenario, the server end 170 establishes a communication connection in a wired or wireless manner with the gateway 150 through the router 190. For example, this wired or wireless manner includes but is not limited to WIFI, etc., so that the server end 170 and the gateway 150 are deployed in the same local area network, and further enables the collection end 130 and the user terminal 110 to realize the interaction with the server end 170 through the local area network path. This local area network includes but is not limited to: ZIGBEE or Bluetooth. In another application scenario, a communication connection in a wired or wireless manner is established between the server end 170 and the collection end 130 or the user terminal 110. For example, this wired or wireless manner includes but is not limited to 2G, 3G, 4G, 5G, WIFI, etc., so that the server end 170 and the collection end 130 or the user terminal 110 are deployed in the same wide area network, and further enables the server end 170 to realize the interaction with the collection end 130 or the user terminal 110 through the wide area network path.
[0045] Among them, the server end 170 can be a single server, or a server cluster composed of multiple servers, or a cloud, a cloud platform, a cloud computing center, etc. composed of multiple servers, so as to better provide background services. For example, the background services include temporal boundary localization services.
[0046] With the interaction between the collection end 130 and the server end 170, after the collection end 130 collects a video containing actions, it can be transmitted to the server end 170 to request the server end 170 to provide temporal boundary localization services.
[0047] For the server - side 170, it can receive the video sent by the acquisition - side 130, and then locate the temporal boundaries of the actions in the video to obtain the boundary - location result, which can be transmitted to the user terminal 110 for the user to view. Specifically, by introducing two coupled mask - reconstruction branches, namely, the visually - aware text mask - reconstruction branch and the text - aware visual mask - reconstruction branch, after obtaining the video, the text - aware visual mask - reconstruction branch is used as a visual aid to reconstruct the text features with some information masked, and at the same time, the visually - aware text mask - reconstruction branch is used as a language aid to reconstruct the video features with some information masked. Through this symmetric and complementary reconstruction supervision, the robust feature matching between the video and the text can be cooperatively carried out, unifying the temporal language localization and the temporal action localization in a concise manner, and finally achieving relatively accurate temporal boundary localization.
[0048] Please refer to Figure 2 , this embodiment of the present application provides a weakly - supervised temporal boundary - location method, which is applicable to electronic devices. For example, the electronic device can be a desktop computer, a laptop computer, a server, etc.
[0049] In the following method embodiments, for the sake of description, the execution subject of each step of the method is taken as an example of an electronic device for illustration, but it is not specifically limited thereto.
[0050] As Figure 2 shown, the method may include the following steps:
[0051] Step 110, obtain a video, and respectively extract features from each video frame in the video and the text description corresponding to the video to obtain the original features of each video frame and the text features of the text description.
[0052] First of all, it should be noted that the video can be obtained by an image - acquisition device shooting and collecting the actions in a scene. Among them, the image - acquisition device can be an electronic device with an image - acquisition function such as a camera, a smart phone equipped with a camera, etc.
[0053] Regarding the acquisition of the video, the video can be from the video shot and collected by the image - acquisition device in real - time, or it can be a video shot and collected by the image - acquisition device in a historical time period pre - stored in the electronic device. Then, for the electronic device, after the image - acquisition device shoots and collects the video, it can be processed in real - time, or it can be pre - stored and then processed. For example, the video can be processed when the CPU of the electronic device is low, or processed according to the instructions of the staff. Thus, the temporal boundary - location in this embodiment can be for the video obtained in real - time or for the video obtained in a historical time period, and no specific limitation is made here.
[0054] Secondly, the text description is used to describe the action tags corresponding to the video. The action tags are used to indicate the action categories to which the actions in each video frame of the video belong, and the action tags in this embodiment are at the video level. It should be noted here that the action tags at the video level indicate that the actions in each video frame of the video belong to the same action category. It can be understood that when detecting the action category of a whole video, often only some video frames contain actions. For example, in a video of a person climbing stairs, in addition to some video frames containing the action of climbing stairs, there are very likely to be video frames that do not contain the action of climbing stairs, such as random activities before climbing stairs. However, the action tags at the video level mean that the actions of all video frames in the whole video are marked as the action of climbing stairs.
[0055] Step 130: Obtain the positive correlation features and negative correlation features of each video frame according to the correlation between the original features of each video frame and the text features of the text description.
[0056] It can be understood that if the action in the video frame conforms to the text description, then the video frame is a positive instance, which can also be understood as that the video frame has a high positive correlation degree with the text description; otherwise, if the action in the video frame does not conform to the text description, then the video frame is a negative instance, that is, it is considered that the video frame has a high negative correlation degree with the text description. Based on this, the correlation essentially reflects the positive and negative correlation degrees between the video frame and the text description.
[0057] The original features, positive correlation features, and negative correlation features of each video frame are used for the subsequent mask reconstruction process. Among them, the original features are the video features directly obtained by extracting features from the video frame; the positive correlation features are the video features obtained by judging the correlation between the original features of the video frame and the text features of the text description, and are used to represent the positive correlation degree between the video frame and the text description; the negative correlation features are the video features obtained by judging the correlation between the original features of the video frame and the text features of the text description, and are used to represent the negative correlation degree between the video frame and the text description.
[0058] Step 150: Align the video features of the video with the text features of the text description by using mask reconstruction, and obtain the reconstructed text features and reconstructed video features of each video frame respectively.
[0059] Among them, the video features of the video include the original features, positive correlation features, and negative correlation features of each video frame.
[0060] In a possible implementation manner, the mask reconstruction process may include the following steps: Based on mask reconstruction, guide the video features of the video to reconstruct the masked text features to obtain the reconstructed text features of each video frame; based on mask reconstruction, guide the text features of the text description to reconstruct the masked video features to obtain the reconstructed video features of each video frame.
[0061] In a possible implementation, the reconstructed text features of each video frame include: reconstructed text features reconstructed based on the positive correlation features of each video frame, reconstructed text features reconstructed based on the original features of each video frame, and reconstructed text features reconstructed based on the negative correlation features of each video frame.
[0062] In a possible implementation, the reconstructed video features of each video frame include: reconstructed video features related to the positive correlation features of each video frame reconstructed based on the text features of the text description, reconstructed video features related to the original features of each video frame reconstructed based on the text features of the text description, and reconstructed video features related to the negative correlation features of each video frame reconstructed based on the text features of the text description.
[0063] As can be seen from the above, the embodiments of the present application use video features as visual assistance to reconstruct text features with some information masked, and use text features as text assistance to reconstruct video features with some information masked. By complementing the visual mask reconstruction branch and the text mask reconstruction branch, the temporal language localization and the temporal action localization are unified, improving the alignment effect of the text and the video, and thus enabling a better temporal boundary localization effect.
[0064] Step 170: According to the reconstructed text features and reconstructed video features of each video frame, localize the temporal boundary of the action in the video to obtain a boundary localization result.
[0065] Among them, the boundary localization result is used to indicate the starting video frame and the ending video frame of the action in the video. It can also be considered that the segment interval of the action in the video is from the starting interval marked by the starting video frame to the ending interval marked by the ending video frame.
[0066] Specifically, post-process the reconstructed text features and reconstructed video features of each video frame to obtain the positive correlation score of the video, normalize the positive correlation score and generate a waveform, and obtain the starting video frame and the ending video frame where the action in the video is located by circularly searching for the peaks of the waveform, so as to obtain the final boundary localization result, thereby fully ensuring the high precision of the boundary localization result. For example, in Figure 6 Application scenario a, if the user wants to query the video segment of a person running up the stairs, in the temporal localization result of the embodiment DM2 of the present application, the 1.23 temporal starting point of the first picture is located as the starting video frame, and the 9.84 temporal ending point of the fourth picture is located as the ending video frame. It can be clearly observed from each video frame that this video segment from the starting video frame to the ending video frame accurately describes the action of a person running up the stairs.
[0067] In one possible implementation, the boundary positioning result is used to indicate the timing start point and timing end point of the action in the video. Then, the segment interval of the action in the video is from the start interval marked by the timing start point to the end interval marked by the timing end point.
[0068] Through the above process, the visual mask reconstruction branch and the text mask reconstruction branch are complemented, the temporal language localization and temporal action localization are unified, and a temporal boundary localization network is designed to complete the temporal boundary localization of actions in the video. The text and visual reconstruction are used to complement and guide each other, which improves the alignment effect of text features and video features, thereby achieving a better temporal boundary localization effect.
[0069] In an exemplary embodiment, the timing boundary positioning is implemented by calling the timing boundary positioning network. Figure 3 , Figure 4 and Figure 5 The schematic diagram of the structure of the timing boundary positioning network in one embodiment is shown together. The timing boundary positioning network includes Figure 3 The proposal generation module for feature extraction in Figure 4 The text mask reconstruction module for text feature reconstruction in Figure 5 The video mask reconstruction module for video feature reconstruction in . Specifically, Figure 3 In , the proposal generation module includes a video feature extractor, an action classifier, a text feature extractor, and a proposal generator; in Figure 4 In , the text mask reconstruction module includes a visually aware text decoder; Figure 5 In , the video mask reconstruction module includes a text-aware visual decoder.
[0070] Now, combined with the structure of the temporal boundary positioning network, the temporal boundary positioning process is described in detail as follows:
[0071] In an exemplary embodiment, step 110 may include the following steps: performing feature extraction on each video frame in the video to obtain the original features of each video frame; performing action category prediction on the original features of each video frame to obtain the action label corresponding to the video, and converting the action label corresponding to the video into a text description corresponding to the video; performing feature extraction on the text description corresponding to the video to obtain the text features of the text description.
[0072] Specifically, Figure 3 As shown, the original features of each video frame are obtained by extracting features of the input video frames using a video feature extractor, the action label corresponding to the video is obtained by predicting the action category of the original features of each input video frame using an action classifier, and the text features of the text description are obtained by extracting features of the input text description using a text feature extractor.
[0073] In an exemplary embodiment, step 130 may include the following steps: calculating a positive correlation score and a negative correlation score between each video frame and the text description respectively according to the correlation between the original features of each video frame and the text features of the text description; calculating the positive correlation features of each video frame according to the calculated positive correlation score and the original features of each video frame; and calculating the negative correlation features of each video frame according to the calculated negative correlation score and the original features of each video frame.
[0074] As Figure 3 shown, after obtaining the original features of each video frame and the text features of the text description, the proposal generator can be used to calculate the positive correlation score and the negative correlation score based on the correlation between the original features of each video frame and the text features of the text description, and then calculate the positive correlation features of each video frame according to the positive correlation score and the original features of each video frame. Similarly, the negative correlation features of each video frame are calculated according to the negative correlation score and the original features of each video frame.
[0075] Specifically, the calculation formulas for the positive correlation features and negative correlation features of each video frame are as follows:
[0076] h = CrossTrans(v, q)
[0077] w pos = σ[Conv(h, ∼φ p )], v pos = w pos ·v
[0078] w neg = σ[Conv(h, ∼φ n )], v neg = w neg ·v.
[0079] Among them, w pos represents the positive correlation score, w neg represents the negative correlation score, v pos represents the positive correlation features, v neg represents the negative correlation features, v represents the original features, q represents the text features, and h represents the original parameters of the proposal generator.
[0080] Through the above process, by converting the action labels into descriptive statements, the temporal language localization and the temporal action localization can be unified, which is conducive to subsequent alignment of language and video features, and thus the temporal boundary localization can be achieved.
[0081] In an exemplary embodiment, step 150 may include the following steps: for each video frame, perform masked reconstruction on the masked text features according to the original features, positively correlated features, and negatively correlated features of the video frame respectively to obtain multiple reconstructed text features of the video frame; perform masked reconstruction on the original features, positively correlated features, and negatively correlated features of the masked video frame according to the text features of the text description to obtain multiple reconstructed video features of the video frame.
[0082] As Figure 4 shown, based on the text masked reconstruction module, input the text features of the masked text description, as well as the original features, positively correlated features, and negatively correlated features of each video frame, into the text decoder of visual perception to obtain the reconstructed text features reconstructed from the positively correlated features of each video frame, the reconstructed text features reconstructed from the original features of each video frame, and the reconstructed text features reconstructed from the negatively correlated features of each video frame respectively.
[0083] Similarly, as Figure 5 shown, based on the video masked reconstruction module, input the text features of the text description, as well as the original features, positively correlated features, and negatively correlated features of each masked video frame, into the visual decoder of text perception to obtain the reconstructed video features related to the positively correlated features reconstructed from the text features of the text description, the reconstructed video features related to the original features reconstructed from the text features of the text description, and the reconstructed video features related to the negatively correlated features reconstructed from the text features of the text description respectively.
[0084] Specifically, the calculation formulas for the above-mentioned multiple reconstructed text features are as follows:
[0085]
[0086]
[0087]
[0088] where, v pos represents the positively correlated feature, v neg represents the negatively correlated feature, v represents the original feature, represents the reconstructed text feature reconstructed under the guidance of the positively correlated feature, represents the reconstructed text feature reconstructed under the guidance of the original feature, represents the reconstructed text feature reconstructed under the guidance of the negatively correlated feature, and CrossTrans represents the text decoder of visual perception.
[0089] Similarly, the calculation formulas for the above-mentioned multiple reconstructed video features are as follows:
[0090]
[0091]
[0092]
[0093] Among them, q represents the text feature of the text description. represents the reconstructed video feature related to the positive correlation feature obtained by text feature-guided reconstruction. represents the reconstructed video feature related to the original feature obtained by text feature-guided reconstruction. represents the reconstructed video feature related to the negative correlation feature obtained by text feature-guided reconstruction. CrossTrans represents the text-aware visual decoder.
[0094] Under the action of the above embodiments, two coupled masked reconstruction branches are introduced, namely the visually aware text masked reconstruction branch and the text-aware visual masked reconstruction branch. The video feature is used as a visual aid to reconstruct the text feature with partially masked information, and the text feature is used as a text aid to reconstruct the video feature with partially masked information. By complementing the visual masked reconstruction branch and the text masked reconstruction branch, the temporal language localization and the temporal action localization are unified, the alignment effect of the text and the video is improved, and thus a better temporal boundary localization effect can be achieved.
[0095] In an exemplary embodiment, the temporal boundary localization network is a machine learning model that is trained and has the ability to locate the temporal boundaries of actions in a video.
[0096] Specifically, the training process of the temporal boundary localization network includes the following steps:
[0097] In the text masked reconstruction module, the text ranking loss and the text reconstruction loss are calculated according to the reconstructed text feature of the video frame, and the text reconstruction loss of the video frame with respect to the reconstructed text feature is constrained by the text ranking loss; in the video masked reconstruction module, the video ranking loss and the video reconstruction loss are calculated according to the reconstructed video feature of the video frame, and the video reconstruction loss of the video frame with respect to the reconstructed video feature is constrained by the video ranking loss; if the text reconstruction loss indicates that the reconstruction effect of the reconstructed text feature reconstructed from the positive correlation feature of the video frame is the best, and the video reconstruction loss indicates that the reconstruction effect of the reconstructed video feature related to the positive correlation feature is the best, the training is completed, and the temporal boundary localization network is obtained.
[0098] Of course, in other embodiments, the training completion condition is not limited to the best reconstruction effect of the reconstructed text features obtained by reconstructing the positive correlation features of the video frames and the best reconstruction effect of the reconstructed video features related to the positive correlation features obtained by reconstruction. It can also be that the reconstruction effect of the reconstructed text features obtained by reconstructing the positive correlation features of the video frames is the best, the reconstruction effect of the reconstructed text features obtained by reconstructing the original features is the second best, the reconstruction effect of the reconstructed text features obtained by reconstructing the negative correlation features is the worst, and the reconstruction effect of the reconstructed video features related to the positive correlation features obtained by reconstruction is the best, the reconstruction effect of the reconstructed video features related to the original features is the second best, and the reconstruction effect of the reconstructed video features related to the negative correlation features is the worst. This is not a specific limitation here.
[0099] Specifically, the calculation formula of the text reconstruction loss is as follows:
[0100]
[0101] Among them, q represents the text features of the text description, represents the reconstructed text features obtained by guiding the reconstruction with positive correlation features, represents the reconstructed text features obtained by guiding the reconstruction with the original features, L ce represents the reconstruction loss of each reconstructed video feature, represents the text reconstruction loss.
[0102] Furthermore, the calculation formula of the text sorting loss is as follows:
[0103]
[0104] Among them, represents the text sorting loss, and α1 and α2 represent the parameters of the text decoder for visual perception.
[0105] Furthermore, the calculation formula of the video reconstruction loss is as follows:
[0106]
[0107] Among them, v represents the original features, represents the reconstructed video features related to the positive correlation features obtained by guiding the reconstruction with the text features, represents the reconstructed video features related to the original features obtained by guiding the reconstruction with the text features, L mse represents the reconstruction loss of each reconstructed text feature, represents the video reconstruction loss.
[0108] Furthermore, the calculation formula of the video sorting loss is as follows:
[0109]
[0110] Among them, represents the video sorting loss, and β1 and β2 represent the parameters of the text-aware visual decoder.
[0111] Through the cooperation of the above embodiments, video features are used as visual assistance to reconstruct text features with partially masked information. At the same time, text features are used as language assistance to reconstruct video features with partially masked information. Through this symmetric complementary reconstruction supervision, the features between video and text are robustly matched in cooperation, unifying temporal language localization and temporal action localization in a concise manner, and then achieving effective alignment of language and video features, thus solving the problem in the related art that language and video features cannot be aligned and temporal boundary localization cannot be achieved.
[0112] This application has conducted a large number of experiments on a general positioning dataset. Specifically, the performance of this application and each existing technology on the Charades-STA dataset is shown in Table 1, the performance of this application and each existing technology on the AcaptivityNet Captions dataset is shown in Table 2, the performance of this application and each existing technology on the ActivityNet v1.3 dataset is shown in Table 3, and the comparison results of different branches of this application are shown in Table 4.
[0113] Table 1 Performance of this application and each existing technology on the Charades-STA dataset
[0114]
[0115] Table 2 Performance of this application and each existing technology on the AcaptivityNet Captions dataset
[0116]
[0117]
[0118] Table 3 Performance of this application and each existing technology on the ActivityNet v1.3 dataset
[0119]
[0120] Table 4 Comparison results of different branches of this application
[0121]
[0122] According to the performance of the embodiments of the present application and various existing technologies shown in Table 1, Table 2, and Table 3 on the datasets Charades-STA, ActivityNet Captions, and ActivityNet v-1.3, where the masking ratio represents the text masking ratio parameter and the visual masking ratio parameter. The text masking ratio parameter refers to the ratio of the number of masked text features to all words in the language, and the visual masking ratio parameter refers to the ratio of the number of masked video frames to the total number of video frames. In all the localization datasets, the indicators of the present application are the highest. Therefore, the present application has reached the state-of-the-art level in both weakly supervised temporal language localization and temporal action localization.
[0123] According to the performance of the embodiments of the present application in different branch cases shown in Table 4, obviously, when both branches are applied to the temporal boundary localization network, through the complementary guidance of the two masked reconstruction branches, the temporal boundary localization network can achieve the best temporal boundary localization effect.
[0124] Figure 6 The figure shows a comparison diagram of the effects of the embodiments of the present application and various existing technologies in the application scenario. DM2 represents the temporal localization result of the present application. The data represents the temporal start point in the identified starting video frame and the temporal end point in the ending video frame. The others are the temporal localization results of the existing technologies. Application scenario a in the figure is the video segment of a person running up the stairs queried. Obviously, the temporal localization result of the present application is the most accurate; Application scenario b is the video segment of a person pressing the light switch by the door queried. Application scenario c is the video segment of hand washing clothes queried. Application scenario d is the video segment of rope skipping queried. It can be clearly observed that the temporal localization results of the present application are the most accurate in each application scenario.
[0125] The following is an embodiment of the device of the present application, which can be used to execute the weakly supervised temporal boundary localization method involved in the present application. For the details not disclosed in the embodiment of the device of the present application, please refer to the method embodiment of the weakly supervised temporal boundary localization method involved in the present application.
[0126] Please refer to Figure 7 , in an exemplary embodiment, a weakly supervised temporal boundary localization device 700.
[0127] The device 700 includes but is not limited to: an original feature extraction module 710, a correlation feature extraction module 730, a masked reconstruction module 750, and a boundary localization module 770.
[0128] Among them, the original feature extraction module 710 is used to obtain a video and respectively extract features from each video frame in the video and the text description corresponding to the video, so as to obtain the original features of each video frame and the text features of the text description; the text description is used to describe the action label corresponding to the video.
[0129] The correlation feature extraction module 730 is configured to obtain the positive correlation features and negative correlation features of each video frame according to the correlation between the original features of each video frame and the text features of the text description.
[0130] The mask reconstruction module 750 is configured to align the video features of the video with the text features of the text description by using mask reconstruction, and obtain the reconstructed text features and reconstructed video features of each video frame respectively; the video features of the video include the original features, positive correlation features and negative correlation features of each video frame.
[0131] The boundary localization module 770 is configured to localize the temporal boundaries of the actions in the video according to the reconstructed text features and reconstructed video features of each video frame, and obtain the boundary localization result.
[0132] It should be noted that when the weakly supervised temporal boundary localization device provided in the above embodiments performs temporal boundary localization, only the above-mentioned division of each functional module is used for illustration. In practical applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the weakly supervised temporal boundary localization device will be divided into different functional modules to complete all or part of the functions described above.
[0133] In addition, the weakly supervised temporal boundary localization device provided in the above embodiments and the embodiments of the weakly supervised temporal boundary localization method belong to the same concept. The specific manners in which each module performs operations have been described in detail in the method embodiments, and will not be elaborated here.
[0134] Figure 8 The structural schematic of an electronic device shown according to an exemplary embodiment. This electronic device is applicable to Figure 1 the server side 170 in the shown implementation environment.
[0135] It should be noted that this electronic device is only an example adapted to this application, and cannot be considered as providing any limitation to the scope of use of this application. This electronic device cannot be interpreted as requiring dependence on or necessarily having Figure 8 one or more components in the shown exemplary electronic device 2000.
[0136] The hardware structure of the electronic device 2000 may vary greatly due to different configurations or performances. For example, Figure 8 as shown, the electronic device 2000 includes: a power supply 210, an interface 230, at least one memory 250, and at least one central processing unit (CPU) 270.
[0137] Specifically, the power supply 210 is used to provide operating voltage for each hardware device on the electronic device 2000.
[0138] The interface 230 includes at least one wired or wireless network interface 231 for interacting with external devices. For example, in Figure 1 the illustrated implementation environment, the interaction between the server side 170, the acquisition side 130, and the user terminal 110.
[0139] Of course, in other examples adapted to this application, the interface 230 may further include at least one serial-to-parallel conversion interface 233, at least one input / output interface 235, and at least one USB interface 237, etc., as Figure 8 shown, and specific limitations are not imposed herein.
[0140] The memory 250, as a carrier for resource storage, can be a read-only memory, a random access memory, a magnetic disk, or an optical disc, etc. The resources stored thereon include an operating system 251, application programs 253, and data 255, etc., and the storage method can be transient storage or permanent storage.
[0141] Among them, the operating system 251 is used to manage and control each hardware device and application program 253 on the electronic device 2000 to enable the central processing unit 270 to perform operations and processing on the massive data 255 in the memory 250. It can be WindowsServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.
[0142] The application program 253 is a computer program that completes at least one specific task based on the operating system 251. It may include at least one module ( Figure 8 not shown), and each module can separately contain a computer program for the electronic device 2000. For example, the weakly supervised temporal boundary localization device can be regarded as an application program 253 deployed on the electronic device 2000.
[0143] The data 255 can be photos, pictures, etc. stored on a magnetic disk, and can also be video data, etc., stored in the memory 250.
[0144] The central processing unit 270 may include one or more than one processors and is configured to communicate with the memory 250 through at least one communication bus to read the computer programs stored in the memory 250, and then perform operations and processing on the massive data 255 in the memory 250. For example, the weakly supervised temporal boundary localization method is completed in the form of reading a series of computer programs stored in the memory 250 by the central processing unit 270.
[0145] In addition, the present application can also be implemented by a hardware circuit or a combination of a hardware circuit and software. Therefore, the implementation of the present application is not limited to any specific hardware circuit, software, or the combination of both.
[0146] Please refer to Figure 9 , in an embodiment of the present application, an electronic device 4000 is provided, and the electronic device 400 may include: a desktop computer, a laptop computer, a server, etc.
[0147] In Figure 9 , the electronic device 4000 includes at least one processor 4001, at least one communication bus 4002, and at least one memory 4003.
[0148] Among them, the processor 4001 and the memory 4003 are connected, such as through the communication bus 4002. Optionally, the electronic device 4000 may further include a transceiver 4004, and the transceiver 4004 may be used for data interaction between the electronic device and other electronic devices, such as data sending and / or data receiving, etc. It should be noted that in practical applications, the transceiver 4004 is not limited to one, and the structure of the electronic device 4000 does not constitute a limitation to the embodiments of the present application.
[0149] The processor 4001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logical blocks, modules, and circuits described in combination with the disclosure of the present application. The processor 4001 may also be a combination that implements a computing function, such as a combination including one or more microprocessors, a combination of a DSP and a microprocessor, etc.
[0150] The communication bus 4002 may include a path for transmitting information between the above components. The communication bus 4002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. The communication bus 4002 may be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 9It is represented only by a thick line, but it does not mean that there is only one bus or one type of bus.
[0151] The memory 4003 can be a ROM (Read Only Memory), or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory), or other types of dynamic storage devices that can store information and instructions. It can also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.
[0152] A computer program is stored on the memory 4003, and the processor 4001 reads the computer program stored in the memory 4003 through the communication bus 4002.
[0153] When the computer program is executed by the processor 4001, it implements the weakly supervised temporal boundary localization method in the above-mentioned various embodiments.
[0154] In addition, an embodiment of the present application provides a storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the weakly supervised temporal boundary localization method in the above-mentioned various embodiments.
[0155] An embodiment of the present application provides a computer program product, which includes a computer program stored in a storage medium. The processor of the computer device reads the computer program from the storage medium, and the processor executes the computer program, so that the computer device executes the weakly supervised temporal boundary localization method in the above-mentioned various embodiments.
[0156] Compared with the related art, the beneficial effects of the present application are:
[0157] 1. The present application proposes a new weakly-supervised temporal boundary localization method. By reconstructing masked features through two modalities, it can effectively enhance the alignment between video and text and improve the temporal boundary localization ability. By introducing two coupled reconstruction branches, namely the visually-perceptive text mask reconstruction module (C-MTM) and the text-perceptive video mask reconstruction module (T-MCM), it uses video features as visual assistance to reconstruct text features with partially masked information. Meanwhile, it uses text features as language assistance to reconstruct video features with partially masked information. Through this symmetric and complementary reconstruction supervision, it can cooperate to perform robust feature matching between video and text, unify temporal language localization and temporal action localization in a concise manner, and achieve the state-of-the-art level in both weakly-supervised temporal language localization and temporal action localization.
[0158] 2. Under the weakly-supervised setting, the present application minimizes the labor cost of boundary marking as much as possible. The present application designs a dual-branch mask reconstruction network (DM2). Through the method of dual-branch mask reconstruction, it aligns language and video features, well solves the problem of temporal boundary localization, and still obtains good temporal boundary localization results.
[0159] 3. Based on the idea of reconstruction, the present application designs a visually-perceptive text mask reconstruction branch and a text-perceptive visual mask reconstruction branch. Through their complementary guidance, it synergistically utilizes visual and language reconstruction to achieve the alignment between video and text in uncropped videos.
[0160] 4. The present application performs an effective post-processing algorithm to post-process the waveform formed by the generated relevant scores to obtain the predicted temporal start point and temporal end point, ensuring the high precision of the temporal boundary localization result.
[0161] It should be understood that although the steps in the flowchart of the accompanying drawings are shown in sequence according to the indication of the arrows, these steps do not necessarily have to be executed in the order indicated by the arrows. Unless otherwise clearly stated in this application, the execution of these steps has no strict order limit and can be executed in other orders. Moreover, at least some of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages. These sub-steps or stages do not necessarily have to be executed at the same time, but can be executed at different times. Their execution order does not necessarily have to be sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.
[0162] The above content is only a preferred exemplary embodiment of the present application and is not used to limit the implementation of the present application. Those of ordinary skill in the art can easily make corresponding adaptations or modifications according to the main concept and spirit of the present application. Therefore, the protection scope of the present application should be subject to the protection scope required by the claims.
Claims
1. A weakly supervised temporal boundary localization method, characterized in that The method includes: Obtaining a video, and respectively extracting features from each video frame in the video and the text description corresponding to the video to obtain the original features of each video frame and the text features of the text description; the text description is used to describe the action label corresponding to the video; According to the correlation between the original features of each video frame and the text features of the text description, obtaining the positive correlation features and negative correlation features of each video frame; Using mask reconstruction to align the video features of the video with the text features of the text description, respectively obtaining the reconstructed text features and reconstructed video features of each video frame; the video features of the video include the original features, positive correlation features and negative correlation features of each video frame; According to the reconstructed text features and reconstructed video features of each video frame, positioning the temporal boundary of the action in the video to obtain a boundary positioning result.
2. The method according to claim 1, characterized in that, The respectively extracting features from each video frame in the video and the text description corresponding to the video to obtain the original features of each video frame and the text features of the text description includes: Extracting features from each video frame in the video to obtain the original features of each video frame; Predicting the action category of the original features of each video frame to obtain the action label corresponding to the video, and converting the action label corresponding to the video into the text description corresponding to the video; Extracting features from the text description corresponding to the video to obtain the text features of the text description.
3. The method according to claim 1, wherein The obtaining the positive correlation features and negative correlation features of each video frame according to the correlation between the original features of each video frame and the text features of the text description includes: According to the correlation between the original features of each video frame and the text features of the text description, respectively calculating the positive correlation score and negative correlation score of each video frame and the text description; According to the calculated positive correlation score and the original features of each video frame, calculating the positive correlation features of each video frame; According to the calculated negative correlation score and the original features of each video frame, calculating the negative correlation features of each video frame.
4. The method according to claim 1, characterized in that, The using mask reconstruction to align the video features of the video with the text features of the text description, respectively obtaining the reconstructed text features and reconstructed video features of each video frame includes: Based on mask reconstruction, guiding the video features of the video to reconstruct the masked text features to obtain the reconstructed text features of each video frame; Based on mask reconstruction, guiding the text features of the text description to reconstruct the masked video features to obtain the reconstructed video features of each video frame.
5. The method according to claim 4, characterized in that, The based on mask reconstruction, guiding the video features of the video to reconstruct the masked text features to obtain the reconstructed text features of each video frame, and based on mask reconstruction, guiding the text features of the text description to reconstruct the masked video features to obtain the reconstructed video features of each video frame includes: For each video frame, respectively performing mask reconstruction on the masked text features according to the original features, positive correlation features and negative correlation features of the video frame to obtain multiple reconstructed text features of the video frame; Mask reconstruction is performed on the original features, positive correlation features, and negative correlation features of the masked video frame respectively according to the text features described in the text, to obtain multiple reconstructed video features of the video frame.
6. The method according to any one of claims 1 to 5, characterized in that, The temporal boundary localization is implemented by invoking a temporal boundary localization network, which is a machine learning model that has been trained and has the ability to locate the temporal boundaries of actions in the video; wherein, The temporal boundary localization network includes a proposal generation module for feature extraction, a text mask reconstruction module for text feature reconstruction, and a video mask reconstruction module for video feature reconstruction.
7. The method according to claim 6, wherein The training process of the temporal boundary localization network includes: In the text mask reconstruction module, calculate the text ranking loss and the text reconstruction loss according to the reconstructed text features of the video frame, and use the text ranking loss to constrain the text reconstruction loss of the video frame with respect to the reconstructed text features; In the video mask reconstruction module, calculate the video ranking loss and the video reconstruction loss according to the reconstructed video features of the video frame, and use the video ranking loss to constrain the video reconstruction loss of the video frame with respect to the reconstructed video features; If the text reconstruction loss indicates that the reconstruction effect of the reconstructed text features reconstructed from the positive correlation features of the video frame is the best, and the video reconstruction loss indicates that the reconstruction effect of the reconstructed video features related to the positive correlation features is the best, then the training is completed to obtain the temporal boundary localization network.
8. A weakly supervised temporal boundary localization device, characterized in that, The device includes: An original feature extraction module, configured to obtain a video, and perform feature extraction on each video frame in the video and the text description corresponding to the video respectively, to obtain the original features of each video frame and the text features of the text description; the text description is used to describe the action label corresponding to the video; A correlation feature extraction module, configured to obtain the positive correlation features and negative correlation features of each video frame according to the correlation between the original features of each video frame and the text features of the text description; A mask reconstruction module, configured to align the video features of the video and the text features of the text description by using mask reconstruction, to obtain the reconstructed text features and reconstructed video features of each video frame respectively; the video features of the video include the original features, positive correlation features, and negative correlation features of each video frame; A boundary localization module, configured to locate the temporal boundaries of actions in the video according to the reconstructed text features and reconstructed video features of each video frame, to obtain a boundary localization result.
9. An electronic device, characterized in that, Includes: At least one processor, at least one memory, and at least one communication bus, wherein, A computer program is stored on the memory, and the processor reads the computer program in the memory through the communication bus; When the computer program is executed by the processor, it implements the weakly supervised temporal boundary localization method according to any one of claims 1 to 7.
10. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the weakly supervised temporal boundary localization method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Weak supervision video time period retrieval method and system based on two-branch proposal network
CN112417206A
Video retrieval method, and method and apparatus for generating video retrieval mapping relationship
US20210142069A1