Method, device, storage medium and electronic device for acquiring video material

By extracting video frame images and video features from the target video and combining them with the HNSW recall tree to optimize the video material recall process, the problem of poor video material recall effect in the existing technology is solved, and more efficient and accurate video material recall is achieved.

CN114049591BActive Publication Date: 2026-03-27DOUYIN VISION CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-15
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing video material retrieval methods rely on the image features of the cover frame or a single video frame, which are too limited to effectively represent the entire video, resulting in poor retrieval performance.

Method used

Extract video frame images and video features from the target video, combine video features and image features to calculate the similarity of candidate videos, optimize the recall process through HNSW recall tree, and obtain candidate videos of the same type as the target video.

Benefits of technology

It improves the accuracy and efficiency of video material retrieval. By introducing a combination of video and image features, it optimizes the candidate video selection process, reduces unnecessary traversal, and improves the efficiency and accuracy of retrieval.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114049591B_ABST
    Figure CN114049591B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a method and device for obtaining video material, a storage medium and an electronic device, wherein a video frame picture of a target video is extracted, a video feature corresponding to the target video and a picture feature corresponding to the video frame picture are extracted, a feature similarity between each candidate video and the target video is calculated according to the video feature and the picture feature, the candidate videos are sorted based on the feature similarity, a target candidate video belonging to the same type as the target video is determined from the candidate videos according to a sorting result, and the target candidate video is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the field of material recall, in particular, to a method and device for obtaining video material, a storage medium and an electronic device. BACKGROUND

[0002] In the field of advertisement delivery, in order to obtain better delivery effect, the similar materials corresponding to the advertisement type with good delivery effect can be found for creative advertisement production, i.e. recalling similar materials belonging to the same vertical category as the seed material from the material library.

[0003] For video material recall, in the existing similar recall method, the cover frame or a certain video frame is usually used for retrieval, but the picture features corresponding to a certain video frame are relatively single, and the cover frame or a certain video frame cannot represent the whole video, resulting in poor recall effect of video material. SUMMARY

[0004] This summary is provided to introduce a selection of concepts, which will be described with greater specificity below in the detailed description section. This summary is not intended to identify key or essential features of the claimed technology, nor is it intended to limit the scope of the claimed technology.

[0005] In a first aspect, the present disclosure provides a method for obtaining video material, the method comprising:

[0006] extracting a video frame picture of a target video;

[0007] extracting video features corresponding to the target video and picture features corresponding to the video frame picture;

[0008] calculating a feature similarity between each candidate video and the target video according to the video features and the picture features;

[0009] sorting the candidate videos based on the feature similarity, and determining a target candidate video belonging to the same type as the target video from a plurality of the candidate videos according to a sorting result;

[0010] obtaining the target candidate video.

[0011] In a second aspect, the present disclosure provides a device for obtaining video material, the device comprising:

[0012] an extraction module configured to extract a video frame picture of a target video;

[0013] an extraction module configured to extract video features corresponding to the target video and picture features corresponding to the video frame picture;

[0014] The similarity calculation module is configured to calculate a feature similarity between each candidate video and the target video according to the video feature and the picture feature;

[0015] The determination module is configured to sort the candidate videos based on the feature similarity, and determine a target candidate video belonging to the same type as the target video from the plurality of candidate videos according to a sorting result.

[0016] The material acquisition module is configured to acquire the target candidate video.

[0017] In a third aspect, the present disclosure provides a computer readable medium having a computer program stored thereon, which, when executed by a processing device, implements the steps of the method of the first aspect of the present disclosure.

[0018] In a fourth aspect, an electronic device is provided, comprising:

[0019] A storage device having a computer program stored thereon;

[0020] A processing device configured to execute the computer program in the storage device to implement the steps of the method of the first aspect of the present disclosure.

[0021] According to the above technical solution, the video frame pictures of the target video are extracted, the video feature corresponding to the target video and the picture feature corresponding to the video frame pictures are extracted, the feature similarity between each candidate video and the target video is calculated according to the video feature and the picture feature, the candidate videos are sorted based on the feature similarity, and a target candidate video belonging to the same type as the target video is determined from the plurality of candidate videos according to a sorting result. Since the video feature can better represent the video content of the entire video, by introducing the video feature of the target video, the video feature and the picture feature are used together as the recall basis of the target candidate video, which can improve the accuracy and recall efficiency of video material recall.

[0022] Other features and advantages of the present disclosure will be described in detail in the following detailed description section. BRIEF DESCRIPTION OF DRAWINGS

[0023] The above and other features, advantages, and aspects of embodiments of the present disclosure will become more apparent by describing in detail exemplary embodiments thereof with reference to the attached drawings in which:

[0024] Figure 1 is a flowchart of a method for acquiring video material according to an exemplary embodiment;

[0025] Figure 2 is a flowchart of a method of acquiring video material according to an example embodiment;

[0026] Figure 3 is a flowchart of a method of acquiring video material according to an example embodiment;

[0027] Figure 4 is a block diagram of an apparatus for acquiring video material according to an example embodiment;

[0028] Figure 5 is a block diagram of an apparatus for acquiring video material according to an example embodiment;

[0029] Figure 6 is a block diagram of an apparatus for acquiring video material according to an example embodiment;

[0030] Figure 7 is a block diagram of an apparatus for acquiring video material according to an example embodiment;

[0031] Figure 8 is a block diagram of an electronic device according to an example embodiment. DETAILED DESCRIPTION

[0032] Embodiments of the present disclosure will be described more fully hereinafter with reference to the accompanying drawings, in which some, but not all embodiments of the present disclosure are shown. Understanding that these drawings depict only some embodiments of the present disclosure and are not therefore to be considered to be limiting the scope of the present disclosure, the embodiments will be described and explained with additional specificity and detail through the use of the accompanying drawings in which:

[0033] It should be understood that each of the steps of the method embodiments of the present disclosure can be performed in a different order, and / or in parallel. Furthermore, the method embodiments can include additional steps and / or omit performing the steps shown. The scope of the present disclosure is not limited in this respect.

[0034] The term "comprising" and variations thereof as used herein are used inclusively, i.e., "comprising but not limited to." The term "based on" is "based at least in part on." The term "one embodiment" means "at least one embodiment." The term "another embodiment" means "at least one additional embodiment." The term "some embodiments" means "at least some embodiments." Related definitions are given below in the description of the other terms.

[0035] It should be noted that the terms "first", "second", and the like in the present disclosure are only used to distinguish different devices, modules or units, and do not limit the order or interdependence of the functions performed by these devices, modules or units.

[0036] It should be noted that the terms "one", "multiple" in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that "one or more" should be understood unless otherwise explicitly indicated in the context.

[0037] The names of the messages or information exchanged between the devices in the embodiments of the present disclosure are only for illustrative purposes, and are not used to limit the scope of the messages or information.

[0038] The specific embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.

[0039] Figure 1 is a flowchart of a method for obtaining video material according to an exemplary embodiment, as shown in Figure 1 The method comprises the following steps:

[0040] In step S101, the video frame pictures of the target video are extracted.

[0041] The target video refers to the video material as the seed material, and the video frame pictures can include the cover frame picture of the target video or the preset video frame picture (such as the first frame of the target video) other than the cover frame picture. In the actual material recall scene, the user can select the cover frame picture to extract the picture feature according to the actual business demand, or select the preset video frame picture other than the cover frame picture to extract the picture feature.

[0042] In determining the cover frame picture of the target video, in one possible implementation, the picture feature of each video frame picture of the target video can be extracted, and the video frame picture is scored according to the picture feature. The highest score is the cover frame picture. The specific implementation details can be referred to the description in the related literature, which will not be repeated here.

[0043] In step S102, the video features corresponding to the target video and the picture features corresponding to the video frame pictures are extracted.

[0044] The video features refer to the features extracted for each video segment of the target video.

[0045] In an actual application scenario, for a short video, a cover frame picture can express the general content of the video, for example, a short video of a beautiful woman or a cute pet, and the content of the video can be understood through the cover frame picture. Therefore, for the material recall of a short video type, the picture feature of the cover frame picture can be extracted for recall, and a good material recall effect can be achieved. For the material recall of a video type that emphasizes action such as dancing or horse riding, the accuracy of material recall using only the picture feature is poor. In the present disclosure, the material recall can be performed based on the picture feature in combination with the video feature of the target video. Therefore, in the present step, the video feature corresponding to the target video and the picture feature corresponding to the video frame picture can be extracted.

[0046] In the process of extracting the picture feature corresponding to the video frame picture of the target video, the cover frame picture in the target video or a preset video frame other than the cover frame picture can be input into a pre-trained picture feature extraction model to obtain the picture feature.

[0047] Here, the picture feature extraction model may, for example, include a resnet-50 model or a resnet-101 model.

[0048] For example, the cover frame picture of the target video can be input into a pre-trained resnet-50 model. Considering that the resnet-50 model is a classification model, the output layer (generally the last layer of the model structure) of the model outputs a classification result. In the present disclosure, the picture feature of the target video is to be extracted based on the picture feature extraction model. Therefore, the network output of the second-to-last layer of the resnet-50 model structure can be used as the extracted picture feature. This is only an example and the present disclosure is not limited in this regard.

[0049] In the process of extracting the video feature corresponding to the target video, the target video can be divided into a plurality of video segments according to a preset step length. For each video segment, the video segment can be input into a pre-trained video feature extraction model to obtain the video feature corresponding to each video segment.

[0050] For example, the video feature extraction model may, for example, include a resnet-3d model or a slowfast model. Taking the resnet-3d model as an example, the resnet-3d model is also a classification model. Therefore, in the process of extracting the video feature of the target video based on the resnet-3d model, the network output of the second-to-last layer in the model structure can be used as the video feature.

[0051] In addition, in the process of dividing the target video into a plurality of video clips according to a preset step length, in order to ensure that complete and sufficient video features can be extracted, each two adjacent video clips can have a preset time length overlap, for example, 4 video frames per second, and 16 frames can be taken as a video clip, the preset step length is 2 seconds, that is, a video clip is divided every 2 seconds, and each two video clips have a 2-second overlap. In this way, the first video clip includes the 1st to 16th frames of the target video, the second video clip includes the 9th to 24th frames of the target video, the third video clip includes the 17th to 32nd frames of the target video, and so on. The target video can be divided into a plurality of video clips. Here, only an example is given, and the present disclosure is not limited in this regard.

[0052] In step S103, a feature similarity of each candidate video to the target video is calculated according to the video features and the picture features.

[0053] The feature similarity can include a picture feature similarity calculated based on the picture features, a video feature similarity calculated based on the video features, and a fusion feature similarity calculated based on fusion features (i.e., features obtained by fusing the video features and the picture features of the same video material).

[0054] In step S104, the candidate videos are sorted based on the feature similarity, and a target candidate video belonging to the same type as the target video is determined from the plurality of candidate videos according to the sorting result.

[0055] Here, the target candidate video belonging to the same type as the target video refers to a video material belonging to the same vertical category as the target video.

[0056] In one possible implementation, the plurality of candidate videos can be all preset candidate videos in a video material library. In this case, the target candidate video can be determined from the plurality of candidate videos in the following manner: picture feature similarities of each preset candidate video to the target video are calculated according to the picture features, and the preset candidate videos are sorted in descending order of the picture feature similarities, and a first preset number of preset candidate videos in front are selected as a first candidate video set according to the sorting result; video feature similarities of each preset candidate video to the target video are calculated according to the video features, and the preset candidate videos are sorted in descending order of the video feature similarities, and a second preset number of preset candidate videos in front are selected as a second candidate video set according to the sorting result; and the target candidate video is determined from the first candidate video set and the second candidate video set.

[0057] In the present disclosure, the feature similarity, i.e. including the picture feature similarity or the video feature similarity, between each candidate video and the target video can be calculated by the cosine distance formula, i.e. formula (1) or the L2 distance formula, i.e. formula (2) as shown below:

[0058]

[0059] similarity(s,c)=||f s -f c ||2 (2)

[0060] Wherein, s represents the target video as the seed material, c represents any candidate video, f s represents the picture feature corresponding to the target video, f c represents the picture feature corresponding to the candidate video, f s represents the video feature corresponding to the target video, f c represents the video feature corresponding to the candidate video.

[0061] Next, the target candidate video can be determined from the first candidate video set and the second candidate video set by the following steps: sorting each candidate video in the first candidate video set and the second candidate video set in order of feature similarity from high to low, the feature similarity including the picture feature similarity or the video feature similarity; selecting the top fourth preset number of candidate videos as the target candidate video according to the sorting result.

[0062] For example, it is assumed that the first candidate video set includes six candidate videos a, b, c, d, e, and f, and the picture feature similarity of each candidate video and its corresponding picture feature can be represented as (a, 0.8), (b, 0.75), (c, 0.6), (d, 0.92), (e, 0.95), and (f, 0.85), respectively. The second candidate video set includes six candidate videos g, h, i, j, k, and l, and the video feature similarity of each candidate video and its corresponding video feature can be represented as (g, 0.99), (h, 0.96), (i, 0.83), (j, 0.97), (k, 0.72), and (l, 0.65), respectively. The candidate videos in the first candidate video set and the second candidate video set are sorted in descending order of feature similarity as follows: video g, video j, video h, video e, video d, video f, video i, video a, video b, video k, video l, and video c. The top five (i.e., the fourth preset number) candidate videos, i.e., video g, video j, video h, video e, and video d, are selected as the target candidate videos. The above example is only illustrative, and the present disclosure is not limited in this regard.

[0063] In view of the fact that the above manner needs to traverse all preset candidate videos in the video material library, that is, the similarity of each preset candidate video with the target video needs to be calculated, the recall cost is high and the recall efficiency is low, in order to improve the recall efficiency of the video material, the plurality of candidate videos can also be the traversed candidate videos determined by traversing a pre-established HNSW (Hierarchical Navigable Small World, hierarchical navigable small world) recall tree according to a preset recall strategy, wherein the HNSW recall tree includes a first HNSW recall tree or a second HNSW recall tree, the first HNSW recall tree is a HNSW recall tree pre-established according to the picture feature similarity between each two candidate videos; and the second HNSW recall tree is a HNSW recall tree pre-established according to the video feature similarity between each two candidate videos. In this way, in another possible implementation manner of the present step, the target candidate video can also be determined from the plurality of candidate videos in the following manner: the picture feature similarity between each traversed candidate video and the target video is calculated based on the preset recall strategy through the first HNSW recall tree according to the picture feature, and a first candidate video set is determined from the plurality of candidate videos according to the picture feature similarity; then the video feature similarity between each traversed candidate video and the target video is calculated based on the preset recall strategy through the second HNSW recall tree according to the video feature, and a second candidate video set is determined from the plurality of traversed candidate videos according to the video feature similarity; and finally, the target candidate video can be determined from the first candidate video set and the second candidate video set.

[0064] The first HNSW recall tree and the second HNSW recall tree each include a plurality of nodes, and the preset recall strategy specifically includes:

[0065] The preset traversal step is performed until the target HNSW recall tree is traversed completely;

[0066] The preset traversal step includes: obtaining a preset node corresponding to the target HNSW recall tree; determining a target node from the target HNSW recall tree according to the preset node, the target node including the preset node and nodes connected with the preset node; calculating the feature similarity between each candidate video to be determined and the target video according to a target feature, the candidate video to be determined being a candidate video corresponding to the target node; taking the node corresponding to the candidate video with the highest feature similarity as a new preset node; and determining whether there is a node connected with the new preset node in the target HNSW recall tree;

[0067] In a case where it is determined that there is a node connected with the new preset node in the target HNSW recall tree, the preset traversal step is re-executed, in a case where it is determined that there is no node connected with the new preset node in the target HNSW recall tree, it is determined that the target HNSW recall tree is traversed completely; in a case where the target HNSW recall tree is traversed completely, the traversed candidate videos corresponding to the traversed nodes are sorted according to the feature similarity in descending order, and the first third preset number of traversed candidate videos are selected as the specific candidate video set according to the sorting result;

[0068] In a case where the target feature is the picture feature, the target HNSW recall tree is the first HNSW recall tree, the feature similarity is a picture feature similarity, and the specific candidate video set is the first candidate video set; in a case where the target feature is the video feature, the target HNSW recall tree is the second HNSW recall tree, the feature similarity is a video feature similarity, and the specific candidate video set is the second candidate video set.

[0069] In this way, each feature similarity can also be calculated by the above formula (1) or formula (2).

[0070] The following describes a specific implementation process of determining a first candidate video set from a plurality of candidate videos according to a picture feature through a pre-established first HNSW recall tree in an example manner:

[0071] Assuming that the first HNSW recall tree, the preset node randomly generated by the HNSW algorithm is node 1, and the node 1 is directly connected with four nodes 2, 3, 4 and 5 on the first HNSW recall tree, at this time, the picture feature similarity of the candidate video corresponding to the nodes 1, 2, 3, 4 and 5 and the target video can be calculated according to the picture feature of the candidate video, and then the node with the highest picture feature similarity is selected from the nodes 2, 3, 4 and 5 as a new preset node (assuming it is node 3), and then the first HNSW recall tree is continued to be traversed, the nodes directly connected with the new preset node are determined from the first HNSW recall tree, assuming they are nodes 6, 7 and 8, and the picture feature similarity of the candidate video corresponding to each of the nodes 6, 7 and 8 and the target video is recalculated, and the node with the highest picture feature similarity is selected from the nodes 6, 7 and 8 as a new preset node according to the picture feature similarity, in the case that there is a node connected with the new preset node in the first HNSW recall tree, the first traversal step is re-executed, until in the case that there is no node connected with the new preset node in the first HNSW recall tree, it is determined that the first HNSW recall tree has been traversed; At this time, the candidate videos corresponding to the traversed nodes can be sorted in order of the picture feature similarity from high to low; The first candidate video set is composed of the first preset number of candidate videos selected according to the sorting result, for example, after the first HNSW recall tree is traversed, it is determined that the traversed nodes include 100 nodes 1, 2,..., 100, wherein after sorting in order of the picture feature similarity from high to low, it is determined that the nodes in the top 5 positions include node 12, node 15, node 23, node 50 and node 55, at this time, it can be determined that the first candidate video set includes the candidate videos corresponding to the nodes 12, 15, 23, 50 and 55 in the plurality of candidate videos, the above example is only illustrative, and the present disclosure is not limited thereto.

[0072] In addition, the second HNSW recall tree also includes a plurality of nodes, each node corresponds to a candidate video one-to-one, and the specific implementation of determining the second candidate video set from the candidate videos based on the second HNSW recall tree is similar to the specific implementation of determining the first candidate video set from the candidate videos based on the first HNSW recall tree, which will not be described here.

[0073] It can be understood that the similarity between two videos is a number greater than 0 and less than or equal to 1, if the similarity is greater than 1, the similarity has no practical significance and cannot be used to compare the similarity between two videos, therefore, in order to avoid the picture feature similarity or the video feature similarity calculated based on the above formula (1) or formula (2) being greater than 1, in a possible implementation of the present disclosure, a preset weight (the preset weight is greater than 0 and less than 1) can be set for the similarity calculated for each type of feature (including picture feature, video feature, and text feature and audio feature mentioned later), and then for each type of feature, the similarity calculated based on the feature is multiplied by the corresponding preset weight, which is the final similarity corresponding to the feature.

[0074] After obtaining the first candidate video set and the second candidate video set, the target candidate video to be recalled can be determined from the first candidate video set and the second candidate video set, specifically, each candidate video in the first candidate video set and the second candidate video set can be sorted in order of feature similarity from high to low, the feature similarity includes the picture feature similarity or the video feature similarity; according to the sorting result, the first fourth preset number of candidate videos are selected as the target candidate video.

[0075] For example, it is assumed that the first candidate video set includes six candidate videos a, b, c, d, e, and f, and each candidate video and its corresponding picture feature similarity can be represented as (a, 0.8), (b, 0.75), (c, 0.6), (d, 0.92), (e, 0.95), and (f, 0.85), respectively, and the second candidate video set includes six candidate videos g, h, i, j, k, and l, and each candidate video and its corresponding video feature similarity can be represented as (g, 0.99), (h, 0.96), (i, 0.83), (j, 0.97), (k, 0.72), and (l, 0.65), respectively, the candidate videos in the first candidate video set and the second candidate video set are sorted in order of feature similarity from high to low as: video g, video j, video h, video e, video d, video f, video i, video a, video b, video k, video l, and video c, and the first five (i.e., the fourth preset number) candidate videos are selected as the target candidate video, i.e., video g, video j, video h, video e, and video d are the target candidate video, the above example is only illustrative, and the present disclosure is not limited thereto.

[0076] In addition, considering that the first candidate video set and the second candidate video set can include one or more same candidate videos in an actual application scenario, the disclosure can take the maximum similarity between the picture feature similarity and the video feature similarity corresponding to each same candidate video in the first candidate video set and the second candidate video set as the feature similarity corresponding to the same candidate video.

[0077] It should be noted that the first HNSW recall tree is a recall tree constructed in advance according to the picture feature similarity between candidate videos, and the picture feature similarity between two candidate videos corresponding to two nodes connected to each other is high. The second HNSW recall tree is a recall tree constructed in advance according to the video feature similarity between candidate videos, and the video feature similarity between two candidate videos corresponding to two nodes connected to each other is high. Therefore, in actual material recall, the disclosure enters the HNSW recall tree from a fixed preset node through the pre-constructed HNSW recall tree, and quickly finds the first candidate video set or the second candidate video set along the edges between nodes, without traversing all preset candidate videos in the entire material library. This can significantly improve the efficiency and accuracy of material recall.

[0078] In addition, in another possible implementation manner of the present step, in the case where the dimensions of the video feature and the picture feature are the same, the picture feature and the video feature can also be fused, and then the target candidate video is determined from the plurality of candidate videos based on the fused fusion feature. That is, the video feature and the picture feature can be fused to obtain a fusion feature. Then, a fusion feature similarity between each candidate video and the target video is calculated according to the fusion feature, and the candidate videos are sorted according to the fusion feature similarity, and a target candidate video belonging to the same type as the target video is determined from the plurality of candidate videos according to the sorting result.

[0079] Among them, the feature fusion can be performed in the following two ways:

[0080] Method one, directly adding the video feature and the picture feature of the target video to obtain the fusion feature. Specifically, for each first feature element of the video feature, the first feature element and the second feature element corresponding to the first feature element in the picture feature are added to obtain the fusion feature.

[0081] For example, assuming that the video feature of the target video corresponds to a feature vector A, A=(a1, a2, …, an), and the picture feature of the target video corresponds to a feature vector B, B=(b1, b2, …, bn), the video feature and the picture feature are directly added, and a feature vector corresponding to the fusion feature is obtained: A+B=(a1+b1, a2+b2, …, an+bn). The above example is only illustrative, and the present disclosure is not limited in this regard.

[0082] It is considered that the features that can represent the main content of videos corresponding to different video types are also different. For example, for short videos of beautiful women, cute pets, etc., the cover frame picture can generally represent the content of the video. Therefore, for short video types, the feature that can represent the main content of the video is a picture feature. For video types that emphasize action, such as dancing and horseback riding, the accuracy of material recall using only the picture feature is poor, and the video feature can represent the main content of the entire video. Therefore, for video types that emphasize action, the feature that can represent the main content of the video is a video feature. That is, for the recall of video materials of different types, the feature that represents the main content of the video is also different, that is, the weight corresponding to each feature is also different. Therefore, the present disclosure can also perform feature fusion in the following way two:

[0083] Way two, identifying the video type of the target video through a material type identification model; determining a first preset weight corresponding to the video feature and a second preset weight corresponding to the picture feature according to the video type. In this way, the video feature and the picture feature can be weighted and summed according to the first preset weight and the second preset weight to obtain the fusion feature.

[0084] The model structure of the material type identification model can be the same as the model structure of the video feature extraction model described above, that is, the material type identification model can include a resnet-3d model or a slowfast model, for example. The video type can include any type such as beautiful women, cute pets, dancing, horseback riding, etc., or the video type can include short videos or action videos.

[0085] In a possible implementation, the output layer of the material type identification model can output probability values corresponding to a plurality of preset material types, respectively. In this way, in the process of inputting the target video into the material type identification model and identifying the video type based on the material type identification model, the preset material type corresponding to the maximum probability value output can be determined as the video type corresponding to the target video.

[0086] After determining the video type of the target video, the first preset weight corresponding to the video feature and the second preset weight corresponding to the picture feature of the target video can be determined according to the corresponding relationship between the video type and the preset weight, and then the video feature and the picture feature can be weighted and summed according to the first preset weight and the second preset weight to obtain the fusion feature.

[0087] For example, the corresponding relationship between a possible video type and a preset weight is as follows: when the video type is a short video, the corresponding first preset weight is 0.2 and the second preset weight is 0.8; when the video type is an action video, the corresponding first preset weight is 0.8 and the second preset weight is 0.2. If it is determined that the video type of the target video is an action video, it can be determined that the first preset weight corresponding to the video feature of the target video is 0.8 and the second preset weight corresponding to the picture feature of the target video is 0.2. Assuming that the video feature of the target video is represented as A and the picture feature of the target video is represented as B, the fusion feature obtained after the video feature and the picture feature are fused is 0.8A+0.2B. The above example is only illustrative, and the present disclosure is not limited in this regard.

[0088] In addition, when the fusion feature similarity between each candidate video and the target video is calculated according to the fusion feature, the candidate videos are sorted based on the fusion feature similarity, and the target candidate video belonging to the same type as the target video is determined from the plurality of candidate videos according to the sorting result, the following two ways can also be used:

[0089] Method one, the plurality of candidate videos include all preset candidate videos in a video material library. In this way, the fusion feature similarity between each preset candidate video and the target video can be calculated according to the fusion feature, and the preset candidate videos can be sorted in order from high to low according to the fusion feature similarity, and the first sixth preset number of preset candidate videos can be selected as the target candidate video according to the sorting result.

[0090] Similarly, the fusion feature similarity between each preset candidate video and the target video can also be calculated by the above formula (1) or formula (2).

[0091] The second mode, in the video material recall based on the fusion feature, if the target candidate video is determined according to the first mode, it is also necessary to traverse all the preset candidate videos, that is, the similarity of the fusion feature between each preset candidate video and the target video needs to be calculated, which not only has high recall cost, but also has low recall efficiency. Therefore, in order to improve the recall efficiency of the video material, similar to the above-mentioned mode of determining the candidate video set according to the picture feature and the video feature respectively, the number of candidate videos traversed can also be reduced based on the idea of HNSW recall tree to improve the recall efficiency of the video material. Therefore, the HNSW recall tree includes a third HNSW recall tree established in advance, and the third HNSW recall tree is an HNSW recall tree established in advance according to the similarity of the fusion feature between each two candidate videos. A plurality of candidate videos include traversed candidate videos determined after traversing the third HNSW recall tree according to the preset recall strategy, and the feature similarity includes the similarity of the fusion feature. In this way, the target candidate video can be determined from a plurality of candidate videos in the following manner: the third HNSW recall tree is traversed by performing the preset traversal step, and the similarity of the fusion feature between each traversed node corresponding to the candidate video (i.e. the traversed candidate video) and the target video is calculated during the traversal process. After traversing the third HNSW recall tree, the candidate videos corresponding to the traversed nodes are sorted in descending order of the similarity of the fusion feature, and the top fifth preset number of traversed candidate videos are selected as the target candidate video according to the sorting result.

[0092] For example, assuming that the preset node randomly generated by the HNSW algorithm on the third HNSW recall tree is node 1, and the node 1 is directly connected with four nodes 2, 3, 4 and 5 on the third HNSW recall tree, at this time, the fusion feature similarity between the candidate videos corresponding to the nodes 1, 2, 3, 4 and 5 and the target video can be calculated, and then the node with the highest fusion feature similarity is selected from the nodes 2, 3, 4 and 5 as a new preset node (assuming that it is node 3). Then the third HNSW recall tree is continuously traversed, and the nodes directly connected with the new preset node are determined from the third HNSW recall tree, assuming that they are nodes 6, 7 and 8. The fusion feature similarity between the candidate videos corresponding to the nodes 6, 7 and 8 and the target video is recalculated according to the fusion feature, and the node with the highest fusion feature similarity is selected from the nodes 6, 7 and 8 as a new preset node. In the case where it is determined that there is a node connected with the new preset node in the third HNSW recall tree, the first traversal step is re-executed until it is determined that there is no node connected with the new preset node in the third HNSW recall tree, and it is determined that the traversal of the third HNSW recall tree is completed. At this time, the candidate videos corresponding to the traversed nodes can be sorted in descending order of the fusion feature similarity. According to the sorting result, the first sixth preset number of candidate videos are selected as the target candidate videos. For example, after the traversal of the third HNSW recall tree is completed, it is determined that the traversed nodes include nodes 1, 2,..., 100, and after the sorting in descending order of the fusion feature similarity, it is determined that the nodes in the top five positions include nodes 12, 15, 23, 50 and 55. At this time, it can be determined that the target candidate videos include the candidate videos corresponding to the nodes 12, 15, 23, 50 and 55 in the plurality of candidate videos. The above example is only illustrative, and the present disclosure is not limited in this regard.

[0093] Similarly, the third HNSW recall tree is a recall tree constructed in advance according to the fusion feature similarity between the candidate videos. The fusion feature similarity between the two candidate videos corresponding to the two nodes connected with each other is high. Therefore, in the actual material recall, the HNSW recall tree constructed in advance is used to enter the HNSW recall tree from a fixed preset node, and the target candidate video is quickly found along the edges between the nodes, without traversing all the preset candidate videos in the entire material library. This can significantly improve the efficiency and accuracy of material recall.

[0094] In step S105, the target candidate video is obtained.

[0095] In this step, the target candidate video can be recalled from the material library to make an advertisement with better delivery effect according to the target candidate video.

[0096] By using the above method, by introducing the video features of the target video as seed materials, the video features and the picture features are used as the basis for recalling the target candidate video, which can improve the accuracy of video material recall. At the same time, in the actual material recall, the HNSW recall tree is pre-constructed, the first candidate video set or the second candidate video set is quickly found by entering the HNSW recall tree from a fixed preset node and along the edges between nodes, without traversing all candidate videos in the entire material library, which can significantly improve the efficiency of material recall.

[0097] Figure 2 According to the method for obtaining video materials shown in the embodiment shown in the flowchart of the method for obtaining video materials shown in Figure 1 Before step S103 is executed, the method further includes the following steps: Figure 2

[0098] In step S106, the text features of the target video are obtained by using the pre-trained text feature extraction model, and / or the audio features of the target video are obtained by using the pre-trained audio feature extraction model.

[0099] The text features can include the text content appearing in the target video, the speech content in the target video, or the classification label of the target video. The text feature extraction model can include a BERT model. The audio features can include the background music and the speech content corresponding to the target video. The audio feature extraction model can include a VGGish model.

[0100] In the actual video material recall scenario, using picture features and video features can ensure that the two videos have high picture similarity, but the content of the videos may still be different. Therefore, in this step, the text features and / or audio features of the target video can be extracted, so that the video material can be recalled according to the picture features, video features, text features, and audio features of the target video, and the accuracy of video material recall can be further improved.

[0101] In the process of extracting the text features of the target video, the text content in the target video can be encoded first. Specifically, the text content can be sentence split, and then the split words are mapped to a vocabulary library to realize vocabulary encoding. Then, the encoded data can be input into a pre-trained BERT model, and the network output of the third layer from the bottom of the model structure is used as the extracted text features.

[0102] ​In the process of extracting the audio features of the target video, the audio data in the target video can be extracted first, and then the audio data can be input into the pre-trained VGGish model to extract the audio features.

[0103] After obtaining the text feature and the audio feature, when performing step S103, the feature similarity between each candidate video and the target video can be calculated based on the image feature, the video feature, and the specified feature, whereby the specified feature includes the text feature and / or the audio feature.

[0104] The specific implementation method for determining the target candidate video based on the above four features is similar to the implementation method for determining the target candidate video based on image features and video features described above. It can calculate a feature similarity based on each of the four features, and then determine the target candidate video by comprehensively ranking the feature similarities calculated based on the four features. Alternatively, in order to improve the efficiency of material retrieval, the video material can be retrieved based on the idea of ​​HNSW retrieval tree. Or, material retrieval can be performed after calculating the feature similarity based on the fusion feature of the four features. For specific implementation methods, please refer to the relevant descriptions above, which will not be repeated here.

[0105] Figure 3 It is based on Figure 1 The illustrated embodiment presents a flowchart of a method for acquiring video footage, as shown below. Figure 3 As shown, before performing step S103, the method further includes the following steps:

[0106] In step S107, the image features and the video features are subjected to feature dimensionality reduction to obtain the dimensionality-reduced target image features and target video features.

[0107] To support the retrieval of tens of millions of videos in the media library and reduce storage and computational overhead, this disclosure can perform dimensionality reduction on the extracted feature vectors. Specifically, for image feature dimensionality reduction, SVD (Singular Value Decomposition) can be used, and for video feature dimensionality reduction, PQ (Product Quantization) can be used to map high-dimensional floating-point feature vectors into low-dimensional discrete vectors.

[0108] For example, assuming the video feature extraction model extracts 512-dimensional video features, the video features can be divided into groups of 4 dimensions, resulting in 128 sub-vectors. Each sub-vector can be clustered into 128 cluster centers using K-Means, thereby encoding any visual feature into a 128-dimensional vector.

[0109] In this way, the feature similarity of each candidate video and the target video can be calculated according to the target picture feature and the target video feature obtained after dimension reduction, so that the target candidate video can be determined more efficiently, and the calculation cost and the storage cost are reduced.

[0110] In addition, before the material recall based on the four types of features including the picture feature, the video feature, the text feature and the audio feature, the SVD method can be used to perform dimension reduction processing on the text feature and the audio feature.

[0111] After the dimension reduction processing, the dimensions of the feature vectors of each type of feature can be made the same.

[0112] Figure 4 is a block diagram of an apparatus for obtaining video material according to an example embodiment, as shown in Figure 4 The apparatus includes:

[0113] The extraction module 401 is configured to extract video frame pictures of a target video.

[0114] The extraction module 402 is configured to extract video features corresponding to the target video and picture features corresponding to the video frame pictures.

[0115] The similarity calculation module 403 is configured to calculate the feature similarity of each candidate video and the target video according to the video features and the picture features.

[0116] The determination module 404 is configured to sort the candidate videos based on the feature similarity, and determine a target candidate video belonging to the same type as the target video from the candidate videos according to the sorting result.

[0117] The material obtaining module 405 is configured to obtain the target candidate video.

[0118] Optionally, the candidate videos include preset candidate videos, and the similarity calculation module 403 and the determination module 404 are configured to calculate the picture feature similarity of each preset candidate video and the target video according to the picture features, sort the preset candidate videos in descending order of the picture feature similarity, and select the first preset number of preset candidate videos as a first candidate video set according to the sorting result; calculate the video feature similarity of each preset candidate video and the target video according to the video features, sort the preset candidate videos in descending order of the video feature similarity, and select the second preset number of preset candidate videos as a second candidate video set according to the sorting result; and determine the target candidate video from the first candidate video set and the second candidate video set.

[0119] Optionally, the plurality of candidate videos comprises traversed candidate videos determined by traversing a pre-established hierarchical navigational small world (HNSW) recall tree according to a preset recall strategy, the HNSW recall tree comprises a first HNSW recall tree or a second HNSW recall tree, the similarity calculation module 403 and the determination module 404 are configured to calculate, according to the picture features, the picture feature similarity between each traversed candidate video and the target video based on the first HNSW recall tree according to the preset recall strategy, and determine a first candidate video set from the plurality of traversed candidate videos according to the picture feature similarity, the first HNSW recall tree being a HNSW recall tree pre-established according to the picture feature similarity between each two candidate videos;

[0120] calculate, according to the video features, the video feature similarity between each traversed candidate video and the target video based on the second HNSW recall tree according to the preset recall strategy, and determine a second candidate video set from the plurality of traversed candidate videos according to the video feature similarity, the second HNSW recall tree being a HNSW recall tree pre-established according to the video feature similarity between each two candidate videos;

[0121] determine the target candidate video from the first candidate video set and the second candidate video set.

[0122] Optionally, the first HNSW recall tree and the second HNSW recall tree each comprise a plurality of nodes, each node corresponding to one of the candidate videos;

[0123] the preset recall strategy comprises:

[0124] performing a preset traversal step until traversal of the target HNSW recall tree is completed; the preset traversal step comprises: obtaining a preset node corresponding to the target HNSW recall tree; determining a target node from the target HNSW recall tree according to the preset node, the target node comprising the preset node and a node connected to the preset node; calculating a feature similarity between each to-be-determined candidate video and the target video according to a target feature, the to-be-determined candidate video being a candidate video corresponding to the target node; taking a node corresponding to a candidate video with the highest feature similarity as a new preset node; determining whether there is a node connected to the new preset node in the target HNSW recall tree; in a case where it is determined that there is a node connected to the new preset node in the target HNSW recall tree, re-performing the preset traversal step, in a case where it is determined that there is no node connected to the new preset node in the target HNSW recall tree, determining that the target HNSW recall tree is traversed completely; in a case where the target HNSW recall tree is traversed completely, sorting the traversed candidate videos corresponding to the traversed nodes in an order from high to low according to the feature similarity, and selecting a top third preset number of the traversed candidate videos as a specific candidate video set according to a sorting result.

[0125] In a case where the target feature is the picture feature, the target HNSW recall tree is the first HNSW recall tree, the feature similarity is a picture feature similarity, and the specific candidate video set is the first candidate video set; in a case where the target feature is the video feature, the target HNSW recall tree is the second HNSW recall tree, the feature similarity is a video feature similarity, and the specific candidate video set is the second candidate video set.

[0126] Optionally, the determining module 404 is configured to sort each candidate video in the first candidate video set and the second candidate video set in an order from high to low according to a feature similarity, the feature similarity comprising the picture feature similarity or the video feature similarity; and select a top fourth preset number of the candidate videos as the target candidate video according to a sorting result.

[0127] Optionally, the similarity calculation module 403 is configured to perform feature fusion on the video feature and the picture feature to obtain a fusion feature; and calculate a fusion feature similarity between each candidate video and the target video according to the fusion feature; and the determining module 404 is configured to sort the candidate videos based on the fusion feature similarity, and determine a target candidate video belonging to a same type as the target video from the plurality of candidate videos according to a sorting result.

[0128] Optionally, the similarity calculation module 403 is configured to, for each first feature element of the video feature, add the first feature element and a second feature element corresponding to the first feature element in the picture feature to obtain the fusion feature.

[0129] Optionally, Figure 5 is a block diagram of an apparatus for obtaining video material according to the embodiment shown in Figure 4 As shown in the figure, the apparatus further comprises: Figure 5

[0130] The weight determination module 406 is configured to identify a video type of the target video by a material type identification model, and determine a first preset weight corresponding to the video feature and a second preset weight corresponding to the picture feature according to the video type.

[0131] The similarity calculation module 403 is configured to perform weighted summation on the video feature and the picture feature according to the first preset weight and the second preset weight to obtain the fusion feature.

[0132] Optionally, the HNSW recall tree comprises a third HNSW recall tree established in advance, the plurality of candidate videos comprises traversed candidate videos determined after traversing the third HNSW recall tree according to the preset traversal strategy, the similarity calculation module 403 is configured to traverse the third HNSW recall tree by performing the preset traversal step, and calculate the fusion feature similarity between each traversed candidate video and the target video according to the fusion feature in the traversal process, and the third HNSW recall tree is an HNSW recall tree established in advance according to the fusion feature similarity between each two candidate videos; and the determination module 404 is configured to, after traversing the third HNSW recall tree, sort the traversed candidate videos corresponding to the nodes in the third HNSW recall tree in descending order of the fusion feature similarity, and select the first fifth preset number of traversed candidate videos as the target candidate video according to the sorting result.

[0133] Optionally, the plurality of candidate videos comprises preset candidate videos, and the similarity calculation module 403 is configured to calculate the fusion feature similarity between each preset candidate video and the target video according to the fusion feature, and the determination module 404 is configured to sort the preset candidate videos in descending order of the fusion feature similarity, and select the first sixth preset number of preset candidate videos as the target candidate video according to the sorting result.

[0134] Optionally, Figure 6 is a block diagram of an apparatus for obtaining video material according to the embodiment shown in Figure 4 As shown in the figure, the apparatus further comprises: Figure 6 ​As shown, the device further comprises:

[0135] The acquisition module 407 is configured to acquire the text feature of the target video by using the pre-trained text feature extraction model, and / or acquire the audio feature of the target video by using the pre-trained audio feature extraction model.

[0136] The similarity calculation module 403 is configured to calculate the feature similarity between each candidate video and the target video according to the picture feature, the video feature, and the specified feature, wherein the specified feature comprises the text feature and / or the audio feature.

[0137] Optionally, Figure 7 According to Figure 4 As shown in the block diagram of a device for acquiring video materials according to an embodiment of the present disclosure, the device comprises: Figure 7 As shown, the device further comprises:

[0138] The dimension reduction module 408 is configured to perform feature dimension reduction on the picture feature and the video feature to obtain the target picture feature and the target video feature after dimension reduction.

[0139] The similarity calculation module 403 is configured to calculate the feature similarity between each candidate video and the target video according to the target picture feature and the target video feature.

[0140] Optionally, the video frame picture comprises a cover frame picture of the target video or a preset video frame picture other than the cover frame picture.

[0141] Optionally, the extraction module 402 is configured to divide the target video into a plurality of video segments according to a preset step length; and input each video segment into a pre-trained video feature extraction model to obtain the video feature corresponding to each video segment.

[0142] By using the above device, the video feature of the target video as the seed material is introduced, and the video feature and the picture feature are used together as the recall basis of the target candidate video, so that the accuracy and recall efficiency of video material recall can be improved.

[0143] Reference will be made to Figure 8 which shows a structural schematic diagram of an electronic device (such as a terminal device) 800 suitable for implementing the embodiments of the present disclosure. The terminal device in the embodiments of the present disclosure can include but is not limited to mobile terminals such as mobile phones, notebook computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablets), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), and the like, and fixed terminals such as digital TVs, desktop computers, and the like. Figure 8The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0144] like Figure 8 As shown, the electronic device 800 may include a processing device (e.g., a central processing unit, a graphics processor, etc.) 801, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 802 or a program loaded from a storage device 808 into a random access memory (RAM) 803. The RAM 803 also stores various programs and data required for the operation of the electronic device 800. The processing device 801, ROM 802, and RAM 803 are interconnected via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0145] Typically, the following devices can be connected to I / O interface 805: input devices 806 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 807 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 808 including, for example, magnetic tapes, hard disks, etc.; and communication devices 809. Communication device 809 allows electronic device 800 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 8 An electronic device 800 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0146] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 809, or installed from a storage device 808, or installed from a ROM 802. When the computer program is executed by a processing device 801, it performs the functions defined in the methods of embodiments of this disclosure.

[0147] It is noted that the aforementioned computer-readable medium of the present disclosure can be a computer-readable signal medium or a computer-readable storage medium or any combination thereof. The computer-readable storage medium can be, for example and without limitation, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the computer-readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program used by or in connection with an instruction execution system, apparatus, or device. In the present disclosure, the computer-readable signal medium can include a computer-readable program code transmitted by a computer-readable storage medium or carried by a carrier wave in a baseband or as part of a carrier wave. Such a propagated computer-readable signal medium can take various forms, including but not limited to electro-magnetic, optical, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium that can be used to carry or store a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained in the computer-readable medium can be transmitted by any suitable medium, including but not limited to wire, cable, RF (radio frequency), or the like, or any suitable combination of the foregoing.

[0148] In some embodiments, the client can communicate using any currently known or future developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communications of any form or medium (e.g., a communications network). Examples of communications networks include local area networks ("LANs"), wide area networks ("WANs"), internetworks (e.g., the Internet), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future developed networks.

[0149] The aforementioned computer-readable medium can be included in the aforementioned electronic device; or can exist separately from the electronic device and not be assembled into the electronic device.

[0150] The computer readable medium described above carries one or more programs, when the one or more programs are executed by the electronic device, cause the electronic device to: extract a video frame picture of a target video; extract a video feature corresponding to the target video and a picture feature corresponding to the video frame picture; calculate a feature similarity between each candidate video and the target video according to the video feature and the picture feature, sort the candidate videos based on the feature similarity, and determine a target candidate video belonging to the same type as the target video from the plurality of candidate videos according to a sorting result; and obtain the target candidate video.

[0151] Computer program code for carrying out operations of the present disclosure can be written in any one or more of a variety of programming languages or combinations of languages including an object oriented programming language such as Java, Smalltalk, C++ or the like and conventional procedural programming languages such as "C" or the like. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).

[0152] The computer program code can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce the operations described.

[0153] The modules described in the embodiments of the present disclosure can be implemented in the form of software, or can be implemented in the form of hardware. In some cases, the name of the module does not constitute a limitation on the module itself, for example, the extraction module can also be described as a "module for extracting video frame pictures".

[0154] The functions described above herein can be performed at least in part by one or more hardware logic components. For example, and without limitation, example types of hardware logic components that can be used include field programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip systems (SOCs), complex programmable logic devices (CPLDs), etc.

[0155] In the context of the present disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable storage media can include, without limitation, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media can include one or more lines of electrical connections, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical storage devices, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0156] According to one or more embodiments of the present disclosure, example 1 provides a method for obtaining video material, comprising:

[0157] extracting a video frame picture of a target video;

[0158] extracting a video feature corresponding to the target video and a picture feature corresponding to the video frame picture;

[0159] calculating a feature similarity between each candidate video and the target video according to the video feature and the picture feature;

[0160] ranking the candidate videos based on the feature similarity, and determining a target candidate video belonging to the same type as the target video from the plurality of candidate videos according to the ranking result;

[0161] obtaining the target candidate video.

[0162] According to one or more embodiments of the present disclosure, example 2 provides the method of example 1, wherein the plurality of candidate videos comprises preset candidate videos, and the calculating of the feature similarity between each candidate video and the target video based on the video feature and the picture feature comprises:

[0163] calculating a picture feature similarity between each preset candidate video and the target video based on the picture feature, and sorting the preset candidate videos in descending order of the picture feature similarity, and selecting the first preset number of preset candidate videos as a first candidate video set according to the sorting result;

[0164] calculating a video feature similarity between each preset candidate video and the target video based on the video feature, and sorting the preset candidate videos in descending order of the video feature similarity, and selecting the second preset number of preset candidate videos as a second candidate video set according to the sorting result;

[0165] determining the target candidate video from the first candidate video set and the second candidate video set.

[0166] According to one or more embodiments of the present disclosure, example 3 provides the method of example 1, wherein the plurality of candidate videos comprises traversed candidate videos determined by traversing a pre-established hierarchical navigation small world (HNSW) recall tree according to a preset recall strategy, the HNSW recall tree comprises a first HNSW recall tree or a second HNSW recall tree, and the calculating of the feature similarity between each candidate video and the target video based on the video feature and the picture feature comprises:

[0167] calculating a picture feature similarity between each traversed candidate video and the target video based on the picture feature through the first HNSW recall tree according to the preset recall strategy, and determining a first candidate video set from the plurality of traversed candidate videos according to the picture feature similarity, wherein the first HNSW recall tree is a HNSW recall tree pre-established based on a picture feature similarity between each two candidate videos;

[0168] According to the video features, the second HNSW recall tree calculates the video feature similarity between each traversed candidate video and the target video based on the preset recall strategy, and determines a second candidate video set from the plurality of traversed candidate videos according to the video feature similarity, the second HNSW recall tree being a HNSW recall tree pre-established according to the video feature similarity between each two candidate videos;

[0169] The target candidate video is determined from the first candidate video set and the second candidate video set.

[0170] According to one or more embodiments of the present disclosure, example 4 provides the method of example 3, the first HNSW recall tree and the second HNSW recall tree each comprising a plurality of nodes, each node corresponding to one candidate video;

[0171] The preset recall strategy comprises:

[0172] A preset traversal step is performed until the target HNSW recall tree is traversed;

[0173] The preset traversal step comprises:

[0174] A preset node corresponding to the target HNSW recall tree is obtained;

[0175] A target node is determined from the target HNSW recall tree according to the preset node, the target node comprising the preset node and nodes connected to the preset node;

[0176] The feature similarity between each to-be-determined candidate video and the target video is calculated according to a target feature, the to-be-determined candidate video being a candidate video corresponding to the target node;

[0177] The node corresponding to the candidate video with the highest feature similarity is taken as a new preset node;

[0178] It is determined whether there is a node connected to the new preset node in the target HNSW recall tree;

[0179] In a case where it is determined that there is a node connected to the new preset node in the target HNSW recall tree, the preset traversal step is re-executed, and in a case where it is determined that there is no node connected to the new preset node in the target HNSW recall tree, it is determined that the target HNSW recall tree is traversed;

[0180] In a case where the target HNSW recall tree is traversed, the traversed candidate videos corresponding to the traversed nodes are sorted in descending order of the feature similarity, and the top third preset number of traversed candidate videos are selected as a specific candidate video set according to the sorting result.

[0181] wherein the target feature comprises the picture feature or the video feature, in a case where the target feature is the picture feature, the target HNSW recall tree is the first HNSW recall tree, the feature similarity is a picture feature similarity, and the specific candidate video set is the first candidate video set; in a case where the target feature is the video feature, the target HNSW recall tree is the second HNSW recall tree, the feature similarity is a video feature similarity, and the specific candidate video set is the second candidate video set.

[0182] According to one or more embodiments of the present disclosure, example 5 provides the method of any one of examples 2-4, and the determining the target candidate video from the first candidate video set and the second candidate video set comprises:

[0183] ordering each candidate video in the first candidate video set and the second candidate video set in a descending order of feature similarity, the feature similarity comprising the picture feature similarity or the video feature similarity;

[0184] selecting a top fourth preset number of candidate videos according to the ordering result as the target candidate video.

[0185] According to one or more embodiments of the present disclosure, example 6 provides the method of example 4, and the calculating the feature similarity of each candidate video with the target video according to the video feature and the picture feature; ordering the candidate videos based on the feature similarity, and determining a target candidate video belonging to the same type as the target video from the plurality of candidate videos according to the ordering result comprises:

[0186] performing feature fusion on the video feature and the picture feature to obtain a fusion feature;

[0187] calculating the feature similarity of each candidate video with the target video according to the fusion feature;

[0188] ordering the candidate videos based on the fusion feature similarity, and determining a target candidate video belonging to the same type as the target video from the plurality of candidate videos according to the ordering result.

[0189] According to one or more embodiments of the present disclosure, example 7 provides the method of example 6, and the performing feature fusion on the video feature and the picture feature to obtain a fusion feature comprises:

[0190] for each first feature element of the video feature, adding the first feature element to a second feature element corresponding to the first feature element in the picture feature to obtain the fusion feature.

[0191] According to one or more embodiments of the present disclosure, example 8 provides the method of example 6, before the feature fusion of the video features and the picture features to obtain the fusion features, the method further comprises:

[0192] identifying a video type of the target video through a material type identification model;

[0193] determining a first preset weight corresponding to the video features and a second preset weight corresponding to the picture features according to the video type respectively;

[0194] the feature fusion of the video features and the picture features to obtain the fusion features comprises:

[0195] weighting and summing the video features and the picture features according to the first preset weight and the second preset weight to obtain the fusion features.

[0196] According to one or more embodiments of the present disclosure, example 9 provides the method of example 6, the HNSW recall tree comprises a third HNSW recall tree established in advance, a plurality of the candidate videos comprises traversed candidate videos determined after traversing the third HNSW recall tree according to the preset recall strategy, the calculation of the fusion feature similarity between each candidate video and the target video according to the fusion features; sorting the candidate videos based on the fusion feature similarity, and determining a target candidate video belonging to the same type as the target video from the plurality of candidate videos according to the sorting result comprises:

[0197] traversing the third HNSW recall tree by performing the preset traversal step, and calculating the fusion feature similarity between each traversed candidate video and the target video according to the fusion features during the traversal process, the third HNSW recall tree being a HNSW recall tree established in advance according to the fusion feature similarity between each two candidate videos;

[0198] in the case where the third HNSW recall tree is traversed, sorting the traversed candidate videos corresponding to the nodes according to the fusion feature similarity from high to low, and selecting the first fifth preset number of traversed candidate videos as the target candidate video according to the sorting result.

[0199] According to one or more embodiments of the present disclosure, example 10 provides the method of example 6, wherein the plurality of candidate videos comprises preset candidate videos, and the calculating of the fusion feature similarity between each candidate video and the target video based on the fusion features comprises:

[0200] The fusion feature similarity between each preset candidate video and the target video is calculated based on the fusion features respectively, and the preset candidate videos are sorted in descending order of the fusion feature similarity, and the top sixth preset number of preset candidate videos are selected as the target candidate videos according to the sorting result.

[0201] According to one or more embodiments of the present disclosure, example 11 provides the method of example 1, wherein before the calculating of the feature similarity between each candidate video and the target video based on the video features and the picture features, the method further comprises:

[0202] obtaining the text features of the target video through a pre-trained text feature extraction model; and / or,

[0203] obtaining the audio features of the target video through a pre-trained audio feature extraction model;

[0204] The calculating of the feature similarity between each candidate video and the target video based on the video features and the picture features comprises:

[0205] The calculating of the feature similarity between each candidate video and the target video based on the picture features, the video features and the specified features, wherein the specified features comprise the text features and / or the audio features.

[0206] According to one or more embodiments of the present disclosure, example 12 provides the method of example 1, wherein before the calculating of the feature similarity between each candidate video and the target video based on the video features and the picture features, the method further comprises:

[0207] performing feature dimension reduction on the picture features and the video features to obtain target picture features and target video features after dimension reduction;

[0208] The calculating of the feature similarity between each candidate video and the target video based on the video features and the picture features comprises:

[0209] The calculating of the feature similarity between each candidate video and the target video based on the target picture features and the target video features.

[0210] According to one or more embodiments of the present disclosure, example 13 provides the method of example 1, wherein the video frame picture comprises a cover frame picture of the target video or a preset video frame picture other than the cover frame picture.

[0211] According to one or more embodiments of the present disclosure, example 14 provides the method of example 1, wherein the extracting the video feature corresponding to the target video comprises:

[0212] dividing the target video into a plurality of video segments according to a preset step length;

[0213] for each of the video segments, inputting the video segment into a pre-trained video feature extraction model to obtain the video feature corresponding to the video segment.

[0214] According to one or more embodiments of the present disclosure, example 15 provides a device for obtaining video material, the device comprising:

[0215] an extraction module configured to extract a video frame picture of a target video;

[0216] an extraction module configured to extract a video feature corresponding to the target video and a picture feature corresponding to the video frame picture;

[0217] a similarity calculation module configured to calculate a feature similarity between each candidate video and the target video according to the video feature and the picture feature;

[0218] a determination module configured to sort the candidate videos based on the feature similarity and determine a target candidate video belonging to the same type as the target video from the plurality of candidate videos according to a sorting result;

[0219] a material obtaining module configured to obtain the target candidate video.

[0220] According to one or more embodiments of the present disclosure, example 16 provides a computer readable medium having a computer program stored thereon, the program being executed by a processing device to implement the steps of the method of any one of examples 1-14.

[0221] According to one or more embodiments of the present disclosure, example 17 provides an electronic device comprising:

[0222] a storage device having a computer program stored thereon;

[0223] a processing device configured to execute the computer program in the storage device to implement the steps of the method of any one of examples 1-14.

[0224] The above description merely illustrates the preferred embodiment of the disclosure and a principle of applied technologies. It should be understood by those skilled in the art that the disclosed range of the disclosure is not limited to the technical solutions formed by the specific combinations of the technical features described above, and should also cover other technical solutions formed by the combinations of the technical features described above or their equivalent features without departing from the disclosed concept. For example, the technical solutions formed by the mutual replacement of the above-described features and the technical features with similar functions disclosed in the disclosure (but not limited to) can be formed.

[0225] Furthermore, although operations are depicted in a particular, sequential order, this should not be understood as requiring or implying that the operations are performed in the order illustrated or sequentially. In certain circumstances, multitasking and parallel processing can be advantageous. Likewise, although specific implementation details are contained in the above discussion, these should not be construed as limiting the scope of the disclosure. Certain features described in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable sub-combination.

[0226] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely illustrative of specific forms of implementing the claims. With respect to the devices in the above-described embodiments, the specific manner in which the various modules perform operations has been described in detail in the embodiments related to the method, and will not be described here in detail.

Claims

1. A method of acquiring video material, characterized by, The method includes: Extract video frame images from the target video; Extract video features corresponding to the target video and image features corresponding to the video frame images; the video features include features extracted for each video segment of the target video. The feature similarity between each candidate video and the target video is calculated based on the video features and the image features; the multiple candidate videos include traversed candidate videos determined after traversing a pre-established HNSW recall tree according to a preset recall strategy, and the HNSW recall tree includes a first HNSW recall tree and a second HNSW recall tree. The candidate videos are sorted based on the feature similarity, and a target candidate video belonging to the same type as the target video is determined from the multiple candidate videos according to the sorting results; The target candidate video is obtained; the target candidate video is determined from the first candidate video set and the second candidate video set; The first candidate video set is determined in the following way: based on the image features, the image feature similarity between each traversed candidate video and the target video is calculated using a first HNSW recall tree based on a preset recall strategy, and the first candidate video set is determined from multiple candidate videos based on the image feature similarity. The second candidate video set is determined as follows: based on the video features, the video feature similarity between each traversed candidate video and the target video is calculated using the second HNSW recall tree based on the preset recall strategy, and the second candidate video set is determined from multiple traversed candidate videos based on the video feature similarity.

2. The method according to claim 1, characterized in that, The plurality of candidate videos includes preset candidate videos, and the feature similarity between each candidate video and the target video is calculated based on the video features and the image features; The candidate videos are ranked based on the feature similarity, and the target candidate videos belonging to the same type as the target video are determined from the multiple candidate videos according to the ranking results, including: Calculate the image feature similarity between each preset candidate video and the target video based on the image features, sort the preset candidate videos in descending order of image feature similarity, and determine the first preset number of preset candidate videos as the first candidate video set based on the sorting results; Calculate the video feature similarity between each preset candidate video and the target video based on the video features, sort the preset candidate videos in descending order of video feature similarity, and determine the first second preset number of preset candidate videos as the second candidate video set based on the sorting results; The target candidate video is determined from the first candidate video set and the second candidate video set.

3. The method according to claim 1, characterized in that, The first HNSW recall tree is an HNSW recall tree pre-built based on the image feature similarity between each two candidate videos; The second HNSW recall tree is an HNSW recall tree pre-built based on the video feature similarity between each pair of candidate videos.

4. The method according to claim 3, characterized in that, Both the first HNSW recall tree and the second HNSW recall tree include multiple nodes, and each node corresponds one-to-one with the candidate video; The preset recall strategy includes: Execute the preset traversal steps until the target HNSW recall tree has been traversed; The preset traversal steps include: Obtain the preset node corresponding to the target HNSW recall tree; A target node is determined from the target HNSW recall tree based on the preset node, wherein the target node includes the preset node and nodes connected to the preset node; Calculate the feature similarity between each candidate video to be determined and the target video based on the target features, wherein the candidate video to be determined is the candidate video corresponding to the target node; The node corresponding to the candidate video with the highest feature similarity is used as the new preset node; Determine whether there is a node in the target HNSW recall tree that is connected to the new preset node; If it is determined that there is a node in the target HNSW recall tree that is connected to the new preset node, the preset traversal step is re-executed. If it is determined that there is no node in the target HNSW recall tree that is connected to the new preset node, the target HNSW recall tree is traversed. After traversing the target HNSW recall tree, the candidate videos corresponding to the traversed nodes are sorted in descending order of feature similarity, and the top three preset number of candidate videos are selected as a specific candidate video set based on the sorting results. Wherein, the target feature includes the image feature or the video feature. When the target feature is the image feature, the target HNSW recall tree is the first HNSW recall tree, the feature similarity is the image feature similarity, and the specific candidate video set is the first candidate video set; when the target feature is the video feature, the target HNSW recall tree is the second HNSW recall tree, the feature similarity is the video feature similarity, and the specific candidate video set is the second candidate video set.

5. The method according to any one of claims 1-4, characterized in that, Determining the target candidate video from the first candidate video set and the second candidate video set includes: Each candidate video in the first candidate video set and the second candidate video set is sorted in descending order of feature similarity, where feature similarity includes image feature similarity or video feature similarity. Based on the sorting results, the top four preset number of candidate videos are selected as the target candidate videos.

6. The method according to claim 1, characterized in that, The step of calculating the feature similarity between each candidate video and the target video based on the video features and the image features; ranking the candidate videos based on the feature similarity; and determining the target candidate video belonging to the same type as the target video from the multiple candidate videos according to the ranking results includes: The video features and the image features are fused to obtain fused features; Calculate the similarity of the fusion features between each candidate video and the target video based on the fusion features; The candidate videos are sorted based on the fusion feature similarity, and a target candidate video belonging to the same type as the target video is determined from the multiple candidate videos according to the sorting results.

7. The method according to claim 6, characterized in that, The feature fusion of the video features and the image features to obtain the fused features includes: For each first feature element of the video features, the first feature element is added to the second feature element of the image features corresponding to the first feature element to obtain the fused feature.

8. The method according to claim 6, characterized in that, Before fusing the video features and the image features to obtain the fused features, the method further includes: The video type of the target video is identified using a material type recognition model; Based on the video type, determine the first preset weight corresponding to the video feature and the second preset weight corresponding to the image feature; The feature fusion of the video features and the image features to obtain the fused features includes: The video features and the image features are weighted and summed according to the first preset weight and the second preset weight to obtain the fused features.

9. The method according to claim 6, characterized in that, The HNSW recall tree includes a pre-established third HNSW recall tree, and the multiple candidate videos include traversed candidate videos determined after traversing the third HNSW recall tree according to the preset recall strategy. The similarity of the fusion features between each candidate video and the target video is calculated based on the fusion features. The candidate videos are ranked based on the fusion feature similarity, and target candidate videos belonging to the same type as the target video are determined from the multiple candidate videos according to the ranking results, including: The third HNSW recall tree is traversed by executing the preset traversal steps, and the similarity of the fusion features between each traversed candidate video and the target video is calculated based on the fusion features during the traversal process. The third HNSW recall tree is an HNSW recall tree pre-built based on the similarity of the fusion features between every two candidate videos. After traversing the third HNSW recall tree, the candidate videos corresponding to the traversed nodes are sorted in descending order of the similarity of the fused features, and the top five preset number of candidate videos are selected as the target candidate videos based on the sorting results.

10. The method according to claim 6, characterized in that, The plurality of candidate videos includes preset candidate videos. The step of calculating the fusion feature similarity between each candidate video and the target video based on the fusion feature; ranking the candidate videos based on the fusion feature similarity; and determining the target candidate video belonging to the same type as the target video from the plurality of candidate videos based on the ranking result includes: The similarity of the fusion features between each preset candidate video and the target video is calculated based on the fusion features. The preset candidate videos are then sorted in descending order of the fusion feature similarity. Based on the sorting results, the top six preset candidate videos are selected as the target candidate videos.

11. The method according to claim 1, characterized in that, Before calculating the feature similarity between each candidate video and the target video based on the video features and the image features, the method further includes: The text features of the target video are obtained by a pre-trained text feature extraction model; and / or, the audio features of the target video are obtained by a pre-trained audio feature extraction model. The step of calculating the feature similarity between each candidate video and the target video based on the video features and the image features includes: The feature similarity between each candidate video and the target video is calculated based on the image features, the video features, and specified features, wherein the specified features include the text features and / or the audio features.

12. The method according to claim 1, characterized in that, Before calculating the feature similarity between each candidate video and the target video based on the video features and the image features, the method further includes: The image features and the video features are subjected to feature dimensionality reduction to obtain the dimensionality-reduced target image features and target video features; The step of calculating the feature similarity between each candidate video and the target video based on the video features and the image features includes: The feature similarity between each candidate video and the target video is calculated based on the features of the target image and the features of the target video.

13. The method according to claim 1, characterized in that, The video frame images include the cover frame image of the target video or preset video frame images other than the cover frame image.

14. The method according to claim 1, characterized in that, The extraction of video features corresponding to the target video includes: The target video is divided into multiple video segments according to a preset step size; For each video segment, the video segment is input into a pre-trained video feature extraction model to obtain the video features corresponding to each video segment.

15. An apparatus for acquiring video footage, characterized in that, The device includes: The extraction module is used to extract video frame images from the target video. An extraction module is used to extract video features corresponding to the target video and image features corresponding to the video frame images; the video features include features extracted for each video segment of the target video. The similarity calculation module is used to calculate the feature similarity between each candidate video and the target video based on the video features and the image features; the multiple candidate videos include traversed candidate videos determined after traversing a pre-established HNSW recall tree according to a preset recall strategy, and the HNSW recall tree includes a first HNSW recall tree and a second HNSW recall tree. The determining module is used to sort the candidate videos based on the feature similarity, and determine the target candidate video that belongs to the same type as the target video from the multiple candidate videos according to the sorting result; The material acquisition module is used to acquire the target candidate video; the target candidate video is determined from a first candidate video set and a second candidate video set; the first candidate video set is determined by: calculating the image feature similarity between each traversed candidate video and the target video based on the image features using a first HNSW recall tree and a preset recall strategy, and determining the first candidate video set from multiple candidate videos based on the image feature similarity; the second candidate video set is determined by: calculating the video feature similarity between each traversed candidate video and the target video based on the video features using a second HNSW recall tree and the preset recall strategy, and determining the second candidate video set from multiple traversed candidate videos based on the video feature similarity.

16. A computer-readable medium having a computer program stored thereon, characterized in that, When the program is executed by the processing device, it implements the steps of the method described in any one of claims 1-14.

17. An electronic device, characterized in that, include: A storage device on which computer programs are stored; A processing device for executing the computer program in the storage device to implement the steps of the method according to any one of claims 1-14.

Citation Information

Patent Citations

  • Video positioning method and device, storage medium and electronic device

    CN110134829A

  • Video retrieval method and device, electronic equipment and storage medium

    CN111241345A