Video feature extraction method, video sampling method, model training method and video retrieval method

By extracting the frame features of video frames and dividing the video frame groups for sampling, the problem of insufficient or redundant video frame sampling in the prior art is solved, and efficient and accurate video feature extraction is achieved.

CN119942396APending Publication Date: 2025-05-06ALIBABA CLOUD COMPUTING CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411758662.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-02
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently and accurately sample video frames and generate video features, especially when processing a large number of video frames, there is a problem of insufficient sampling or redundancy in uniform sampling and random sampling.

Method used

By extracting the frame characteristics of the video frame, dividing the video frame into multiple video frame groups based on the frame characteristics, and sampling each video frame group to obtain the target video frame. Then, the frame features of the target video frame are aggregated to generate the video features of the video.

Benefits of technology

Efficient and accurate video frame sampling is achieved, ensuring the diversity and depth of sampling, and the generated video features can accurately reflect the video content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942396A_ABST
    Figure CN119942396A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a video feature extraction method, a video sampling method, a model training method, a video retrieval method, computing equipment, a computer storage medium and a computer program product. The method comprises the following steps: extracting frame features of video frames in a video; dividing video frames in the video into a plurality of video frame groups based on frame features; respectively sampling the plurality of video frame groups to obtain a plurality of target video frames; and carrying out aggregation processing on the frame features of the plurality of target video frames to obtain video features of the video. According to the technical scheme provided by the embodiment of the invention, layered sampling is carried out on the video frames, so that the sampling diversity and the sampling depth are ensured, efficient and accurate video frame sampling is realized, then the frame features of the video frames obtained by sampling are subjected to aggregation processing to obtain the video features, and the accuracy of the video features is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of video processing technology, and in particular to a video feature extraction method, a video sampling method, a model training method, a video retrieval method, a computing device, a computer storage medium, and a computer program product. Background Art

[0002] With the development of Internet technology, video content has exploded and there are more and more video applications, such as video search, video recommendation, video question and answer, etc. These video applications often involve video retrieval operations and are implemented based on video features. Since videos usually contain a large number of video frames, how to efficiently and accurately sample video frames and generate video features based on them has become a technical problem that needs to be solved. Summary of the invention

[0003] Embodiments of the present application provide a video feature extraction method, a video sampling method, a model training method, a video retrieval method, a computing device, a computer storage medium, and a computer program product.

[0004] In a first aspect, an embodiment of the present application provides a video feature extraction method, comprising:

[0005] Extracting frame features of video frames in the video;

[0006] Based on the frame features, dividing the video frames in the video into a plurality of video frame groups;

[0007] Sampling the multiple video frame groups respectively to obtain multiple target video frames;

[0008] Aggregate the frame features of the multiple target video frames to obtain video features of the video.

[0009] In a second aspect, an embodiment of the present application provides a model training method, comprising:

[0010] Acquire training data; the training data includes a sample video and a sample text matching the sample video;

[0011] Extracting frame features of video frames in the sample video using a feature processing model;

[0012] Based on the frame features, dividing the video frames in the sample video into a plurality of sample frame groups;

[0013] Sampling the multiple sample frame groups respectively to obtain multiple sample frames;

[0014] Aggregating the frame features of the plurality of sample frames to obtain video features of the sample video;

[0015] Extracting text features of the sample text using the feature processing model;

[0016] The feature processing model is trained based on the video features of the sample video and the text features of the sample text; the feature processing model is used to extract the text features of the target text and the frame features of the video frames in the video to be matched; the frame features of the multiple video frames acquired by the video to be matched are aggregated to obtain the video features of the video to be matched.

[0017] In a third aspect, an embodiment of the present application provides a video retrieval method, comprising:

[0018] Responding to the video retrieval request, obtaining retrieval information;

[0019] Extracting text features from the search information;

[0020] At least one target video is determined based on feature similarities between text features and video features of multiple videos to be matched; wherein the video features are obtained by aggregating frame features of multiple target video frames in the video; the multiple target video frames are obtained by sampling from multiple video frame groups corresponding to the video; the multiple video frame groups are obtained by hierarchically processing video frames in the video based on frame features.

[0021] In a fourth aspect, an embodiment of the present application provides a video sampling method, including:

[0022] Extracting frame features of video frames in the video;

[0023] Based on the frame features, dividing the video frames in the video into a plurality of video frame groups;

[0024] The multiple video frame groups are sampled respectively to obtain multiple target video frames.

[0025] In a fifth aspect, an embodiment of the present application provides a computing device, including a processing component and a storage component;

[0026] The storage component stores one or more computer instructions; the one or more computer instructions are used to be called and executed by the processing component to implement the video feature extraction method as described in the first aspect above, or to implement the model training method as described in the second aspect above, or to implement the video retrieval method as described in the third aspect above, or to implement the video sampling method as described in the fourth aspect above.

[0027] In a fifth aspect, a computer storage medium is provided in an embodiment of the present application, storing a computer program, which, when executed by a computer, implements the video feature extraction method as described in the first aspect above, or implements the model training method as described in the second aspect above, or implements the video retrieval method as described in the third aspect above, or implements the video sampling method as described in the fourth aspect above.

[0028] In a sixth aspect, a computer program product is provided in an embodiment of the present application, wherein the computer program product includes a computer program code, and when the computer program code is executed by a computer, the video feature extraction method as described in the first aspect above is implemented, or the model training method as described in the second aspect above is implemented, or the video retrieval method as described in the third aspect above is implemented, or the video sampling method as described in the fourth aspect above is implemented.

[0029] In the embodiment of the present application, the frame features of the video frames in the video are extracted; based on the frame features, the video frames in the video are divided into multiple video frame groups; the multiple video frame groups are sampled respectively to obtain multiple target video frames; the frame features of the multiple target video frames are aggregated to obtain the video features of the video. The technical solution is to first group the video frames according to the frame features, then sample the grouped video frame groups to obtain the target video frames, and then aggregate the frame features of the target video frames to obtain the video features. By layering the video frames, the diversity and depth of the sampling are guaranteed, and efficient and accurate video frame sampling is achieved. And the video features generated based on the frame features of the target video frames can accurately reflect the video content.

[0030] These and other aspects of the present application will become more clearly understood in the description of the following embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0032] Figure 1 A flowchart of an embodiment of a video feature extraction method provided by the present application is shown;

[0033] Figure 2 A schematic diagram of a video feature extraction process in a practical application of an embodiment of the present application is shown;

[0034] Figure 3A flow chart of an embodiment of a model training method provided by the present application is shown;

[0035] Figure 4 A flowchart of an embodiment of a video retrieval method provided by the present application is shown;

[0036] Figure 5 A schematic diagram of implementing a video retrieval process in a practical application of an embodiment of the present application is shown;

[0037] Figure 6 A system architecture diagram is shown in which the technical solution of the embodiment of the present application is applied;

[0038] Figure 7 A flowchart of an embodiment of a video sampling method provided by the present application;

[0039] Figure 8 A block diagram of an embodiment of a video feature extraction device provided by the present application is shown;

[0040] Fig. 9 A block diagram of an embodiment of a model training device provided by the present application is shown;

[0041] Fig.10 A block diagram of an embodiment of a video retrieval device provided by the present application is shown;

[0042] Fig.11 A block diagram of an embodiment of a video sampling device provided by the present application is shown;

[0043] Fig.12 A block diagram of an embodiment of a computing device provided by the present application is shown. DETAILED DESCRIPTION

[0044] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application.

[0045] In some of the processes described in the specification and claims of this application and the above-mentioned figures, multiple operations that appear in a specific order are included, but it should be clearly understood that these operations may not be executed in the order in which they appear in this article or executed in parallel. The serial numbers of the operations, such as 101, 102, etc., are only used to distinguish between different operations, and the serial numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions of "first", "second", etc. in this article are used to distinguish different messages, devices, modules, etc., do not represent the order of precedence, and do not limit the "first" and "second" to be different types.

[0046] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0047] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of this application.

[0048] With the development of Internet technology, video content has exploded and there are more and more video applications, such as video search, video recommendation, video question and answer, etc. These video applications often involve video retrieval operations and the matching of video and text. The matching of video and text is based on video features. Since videos usually contain a large number of video frames, in order to accurately extract video features that express the video content and avoid a large amount of redundant information, the video can be sampled and video features can be generated based on the adopted video frames.

[0049] In the process of realizing the concept of the present application, the inventor found that if sampling is performed in a uniform sampling manner, that is, video frames are selected from the video at fixed intervals; for example, if 1 frame is selected every 5 frames, then the 1st frame, the 6th frame, the 11th frame, etc. of the original video will be selected. Although this method can obtain the pictures of different time periods of the video more evenly, since it is mechanically sampled at fixed intervals, the sampling strategy cannot be adjusted according to the importance or degree of change of the video content, so the coverage of key content may be missing in diversity. For example, in a video containing a fast action scene (such as a wonderful moment in a sports game), uniform sampling may be sampled in the intervals where these key actions occur, resulting in the collected video frames not being able to well reflect the details of these key actions. In addition, uniform sampling may collect some redundant video frames in some cases. For example, in a part where the video content is relatively stable and there is not much change, uniform sampling will still select video frames at fixed intervals. The difference between these frames may be very small. For subsequent analysis, it may not provide much new information, but it increases the amount of data processing, thereby affecting the sampling efficiency. If the random sampling method is used, that is, the video frames are randomly selected, this method will cause the information to be incoherent or repeated. The inventor also thought that although key information can be retained if only key frames are selected, important information in transition frames may be lost. If the action information in the video is identified and video frames are selected according to the action changes, the method is less universal and is not suitable for videos of static scenes.

[0050] In order to efficiently and accurately sample video frames and generate accurate video features based on them, the inventor has proposed the technical solution of the present application after a series of studies. In the embodiment of the present application, the frame features of the video frames in the video are first extracted; then, based on the frame features, the video frames in the video are divided into multiple video frame groups; then, the multiple video frame groups are sampled respectively to obtain multiple target video frames; the frame features of the multiple target video frames are aggregated to obtain the video features of the video. By grouping the video frames according to the frame features, and then sampling the grouped video frame groups to obtain the target video frames, the sampling diversity and sampling depth can be guaranteed, so that the video content can be fully captured. Then, the frame features of the target video frames are aggregated to obtain the video features, and efficient and accurate video frame sampling is achieved. And the video features generated based on the frame features of the target video frames can accurately reflect the video content.

[0051] The implementation details of the technical solution of the embodiment of the present application are elaborated in detail below.

[0052] Figure 1 A flowchart of a video feature extraction method provided in one embodiment of the present application. The technical solution of this embodiment can be executed by the server. The method may include the following steps:

[0053] 101: Extract frame features of video frames in the video.

[0054] Among them, the frame features of the video frames can characterize the content characteristics of each video frame in an informative and representative manner. By extracting the frame features of the video frames, it can provide a basis for the grouping, sampling and acquisition of video features of the video frames.

[0055] In an optional method of the present application, a visual feature extraction model may be used to extract frame features of video frames in a video. The visual feature extraction model may be, for example, a visual encoder, etc. The visual encoder may be constructed and generated based on, for example, a CNN (Convolutional Neural Network), an RNN (Recurrent Neural Network), a Transformer architecture, etc.

[0056] In another optional mode of the present application, in a scenario involving video and text matching, a feature processing model may be used to extract frame features of video frames in a video. The feature processing model may be, for example, a multimodal model, which may include a visual feature extraction model such as a visual encoder and a text feature extraction model such as a text encoder, and the visual encoder in the feature processing model may be used to extract frame features of video frames in a video.

[0057] In addition, color features, texture features, shape features, etc. of the video frame may also be obtained through computational analysis as frame features of the video frame.

[0058] 102: Divide video frames in the video into multiple video frame groups based on frame features.

[0059] In an embodiment of the present application, the frame features extracted through the above step 101 can reflect the picture content of the video frame, so that the video frames can be divided into different video frame groups based on the similarities or differences between these frame features, so that the video frames in the same video frame group have relatively similar features.

[0060] In an optional method of the present application, for example, a clustering algorithm can be used to cluster video frames in a video so as to divide the video frames into multiple video frame groups. For example, the clustering algorithm can include a K-Means clustering algorithm, a hierarchical clustering algorithm, a split hierarchical clustering algorithm, etc. Of course, other methods can also be used to implement it, which will be described in detail in the following embodiments.

[0061] 103: Sample the multiple video frame groups respectively to obtain multiple target video frames.

[0062] In an embodiment of the present application, the video frames of a video are divided into a plurality of video frame groups according to frame features, so that each video frame group has unique content and features. For example, the content shown in the video may include: a person first walks in the park, then sits on a bench to look at the scenery, and finally gets up and leaves. By performing feature extraction on the video frames, the video frames can be divided into a plurality of video frame groups according to the extracted frame features. For example, in this example, the video frames can be divided into three video frame groups, wherein the plurality of video frames contained in the first video frame group can be used to characterize the walking stage in the video, the plurality of video frames contained in the second video frame group can be used to characterize the sitting and looking at the scenery stage in the video, and the plurality of video frames in the third video frame group can be used to characterize the getting up and leaving stage in the video.

[0063] On this basis, multiple video groups can be sampled respectively. Since each video frame group corresponds to a specific content feature of the video, the multiple target video frames obtained through sampling can well represent the characteristics of each video group.

[0064] In the embodiments of the present application, uniform sampling, random sampling, key frame sampling and other sampling methods can be used to perform sparse sampling on multiple video frame groups respectively to improve sampling efficiency. Among them, uniform sampling can refer to selecting video frames as target video frames at fixed intervals in each video frame group. For example, it is set to select 1 frame every n frames, so that some video frames can be regularly obtained from the group; random sampling can refer to randomly selecting several video frames as target video frames in each video frame group; thereby, data deviation can be avoided, video frames at different times and under different circumstances can be obtained, and the characteristics of the video frames in the group can be more comprehensively reflected; key frame sampling can refer to selecting key frames in the video frame group as target video frames; key frames usually include frames that are representative of changes in the content of video frames in the video group, such as scene switching frames, frames with significant changes in actions, etc. There are many ways to determine key frames, such as by analyzing the degree of difference between video frames in the group, when the difference exceeds a certain threshold, the frame can be determined as a key frame. Of course, other sampling methods can also be used, and different grouping methods can also use different sampling methods, which will be introduced in detail in the following embodiments.

[0065] 104: Aggregate frame features of multiple target video frames to obtain video features of the video.

[0066] After sampling multiple video groups to obtain multiple target video frames, frame features of these target video frames may be aggregated to obtain video features that can represent the characteristics of the entire video content.

[0067] In an optional method of the present application, for frame features represented in vector form, the frame features of multiple target video frames can be aggregated in an average aggregation manner to obtain video features. Average aggregation can add the frame features of multiple target video frames by vector, and then divide them by the number of target video frames, and the result is used as the video feature of the video.

[0068] In another optional method of the present application, the frame features of multiple target video frames can be aggregated to obtain video features by weighted average aggregation. Weighted average aggregation can assign a weight to the frame features of each target video frame according to factors such as the importance of the target video frame in the group and the sampling method, and then add them according to the weights and divide them by the sum of the weights to obtain the video features of the video.

[0069] In another optional method of the present application, a feature aggregation model can be used to aggregate the frame features of multiple target video frames to obtain video features of the video. The feature aggregation model can be, for example, a recurrent neural network, a long short-term memory network, etc. The frame features of multiple target video frames are input into the model in sequence, and a video feature that can represent the characteristics of the entire video content is automatically generated through the processing mechanism within the model.

[0070] In the embodiment of the present application, by adopting: extracting frame features of video frames in a video; dividing the video frames in a video into multiple video frame groups based on the frame features; sampling the multiple video frame groups respectively to obtain multiple target video frames; aggregating the frame features of the multiple target video frames to obtain the technical solution of the video features of the video, first grouping the video frames according to the frame features, then sampling the grouped video frame groups to obtain the target video frames, and then aggregating the frame features of the target video frames to obtain the video features, by layering the video frames, the diversity and sampling depth of the sampling are guaranteed, and efficient and accurate video frame sampling is achieved. And the video features generated based on the frame features of the target video frames can accurately reflect the video content.

[0071] In some embodiments, extracting frame features of video frames in a video can be specifically implemented as follows: extracting features of video frames in the video to obtain initial features of the video frames; determining position information of the video frames in the video; and fusing the position information into the initial features to obtain frame features of the video frames.

[0072] Among them, the position information of the video frame in the video can be used to characterize the positional relationship between the video frames, such as the continuity between adjacent video frames, the changes of the video content in different time periods, etc.

[0073] In addition, in an optional method of the present application, the position information can be determined based on the frame number. For example, in a video with a frame rate of 30fps, the 1st frame, the 2nd frame, the 3rd frame... are arranged in sequence, and the relative position of each video frame in the video can be determined by the frame number.

[0074] In another optional method of the present application, the position information can be determined according to the timestamp corresponding to the video frame, so that by comparing the timestamps of different video frames, the time sequence or interval time of the video frames in the video can be determined.

[0075] Among them, since the video frames in the video can constitute a sequence, the position information can be represented by sinusoidal position embedding. Sinusoidal position embedding can be generated using sine and cosine functions, can uniquely identify the position in the sequence, and has smooth periodicity.

[0076] By fusing the position information of the video frame with the extracted initial features, the fused frame features not only contain the visual information of the video frame itself, but also its position information in the video, so that the obtained frame features can more comprehensively reflect the situation of the video frame.

[0077] In an optional method of the present application, the position information can be fused into the initial features by splicing and fusion to obtain frame features. For example, the position information can be represented in the form of a vector (for example, the position information based on the frame number can be represented by an integer, and the position information based on the timestamp can be represented by a vector), and then the position information vector and the initial feature vector are spliced ​​in dimension in a certain order. For example, the initial feature vector is a vector with a dimension of n, and the position information vector is a vector with a dimension of m, then the new vector after splicing is a vector with a dimension of n+m, and this new vector is used as the frame feature of the fused video frame.

[0078] In another optional method of the present application, the position information can be fused into the initial features by weighted summation fusion to obtain frame features. For example, weights can be assigned to the position information and the initial features respectively, and then the position information and the initial features are summed according to their respective weights to obtain the frame features of the fused video frame. In an embodiment of the present application, the weight can be adjusted according to specific application requirements. For example, if the position information is expected to account for a larger proportion of the fused features, the weight of the position information can be appropriately increased; if more attention is paid to the visual information of the initial features themselves, the weight of the initial features can be appropriately increased.

[0079] In some embodiments, dividing the video frames in the video into a plurality of video frame groups may be specifically implemented as follows: based on frame features, clustering the video frames in the video to obtain a plurality of video frame groups.

[0080] In some embodiments, based on frame features, clustering video frames in a video to obtain multiple video frame groups can be specifically implemented as follows: determining multiple cluster centers from multiple video frames; for any cluster center, determining the similarity distances between the multiple video frames and the cluster center frame based on video frame feature data corresponding to the multiple video frames respectively; clustering at least one video frame that meets the similarity distance condition to the cluster center frame to generate multiple video frame groups.

[0081] In one embodiment of the present application, when clustering video frames in a video, the number of video frame groups to be divided into, K, can be determined first, where K is a positive integer, and the value of the number of groups K can be determined based on experience or through some experiments. For example, if it is roughly known that there are several obviously different scenes in the video, the K value can be set according to the number of scenes.

[0082] After determining the number of groups K, K cluster centers can be randomly initialized. These cluster centers can represent the initial cluster centers of different video frame groups in the feature space. The initial positions of the cluster centers can be randomly selected, but will be continuously updated as the algorithm runs.

[0083] The frame feature of each video frame can be regarded as a point in the feature space, and then the distance from this point to K cluster centers is calculated. In an embodiment of the present application, the Euclidean distance can be used as a distance metric, that is, the distance is determined by calculating the square root of the sum of the squares of the coordinate differences of the two points. According to the calculated distance, each video frame can be assigned to the group where the cluster center closest to it is located. For example, if the distance from the frame feature of video frame A to cluster center 1 is shorter than the distance to other cluster centers, then video frame A will be assigned to the group represented by cluster center 1.

[0084] After all video frames are assigned to corresponding groups, the cluster center of each group can be recalculated. When recalculating the cluster center of each group, the new cluster center can be obtained by calculating the mean of the frame features of all video frames in the group. In other words, the frame feature vectors of all video frames in the group can be added together and then divided by the number of video frames in the group, and the result is the new cluster center. In this way, the cluster center will be adjusted according to the actual distribution of the video frames in the group.

[0085] Furthermore, the above steps of allocating video frames and updating cluster centers can be repeated until the cluster centers no longer change significantly, that is, the convergence condition is reached. Usually, a threshold can be set. For example, when the change in the position of the cluster center in two consecutive iterations is less than the threshold, the algorithm is considered to have converged, and the operation of dividing the video frames into K video groups is completed.

[0086] In another implementation of the present application, when clustering video frames in a video, each video frame can be first regarded as a separate group, that is, there are the same number of video frame groups as the number of video frames. Then, the two video frame groups with the closest distance can be merged together. In an embodiment of the present application, for example, a single connection method, a full connection method, and an average connection method can be used to calculate the distance between video groups. Among them, the single connection method can determine the distance between video groups based on the distance between the two closest video frames in the two video groups; the full connection method can determine the distance between groups based on the maximum value of the distance between all video frames in the two groups; the average connection method can determine the distance between groups based on the average value of the distance between all video frames in the two groups. After determining the distance, the two video groups with the closest distance can be selected for merging.

[0087] In the process of continuous merging, the merging can be stopped when the stop condition is met, thereby obtaining multiple video frame groups. For example, the stop condition may include a preset target value of the number of groups, and the merging operation is stopped when the number of groups reaches the target value; or the change of the distance between groups during the merging process is observed, and when the distance between groups suddenly increases a lot, it means that relatively different groups have been merged, and the merging operation can also be stopped at this time.

[0088] In some embodiments, sampling multiple video frame groups separately to obtain multiple target video frames can be specifically implemented as follows: for any video frame group, filtering video frames whose distance from the cluster center of the video frame group is greater than a first threshold; sampling the filtered video frame group to obtain multiple target video frames.

[0089] In an embodiment of the present application, for each video frame in each video frame group, its similarity distance with the cluster center may be determined.

[0090] After setting the first threshold, when the distance between a video frame and the cluster center of its group is greater than the first threshold, it can be determined that the video frame may not represent the overall characteristics of the group well, and retaining them may interfere with subsequent sampling, analysis and other operations based on the group. Therefore, the video frame can be filtered out of the video frame group, thereby removing video frames that are too different in characteristics from other video frames in the group and are relatively "deviated".

[0091] After filtering out video frames that are far from the cluster center, the remaining video frames in the video frame group can be sampled, so that the amount of data can be further reduced on the basis of retaining the key information of the group, and video frames that can represent the characteristics of the group can be obtained as target video frames.

[0092] In some embodiments, sampling the video frame group after base pair filtering can be specifically implemented as follows:

[0093] The filtered video frame group is sampled with a first number of target video frames in ascending order of distance from the cluster center.

[0094] In the embodiments of the present application, for example, a bubble sort or a quick sort algorithm may be used to sort the multiple video frames.

[0095] In an embodiment of the present application, after the video frames in the filtered video frame group are sorted according to the distance, video frames can be selected in sequence from the sorted video frame sequence as target video frames according to the predetermined number of target video frames to be selected, i.e., the first number. For example, if the first number is 10, the first 10 video frames after sorting can be selected as target video frames. Since these selected target video frames are selected in order of distance from the cluster center from small to large, they are closer to the cluster center in terms of features, can better represent the overall features of the group, can accurately reflect the situation of the filtered video frame group, and can only sample the first number of video frames without sampling all video frames, thereby ensuring both accuracy and sampling efficiency.

[0096] In some embodiments, dividing the video frames in the video into multiple video frame groups based on the frame features may be specifically implemented as follows:

[0097] Based on the frame features, the similarity between any two adjacent video frames is determined; when the similarity between any two adjacent video frames is greater than a second threshold, it is determined that there is a group boundary between the two adjacent video frames; based on the group boundary, the video frames of the video are divided into multiple video frame groups.

[0098] In the embodiment of the present application, the second threshold may be used to determine whether there is a difference between video frames sufficient to divide the video frames into different groups.

[0099] When the similarity between two adjacent video frames is greater than the second threshold, it means that the two video frames have obvious differences in features, and this difference is sufficient to classify them into different groups, so at this time it is determined that there is a group boundary between the two video frames. For example, in a video, if the previous video frame shows an indoor scene, its frame features are relatively dark light, indoor object shapes and textures, etc., and the following video frame suddenly switches to an outdoor scene, and its frame features become bright light, outdoor natural scenery shapes and textures, etc., by calculating the similarity between the frame features of the two adjacent video frames, if it is found that it is greater than the second threshold, then it can be determined that there is a group boundary between the two video frames, that is, the switch from the indoor scene group to the outdoor scene group.

[0100] After determining the group boundaries between all adjacent video frames, the video frames of the video can be divided into multiple video frame groups according to these group boundaries. You can start from the start frame of the video and follow the video frame sequence. When you encounter the first group boundary, divide the previous video frames into a group; then continue to divide the video frames before the current group boundary into a new group every time you encounter a new group boundary until the end of the video. For example, assuming that in a video, three group boundaries are determined through the previous steps, then the video frames will be divided into four video frame groups, corresponding to the video content in different stages or scenes.

[0101] Through this division method based on group boundaries, since the similarity between video frames in the same video group does not exceed the second threshold, the video frames in the same group can have high similarity in features, so they are more consistent in visual content, scenes, actions, etc.; and the division between different groups is based on the group boundaries. The similarity between the video frames on both sides of the group boundaries is greater than the second threshold, so the video frames of different video groups have more obvious differences, so the video frames between different groups show obvious changes in features, such as scene switching, action changes, etc.

[0102] In some embodiments, sampling multiple video frame groups respectively to obtain multiple target video frames can be specifically implemented as follows:

[0103] For any video frame of any video frame group, when the similarity between the video frame and its adjacent previous video frame is less than a third threshold, the video frame is sampled as a target video frame according to probability.

[0104] In an embodiment of the present application, each video frame can be sequentially taken from the video frame group as the current candidate video frame, and then its similarity with the previous candidate video frame is calculated. For example, for the first candidate video frame, since it has no previous candidate video frame, its similarity calculation can be skipped first; for the second candidate video frame, its similarity with the first candidate video frame is calculated; for the third candidate video frame, its similarity with the second candidate video frame is calculated, and so on, to obtain the similarity result of each current candidate video frame with the previous candidate video frame.

[0105] If the similarity between the current candidate video frame and the previous candidate video frame is less than a preset third threshold, it indicates that the two candidate video frames are too similar in features. In order to avoid selecting too similar video frames as target video frames, which results in the target video frame set not being able to well reflect the overall features of the video frame group, the current candidate video frame can be discarded in this case. If the similarity result is not less than the preset third threshold, it indicates that the current candidate video frame and the previous candidate video frame have certain differences in features, and the current video frame can be determined as the target video frame.

[0106] In an embodiment of the present application, after determining adjacent video frames whose similarity is less than the third threshold, the video frames may be sampled as target video frames according to probability. For example, the latter of the two adjacent video frames may be randomly assigned a value, and whether to sample the video frame as the target video frame may be determined based on the assignment result.

[0107] For example, when the similarity between a video frame and a previous video frame is less than a third threshold and the value assigned to the video frame is 1, the video frame is used as a target video frame; if the value assigned to the video frame is 0, no sampling is performed.

[0108] Therefore, on the basis of considering the difference between the video frame and the previous frame, a filtering condition can be further added by assigning a value of 0 or 1, so that the sampling process is more sparse, and the video frames that can reflect a certain difference from the previous frame and meet the assignment conditions can be more accurately selected as target video frames, enriching the diversity of the target video frame set and ensuring accuracy. At the same time, random sampling according to probability does not require sampling of all video frames, thereby ensuring both accuracy and sampling efficiency.

[0109] In some embodiments, the method may further include:

[0110] A second number of video frames are sampled from the video to form a frame sequence from the second number of video frames.

[0111] In an embodiment of the present application, the second number may refer to the value of a specific number of video frames selected from the video, and the value may be determined based on actual application requirements. For example, in a video preview application, it may be necessary to select only a small number of video frames (such as dozens of frames) to quickly display the general content of the video so that the user can have a preliminary impression of the video; and in some more sophisticated video analysis tasks, it may be necessary to select a relatively large number of video frames (such as hundreds of frames or even thousands of frames) to ensure that sufficiently detailed video information can be obtained for accurate analysis.

[0112] After the second number of video frames are selected, these video frames may be arranged in order of their occurrence in the video to form a frame sequence.

[0113] The above-mentioned determination of the similarity between any two adjacent video frames may be specifically implemented as follows: calculating the similarity between any two adjacent sampling frames in the frame sequence.

[0114] The above-mentioned dividing the video into a plurality of video frame groups based on the group boundary may be specifically implemented as follows: dividing the frame sequence into a plurality of video frame groups based on the group boundary.

[0115] In a practical application, the technical solution of the embodiment of the present application can be applied to a scenario involving video and text matching, such as a video search scenario involving matching of search keywords and videos. Therefore, in some embodiments, the method may further include:

[0116] Determine the text to be matched corresponding to the video, and extract the text features of the text to be matched.

[0117] Then the above-mentioned determination of the similarity between any two adjacent video frames based on frame features can be specifically implemented as follows:

[0118] Determine any adjacent first video frame and second video frame; fuse the text feature into the frame feature of the first video frame to obtain first feature data; fuse the text feature into the frame feature of the second video frame to obtain second feature data; and use the similarity between the first feature data and the second feature data as the similarity between the first video frame and the second video frame.

[0119] In some embodiments, determining the text to be matched may be specifically implemented as follows:

[0120] In response to a video search request sent by a target user, a text to be matched provided by the target user is obtained.

[0121] In some other embodiments, determining the text to be matched may be specifically implemented as follows:

[0122] In response to a video recommendation event for a target user, attribute information of the target user is obtained, and the attribute information is used as text to be matched.

[0123] In some other embodiments, determining the text to be matched may be specifically implemented as follows:

[0124] In response to a video question-and-answer request sent by a target user, question information provided by the target user is obtained, and the question information is used as text to be matched.

[0125] In an application scenario of the present application, when a video search request is received from a target user, text information provided by the target user to describe the video content that the target user wants to find can be extracted from the video search request. This text information is the text to be matched. Assuming that the target user enters "funny pet videos" in the input box of the video search engine, "funny pet videos" can be the text to be matched. Therefore, the video to be matched can be selected based on the text to be matched.

[0126] In another application scenario of the present application, when a video recommendation event for a target user occurs, the attribute information of the target user can be obtained. The attribute information may include, for example, the user's age, gender, region, hobbies (such as music preferences, sports preferences, etc.), browsing history (previously browsed video types, web page content, etc.), etc. Then, the attribute information can be used as a text to be matched. By using the attribute information as a text to be matched, it is possible to more accurately recommend videos that may be of interest to the user based on the user's own characteristics and preferences. For example, the system learns that the target user is a young woman who lives in a first-tier city, usually likes fashion and fitness content, and has recently browsed many yoga-related videos. Then, the system will use user attribute information such as "young women, first-tier cities, fashion, fitness, yoga" as text to be matched to screen and recommend videos that may be in line with the user's interests and life background, such as fashion fitness course videos, urban yoga life record videos, etc.

[0127] In another application scenario of the present application, when the target user issues a video question-and-answer request, the specific question information input by the user will be received. For example, the user asks a question about the video: "At what minute and second does the protagonist first appear in this video?" or asks about a certain type of video: "How to shoot a landscape video with a movie-like texture?" Thus, the question information provided by these users can be completely obtained and used as the text to be matched.

[0128] In an embodiment of the present application, in a video frame sequence, any two adjacent video frames may be selected and marked as a first video frame and a second video frame, respectively. The two video frames are adjacent in the video playback order, and there may be a certain continuity or change in the presentation of the video content.

[0129] After the text features are fused into the frame features of the adjacent first and second video frames to obtain the first feature data and the second feature data, the text information can be incorporated into the consideration of the video frame similarity by calculating the similarity between the two fused feature data. The similarity obtained in this way not only reflects the similarity of the visual information of the video frames themselves, but also takes into account the influence of the text information related thereto, thereby more comprehensively evaluating the similarity between adjacent video frames.

[0130] In some embodiments, aggregating frame features of multiple target video frames to obtain video features of a video may be specifically implemented as follows:

[0131] The frame features of multiple target video frames are aggregated using a feature aggregation model to obtain video features.

[0132] In an embodiment of the present application, each target video frame obtained by sampling multiple video frame groups has its own frame features, which can describe the visual content, scene characteristics, action status and other aspects of the video frame from different angles. However, the frame features of each target video frame can only reflect the local situation of the video. Through the feature aggregation model, these scattered local information can be summarized to form a comprehensive representation that can cover the overall characteristics of the video, that is, the video feature. Assuming that the video can show the scene of an outdoor concert, by dividing the video frame groups and sampling the video frame groups, multiple target video frames can be obtained, some of which may focus on the performance movements and expressions of the singers on the stage, and others may focus on the reactions of the audience and the lighting effects of the scene. The feature aggregation model can integrate these frame features with different focuses, so that the final generated video features can not only reflect the singer's performance, but also reflect the overall scene elements such as the atmosphere of the audience and lighting, thereby comprehensively describing the characteristics of this outdoor concert video.

[0133] In an embodiment of the present application, the feature aggregation model may have learning capabilities, which can automatically adjust the aggregation method and weight distribution according to factors such as the content characteristics of different videos, the distribution of target video frames, and specific application requirements, so as to more accurately reflect the overall characteristics of the video.

[0134] In an embodiment of the present application, the feature aggregation model can be trained to generate more accurate video features.

[0135] In some embodiments, determining the similarity between any two adjacent video frames may be specifically implemented as follows:

[0136] The feature calculation model is used to calculate the similarity between any two adjacent video frames.

[0137] In an embodiment of the present application, the feature calculation model may have learning capabilities, may deeply analyze the frame features of the video frames, and calculate the similarity based on the intrinsic relationship between the frame features, thereby obtaining a result that is more consistent with the actual similarity of the video frames.

[0138] The feature calculation model can continuously adjust its parameters through training to adapt to the feature combination and similarity relationship of different types of video frames, and then calculate a more accurate similarity. With the increase of training data and training rounds, the accuracy of the feature calculation model will continue to improve, and it can better cope with similarity calculation tasks under various complex video content and scene changes.

[0139] In some embodiments, aggregating frame features of multiple target video frames to obtain video features of a video may be specifically implemented as follows:

[0140] A sampling sequence is formed by a plurality of target video frames;

[0141] For any target video frame, determine the position information of the target video frame in the sampling sequence;

[0142] Fusion of location information into frame features of target video frames;

[0143] The fused frame features corresponding to multiple target video frames are aggregated to obtain the video features of the video.

[0144] In the embodiment of the present application, by integrating the position information of the target video frame in the sampling sequence into its frame feature, the frame feature can contain not only the visual information of the video frame itself, but also reflect its position context information in the sequence. In this way, in the subsequent aggregation processing, the position factor of the video frame can be comprehensively considered, so that the final generated video feature can more accurately reflect the overall structure and content changes of the video.

[0145] In some embodiments, extracting text features of the text to be matched may be specifically implemented as follows:

[0146] The feature processing model is used to extract text features of the text to be matched. The feature processing model may include a visual encoder and a text encoder, which may be implemented as a multimodal model. A multimodal model is a machine learning model that can process and integrate multiple types of data (such as images, text, audio, video, sensor data, etc.), which can learn the joint embedding representation between video frames and text, so that the extracted video features and text features are more accurate.

[0147] In some embodiments, the feature processing model can be trained and generated by the following operations:

[0148] Acquire training data; the training data includes sample videos and sample texts matching the sample videos;

[0149] Extracting frame features of video frames in sample videos using a feature processing model;

[0150] Based on the frame features, the video frames in the sample video are divided into a plurality of sample frame groups;

[0151] Sampling the multiple sample frame groups respectively to obtain multiple sample frames;

[0152] Aggregate frame features of multiple sample frames to obtain video features of a sample video;

[0153] Extracting text features of sample text using feature processing model;

[0154] A feature processing model is trained based on the video features of the sample video and the text features of the sample text; the feature processing model is used to extract the text features of the target text and the frame features of the video frames in the video to be matched; the frame features of multiple video frames obtained by capturing the video to be matched are aggregated to obtain the video features of the video to be matched.

[0155] In an embodiment of the present application, a sample video and a sample text matched therewith may be used as training data for a feature processing model, wherein the sample video may include video clips for displaying various video contents, scenes, actions, etc., and the sample text may be a text description semantically associated with the sample video.

[0156] In the embodiment of the present application, the feature processing model can first be used to process the sample video. The feature processing model can perform feature extraction operations on each video frame in the sample video one by one, and convert the rich visual information in the video frame into a frame feature form that can be understood and processed by the computer. These frame features can describe the video frame from multiple dimensions.

[0157] After obtaining the frame features of each video frame in the sample video, the video frames in the sample video can be divided into multiple different sample frame groups according to the similarities and differences between the frame features. Through such grouping, the video frames in the same group can have high similarity in certain key features, while there are obvious differences between different groups, thereby achieving a preliminary classification and induction of the sample video content.

[0158] Then, the divided multiple sample frame groups can be sampled. By sampling the divided multiple sample frame groups, partial video frames can be obtained from each sample frame group. These video frames are multiple sample frames, which can reflect the characteristics of each sample frame group to a certain extent, while avoiding processing too much redundant data.

[0159] In the embodiments of the present application, for example, sampling methods such as uniform sampling, random sampling, key frame sampling, etc. may be used to sample the divided multiple sample frame groups.

[0160] By aggregating the frame features of multiple sample frames obtained through sampling, video features that can represent the characteristics of the entire sample video content can be obtained. The video features can be used as a condensed and representative expression of the sample video.

[0161] In addition to processing sample videos, feature processing models can also be used to extract text features of sample text. Sample text can be converted into a quantized, processable form for matching with video features of sample videos and subsequent training operations.

[0162] After obtaining the video features of the sample video and the text features of the sample text, the feature processing model can be trained using the video features and text features. By training the feature processing model, the feature processing model can accurately extract the text features of the target text and the frame features of the video frames in the video to be matched. During the training process, the feature processing model can adjust its own parameters based on the video features of the input sample video and the text features of the sample text so that the output of the model is as consistent as possible with the expected results. For example, if the video features predicted by the model are inconsistent with the video features of the actual sample video, or if there is a deviation between the extracted text features and the text features of the actual sample text, the model will adjust its internal parameters through algorithms such as back propagation to improve the accuracy and reliability of the model.

[0163] In some embodiments, aggregating frame features of multiple sample frames to obtain video features of a sample video may be specifically implemented as follows:

[0164] The frame features of multiple sample frames are aggregated using a feature aggregation model to obtain the video features of the sample video.

[0165] In an embodiment of the present application, the feature aggregation model may have learning capabilities, which can automatically adjust the aggregation method and weight distribution according to factors such as the content characteristics of different videos, the distribution of target video frames, and specific application requirements, so as to more accurately reflect the overall characteristics of the video.

[0166] In an embodiment of the present application, the feature aggregation model can be trained to generate more accurate video features.

[0167] The feature aggregation model can be implemented by combining a feedforward layer and an activation function such as ReLU (Linear Rectification function), for example, and the feedforward layer can be two fully connected layers. The feedforward layer and the ReLU activation function can be used to perform nonlinear transformation on the features to finally obtain the aggregated video features. Of course, the present application is not limited to this.

[0168] In some embodiments, based on frame features, dividing the video frames in the sample video to obtain multiple sample frame groups may be specifically implemented as follows:

[0169] The similarity between two adjacent video frames is calculated using a feature calculation model;

[0170] When the similarity between any two adjacent video frames is greater than a second threshold, determining that there is a group boundary between the two adjacent video frames;

[0171] Based on group boundaries, video frames of the sample video are divided into a plurality of video frame groups.

[0172] In an embodiment of the present application, the feature calculation model may have learning capabilities, may deeply analyze the frame features of the video frames, and calculate the similarity based on the intrinsic relationship between the frame features, thereby obtaining a result that is more consistent with the actual similarity of the video frames.

[0173] The feature calculation model can continuously adjust its parameters through training to adapt to the feature combination and similarity relationship of different types of video frames, and then calculate a more accurate similarity. With the increase of training data and training rounds, the accuracy of the feature calculation model will continue to improve, and it can better cope with similarity calculation tasks under various complex video content and scene changes.

[0174] In some embodiments, based on the video features of the sample video and the text features of the sample text, the training feature processing model can be specifically implemented as follows:

[0175] Take any training data as a positive sample, and the sample video in one training data and the sample text in another training data as negative samples;

[0176] Based on the similarity between video features and text features in positive samples and the similarity between video features and text features in negative samples, the feature processing model is trained through comparative learning.

[0177] In an embodiment of the present application, by constructing positive samples and negative samples and adopting a contrastive learning training method to train a feature extraction model, the feature extraction model can be guided to learn a more accurate feature representation through the similarity difference between video features and text features in positive samples and negative samples, so that the feature extraction model can better distinguish between relevant and irrelevant combinations of videos and texts, thereby improving the accuracy of the feature extraction model in extracting target text features and video frames to be matched.

[0178] In an embodiment of the present application, a positive sample may be composed of a set of matching sample videos and sample texts. For example, in the training data, the sample video is a video showing a sunrise at the seaside, and the sample text is "This video shows a beautiful sunrise at the seaside, the sun slowly rises from the sea level, and the sky is dyed orange-red", then the entire training data is regarded as a positive sample. This method of selecting positive samples can be based on the original corresponding relationship between video and text in practice, providing the model with a correct matching example, allowing the model to learn the association pattern between video features and text features in matching situations.

[0179] Unlike positive samples, sample videos and sample texts that originally did not match can be combined to form negative samples. For example, in the training data, the sample video is a video about the night view of the city, and the sample text of another training data is "The trees in this forest are lush and green, and the birds are singing on the branches." Combining the video of the night view of the city and the text describing the forest scene constitutes a negative sample. By constructing a large number of negative samples in this way, the model can clearly recognize which combinations of videos and texts are mismatched, so as to better grasp the matching relationship between video features and text features during the learning process.

[0180] By allowing the feature processing model to learn the differences between positive samples and negative samples, the feature processing model can accurately distinguish between matching and mismatching video and text combinations.

[0181] In some embodiments, based on the video features of the sample video and the text features of the sample text, the training feature processing model can be specifically implemented as follows:

[0182] Based on video features and text features, feature processing models, feature aggregation models, and feature calculation models are trained.

[0183] In a possible implementation of the present application, the feature processing model, the feature aggregation model and the feature calculation model can be trained separately. Of course, the feature processing model, the feature aggregation model and the feature calculation model can also be trained jointly. That is, based on the similarity between the video features and the text features in the positive samples and the similarity between the video features and the text features in the negative samples, the feature processing model, the feature aggregation model and the feature calculation model can be trained by comparative learning, so that the feature processing model, the feature aggregation model and the feature calculation model can work together to better adapt to the needs of video and text related tasks, so that in the subsequent processing of real video data and text data, features can be extracted more accurately, similarities can be calculated, and features can be aggregated to realize application scenarios such as video retrieval, video classification, and video question and answer.

[0184] Figure 2 FIG. 1 is a schematic diagram of a video feature extraction process in a practical application of the embodiment of the present application. Figure 2 As shown in step 201, the video frame may be encoded first, and the frame features of the video frame in the video may be extracted using the visual encoder in the feature processing model.

[0185] After the frame features are extracted, the video frames can be sparsely layered sampled based on the frame features to obtain the target video frames. The embodiment of the present application first performs grouping, and samples different video frame groups to achieve layered sampling. Layered sampling can ensure the diversity of sampled video frames, improve sampling accuracy, and ensure the accuracy of the final generated video features. In addition, a certain number of video frames can be sampled for each video frame group according to the above implementation method to achieve sparse sampling. Sparse sampling can reduce the number of samples and improve sampling efficiency.

[0186] Based on the frame features, for example, as shown in step 202, sparse layered sampling of video frames can be implemented offline. However, it is not limited thereto, as shown in step 203, sparse layered sampling of video frames can also be implemented online.

[0187] In one implementation of the present application, offline sparse stratified sampling of video frames can be achieved as follows: based on frame features, the video frames in the video are clustered to obtain multiple video groups, and then for any video group, the video frames whose distance from the cluster center of the video group is greater than a first threshold are filtered out, and then the filtered video frame groups are sampled in ascending order of distance from the cluster center to obtain a first number of target video frames.

[0188] In another implementation of the present application, the sparse layered sampling of video frames can be implemented online as follows: based on frame features, the similarity between two adjacent video frames is calculated using a feature calculation model, and when the similarity between any two adjacent video frames is greater than a second threshold, it is determined that there is a group boundary between the two adjacent video frames, and based on the group boundary, the video frames of the video are divided into multiple video frame groups, and then for any video frame of any video frame group, when the similarity between the video frame and its adjacent previous video frame is less than a third threshold, the video frame is sampled as a target video frame according to probability. The detailed sampling implementation method can be found in the above description and will not be repeated here.

[0189] As shown in step 204, after sampling to obtain a plurality of target video frames, a feature aggregation model may be used to aggregate frame features of the plurality of target video frames to generate video features.

[0190] Further, as shown in step 205, the text to be matched corresponding to the video can be obtained, and the text features of the text to be matched can be extracted using the text encoder in the feature processing model. Then, as shown in step 206, for example, the text features can be matched with the video features by using the feature calculation model to calculate the similarity between the text features and the video features to determine the video features that match the text features.

[0191] In an embodiment of the present application, as in step 207, a feature processing model, a feature calculation model, a feature aggregation model, etc. may be trained by comparative learning.

[0192] Figure 3 A flowchart of a model training method provided in one embodiment of the present application is as follows: Figure 3 As shown, the model training method may include the following steps:

[0193] 301: Obtain training data; the training data includes sample videos and sample texts matching the sample videos.

[0194] 302: Extract frame features of video frames in the sample video using a feature processing model.

[0195] 303: Divide the video frames in the sample video into a plurality of sample frame groups based on the frame features.

[0196] 304: Sampling the multiple sample frame groups respectively to obtain multiple sample frames.

[0197] 305: Aggregate frame features of multiple sample frames to obtain video features of the sample video.

[0198] 306: Extract text features of the sample text using the feature processing model.

[0199] 307: Based on the video features of the sample video and the text features of the sample text, a feature processing model is trained; the feature processing model is used to extract the text features of the target text and the frame features of the video frames in the to-be-matched video; the frame features of the multiple video frames acquired by the to-be-matched video are aggregated to obtain the video features of the to-be-matched video.

[0200] In an embodiment of the present application, a sample video and a sample text matched therewith may be used as training data for a feature processing model, wherein the sample video may include video clips for displaying various video contents, scenes, actions, etc., and the sample text may be a text description semantically associated with the sample video.

[0201] In the embodiment of the present application, the feature processing model can first be used to process the sample video. The feature processing model can perform feature extraction operations on each video frame in the sample video one by one, and convert the rich visual information in the video frame into a frame feature form that can be understood and processed by the computer. These frame features can describe the video frame from multiple dimensions.

[0202] After obtaining the frame features of each video frame in the sample video, the video frames in the sample video can be divided into multiple different sample frame groups according to the similarities and differences between the frame features. Through such grouping, the video frames in the same group can have high similarity in certain key features, while there are obvious differences between different groups, thereby achieving a preliminary classification and induction of the sample video content.

[0203] Then, the divided multiple sample frame groups can be sampled. By sampling the divided multiple sample frame groups, partial video frames can be obtained from each sample frame group. These video frames are multiple sample frames, which can reflect the characteristics of each sample frame group to a certain extent, while avoiding processing too much redundant data.

[0204] In the embodiments of the present application, for example, sampling methods such as uniform sampling, random sampling, key frame sampling, etc. may be used to sample the divided multiple sample frame groups.

[0205] By aggregating the frame features of multiple sample frames obtained through sampling, video features that can represent the characteristics of the entire sample video content can be obtained. The video features can be used as a condensed and representative expression of the sample video.

[0206] In addition to processing sample videos, feature processing models can also be used to extract text features of sample text. Sample text can be converted into a quantized, processable form for matching with video features of sample videos and subsequent training operations.

[0207] After obtaining the video features of the sample video and the text features of the sample text, the feature processing model can be trained using the video features and text features. By training the feature processing model, the feature processing model can accurately extract the text features of the target text and the frame features of the video frames in the video to be matched. During the training process, the feature processing model can adjust its own parameters based on the video features of the input sample video and the text features of the sample text so that the output of the model is as consistent as possible with the expected results. For example, if the video features predicted by the model are inconsistent with the video features of the actual sample video, or if there is a deviation between the extracted text features and the text features of the actual sample text, the model will adjust its internal parameters through algorithms such as back propagation to improve the accuracy and reliability of the model.

[0208] In some embodiments, aggregating frame features of multiple sample frames to obtain video features of a sample video may be specifically implemented as follows:

[0209] The frame features of multiple sample frames are aggregated using a feature aggregation model to obtain the video features of the sample video.

[0210] In an embodiment of the present application, the feature aggregation model may have learning capabilities, which can automatically adjust the aggregation method and weight distribution according to factors such as the content characteristics of different videos, the distribution of target video frames, and specific application requirements, so as to more accurately reflect the overall characteristics of the video.

[0211] In an embodiment of the present application, the feature aggregation model can be trained to generate more accurate video features.

[0212] In some embodiments, based on frame features, dividing the video frames in the sample video to obtain multiple sample frame groups may be specifically implemented as follows:

[0213] The similarity between two adjacent video frames is calculated using a feature calculation model; when the similarity between any two adjacent video frames is greater than a second threshold, it is determined that there is a group boundary between the two adjacent video frames; based on the group boundary, the video frames of the sample video are divided into a plurality of video frame groups.

[0214] In an embodiment of the present application, the feature calculation model may have learning capabilities, may deeply analyze frame features of video frames, and calculate similarities based on the intrinsic relationships between frame features, thereby obtaining results that are more consistent with the actual similarities of video frames.

[0215] The feature calculation model can continuously adjust its parameters through training to adapt to the feature combination and similarity relationship of different types of video frames, and then calculate a more accurate similarity. With the increase of training data and training rounds, the accuracy of the feature calculation model will continue to improve, and it can better cope with similarity calculation tasks under various complex video content and scene changes.

[0216] In another possible implementation of the present application, a clustering algorithm may be used to divide the video frames of the sample video into multiple video frame groups. The specific implementation of using the clustering algorithm to divide the video frames of the sample video into multiple video frame groups can be referred to. Figure 1 The specific implementation methods provided in the relevant embodiments will not be repeated here.

[0217] In some embodiments, based on the video features of the sample video and the text features of the sample text, the training feature processing model can be specifically implemented as follows:

[0218] Any training data is taken as a positive sample, and a sample video in one training data and a sample text in another training data constitute a negative sample; based on the similarity between the video features and the text features in the positive sample and the similarity between the video features and the text features in the negative sample, the feature processing model is trained by comparative learning.

[0219] In an embodiment of the present application, by constructing positive samples and negative samples and adopting a contrastive learning training method to train a feature extraction model, the feature extraction model can be guided to learn a more accurate feature representation through the similarity difference between video features and text features in positive samples and negative samples, so that the feature extraction model can better distinguish between relevant and irrelevant combinations of videos and texts, thereby improving the accuracy of the feature extraction model in extracting target text features and video frames to be matched.

[0220] In an embodiment of the present application, a positive sample may be composed of a set of matching sample videos and sample texts. For example, in the training data, the sample video is a video showing a sunrise at the seaside, and the sample text is "This video shows a beautiful sunrise at the seaside, the sun slowly rises from the sea level, and the sky is dyed orange-red", then the entire training data is regarded as a positive sample. This method of selecting positive samples can be based on the original corresponding relationship between video and text in practice, providing the model with a correct matching example, allowing the model to learn the association pattern between video features and text features in matching situations.

[0221] Unlike positive samples, sample videos and sample texts that originally did not match can be combined to form negative samples. For example, in the training data, the sample video is a video about the night view of the city, and the sample text of another training data is "The trees in this forest are lush and green, and the birds are singing on the branches." Combining the video of the night view of the city and the text describing the forest scene constitutes a negative sample. By constructing a large number of negative samples in this way, the model can clearly recognize which combinations of videos and texts are mismatched, so as to better grasp the matching relationship between video features and text features during the learning process.

[0222] By allowing the feature processing model to learn the differences between positive samples and negative samples, the feature processing model can accurately distinguish between matching and mismatching video and text combinations.

[0223] In some embodiments, based on the video features of the sample video and the text features of the sample text, the training feature processing model can be specifically implemented as follows:

[0224] Based on video features and text features, feature processing models, feature aggregation models, and feature calculation models are trained.

[0225] In a possible implementation of the present application, the feature processing model, feature aggregation model and feature calculation model can be trained separately. Of course, the feature processing model, feature aggregation model and feature calculation model can also be trained jointly, that is, based on the similarity between the video features and the text features in the positive samples and the similarity between the video features and the text features in the negative samples, the feature processing model, feature aggregation model and feature calculation model can be trained by comparative learning. The feature processing model, feature aggregation model and feature calculation model can work together to better meet the needs of video and text related tasks, so that when processing real video data and text data in the future, features can be extracted more accurately, similarities can be calculated and features can be aggregated to achieve application scenarios such as video retrieval, video classification, and video question and answer.

[0226] Figure 4 A flowchart of a video retrieval method provided in one embodiment of the present application is as follows: Figure 4 As shown, the video retrieval method may specifically include the following steps:

[0227] 401: Respond to the video retrieval request and obtain retrieval information.

[0228] 402: Extract text features from the search information.

[0229] 403: Determine at least one target video based on feature similarities between text features and video features of multiple videos to be matched; wherein the video features are obtained by aggregating frame features of multiple target video frames in the video; the multiple target video frames are obtained by sampling from multiple video frame groups corresponding to the video; and the multiple video frame groups are obtained by layering video frames in the video based on the frame features.

[0230] In some embodiments, the method may further include:

[0231] For any video to be matched, frame features of video frames in the video are extracted; based on the frame features, the video frames in the video are divided into multiple video frame groups; the multiple video frame groups are sampled respectively to obtain multiple target video frames; the frame features of the multiple target video frames are aggregated to obtain video features of the video.

[0232] In some embodiments, in response to a video retrieval request, obtaining retrieval information may be specifically implemented as follows:

[0233] In response to a video search request sent by a target user, a search keyword provided by the target user is obtained.

[0234] In some other embodiments, in response to a video search request, obtaining search information may be specifically implemented as follows:

[0235] In response to a video recommendation event for a target user, attribute information of the target user is obtained.

[0236] In some other embodiments, in response to a video search request, obtaining search information may be specifically implemented as follows:

[0237] In response to a video question-and-answer request sent by a target user, question information provided by the target user is obtained.

[0238] In some embodiments, the method may further include: recommending at least one target video to the target user.

[0239] In some embodiments, when the retrieval information is question information provided by a target user, recommending at least one target video to the target user can be specifically implemented by: generating reply information corresponding to the question information based on the at least one target video; and providing the reply information to the target user.

[0240] In an embodiment of the present application, the reply information can be generated using a large model. Among them, a large model refers to a machine learning model with a large number of parameters and complex structure, which can process massive data and complete various complex tasks, such as natural language processing, computer vision, speech recognition, etc. It is an AI (artificial intelligence) model. Among them, the large model can be implemented, for example, using a large language model (Large Language Model, LLM) or a multimodal large model (Multimodal Large Model, MLM). This application does not limit this.

[0241] Among them, the implementation of the relevant technical solutions in the video retrieval method can refer to Figure 2 The video feature extraction method shown is not repeated here.

[0242] Figure 5 A system architecture diagram is shown in which the technical solution of an embodiment of the present application can be applied. The system architecture may include a user end 501 and a server end 502.

[0243] The user terminal 501 and the server terminal 502 are connected via a network. The network provides a medium for a communication link between the user terminal 501 and the server terminal 502. The network may include various connection types, such as wired or wireless communication links or optical fiber cables.

[0244] The client 501 can interact with the server 502 through the network to receive or send messages, etc.

[0245] Among them, the user terminal 501 can be a browser, an APP (Application), or a web application such as an H5 (HyperText Markup Language 5, Hypertext Markup Language Version 5) application, or a light application (also known as a mini-program, a lightweight application) or a cloud application, etc. The user terminal 501 can be deployed in an electronic device and needs to rely on the device to run or some apps in the device to run. For example, the electronic device can have a display screen and support information browsing, such as a personal mobile terminal such as a mobile phone, a tablet computer, a personal computer, etc. For ease of understanding, Figure 5 In the text, the user end is mainly represented by the device image. Various other types of applications can usually be configured in electronic devices, such as human-computer dialogue applications, model training applications, text processing applications, web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.

[0246] The server 502 may include servers that provide various services, such as a server for background training that provides support for the model used on the user terminal 501, or a server that processes interactive information sent by the user terminal.

[0247] It should be noted that the server 502 can be implemented as a distributed server cluster consisting of multiple servers, or as a single server. The server can also be a server of a distributed system, or a server combined with a blockchain. The server can also be a cloud server, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology.

[0248] It should be noted that the video retrieval method and video sampling method provided in the embodiments of the present application are generally executed by the server 502, and the corresponding video retrieval device is generally set in the server 502. However, in other embodiments of the present application, the user end 501 may also have similar functions to the server 502, so as to execute the video feature extraction method, model training method, video retrieval method, and video sampling method provided in the embodiments of the present application. In other embodiments, the video feature extraction method, model training method, video retrieval method, and video sampling method provided in the embodiments of the present application may also be jointly executed by the user end 501 and the server 502,

[0249] It should be understood that Figure 5 The number of the user terminals and the server terminals in the figure is only for illustration. Any number of the user terminals and the server terminals may be provided according to the implementation requirements.

[0250] In actual applications, the server 502 can implement the above Figure 1 The video feature extraction method described in the embodiment shown, Figure 3 The model training method described in the embodiment shown or the above Figure 4 The video retrieval method of the illustrated embodiment, etc.

[0251] Combination Figure 5 The system architecture diagram shown in Figure 6 As shown, a schematic diagram of the video retrieval process implementation of an embodiment of the present application in a practical application is shown.

[0252] For example, the user 601 may use the user terminal 501 to trigger a video search request 602 , whereby the server 502 may obtain the search information 603 carried in the video search request 602 .

[0253] Then, the server 502 may perform feature extraction on the search information 603 to obtain text features 604 .

[0254] The server 502 may store multiple videos to be matched. After obtaining the text feature 604, similarity calculation may be performed based on the text feature 604 and the video features of the multiple videos to be matched, thereby determining at least one target video that matches the text feature 604. In this example, for example, the target video that matches the text feature 604 may be video 605.

[0255] Furthermore, in the case where the video retrieval request 602 is a video search request, the video 605 can be returned to the user terminal 501; in the case where the video retrieval request 602 is a video recommendation event, the video 605 can be returned to the user terminal 501; in the case where the video retrieval request 602 is a video question and answer request, reply information can be generated based on the video 605 and then returned to the user terminal 501.

[0256] In addition to the video sampling involved in the above video retrieval scenarios, video sampling is also involved in practical applications, such as video editing, video compression and other video processing scenarios. Efficient and accurate video sampling can ensure the accuracy of video processing. Therefore, Figure 7 As shown, a flow chart of an embodiment of a video sampling method provided by the present application is shown. Figure 7 As shown, the video sampling method may specifically include the following steps:

[0257] 701: Extract frame features of video frames in the video.

[0258] 702: Divide video frames in the video into multiple video frame groups based on frame features.

[0259] 703: Sample the multiple video frame groups respectively to obtain multiple target video frames.

[0260] The implementation of the relevant technical solutions of steps 701 to 703 can refer to Figure 1 As shown in steps 101 to 103 in the illustrated embodiment, they will not be described in detail here.

[0261] Figure 8 A block diagram of a video feature extraction device provided by an embodiment of the present application is shown. Figure 8 As shown, the video feature extraction device may specifically include:

[0262] A first extraction module 801 is used to extract frame features of video frames in a video;

[0263] A first division module 802, configured to divide the video frames in the video into a plurality of video frame groups based on the frame features;

[0264] A first sampling module 803 is used to sample the multiple video frame groups respectively to obtain multiple target video frames;

[0265] The first aggregation module 804 is used to aggregate the frame features of the multiple target video frames to obtain the video features of the video.

[0266] In some embodiments, the first division module 802 is specifically used to: cluster the video frames in the video based on the frame features to obtain multiple video frame groups.

[0267] In some embodiments, the first sampling module 803 includes:

[0268] A filtering submodule, configured to filter, for any video frame group, video frames whose distance from a cluster center of the video frame group is greater than a first threshold;

[0269] The sampling submodule is used to sample the filtered video frame group to obtain multiple target video frames.

[0270] In some embodiments, the sampling submodule is specifically used for:

[0271] The filtered video frame group is sampled with a first number of target video frames in ascending order of distance from the cluster center.

[0272] In some embodiments, the first division module 802 is specifically used to: determine the similarity between any two adjacent video frames based on the frame features; determine that there is a group boundary between the two adjacent video frames when the similarity between any two adjacent video frames is greater than a second threshold; and divide the video frames of the video into multiple video frame groups based on the group boundary.

[0273] In some embodiments, the first sampling module 803 includes:

[0274] For any video frame of any video frame group, when the similarity between the video frame and its adjacent previous video frame is less than a third threshold, the video frame is sampled as a target video frame according to probability.

[0275] In some embodiments, the device may further include:

[0276] The sequence construction module is used to sample a second number of video frames from the video to form a frame sequence from the second number of video frames.

[0277] In some embodiments, the first division module 802 is specifically used to: calculate the similarity between any two adjacent sampling frames in the frame sequence; and dividing the video into multiple video frame groups based on the group boundary includes: dividing the frame sequence into multiple video frame groups based on the group boundary.

[0278] In some embodiments, the device may further include:

[0279] A text determination module, used to determine the text to be matched corresponding to the video and extract text features of the text to be matched;

[0280] In some embodiments, the first division module 802 is specifically used to: determine any adjacent first video frame and second video frame; fuse the text feature into the frame feature of the first video frame to obtain first feature data; fuse the text feature into the frame feature of the second video frame to obtain second feature data; and use the similarity between the first feature data and the second feature data as the similarity between the first video frame and the second video frame.

[0281] In some embodiments, the first aggregation module 804 is specifically used to: form a sampling sequence from the multiple target video frames; for any target video frame, determine the position information of the target video frame in the sampling sequence; fuse the position information into the frame features of the target video frame; and aggregate the fused frame features corresponding to the multiple target video frames to obtain video features of the video.

[0282] Figure 8 The video feature extraction device can perform Figure 1 The implementation principle and technical effect of the video feature extraction method described in the illustrated embodiment will not be described in detail. The specific manner in which each module and unit performs operations in the video feature extraction device in the above embodiment has been described in detail in the embodiment of the method, and will not be described in detail here.

[0283] Fig. 9 A block diagram of a model training device provided by an embodiment of the present application is shown as follows: Figure 8 As shown, the device may specifically include:

[0284] The data acquisition module 901 is used to acquire training data; the training data includes a sample video and a sample text matching the sample video;

[0285] A second extraction module 902, configured to extract frame features of video frames in the sample video using a feature processing model;

[0286] A second division module 903, configured to divide the video frames in the sample video into a plurality of sample frame groups based on the frame features;

[0287] A second sampling module 904 is used to sample the multiple sample frame groups respectively to obtain multiple sample frames;

[0288] A second aggregation module 905 is used to aggregate the frame features of the multiple sample frames to obtain the video features of the sample video;

[0289] A third extraction module 906, configured to extract text features of the sample text using the feature processing model;

[0290] The training module 907 is used to train the feature processing model based on the video features of the sample video and the text features of the sample text; the feature processing model is used to extract the text features of the target text and the frame features of the video frames in the video to be matched; the frame features of the multiple video frames acquired by the video to be matched are aggregated to obtain the video features of the video to be matched.

[0291] In some embodiments, the second aggregation module 905 is specifically used to:

[0292] The frame features of the multiple sample frames are aggregated using a feature aggregation model to obtain video features of the sample video.

[0293] In some embodiments, the second division module 903 is specifically used to:

[0294] The similarity between two adjacent video frames is calculated using a feature calculation model; when the similarity between any two adjacent video frames is greater than a second threshold, it is determined that there is a group boundary between the two adjacent video frames; based on the group boundary, the video frames of the sample video are divided into a plurality of video frame groups.

[0295] In some embodiments, the training module 907 is specifically used to:

[0296] Based on the video features and the text features, the feature processing model, the feature aggregation model and the feature calculation model are trained.

[0297] Fig. 9 The model training device can perform Figure 3 The implementation principle and technical effect of the model training method described in the illustrated embodiment will not be described in detail. The specific manner in which each module and unit performs operations in the model training device in the above embodiment has been described in detail in the embodiment of the method, and will not be described in detail here.

[0298] Fig.10 A block diagram of a video retrieval device provided by an embodiment of the present application is shown as follows: Fig.10 As shown, the device may specifically include:

[0299] The information acquisition module 1001 is used to obtain the search information in response to the video search request;

[0300] The fourth extraction module 1002 is used to extract text features in the search information;

[0301] The target video determination module 1003 is used to determine at least one target video based on the feature similarity between the text features and the video features of multiple videos to be matched; wherein the video features are obtained by aggregating the frame features of multiple target video frames in the video; the multiple target video frames are obtained by sampling from multiple video frame groups corresponding to the video; and the multiple video frame groups are obtained by layering the video frames in the video based on the frame features.

[0302] In some embodiments, the video retrieval apparatus further comprises:

[0303] A sixth extraction module, for extracting frame features of video frames in any video to be matched;

[0304] a fourth division module, configured to divide the video frames in the video into a plurality of video frame groups based on the frame features;

[0305] A fourth sampling module, used to sample the multiple video frame groups respectively to obtain multiple target video frames;

[0306] The target aggregation module is used to aggregate the frame features of the multiple target video frames to obtain the video features of the video.

[0307] In some embodiments, the information acquisition module 1001 is specifically used to:

[0308] In response to a video search request sent by a target user, a search keyword provided by the target user is obtained.

[0309] In some other embodiments, the information acquisition module 1001 is specifically used for:

[0310] In response to a video recommendation event for a target user, attribute information of the target user is obtained.

[0311] In some other embodiments, the information acquisition module 1001 is specifically used for:

[0312] In response to a video question-and-answer request sent by a target user, obtaining question information provided by the target user;

[0313] In some embodiments, the device further comprises:

[0314] A recommendation module is used to recommend the at least one target video to the target user.

[0315] Fig.10 The video retrieval device can execute Figure 4The implementation principle and technical effect of the video retrieval method described in the embodiment are not described in detail. The specific way in which each module and unit performs operations in the video retrieval device in the above embodiment has been described in detail in the embodiment of the method, and will not be described in detail here.

[0316] Fig.11 A block diagram of a video sampling device provided by an embodiment of the present application is shown as follows: Fig.11 As shown, the device may specifically include:

[0317] A fifth extraction module 1101 is used to extract frame features of video frames in the video;

[0318] A third division module 1102 is used to divide the video frames in the video into a plurality of video frame groups based on the frame features;

[0319] The third sampling module 1103 is used to sample the multiple video frame groups respectively to obtain multiple target video frames.

[0320] Fig.11 The video sampling device can perform Figure 7 The implementation principle and technical effect of the video sampling method described in the embodiment are not described in detail. The specific way in which each module and unit performs operations in the video sampling device in the above embodiment has been described in detail in the embodiment of the method, and will not be described in detail here.

[0321] In a possible design, the video feature extraction device, model training device, video retrieval device, and video sampling device provided in the embodiments of the present application can be implemented as a computing device, such as Fig.12 As shown, the computing device may include a storage component 1201 and a processing component 1202;

[0322] The storage component 1201 stores one or more computer instructions, wherein the one or more computer instructions are called and executed by the processing component 1202 to implement the video feature extraction method, model training method, video retrieval method, and video sampling method provided in the embodiments of the present application.

[0323] Of course, the computing device may also include other components, such as input / output interfaces, communication components, etc. The input / output interface provides an interface between the processing component and the peripheral interface module, which may be an output device, an input device, etc. The communication component is configured to facilitate wired or wireless communication between the computing device and other devices.

[0324] Among them, the computing device can be a physical device or an elastic computing host provided by a cloud computing platform, etc. In this case, the computing device can refer to a cloud server, and the above-mentioned processing components, storage components, etc. can be basic server resources rented or purchased from the cloud computing platform.

[0325] When the computing device is a physical device, it can be implemented as a distributed cluster consisting of multiple servers or terminal devices, or as a single server or a single terminal device.

[0326] The embodiment of the present application also provides a computer-readable storage medium storing a computer program. When the computer program is executed by a computer, the video feature extraction method, model training method, video retrieval method, and video sampling method provided in the embodiment of the present application can be implemented.

[0327] The embodiments of the present application also provide a computer program product, including a computer program. When the computer program is executed by a computer, it can implement the video feature extraction method, model training method, video retrieval method, and video sampling method provided in the embodiments of the present application.

[0328] The processing components in the above corresponding embodiments may include one or more processors to execute computer instructions to complete all or part of the steps in the above method. Of course, the processing components may also be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors or other electronic components to perform the above method.

[0329] The storage component is configured to store various types of data to support operations in the device. The storage component can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.

[0330] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0331] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.

[0332] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0333] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A video feature extraction method, characterized in that: include: Extracting frame features of video frames in the video; Based on the frame features, dividing the video frames in the video into a plurality of video frame groups; Sampling the multiple video frame groups respectively to obtain multiple target video frames; Aggregate the frame features of the multiple target video frames to obtain video features of the video.

2. The method according to claim 1, characterized in that The dividing the video frames in the video into a plurality of video frame groups based on the frame features comprises: Based on the frame features, the video frames in the video are clustered to obtain a plurality of video frame groups.

3. The method according to claim 2, characterized in that The sampling of the plurality of video frame groups respectively to obtain a plurality of target video frames comprises: For any video frame group, filter the video frames whose distances from the cluster center of the video frame group are greater than a first threshold; The filtered video frame group is sampled to obtain a plurality of target video frames.

4. The method according to claim 3, characterized in that The sampling of the filtered video frame group comprises: The filtered video frame group is sampled with a first number of target video frames in ascending order of distance from the cluster center.

5. The method according to claim 1, characterized in that The dividing the video frames in the video into a plurality of video frame groups based on the frame features comprises: Based on the frame features, determining the similarity between any two adjacent video frames; When the similarity between any two adjacent video frames is greater than a second threshold, determining that there is a group boundary between the two adjacent video frames; Based on the group boundaries, video frames of the video are divided into a plurality of video frame groups.

6. The method according to claim 5, characterized in that The sampling of the plurality of video frame groups respectively to obtain a plurality of target video frames comprises: For any video frame of any video frame group, when the similarity between the video frame and its adjacent previous video frame is less than a third threshold, the video frame is sampled as a target video frame according to probability.

7. The method according to claim 5, characterized in that Also includes: Sampling a second number of video frames from the video to form a frame sequence from the second number of video frames; Determining the similarity between any two adjacent video frames includes: Calculating the similarity between any two adjacent sampling frames in the frame sequence; The dividing the video into a plurality of video frame groups based on the group boundary comprises: The frame sequence is divided into a plurality of video frame groups based on the group boundaries.

8. The method according to claim 1, characterized in that Also includes: Determine the text to be matched corresponding to the video, and extract text features of the text to be matched; The determining the similarity between any two adjacent video frames based on the frame features includes: Determine any adjacent first video frame and second video frame; Fusion the text feature into the frame feature of the first video frame to obtain first feature data; Fusion the text feature into the frame feature of the second video frame to obtain second feature data; The similarity between the first feature data and the second feature data is used as the similarity between the first video frame and the second video frame.

9. The method according to claim 7, characterized in that: The aggregating the frame features of the plurality of target video frames to obtain the video features of the video comprises: The plurality of target video frames form a sampling sequence; For any target video frame, determining position information of the target video frame in the sampling sequence; fusing the position information into the frame features of the target video frame; Aggregate the fused frame features corresponding to the multiple target video frames to obtain video features of the video.

10. A model training method, characterized in that: include: Acquire training data; the training data includes a sample video and a sample text matching the sample video; Extracting frame features of video frames in the sample video using a feature processing model; Based on the frame features, dividing the video frames in the sample video into a plurality of sample frame groups; Sampling the multiple sample frame groups respectively to obtain multiple sample frames; Aggregating the frame features of the multiple sample frames to obtain video features of the sample video; Extracting text features of the sample text using the feature processing model; Training the feature processing model based on the video features of the sample video and the text features of the sample text; The feature processing model is used to extract text features of the target text and frame features of video frames in the video to be matched; the frame features of multiple video frames acquired by the video to be matched are aggregated to obtain the video features of the video to be matched.

11. The method according to claim 10, characterized in that The aggregating the frame features of the plurality of sample frames to obtain the video features of the sample video comprises: The frame features of the multiple sample frames are aggregated using a feature aggregation model to obtain video features of the sample video.

12. The method according to claim 11, characterized in that The dividing the video frames in the sample video to obtain a plurality of sample frame groups based on the frame features comprises: The similarity between two adjacent video frames is calculated using a feature calculation model; When the similarity between any two adjacent video frames is greater than a second threshold, determining that there is a group boundary between the two adjacent video frames; Based on the group boundaries, the video frames of the sample video are divided into a plurality of video frame groups.

13. The method according to claim 12, characterized in that The training of the feature processing model based on the video features of the sample video and the text features of the sample text includes: Based on the video features and the text features, the feature processing model, the feature aggregation model and the feature calculation model are trained.

14. A video retrieval method, characterized in that: include: Responding to the video retrieval request, obtaining retrieval information; Extracting text features from the search information; At least one target video is determined based on feature similarities between text features and video features of multiple videos to be matched; wherein the video features are obtained by aggregating frame features of multiple target video frames in the video; the multiple target video frames are obtained by sampling from multiple video frame groups corresponding to the video; the multiple video frame groups are obtained by hierarchically processing video frames in the video based on frame features.

15. The method according to claim 14, characterized in that Also includes: For any video to be matched, extract frame features of video frames in the video; Based on the frame features, dividing the video frames in the video into a plurality of video frame groups; Sampling the multiple video frame groups respectively to obtain multiple target video frames; Aggregate the frame features of the multiple target video frames to obtain video features of the video.

16. The method according to claim 14, characterized in that In response to the video retrieval request, obtaining retrieval information includes: In response to a video search request sent by a target user, obtaining a search keyword provided by the target user; In response to a video recommendation event for a target user, acquiring attribute information of the target user; or, In response to a video question-and-answer request sent by a target user, obtaining question information provided by the target user; The method further comprises: The at least one target video is recommended to the target user.

17. A video sampling method, characterized in that: include: Extracting frame features of video frames in the video; Based on the frame features, dividing the video frames in the video into a plurality of video frame groups; The multiple video frame groups are sampled respectively to obtain multiple target video frames.

18. A computing device, characterized in that including a processing component and a storage component; The storage component stores one or more computer instructions; the one or more computer instructions are used to be called and executed by the processing component to implement the video feature extraction method as described in any one of claims 1 to 9, or to implement the model training method as described in any one of claims 10 to 13, or to implement the video retrieval method as described in any one of claims 14 to 16, or to implement the video sampling method as described in claim 17.

19. A computer storage medium, characterized in that: A computer program is stored, and when the computer program is executed by a computer, it implements the video feature extraction method as described in any one of claims 1 to 9, or implements the model training method as described in any one of claims 10 to 13, or implements the video retrieval method as described in any one of claims 14 to 16, or implements the video sampling method as described in claim 17.

20. A computer program product, characterized in that The computer program product includes a computer program code, and when the computer program code is executed by a computer, it implements the video feature extraction method as described in any one of claims 1 to 9, or implements the model training method as described in any one of claims 10 to 13, or implements the video retrieval method as described in any one of claims 14 to 16, or implements the video sampling method as described in claim 17.