Model training method, video processing method, computer device and medium

By combining visual consistency and thematic consistency in feature extraction model training and adopting self-supervised comparative learning, the problems of low efficiency and poor accuracy of feature extraction model training in existing technologies are solved, and efficient and accurate video feature extraction is achieved.

CN114677623BActive Publication Date: 2025-10-10ALIBABA DAMO (HANGZHOU) TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210249167.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-14
Publication Date
2025-10-10
Estimated Expiration
2042-03-14

AI Technical Summary

Technical Problem

Existing feature extraction models have difficulty accurately extracting feature vectors of videos during training, especially for long videos. In addition, existing training methods require a lot of manual labeling costs and are time-consuming.

Method used

During the model training process, multiple video clips with visual consistency and multiple video clips with thematic consistency are prepared. A self-supervised comparative learning method is adopted, focusing on visual consistency and thematic consistency, and learning video representation from multiple levels to reduce manual labeling costs and improve training efficiency.

Benefits of technology

The trained feature extraction model can accurately extract the feature vectors of the video, has good generalization performance, is suitable for short and long videos, and has high model training efficiency and does not require a large amount of manual labeling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114677623B_ABST
    Figure CN114677623B_ABST
Patent Text Reader

Abstract

The embodiment of the application provides a model training method, a video processing method, a computer device and a medium. In the embodiment of the application, a plurality of video clips with visual consistency are prepared, and a plurality of video clips with theme consistency are prepared. In the model training process of the feature extraction model, a self-supervised contrast learning mode is used to pay attention to whether different video clips have the same visual feature dimension of visual consistency and whether different video clips have the same theme dimension of theme consistency, so that video representation is learned based on the self-supervised contrast learning mode from multiple levels. The model performance of the trained feature extraction model is good, the feature vector of the video can be accurately extracted, the generalization performance is good, and the feature extraction can be applied to short videos and long videos. In addition, the model training process does not need to introduce a large amount of artificial labeling cost, the time consumption is less, and the model training efficiency is high.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a model training method, a video processing method, a computer device, and a medium. Background Art

[0002] Video is multimedia data composed of several image frames arranged in chronological order. It has an additional dimension of temporal information compared to image frames. Because video carries richer information, it has become a common form of information carrier, and video-based recognition processing has also become more and more common. The process of video-based recognition processing is roughly to use a feature extraction model to extract features from the video and obtain feature vectors that represent the video; then, based on the feature vectors of the video, recognition processing such as behavior recognition, behavior location, and anomaly analysis is performed. Therefore, it can be seen that the model performance of the feature extraction model is directly related to whether the extracted video feature vectors can accurately represent the video, which directly affects the recognition effect of the video-based recognition processing method.

[0003] In related technologies, self-supervised contrastive learning methods are commonly used to train feature extraction models for video feature extraction. This training method primarily involves randomly sampling multiple different video clips from the video during each training round, with the goal of training the feature extraction model to extract the same feature vectors from these multiple different video clips. However, feature extraction models trained with this method struggle to accurately extract the feature vectors of the video, which in turn affects the recognition performance of video-based recognition processing methods. Summary of the Invention

[0004] Various aspects of the present application provide a model training method, a video processing method, a computer device, and a medium for improving the model performance of a feature extraction model.

[0005] An embodiment of the present application provides a model training method, including: obtaining a plurality of sample videos, each corresponding to a group of video clips, each group of video clips including two video clips with visual consistency, and other video clips with no limitation on visual consistency, and the video clips in the same group of video clips have thematic consistency; using the plurality of groups of video clips to perform a current round of model training on a feature extraction model to obtain a plurality of groups of feature vectors generated by the feature extraction model in the current round of model training; generating a first loss value corresponding to the visual consistency dimension of the current round of model training based on the plurality of groups of feature vectors, and generating a second loss value corresponding to the thematic consistency dimension of the current round of model training based on the thematic consistency prediction model and the plurality of groups of feature vectors; when it is determined that the model training end condition is not met based on the first loss value and the second loss value, continuing the next round of model training on the feature extraction model.

[0006] An embodiment of the present application also provides a video processing method, including: obtaining a target video to be identified; using a feature extraction model to extract features from the target video to obtain a feature vector of the target video; performing identification processing on the target video based on the feature vector of the target video to obtain an identification result; wherein the feature extraction model is a model trained according to the model training method provided in an embodiment of the present application.

[0007] An embodiment of the present application also provides a computer device, comprising: a memory and a processor; the memory is used to store a computer program; the processor is coupled to the memory, and is used to execute the computer program to execute a model training method or a video processing method.

[0008] An embodiment of the present application also provides a computer storage medium storing a computer program. When the computer program is executed by a processor, the processor is enabled to implement a model training method or a video processing method.

[0009] In an embodiment of the present application, multiple video clips with visual consistency and multiple video clips with thematic consistency are prepared. In this way, during the model training process of the feature extraction model, a self-supervised comparative learning method is adopted, which focuses on both the visual consistency dimension of whether different video clips have the same visual features and the thematic consistency dimension of whether different video clips have the same themes, thereby achieving learning of video representations based on a self-supervised comparative learning method from multiple levels. The feature extraction model trained in this way has good model performance, can accurately extract the feature vectors of the video, has good generalization performance, and can be applied to feature extraction of short videos and long videos. In addition, the model training process does not require the introduction of a large amount of manual labeling costs, is less time-consuming, and has high model training efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0011] Figure 1 A flowchart of a model training method provided in an embodiment of the present application;

[0012] Figure 2 An application scenario diagram applicable to a model training method provided in an embodiment of the present application;

[0013] Figure 3 A flowchart of a video processing method provided in an embodiment of the present application;

[0014] Figure 4 A schematic diagram of the structure of a model training device provided in an embodiment of the present application;

[0015] Figure 5 A schematic diagram of the structure of a video processing device provided in an embodiment of the present application;

[0016] Figure 6 A schematic diagram of the structure of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0017] To make the purpose, technical solutions, and advantages of this application more clear, the technical solutions of this application will be clearly and completely described below in conjunction with the specific embodiments of this application and the corresponding drawings. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0018] First, the terms involved in the embodiments of this application are explained:

[0019] Visual consistency: For multiple video clips sampled from long or short videos with a close temporal distance (i.e., time interval), there is a high probability that the multiple video clips have the same visual information, for example, they are all describing the same thing. Therefore, there is visual consistency between these multiple video clips.

[0020] Thematic consistency: For multiple video clips sampled from a long video or a short video, since the multiple video clips are from the same long video or short video, the multiple video clips are likely to have the same theme that reflects the main tone of the video content.

[0021] It is worth noting that if there is visual consistency between multiple video clips, then there is also thematic consistency between the multiple video clips. However, if there is thematic consistency between multiple video clips, there is not necessarily visual consistency between the multiple video clips.

[0022] Self-supervised contrastive learning: This refers to the process of completing self-supervised learning through contrastive learning. In self-supervised contrastive learning, no labels are used. To train the network, each image or video is subjected to multiple independent random data augmentations, generating multiple visually distinct but semantically unchanged samples. The network extracts features from these samples, requiring that the feature distances between samples from the same image or video are close (i.e., the feature vectors have high similarity), while the feature distances between samples from different images or videos are far (i.e., the feature vectors have low similarity).

[0023] Whether it is a long video or a short video, the feature extraction model trained by the existing model training method has difficulty in accurately extracting the feature vectors of the video, and thus it is difficult to guarantee the recognition effect of the video-based recognition processing method. In particular, for the feature extraction of long videos, the feature extraction model trained by the existing model training method has even more difficulty in accurately extracting the feature vectors of long videos. This is because long videos are long in duration, and the visual information of multiple video clips extracted from long videos is not consistent. Therefore, it is not feasible to use the feature extraction model to be able to extract the same feature vectors for these multiple different video clips as the training goal. In addition, the existing model training method requires manual labeling of the video, which will introduce a large amount of manual labeling costs, is time-consuming, and has low model training efficiency.

[0024] In response to the above technical problems, the embodiments of the present application provide a model training method, a video processing method, a computer device and a medium. In the embodiments of the present application, a plurality of video clips with visual consistency and a plurality of video clips with thematic consistency are prepared. In this way, in the model training process of the feature extraction model, a self-supervised comparative learning method is adopted, which not only focuses on the visual consistency dimension of whether different video clips have the same visual features, but also focuses on the thematic consistency dimension of whether different video clips have the same themes. This realizes learning video representation from multiple levels based on the self-supervised comparative learning method. The feature extraction model trained in this way has good model performance, can accurately extract the feature vectors of the video, has good generalization performance, and can be applied to the feature extraction of short videos and long videos. In addition, the model training process does not require the introduction of a large amount of manual labeling costs, consumes less time, and has high model training efficiency.

[0025] The following describes in detail the technical solutions provided by various embodiments of the present application in conjunction with the accompanying drawings.

[0026] Figure 1 This is a flow chart of a model training method provided in an embodiment of the present application. Figure 1 , the method may include the following steps:

[0027] 101. Obtain a plurality of sample videos, each corresponding to a group of video clips. Each group of video clips includes two video clips with visual consistency and other video clips without limitation on visual consistency. The video clips in the same group of video clips have theme consistency.

[0028] 102. Use multiple groups of video clips to perform current round model training on the feature extraction model to obtain multiple groups of feature vectors generated by the feature extraction model in this round model training.

[0029] 103. Based on the multiple sets of feature vectors, a first loss value corresponding to the visual consistency dimension of this round of model training is generated, and based on the topic consistency prediction model and the multiple sets of feature vectors, a second loss value corresponding to the topic consistency dimension of this round of model training is generated.

[0030] 104. When it is determined according to the first loss value and the second loss value that the model training end condition is not met, continue to perform the next round of model training on the feature extraction model.

[0031] In this embodiment, a plurality of different sample videos are first prepared. The plurality of sample videos need to include at least a plurality of long videos with a longer duration, and may also include short videos with a shorter duration. Of course, the duration conditions for dividing long videos and short videos can be flexibly set according to actual application requirements. For example, the plurality of sample videos include the following videos: a long video with a duration of 5 minutes, a long video with a duration of 3 minutes, and a short video with a duration of 30 seconds. It is worth noting that, due to the long duration of the long video, it is usually very likely that the video content of multiple actions or multiple events will appear in the long video. Due to the short duration of the short video, it is usually often only the video content of one action or event.

[0032] In this embodiment, to ensure that model training can balance visual consistency and thematic consistency, for each round of model training, two video clips with visual consistency, as well as other video clips with no restrictions on visual consistency, are sampled from each sample video to form a set of video clips for the sample video. It is worth noting that the two video clips with visual consistency have the same visual information, while the other video clips with no restrictions on visual consistency can have the same visual information as the two video clips with visual consistency, or they can have different visual information from the two video clips with visual consistency, without any restrictions on this.

[0033] Furthermore, when sampling a group of video clips corresponding to each sample video, two visually consistent video clips can be sampled from the sample video at a set sampling interval, and other video clips can be randomly sampled from the sample video. The sampling interval is set based on actual needs and ensures that the two sampled video clips have the same visual information. Unlike the sampling rule for the two visually consistent video clips, the other video clips are randomly sampled from the corresponding sample video, thus exhibiting a random nature.

[0034] In practical applications, the more times the model is trained, the better the performance of the trained model. In this embodiment, the feature extraction model needs to undergo multiple model trainings, and each round of model training can use the same sampling time interval or different sampling time intervals. Further optionally, a progressive sampling strategy can be adopted to gradually increase the time distance between the two video clips during the training process to control the training difficulty of the video clips from easy to difficult, thereby further improving the generalization performance of the feature extraction model. It is worth noting that when the time distance between the two video clips is close, the two video clips are likely to share more visual features, and the training difficulty in the visual consistency dimension and the subject consistency dimension will be relatively small; when the time distance between the two video clips is far, the two video clips share fewer common visual features, and the training difficulty in the visual consistency dimension and the subject consistency dimension will be relatively large.

[0035] Based on the above, further optionally, when determining the sampling time interval used in each round of model training, the sampling time interval used in the current round of model training can be determined according to the training progress and maximum time interval of the current round of model training.

[0036] Further, optionally, when determining the sampling time interval for the current round of model training based on the training progress and maximum time interval of the current round of model training, the sampling time interval for the current round of model training can be continuously increased as the number of model training times increases, thereby increasing the training difficulty. In practical applications, the maximum sampling time interval for the current round of model training can be determined based on the training progress and maximum time interval of the current round of model training, and the sampling time interval for the current round of model training can be set to be less than or equal to the maximum sampling time interval.

[0037] As an example, the maximum sampling time interval used in this round of model training can be calculated according to formula (1).

[0038]

[0039] Among them, δ max Indicates the maximum sampling time interval used in this round of model training. This value is generally small, much smaller than the total length of the video. max For example, 1 second; max Indicates the total number of model training rounds; α indicates the training progress of this round of model training, that is, which round of model training this round of model training is; Δ indicates the maximum time interval set during the model training process.

[0040] In this embodiment, after obtaining multiple groups of video clips from multiple sample videos, the feature extraction model is trained using the multiple groups of video clips. Optionally, the model structure of the feature extraction model may include, but is not limited to, convolutional neural networks (CNN), recurrent neural networks (RNN), and long short-term memory networks (LSTM). Specifically, in each round of model training, the feature extraction model is used to extract a set of feature vectors corresponding to each of the multiple groups of video clips, and visual consistency learning and thematic consistency learning are performed using the set of feature vectors corresponding to each group of video clips.

[0041] In this embodiment, when using a feature extraction model to extract a set of feature vectors corresponding to each group of video clips, each video clip in each group of video clips can be directly input into the feature extraction model for feature extraction, thereby obtaining a feature vector for each video clip as a set of feature vectors. Furthermore, optionally, to improve model performance, each video clip in each group of video clips can be enhanced; each enhanced video clip can be input into the feature extraction model for feature extraction, thereby obtaining a feature vector for each video clip as a set of feature vectors. Specifically, an image enhancement algorithm can be used to enhance the color, contrast, or clarity of each video clip, but this is not limited to this. Of course, depending on actual application needs, image enhancement can also be performed only on other randomly sampled video clips, or only on two video clips with visual consistency. This embodiment does not impose any restrictions on this.

[0042] In this embodiment, when visual consistency learning is performed using a set of feature vectors corresponding to each set of video clips, after extracting multiple sets of feature vectors corresponding to the multiple sets of video clips based on the feature extraction model, a first loss value corresponding to the visual consistency dimension of this round of model training can be generated based on the multiple sets of feature vectors. Furthermore, optionally, before determining the first loss value based on the multiple sets of feature vectors, each feature vector in the multiple sets of feature vectors can be input into a feature space transformation network to reduce the dimensionality of each feature vector. The feature space transformation network can be a fully connected layer for reducing the dimensionality of the feature vectors.

[0043] In this embodiment, the first loss value is calculated based on a loss function associated with the visual consistency dimension. The first loss value can reflect the degree of difference between the feature vector output by the feature extraction model and the actual visual information of the sample video. In practical applications, the loss function associated with the visual consistency dimension can be flexibly set, and this embodiment does not impose any restrictions on this.

[0044] In this embodiment, when using a set of feature vectors corresponding to each group of video clips for topic consistency learning, a second loss value corresponding to the topic consistency dimension of this round of model training can be generated based on the topic consistency prediction model and multiple sets of feature vectors. The second loss value can reflect the degree of difference between the situation of the predicted video clips with consistent topics and the situation of the actual video clips with consistent topics. The topic consistency prediction model can be a multi-layer perceptron, which can predict the topic consistency of two video clips. The topic consistency prediction results include but are not limited to: whether the two video clips have the same topic, or the probability that the two video clips have the same topic, etc. Optionally, the model structure of the topic consistency prediction model can include but is not limited to: Convolutional Neural Networks (CNN), Recurrent Neural Networks (RNN) and Long Short-Term Memory (LSTM).

[0045] In each round of model training, numerical calculations such as averaging, accumulation, or weighted summation are performed on the first loss value corresponding to the visual consistency dimension and the second loss value corresponding to the subject consistency dimension of this round of model training to obtain the total loss value of this round of model training.

[0046] In some application scenarios, the model training termination condition only considers the total loss value of the current round of model training. In this case, if the total loss value of the current round of model training is greater than the preset loss value, the model training termination condition is determined to be unsatisfied, and the model parameters of the feature extraction model and the topic consistency prediction model are adjusted. Steps 101 to 104 are then re-executed to continue the next round of model training for the feature extraction model. If the total loss value of the current round of model training is less than or equal to the preset loss value, the model training termination condition is determined to be met, and the trained feature extraction model can now be used for model inference.

[0047] In other application scenarios, the model training termination condition considers not only the total loss value of the current round of model training, but also the current number of model training cycles. If the current number of model training cycles is less than the preset maximum number of model training cycles, regardless of whether the total loss value of the current round of model training is greater than or equal to the preset loss value, the model training termination condition is determined to be unsatisfied, and the model parameters of the feature extraction model and the topic consistency prediction model are adjusted. Steps 101 to 104 are re-executed to continue the next round of model training for the feature extraction model. If the current number of model training cycles is greater than or equal to the preset maximum number of model training cycles, if the total loss value of the current round of model training is greater than the preset loss value, the model training termination condition is determined to be unsatisfied, and the model parameters of the feature extraction model and the topic consistency prediction model are adjusted. Steps 101 to 104 are re-executed to continue the next round of model training for the feature extraction model. If the total loss value of the current round of model training is less than or equal to the preset loss value, the model training termination condition is determined to be satisfied, and the trained feature extraction model can now be used for model inference. It is worth noting that the preset loss value and the preset maximum number of model training cycles can be flexibly set according to actual application requirements.

[0048] The model training method provided in the embodiment of the present application prepares multiple video clips with visual consistency and multiple video clips with thematic consistency. In this way, in the model training process of the feature extraction model, a self-supervised comparative learning method is adopted, which focuses on both the visual consistency dimension of whether different video clips have the same visual features and the thematic consistency dimension of whether different video clips have the same themes, thereby realizing learning video representation from multiple levels based on the self-supervised comparative learning method. The feature extraction model trained in this way has good model performance, can accurately extract the feature vectors of the video, has good generalization performance, and can be applied to feature extraction of short videos and long videos. In addition, the model training process does not require the introduction of a large amount of manual labeling costs, is less time-consuming, and has high model training efficiency.

[0049] In some embodiments, when generating the first loss value corresponding to the visual consistency dimension of this round of model training based on multiple sets of feature vectors, the loss value corresponding to each set of feature vectors is first calculated. For ease of understanding and distinction, the loss value corresponding to each set of feature vectors is referred to as the third loss value; then, based on the multiple third loss values ​​corresponding to the multiple sets of feature vectors, the first loss value corresponding to the visual consistency dimension of this round of model training is obtained. Further optionally, when obtaining the first loss value corresponding to the visual consistency dimension of this round of model training based on the multiple third loss values ​​corresponding to the multiple sets of feature vectors, numerical calculations such as averaging, weighted summing, or accumulation can be performed on the multiple third loss values ​​corresponding to the multiple sets of feature vectors to obtain the first loss value corresponding to the visual consistency dimension of this round of model training, but the present invention is not limited thereto.

[0050] Further optionally, when determining the third loss value corresponding to each group of feature vectors, the third loss value corresponding to the group of feature vectors can be obtained based on the similarity between the first feature vector and the second feature vector corresponding to the two video clips with visual consistency in the group of feature vectors, and the similarity between the first feature vector and the second feature vector and other feature vectors different from themselves in each group of feature vectors.

[0051] In an optional implementation, when determining the third loss value corresponding to each group of eigenvectors, the fourth loss value can be calculated based on the similarity between the first eigenvector and the second eigenvector, and the similarity between the first eigenvector and other eigenvectors in each group of eigenvectors except the first eigenvector; the fifth loss value can be calculated based on the similarity between the first eigenvector and the second eigenvector, and the similarity between the second eigenvector and other eigenvectors in each group of eigenvectors except the second eigenvector; and the third loss value corresponding to the group of eigenvectors can be determined based on the fourth loss value and the fifth loss value.

[0052] In this embodiment, a loss function for calculating the fourth loss value and the fifth loss value can be designed based on actual application requirements. Assume that there are N sample videos, and each sample video corresponds to a set of video clips including two video clips with the same theme and other randomly sampled video clips. Thus, the N sample videos correspond to 3N video clips, and the 3N video clips correspond to 3N feature vectors, where N is a positive integer.

[0053] As an example, the formula for calculating the loss function of the fourth loss value and the fifth loss value is as follows:

[0054]

[0055] As another example, the formula for calculating the loss function of the fourth loss value and the fifth loss value is as follows:

[0056]

[0057] As another example, the formula for calculating the loss function of the fourth loss value and the fifth loss value is as follows:

[0058]

[0059] As another example, the formula for calculating the loss function of the fourth loss value and the fifth loss value is as follows:

[0060]

[0061] It is worth noting that s i,jrepresents the similarity between the first feature vector and the second feature vector, which may be, but is not limited to, the cosine similarity or the Euclidean distance between the first feature vector and the second feature vector. is 1 when k≠i, otherwise, is 0. τ is a constant coefficient. exp() represents an exponential operation, and log() represents a logarithmic operation.

[0062] For the case that the 2N feature vectors not including the randomly sampled feature vector are selected to participate in the calculation of the fourth loss value or the fifth loss value: when i represents the first feature vector, k represents the other feature vectors in the 2N feature vectors except the first feature vector, s i,k represents the similarity between the first feature vector and the other feature vectors; when i represents the second feature vector, k represents the other feature vectors in the 2N feature vectors except the second feature vector, s i,k represents the similarity between the second feature vector and the other feature vectors.

[0063] For the case that the 3N feature vectors participate in the calculation of the fourth loss value or the fifth loss value: when i represents the first feature vector, k represents the other feature vectors in the 3N feature vectors except the first feature vector, s i,k represents the similarity between the first feature vector and the other feature vectors; when i represents the second feature vector, k represents the other feature vectors in the 3N feature vectors except the second feature vector, s i,k represents the similarity between the second feature vector and the other feature vectors.

[0064] For the convenience of understanding, the calculation of the first loss value corresponding to the visual consistency dimension in the model training is described by taking the loss function for calculating the fourth loss value and the fifth loss value as formula (5) as an example.

[0065] Suppose N sample videos are prepared, and N groups of video clips are sampled from the N sample videos, each group of video clips including two video clips with video consistency and one randomly sampled video clip. As an example, the loss function shown in formula (6) can be used to calculate the first loss value corresponding to the visual consistency dimension in the model training

[0066]

[0067] In formula (6), the 2n-1th eigenvector corresponds to the first eigenvector in each group of eigenvectors, the 2nth eigenvector corresponds to the second eigenvector in each group of eigenvectors, and l(2n-1,2n)+l(2n,2n-1) represents the third loss value corresponding to each group of eigenvectors. l(2n-1,2n) represents the fourth loss value corresponding to each group of eigenvectors, and l(2n,2n-1) represents the fifth loss value corresponding to each group of eigenvectors. When using formula (5) to calculate the fourth loss value, i in formula (5) takes the value of 2n-1, and j in formula (5) takes the value of 2n. When using formula (5) to calculate the fifth loss value, i in formula (5) takes the value of 2n, and j in formula (5) takes the value of 2n-1.

[0068] In some embodiments, when generating the second loss value corresponding to the topic consistency dimension of this round of model training based on the topic consistency prediction model and multiple sets of feature vectors, any two feature vectors in the multiple sets of feature vectors can be vector spliced ​​to obtain multiple splicing vectors; each splicing vector is input into the topic consistency prediction model to perform topic consistency prediction on the two video clips corresponding to each splicing vector; based on the topic consistency prediction results and classification labels corresponding to each of the multiple splicing vectors, the second loss value corresponding to the topic consistency dimension of this round of model training is determined. The classification label may include a first classification label and a second classification label. The first classification label indicates that the topics of the two video clips corresponding to the splicing vector are consistent; the second classification label indicates that the topics of the two video clips corresponding to the splicing vector are inconsistent.

[0069] Optionally, before each concatenated vector is input into the topic consistency prediction model, each concatenated vector may be input into a feature space transformation network to reduce the feature dimension of the concatenated vector. The feature space transformation network may be a fully connected layer for reducing the feature dimension of the feature vector.

[0070] In practical applications, when determining the second loss value corresponding to the topic consistency dimension of this round of model training, the loss value corresponding to each splicing vector can be first determined based on the topic consistency prediction result and classification label corresponding to each splicing vector. For ease of understanding and distinction, the loss value corresponding to each splicing vector is referred to as the sixth loss value. Then, based on the multiple sixth loss values ​​corresponding to the multiple splicing vectors, the second loss value corresponding to the topic consistency dimension of this round of model training is determined. For example, numerical calculations such as averaging, weighted summing, or accumulation can be performed on the multiple sixth loss values ​​corresponding to the multiple splicing vectors to determine the second loss value corresponding to the topic consistency dimension of this round of model training. Further optionally, the multiple sixth loss values ​​with the first classification label are averaged to obtain a first average value, and the multiple sixth loss values ​​with the second classification label are averaged to obtain a second average value; based on the first average value and the second average value, the second loss value corresponding to the topic consistency dimension of this round of model training is determined.

[0071] In an optional implementation, for each splicing vector, if it has a first classification label, and the first classification label indicates that the themes of the two video clips corresponding to the splicing vector are consistent, then the theme consistency prediction result of the splicing vector is used as the input parameter of the target loss function, and the sixth loss value corresponding to the splicing vector is calculated; if it has a second classification label, and the second classification label indicates that the themes of the two video clips corresponding to the splicing vector are inconsistent, then the difference between the theme consistency prediction result of the splicing vector and a specified value is used as the input parameter of the target loss function, and the sixth loss value corresponding to the splicing vector is calculated. Optionally, the target loss function may include, but is not limited to: log loss function, cross-entropy loss function, and focal loss function for solving data imbalance problem.

[0072] Assume that the three video clips sampled from each sample video in N sample videos are v i 、v j 、v k According to formula (7), the feature vector t can be obtained by processing the feature extraction model f and the feature space conversion network h in sequence. i , t j , t k .

[0073] {t i ,t j ,t k}=h(f({v i ,v j ,v k}))……(7)

[0074] In formula (7), f({v i ,v j ,v k} means to convert three video clips v i 、v j 、v k Input them into the feature extraction model f for feature extraction, and get three video clips v i 、v j 、v k The corresponding eigenvectors, h(f({v i ,v j ,v k}) means to convert three video clips v i 、v j 、v k The corresponding eigenvectors f({v i ,v j ,v k}Input into the feature space transformation network h for dimensionality reduction.

[0075] After 3N video clips are processed by the feature extraction model f and the feature space conversion network h, 3N feature vectors can be obtained. By concatenating any two feature vectors from the 3N feature vectors, a 3N×3N concatenated vector can be obtained. The feature set including the 3N×3N concatenated vectors is denoted as U. The feature set U can be expressed as:

[0076]

[0077] It is worth noting that in formula (8), the eigenvector The superscripts 1…N of the eigenvectors represent the numbers of the N sample videos; Represents vector concatenation operation, C T Represents the feature dimension, and the feature dimensions are t i , t j , t k . R represents the set of real numbers, and ∈ represents belonging to a symbol.

[0078] In this real number example, U is input into the topic consistency prediction model, which predicts whether the two video clips corresponding to each splicing vector have the same topic and outputs the topic consistency prediction result M. Where M∈R 3N×3N In order to perform supervised training, define the classification label G for supervision, where G∈R 3N×3N, each classification label in G indicates whether the corresponding two video clips have the same theme. If the two video clips have the same theme, the classification label value is 1, and the classification label with a value of 1 can be called the first classification label; if the two video clips have different themes, the classification label value is 0, and the classification label with a value of 0 can be called the second classification label.

[0079] In order to make the network pay more attention to visual differences and better predict difficult samples in the dimension of topic consistency, the Focal loss function can be used to supervise the topic consistency prediction results.

[0080] As an example, the loss function shown in formula (9) can be used to calculate the second loss value

[0081]

[0082] Among them, γ1 is the first classification label in G (G i,j =1), γ2 is the number of the second classification label in G (G i,j =0). M i,j represents the topic consistency prediction result of any one of the 3N×3N splicing vectors, Indicates M i,j As input parameters and input to the Focal loss function Calculate and obtain the loss value; As input parameters and input to the Focal loss function Calculate and get the loss value.

[0083] For ease of understanding, combined Figure 2 The application scenario diagram shown in FIG. 1 illustrates the model training method provided in the embodiment of the present application. Figure 2 As shown in Figure 1, the entire model architecture includes a feature extraction model f, a topic consistency prediction model Φ, a feature space transformation network g related to visual consistency, and a feature space transformation network h related to topic consistency. The feature vector output by the feature extraction model f undergoes dimensionality reduction processing in the feature space transformation network g before learning visual consistency. The feature vector output by the feature extraction model f undergoes dimensionality reduction processing in the feature space transformation network h before being input into the topic consistency prediction model for topic consistency learning.

[0084] Specifically, assuming that N sample videos are prepared, in each round of model training, according to the training progress α of this round of model training, the maximum time interval Δ, the total number of model training rounds α max The maximum sampling time interval δ used in this round of model training can be calculated max; Then, for each sample video, max The sampling time interval is denoted as v i and v j ; and randomly sample a video clip from the sample video, the randomly sampled video clip is recorded as v k .

[0085] Since there are N sample videos, in order to facilitate understanding and distinction, the N sample videos are numbered 1, 2, 3...N, and the three video clips v corresponding to each sample video are i 、v j 、v k A superscript is added to indicate the code of the sample video to which it is associated. For example, Respectively represent the three video clips of the first sample video; Represent the three video clips of the second sample video respectively; and so on. Respectively represent the three video clips of the nth sample video, where n is any positive integer from 1 to N.

[0086] After sampling N groups of video clips from N sample videos, each group of video clips includes v i and v j , and randomly sampled v k . Each video clip in the N groups of video clips is input into the feature extraction model f for feature extraction, and N groups of feature vectors can be obtained. The feature space conversion network g is used to reduce the dimension of each feature vector in the N groups of feature vectors, and the N groups of feature vectors after the dimension reduction are used for visual consistency learning. The feature space conversion network h is used to reduce the dimension of each feature vector in the N groups of feature vectors, and the N groups of feature vectors after the dimension reduction are input into the topic consistency prediction model for topic consistency learning.

[0087] exist Figure 2 In the example, two video clips from the first sample video v j 1 Marked with gray circles, random video clips from the first sample video Marked with white circles, the video clips from the nth sample video Marked with a grey diamond.

[0088] from Figure 2From the topic consistency prediction result M, it can be seen that for the splicing vectors marked with two gray circles, the corresponding topic consistency prediction result value is 1, and the value 1 indicates that the two video clips corresponding to the splicing vector are derived from the same sample video (both from the first sample video).

[0089] For the splicing vectors marked with a gray circle and a white circle, the corresponding topic consistency prediction result has a value of 1, which means that the two video clips corresponding to the splicing vectors come from the same sample video (both come from the first sample video).

[0090] For a splicing vector marked with a gray circle and a gray diamond, the corresponding topic consistency prediction result is 0. A value of 0 indicates that the two video clips corresponding to the splicing vector come from different sample videos (one comes from the first sample video, and the other comes from the nth sample video);

[0091] For a splicing vector marked with a white circle and a gray diamond, the corresponding topic consistency prediction result is 0. A value of 0 indicates that the two video clips corresponding to the splicing vector come from different sample videos (one comes from the first sample video and the other comes from the nth sample video);

[0092] For the two splicing vectors marked with gray diamonds, the corresponding topic consistency prediction result has a value of 1, which means that the two video clips corresponding to the splicing vectors are derived from the same sample video (both from the nth sample video).

[0093] It is worth noting that Figure 2 Three video clips from the first sample video v j 1 、 and the video clip v from the nth sample video j n Taking 5 video clips as an example, we introduce the combination of splicing vectors and the value of the topic consistency prediction results. Similarly, for splicing vectors from the same sample video, the corresponding topic consistency prediction result is 1; for splicing vectors from different sample videos, the corresponding topic consistency prediction result is 0.

[0094] The feature extraction model trained using the model training method provided in the embodiments of this application can be applied to various application scenarios requiring video processing. For example, face recognition scenarios, behavior recognition scenarios, video recommendation scenarios, and video cover generation scenarios. To this end, the embodiments of this application also provide a video processing method based on the feature extraction model. Figure 3 This is a flow chart of a video processing method provided in an embodiment of the present application. Figure 3 , the method may include the following steps:

[0095] 301. Obtain a target video to be identified.

[0096] 302. Extract features from the target video using a feature extraction model to obtain a feature vector of the target video.

[0097] 303. Perform recognition processing on the target video according to the feature vector of the target video to obtain a recognition result.

[0098] In this embodiment, the target video is the video that requires recognition processing. In different application scenarios, the video content in the target video varies. Because the feature extraction model trained using the model training method provided in the embodiment of the present application has good model performance after training, the feature vector of the target video extracted using the feature extraction model can more accurately represent the visual information of the target video. Of course, performing recognition processing based on the feature vector that more accurately represents the visual information of the target video can obtain a more accurate recognition result.

[0099] It's worth noting that different recognition processing algorithms are used in different application scenarios. For example, face recognition uses a face recognition algorithm, behavior recognition uses a behavior recognition algorithm, video recommendation uses big data recommendation technology, and video cover generation uses video content understanding technology to understand the video content and select the most exciting cover for the understood video content.

[0100] The video processing method provided in the embodiment of the present application uses the feature vector of the target video extracted by the feature extraction model to more accurately represent the visual information of the target video. Recognition processing based on the feature vector that more accurately represents the visual information of the target video can obtain a more accurate recognition result.

[0101] Furthermore, the video processing method can optionally be applied to AR (Augmented Reality) scenes and / or VR (Virtual Reality) scenes. Specifically, based on the recognition results, the AR scene and / or VR scene can be triggered to perform corresponding operations.

[0102] It is worth noting that in different application scenarios, the recognition results may vary, and accordingly, the operations triggered to perform in the augmented reality (AR) scene and / or virtual reality (VR) scene may also vary. Several exemplary application scenarios are described below.

[0103] Scenario 1: Controlling the movement of a real target object in the real world and a virtual model in the virtual world. Specifically, if the recognition result is the motion information of the target object in the target video, triggering the AR scene and / or VR scene to perform the corresponding operation specifically includes: controlling the virtual model corresponding to the target object in the AR scene and / or VR scene to perform the corresponding action based on the motion information of the target object.

[0104] For example, when a target object in the real world performs an action such as running, walking, or sitting, the virtual model corresponding to the target object in the AR scene and / or VR scene also runs, walks, or sits.

[0105] Scenario 2: Controlling the expression linkage between a real target object in the real world and a virtual model in the virtual world. Specifically, if the recognition result is the expression information of the target object in the target video, triggering the AR scene and / or VR scene to perform the corresponding operation specifically includes: controlling the virtual model corresponding to the target object in the AR scene and / or VR scene to display the corresponding expression based on the expression information of the target object.

[0106] For example, when a target object in the real world presents a smiling face or a crying face, the virtual model corresponding to the target object in the AR scene and / or VR scene also presents a smiling face or a crying face.

[0107] Scenario 3: The AR or VR device can be unlocked or locked based on the actions of a real target object in the real world. Specifically, if the recognition result indicates that the target object in the target video has been unlocked, the AR device and / or VR device to which the AR scene belongs will be triggered to unlock; if the recognition result indicates that the target object in the target video has been locked, the AR device and / or VR device to which the AR scene belongs will be triggered to lock.

[0108] For example, when the AR device or VR device is in a locked state, the user can input an unlocking gesture or unlocking voice message associated with an unlocking operation to trigger the AR device or VR device to enter an unlocked state from a locked state. Similarly, when the AR device or VR device is in an unlocked state, the user can input a locking gesture or locking voice message associated with a locking operation to trigger the AR device or VR device to enter a locked state from an unlocked state.

[0109] It is worth noting that the target object may be a person, an animal, a plant, etc.; the virtual model corresponding to the target object includes, but is not limited to, a three-dimensional model of the target object, or a three-dimensional model of other objects associated with the target object.

[0110] It should be noted that the execution entity of each step of the method provided in the above embodiment can be the same device, or the method can be executed by different devices. For example, the execution entity of steps 101 to 104 can be device A; for another example, the execution entity of steps 101 and 102 can be device A, and the execution entity of steps 103 and 104 can be device B; and so on.

[0111] In addition, some of the processes described in the above embodiments and the accompanying drawings include multiple operations that appear in a specific order, but it should be clearly understood that these operations may not be executed in the order in which they appear in this article or may be executed in parallel. The serial numbers of the operations, such as 101, 102, etc., are only used to distinguish between different operations, and the serial numbers themselves do not represent any order of execution. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions of "first", "second", etc. in this article are used to distinguish different messages, devices, modules, etc., and do not represent a sequential order, nor do they limit "first" and "second" to different types.

[0112] Figure 4 This is a schematic diagram of the structure of a model training device provided in an embodiment of the present application. Figure 4 As shown, the model training device may include: an acquisition module 41 and a training module 42.

[0113] Among them, the acquisition module 41 is used to obtain a group of video clips corresponding to each of the multiple sample videos, each group of video clips includes two video clips with visual consistency and other video clips with no limitation on visual consistency, and the video clips in the same group of video clips have thematic consistency.

[0114] The training module 42 uses multiple sets of video clips to perform this round of model training on the feature extraction model to obtain multiple sets of feature vectors generated by the feature extraction model in this round of model training; based on the multiple sets of feature vectors, a first loss value corresponding to the visual consistency dimension of this round of model training is generated, and based on the topic consistency prediction model and the multiple sets of feature vectors, a second loss value corresponding to the topic consistency dimension of this round of model training is generated; when it is judged that the model training end conditions are not met according to the first loss value and the second loss value, the feature extraction model is continued to perform the next round of model training.

[0115] Further optionally, the training module 42 uses multiple groups of video clips to perform this round of model training on the feature extraction model to obtain multiple groups of feature vectors generated by the feature extraction model in this round of model training, which are specifically used for: enhancing each video clip in each group of video clips; inputting each enhanced video clip into the feature extraction model for feature extraction, and obtaining the feature vector of each video clip as a group of feature vectors.

[0116] Further optionally, when the training module 42 generates the first loss value corresponding to the visual consistency dimension of this round of model training based on multiple groups of feature vectors, it is specifically used to: for each group of feature vectors, according to the similarity between the first feature vector and the second feature vector corresponding to the two video clips with visual consistency in the group of feature vectors, and the similarity between the first feature vector and the second feature vector and other feature vectors different from themselves in each group of feature vectors, obtain the third loss value corresponding to the group of feature vectors; according to the multiple third loss values ​​corresponding to the multiple groups of feature vectors, obtain the first loss value corresponding to the visual consistency dimension of this round of model training.

[0117] Further optionally, the training module 42 obtains the third loss value corresponding to each group of feature vectors based on the similarity between the first feature vector and the second feature vector corresponding to the two video clips with visual consistency in the group of feature vectors, and the similarity between the first feature vector and the second feature vector and other feature vectors different from themselves in each group of feature vectors. The training module 42 is specifically used to: calculate the fourth loss value based on the similarity between the first feature vector and the second feature vector, and the similarity between the first feature vector and other feature vectors in each group of feature vectors except the first feature vector; calculate the fifth loss value based on the similarity between the first feature vector and the second feature vector, and the similarity between the second feature vector and other feature vectors in each group of feature vectors except the second feature vector; determine the third loss value corresponding to the group of feature vectors based on the fourth loss value and the fifth loss value.

[0118] Further optionally, when the training module 42 generates a second loss value corresponding to the topic consistency dimension of this round of model training based on the topic consistency prediction model and multiple groups of feature vectors, it is specifically used to: vector splice any two feature vectors in the multiple groups of feature vectors to obtain multiple splicing vectors; input each splicing vector into the topic consistency prediction model to perform topic consistency prediction on the two video clips corresponding to each splicing vector; and determine the second loss value corresponding to the topic consistency dimension of this round of model training based on the topic consistency prediction results and classification labels corresponding to each of the multiple splicing vectors.

[0119] Further optionally, when the training module 42 determines the second loss value corresponding to the subject consistency dimension of this round of model training based on the subject consistency prediction results and classification labels corresponding to each of the multiple splicing vectors, it is specifically used as follows: for each splicing vector, if it has a first classification label, and the first classification label indicates that the themes of the two video clips corresponding to the splicing vector are consistent, then the subject consistency prediction result of the splicing vector is used as the input parameter of the target loss function, and the sixth loss value corresponding to the splicing vector is calculated; if it has a second classification label, and the second classification label indicates that the themes of the two video clips corresponding to the splicing vector are inconsistent, then the difference between the subject consistency prediction result of the splicing vector and a specified value is used as the input parameter of the target loss function, and the sixth loss value corresponding to the splicing vector is calculated; based on the multiple sixth loss values ​​corresponding to the multiple splicing vectors, the second loss value corresponding to the subject consistency dimension of this round of model training is determined.

[0120] Further optionally, when the training module 42 determines the second loss value corresponding to the topic consistency dimension of this round of model training based on the multiple sixth loss values ​​corresponding to the multiple splicing vectors, it is specifically used to: average the multiple sixth loss values ​​with the first classification label to obtain a first average value, and average the multiple sixth loss values ​​with the second classification label to obtain a second average value; determine the second loss value corresponding to the topic consistency dimension of this round of model training based on the first average value and the second average value.

[0121] Further optionally, the target loss function is a Focal loss function used to solve the data imbalance problem.

[0122] Further optionally, before inputting each splicing vector into the topic consistency prediction model, the training module 42 is further configured to: input each splicing vector into a feature space conversion network to perform dimensionality reduction processing on the feature dimension of the splicing vector.

[0123] Further optionally, when the acquisition module 41 obtains multiple sample videos each corresponding to a group of video clips, it is specifically used to: for each sample video, sample two video clips with visual consistency from the sample video according to the sampling time interval used in this round of model training, and randomly sample other video clips from the sample video, and take the two video clips with visual consistency and the other video clips as a group of video clips.

[0124] Further optionally, the acquisition module 41 is further configured to determine a sampling time interval used in the current round of model training according to the training progress and maximum time interval of the current round of model training.

[0125] Figure 4 The model training apparatus shown can be executed Figure 1The implementation principle and technical effects of the model training method in the illustrated embodiment will not be described in detail. The specific manner in which each module and unit performs operations in the model training device in the above embodiment has been described in detail in the embodiment of the method, and will not be elaborated here.

[0126] Figure 5 This is a schematic diagram of the structure of a video processing device provided in an embodiment of the present application. Figure 5 As shown, the video processing device may include: an acquisition module 51, a feature extraction module 52 and a recognition processing module 53.

[0127] The acquisition module 51 is used to acquire the target video to be identified;

[0128] A feature extraction module 52 is configured to extract features from a target video using a feature extraction model to obtain a feature vector of the target video; wherein the feature extraction model is a model trained according to the model training method provided in the above embodiment;

[0129] The recognition processing module 53 is used to perform recognition processing on the target video according to the feature vector of the target video to obtain a recognition result.

[0130] Figure 5 The video processing device shown can perform Figure 3 The video processing method in the embodiment shown, its implementation principle and technical effects are not repeated here. The specific manner in which each module and unit performs operations in the video processing device in the above embodiment has been described in detail in the embodiment of the method, and will not be elaborated here.

[0131] Figure 6 This is a schematic diagram of the structure of a computer device provided in an embodiment of the present application. Figure 6 As shown, the computer device may include: a memory 61 and a processor 62;

[0132] The memory 61 is used to store computer programs and can be configured to store various other data to support operations on the computing platform. Examples of such data include instructions for any application or method operating on the computing platform, contact data, phone book data, messages, pictures, videos, etc.

[0133] The memory 61 can be implemented by any type of volatile or non-volatile memory device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.

[0134] The processor 62 is coupled to the memory 61 and is used to execute the computer program in the memory 61, so as to: obtain a group of video clips corresponding to each of the multiple sample videos, each group of video clips includes two video clips with visual consistency, and other video clips with no limitation on visual consistency, and the video clips in the same group of video clips have thematic consistency; use the multiple groups of video clips to perform current round model training on the feature extraction model to obtain multiple groups of feature vectors generated by the feature extraction model in this round of model training; generate a first loss value corresponding to the visual consistency dimension of the current round model training based on the multiple groups of feature vectors, and generate a second loss value corresponding to the thematic consistency dimension of the current round model training based on the thematic consistency prediction model and the multiple groups of feature vectors; when it is judged that the model training end condition is not met based on the first loss value and the second loss value, continue to perform the next round of model training on the feature extraction model.

[0135] Further optionally, the processor 62 uses multiple groups of video clips to perform this round of model training on the feature extraction model to obtain multiple groups of feature vectors generated by the feature extraction model in this round of model training, which are specifically used for: enhancing each video clip in each group of video clips; inputting each enhanced video clip into the feature extraction model for feature extraction, and obtaining the feature vector of each video clip as a group of feature vectors.

[0136] Further optionally, when the processor 62 generates the first loss value corresponding to the visual consistency dimension of this round of model training based on multiple groups of feature vectors, it is specifically used to: for each group of feature vectors, according to the similarity between the first feature vector and the second feature vector corresponding to the two video clips with visual consistency in the group of feature vectors, and the similarity between the first feature vector and the second feature vector and other feature vectors different from themselves in each group of feature vectors, obtain the third loss value corresponding to the group of feature vectors; according to the multiple third loss values ​​corresponding to the multiple groups of feature vectors, obtain the first loss value corresponding to the visual consistency dimension of this round of model training.

[0137] Further optionally, when the processor 62 obtains the third loss value corresponding to each group of feature vectors based on the similarity between the first feature vector and the second feature vector corresponding to the two video clips with visual consistency in the group of feature vectors, and the similarity between the first feature vector and the second feature vector and other feature vectors different from themselves in each group of feature vectors, the processor 62 is specifically used to: calculate the fourth loss value based on the similarity between the first feature vector and the second feature vector, and the similarity between the first feature vector and other feature vectors in each group of feature vectors except the first feature vector; calculate the fifth loss value based on the similarity between the first feature vector and the second feature vector, and the similarity between the second feature vector and other feature vectors in each group of feature vectors except the second feature vector; determine the third loss value corresponding to the group of feature vectors based on the fourth loss value and the fifth loss value.

[0138] Further optionally, when the processor 62 generates a second loss value corresponding to the topic consistency dimension of this round of model training based on the topic consistency prediction model and multiple groups of feature vectors, it is specifically used to: vector splice any two feature vectors in the multiple groups of feature vectors to obtain multiple splicing vectors; input each splicing vector into the topic consistency prediction model to perform topic consistency prediction on the two video clips corresponding to each splicing vector; and determine the second loss value corresponding to the topic consistency dimension of this round of model training based on the topic consistency prediction results and classification labels corresponding to each of the multiple splicing vectors.

[0139] Further optionally, when the processor 62 determines the second loss value corresponding to the subject consistency dimension of this round of model training based on the subject consistency prediction results and classification labels corresponding to each of the multiple splicing vectors, it is specifically used to: for each splicing vector, if it has a first classification label, the first classification label indicates that the themes of the two video clips corresponding to the splicing vector are consistent, then the subject consistency prediction result of the splicing vector is used as the input parameter of the target loss function, and the sixth loss value corresponding to the splicing vector is calculated; if it has a second classification label, the second classification label indicates that the themes of the two video clips corresponding to the splicing vector are inconsistent, then the difference between the subject consistency prediction result of the splicing vector and a specified value is used as the input parameter of the target loss function, and the sixth loss value corresponding to the splicing vector is calculated; based on the multiple sixth loss values ​​corresponding to the multiple splicing vectors, the second loss value corresponding to the subject consistency dimension of this round of model training is determined.

[0140] Further optionally, when the processor 62 determines the second loss value corresponding to the topic consistency dimension of this round of model training based on multiple sixth loss values ​​corresponding to multiple splicing vectors, it is specifically used to: average the multiple sixth loss values ​​with the first classification label to obtain a first average value, and average the multiple sixth loss values ​​with the second classification label to obtain a second average value; determine the second loss value corresponding to the topic consistency dimension of this round of model training based on the first average value and the second average value.

[0141] Further optionally, the target loss function is a Focal loss function used to solve the data imbalance problem.

[0142] Further optionally, before inputting each splicing vector into the topic consistency prediction model, the processor 62 is further configured to: input each splicing vector into a feature space conversion network to perform dimensionality reduction processing on the feature dimension of the splicing vector.

[0143] Further optionally, when the processor 62 obtains multiple sample videos each corresponding to a group of video clips, it is specifically used to: for each sample video, sample two video clips with visual consistency from the sample video according to the sampling time interval used in this round of model training, and randomly sample other video clips from the sample video, and take the two video clips with visual consistency and the other video clips as a group of video clips.

[0144] Further optionally, the processor 62 is further configured to determine a sampling time interval used in the current round of model training according to the training progress and maximum time interval of the current round of model training.

[0145] The detailed implementation process of the processor executing each action can be found in the relevant description in the aforementioned method embodiment or device embodiment, and will not be repeated here.

[0146] Further, if Figure 6 As shown, the computer device also includes: a communication component 63, a display 64, a power component 65, an audio component 66 and other components. Figure 6 Only some components are shown schematically, which does not mean that the computer equipment only includes Figure 6 In addition, Figure 6 The components in the dotted box are optional components, not mandatory components, and the specific configuration depends on the product form of the production scheduling equipment. The computer device of this embodiment can be implemented as a terminal device such as a desktop computer, a laptop computer, a smart phone or an IOT device, or as a server device such as a conventional server, a cloud server or a server array. If the computer device of this embodiment is implemented as a terminal device such as a desktop computer, a laptop computer, a smart phone, etc., it can include Figure 6If the computer device of this embodiment is implemented as a server device such as a conventional server, a cloud server or a server array, it may not include the components in the dotted box; Figure 6 Components within the dotted box.

[0147] The detailed implementation process of the processor executing each action can be found in the relevant description in the aforementioned method embodiment or device embodiment, and will not be repeated here.

[0148] The embodiment of the present application also provides a computer device, the structure and Figure 6 The computer devices shown are the same, but the processing logic is different. Specifically, the computer device includes: a memory and a processor;

[0149] Memory is used to store computer programs and can be configured to store various other data to support operations on the computing platform. Examples of such data include instructions for any application or method operating on the computing platform, contact data, phone book data, messages, pictures, videos, etc.

[0150] The memory may be implemented by any type of volatile or non-volatile memory device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.

[0151] A processor is coupled to a memory and is used to execute a computer program in the memory to: obtain a target video to be identified; use a feature extraction model to extract features from the target video to obtain a feature vector of the target video; and perform recognition processing on the target video based on the feature vector of the target video to obtain a recognition result; wherein the feature extraction model is a model trained according to the model training method provided in the above embodiment.

[0152] Accordingly, an embodiment of the present application further provides a computer-readable storage medium storing a computer program, which, when executed, can implement the steps that can be executed by a computer device in the above method embodiment.

[0153] Accordingly, an embodiment of the present application also provides a computer program product, including a computer program / instruction. When the computer program / instruction is executed by a processor, the processor is enabled to implement the steps that can be executed by a computer device in the above method embodiment.

[0154] The communication component is configured to facilitate wired or wireless communication between the device on which the communication component is installed and other devices. The device on which the communication component is installed can access a wireless network based on a communication standard, such as WiFi, 2G, 3G, 4G / LTE, 5G, or the like, or a combination thereof. In an example embodiment, the communication component receives a broadcast signal or broadcast related information from an external broadcast management system via a broadcast channel. In an example embodiment, the communication component further includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on Radio Frequency Identification (RFID) technology, infrared Data Association (IrDA) technology, Ultra-Wide Band (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0155] The display includes a screen, which can include a Liquid Crystal Display (LCD) and a Touch Panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive an input signal from a user. The touch panel includes one or more touch sensors to sense a touch, a slide, and a gesture on the touch panel. The touch sensor can not only sense a boundary of a touch or a slide action, but also detect a duration and a pressure associated with a touch or a slide operation.

[0156] The power supply component provides power to various components of the device on which the power supply component is installed. The power supply component can include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power for the device on which the power supply component is installed.

[0157] The audio component can be configured to output and / or input audio signals. For example, the audio component includes a microphone (MIC) that is configured to receive an external audio signal when the device on which the audio component is installed is in a particular mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signal can be further stored in memory or transmitted via the communication component. In some embodiments, the audio component also includes a speaker for outputting audio signals.

[0158] Those skilled in the art will appreciate that embodiments of the present application can be readily used as a method, a system, or a computer program product. Accordingly, the present application can take the form of an entirely hardware embodiment, an entirely software embodiment or an embodiment combining software and hardware aspects. Furthermore, the present application can take the form of a computer program product on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage devices, etc.) embodying computer readable program code thereon.

[0159] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0160] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0161] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 The steps for the function specified in one or more boxes.

[0162] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.

[0163] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.

[0164] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.

[0165] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.

[0166] The above are merely embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should all be included within the scope of the claims of the present application.

Claims

1. A model training method, characterized in that: include: A plurality of sample videos are obtained, each corresponding to a group of video clips, each group of video clips including two video clips with visual consistency and other video clips without limitation on visual consistency, and each video clip in the same group of video clips has thematic consistency; Using multiple sets of video clips to perform current model training on the feature extraction model to obtain multiple sets of feature vectors generated by the feature extraction model in the current model training; Generating a first loss value corresponding to the visual consistency dimension of this round of model training based on the multiple sets of feature vectors, and generating a second loss value corresponding to the thematic consistency dimension of this round of model training based on the thematic consistency prediction model and the multiple sets of feature vectors; When it is determined that the model training end condition is not met according to the first loss value and the second loss value, continuing to perform the next round of model training on the feature extraction model; wherein, generating a first loss value corresponding to the visual consistency dimension of this round of model training based on the multiple groups of feature vectors includes: for each group of feature vectors, obtaining a third loss value corresponding to the group of feature vectors based on the similarity between the first feature vector and the second feature vector corresponding to the two visually consistent video clips in the group of feature vectors, and the similarity between the first feature vector and the second feature vector and other feature vectors different from themselves in each group of feature vectors; and obtaining the first loss value corresponding to the visual consistency dimension of this round of model training based on the multiple third loss values ​​corresponding to the multiple groups of feature vectors; wherein, based on the topic consistency prediction model and the multiple groups of feature vectors, generating a second loss value corresponding to the topic consistency dimension of this round of model training, including: vector splicing any two feature vectors from the multiple groups of feature vectors to obtain multiple splicing vectors; inputting each splicing vector into the topic consistency prediction model to perform topic consistency prediction on the two video clips corresponding to each splicing vector; and determining the second loss value corresponding to the topic consistency dimension of this round of model training based on the topic consistency prediction results and classification labels corresponding to each of the multiple splicing vectors; Among them, according to the topic consistency prediction results and classification labels corresponding to each of the multiple splicing vectors, the second loss value corresponding to the current round of model training on the topic consistency dimension is determined, including: according to the topic consistency prediction results and classification labels corresponding to each splicing vector, the sixth loss value corresponding to each splicing vector is determined; according to the multiple sixth loss values ​​corresponding to the multiple splicing vectors, the second loss value corresponding to the current round of model training on the topic consistency dimension is determined.

2. The method according to claim 1, characterized in that The feature extraction model is trained in this round using multiple sets of video clips to obtain multiple sets of feature vectors generated by the feature extraction model in this round of model training, including: Performing enhancement processing on each video clip in each group of video clips; Each enhanced video segment is input into the feature extraction model for feature extraction, and a feature vector of each video segment is obtained as a group of feature vectors.

3. The method according to claim 1, characterized in that For each group of feature vectors, a third loss value corresponding to the group of feature vectors is obtained based on the similarity between the first feature vector and the second feature vector corresponding to the two visually consistent video clips in the group of feature vectors, and the similarity between the first feature vector and the second feature vector and other feature vectors different from the first feature vector and the second feature vector in each group of feature vectors, including: Calculating a fourth loss value based on a similarity between the first eigenvector and the second eigenvector, and a similarity between the first eigenvector and other eigenvectors in each group of eigenvectors except the first eigenvector; Calculating a fifth loss value based on a similarity between the first eigenvector and the second eigenvector, and a similarity between the second eigenvector and other eigenvectors in each group of eigenvectors except the second eigenvector; Determine a third loss value corresponding to the group of feature vectors based on the fourth loss value and the fifth loss value.

4. The method according to claim 1, wherein Based on the topic consistency prediction results and classification labels corresponding to the multiple concatenated vectors, the second loss value corresponding to the topic consistency dimension of this round of model training is determined, including: For each splicing vector, if it has a first classification label, and the first classification label indicates that the themes of the two video clips corresponding to the splicing vector are consistent, then the theme consistency prediction result of the splicing vector is used as the input parameter of the target loss function, and the sixth loss value corresponding to the splicing vector is calculated; if it has a second classification label, and the second classification label indicates that the themes of the two video clips corresponding to the splicing vector are inconsistent, then the difference between the theme consistency prediction result of the splicing vector and the specified value is used as the input parameter of the target loss function, and the sixth loss value corresponding to the splicing vector is calculated; According to the multiple sixth loss values ​​corresponding to the multiple splicing vectors, the second loss value corresponding to the topic consistency dimension of this round of model training is determined.

5. The method according to claim 4, characterized in that Determining a second loss value corresponding to the topic consistency dimension of this round of model training based on multiple sixth loss values ​​corresponding to the multiple splicing vectors includes: Averaging a plurality of sixth loss values ​​having the first classification label to obtain a first average value, and averaging a plurality of sixth loss values ​​having the second classification label to obtain a second average value; Based on the first average value and the second average value, determine the second loss value corresponding to the topic consistency dimension of this round of model training.

6. The method according to claim 4, characterized in that The objective loss function is a Focal loss function used to solve the data imbalance problem.

7. The method according to claim 4, characterized in that Before each concatenation vector is fed into the topic coherence prediction model, it also includes: Each concatenated vector is input into a feature space transformation network to reduce the feature dimension of the concatenated vector.

8. The method according to any one of claims 1 to 7, characterized in that Get multiple sample videos, each corresponding to a set of video clips, including: For each sample video, two video segments with visual consistency are sampled from the sample video according to the sampling time interval used in this round of model training, and other video segments are randomly sampled from the sample video. The two video segments with visual consistency and the other video segments are taken as a group of video segments.

9. The method according to claim 8, characterized in that Also includes: Determine the sampling time interval used in this round of model training based on the training progress and maximum time interval of this round of model training.

10. A video processing method, characterized in that: include: Obtain the target video to be identified; Extracting features from the target video using a feature extraction model to obtain a feature vector of the target video; Performing recognition processing on the target video according to the feature vector of the target video to obtain a recognition result; The feature extraction model is a model trained according to the model training method according to any one of claims 1 to 9.

11. The method according to claim 10, characterized in that Also includes: According to the recognition result, the augmented reality AR scene and / or virtual reality VR scene is triggered to perform corresponding operations.

12. A computer device, characterized in that: include: memory and processor; The memory is used to store computer programs; The processor is coupled to the memory and configured to execute the computer program to perform the steps of the method according to any one of claims 1 to 9.

13. A computer storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the processor is enabled to implement the steps of the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Training method and device of video feature extraction model and electronic equipment

    CN113378781A

  • Video recognition model training method and device and video recognition method and device

    CN113569740A