Video processing method, model training method, corresponding device and electronic equipment

Through the cross-domain feature alignment method, video and text feature encoder and dimensional transformation model are used to align video features with label features, solving the problems of slow UGC video tag generation speed and difficult label range adjustment, and achieving efficient video tag generation.

CN120526342APending Publication Date: 2025-08-22TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410193032.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-02-20
Publication Date
2025-08-22

AI Technical Summary

Technical Problem

When generating UGC video tags, the prior art faces the slow and difficult generation speed caused by the diversity of labels, and it is difficult to quickly adjust the scope of supporting tags. New tags require a lot of manual annotation and model retraining.

Method used

Through the trained video feature encoder and dimensional transformation model, the video features are converted into predetermined dimensions, and similar tag features are retrieved in the tag feature library. Combined with the text feature encoder and dimensional transformation model, the tag information and video features are aligned to the same space, cross-domain feature retrieval is realized, and tag information is directly obtained.

Benefits of technology

Improves the speed and efficiency of video tag generation, reduces dependence on manual annotation and model retraining, and supports rapid adjustment of tag range.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120526342A_ABST
    Figure CN120526342A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a video processing method, a model training method, a corresponding device and electronic equipment, and relates to the fields of artificial intelligence, computer vision, natural language processing, machine learning, cloud technology, big data, traffic and the like. The method comprises the following steps of: extracting a target video feature of a video to be processed through a trained video feature encoder, and converting the dimension of the target video feature into a predetermined dimension through a trained video feature dimension transformation model; retrieving a target label feature similar to the target video feature of the predetermined dimension in a preset label feature library, wherein the dimension of each label feature in the label feature library is the predetermined dimension; according to the method, the association relationship is established between the tag information corresponding to the target tag feature and the to-be-processed video, that is, the target video feature and the tag feature in the tag feature library are aligned to the same space, the corresponding tag information can be directly obtained through a cross-domain feature retrieval method, and the video tag reasoning speed is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to technical fields such as artificial intelligence, computer vision, natural language processing, machine learning, cloud technology, big data, etc. Specifically, the present disclosure relates to a video processing method, a model training method, corresponding devices and electronic equipment. Background Art

[0002] Video content recognition technology refers to the use of technologies such as computer vision and machine learning to identify and understand the content in the video.

[0003] Video tag generation is a very important part of video content recognition. For example, machines can automatically generate video tags for massive amounts of UGC (User-Generated Content) videos.

[0004] However, due to the diversity of UGC video content, the number of tags used in business scenarios can usually reach hundreds of thousands or even millions, making it difficult to label each video accordingly. Summary of the Invention

[0005] The embodiments of the present disclosure provide a video processing method, a model training method, a corresponding device, and an electronic device that can solve the problem of how to increase the speed of generating video tags. The technical solutions provided by the present disclosure are as follows:

[0006] According to one aspect of an embodiment of the present disclosure, a video processing method is provided, the method comprising:

[0007] Get the video to be processed;

[0008] Extract target video features of the video to be processed through the trained video feature encoder, and convert the dimension of the target video features into a predetermined dimension through the trained video feature dimension transformation model;

[0009] Retrieving target tag features similar to target video features of predetermined dimensions from a preset tag feature library, wherein the dimensions of each tag feature in the tag feature library are predetermined dimensions;

[0010] Establish an association between the label information corresponding to the target label feature and the video to be processed.

[0011] In an optional embodiment, the method further includes:

[0012] Get the description text corresponding to at least one tag information;

[0013] The trained text feature encoder is used to extract the text features of each descriptive text, and the trained text feature dimension transformation model is used to convert each text feature into a predetermined dimension to obtain each label feature.

[0014] Storing each tag feature in a tag feature library;

[0015] Among them, the video feature encoder, video feature dimension transformation model, text feature encoder and text feature dimension transformation model are jointly trained.

[0016] In an optional embodiment, extracting target video features of a video to be processed by using a trained video feature encoder includes:

[0017] Obtain multiple modal information corresponding to the video to be processed;

[0018] For each modal information, the modal features of the modal information are extracted through the modal feature encoder corresponding to the modal information;

[0019] Multiple modal features are fused to obtain the target video features of the video to be processed.

[0020] In an optional embodiment, the multiple modal features include text features, at least one video frame feature, and at least one audio feature, and the text features include tag words. The multiple modal features are fused to obtain target video features of the video to be processed, including:

[0021] For each word feature in the text feature, the word feature is fused with the corresponding position feature and the corresponding type feature to obtain the first feature;

[0022] For each video frame feature, the video frame feature is fused with the corresponding position feature and the corresponding type feature to obtain a second feature;

[0023] For each audio feature, the audio-visual feature is fused with the corresponding position feature and the corresponding type feature to obtain a third feature;

[0024] Each first feature, each second feature, and each third feature are fused to obtain a fused feature corresponding to the marked word, which is used as the target video feature of the video to be processed.

[0025] In an optional implementation, obtaining description text corresponding to at least one piece of tag information includes:

[0026] Obtain the name, at least one level classification, and description information corresponding to at least one tag information;

[0027] For each tag information, the name, at least one level classification, and description information corresponding to the tag information are combined according to a predetermined format to obtain a description text corresponding to the tag information.

[0028] In an optional embodiment, the method further comprises at least one of the following:

[0029] Acquire the tag information to be deleted, determine the first tag feature corresponding to the tag information to be deleted, and delete the first tag feature in the tag feature library;

[0030] Obtain the target description text corresponding to the label information to be added, extract the target text features of the target description text through the trained text feature encoder, and convert the target text features into a predetermined dimension through the trained text feature dimension transformation model to obtain a second label feature, and add the second label feature to the label feature library.

[0031] According to another aspect of the embodiments of the present disclosure, a model training method is provided, the method comprising:

[0032] Obtaining a training sample, where the training sample includes a sample video and at least one annotated label information corresponding to the sample video;

[0033] Extracting a first sample video feature of a sample video using a preset video feature encoder, and classifying the first sample video feature using at least one preset classification network to obtain at least one sample label information; training the preset video feature encoder and the at least one preset classification network based on the sample label information and the annotated label information until a first preset condition is met, thereby obtaining a pre-trained video feature encoder;

[0034] Through a pre-trained video feature encoder, a second sample video feature of the sample video is extracted, and through a preset video feature dimension transformation model, the dimension of the second sample video feature is converted into a predetermined dimension; through a preset text feature encoder, sample label features of each preset label information description text are extracted, and through a preset text feature dimension transformation model, each sample label feature is converted into a predetermined dimension; based on the sample label features of the predetermined dimension, the second sample video features of the predetermined dimension and the annotated label information, the pre-trained video feature encoder, the preset video feature dimension transformation model, the preset text feature encoder and the preset text feature dimension transformation model are trained until the first preset condition is met, thereby obtaining a trained video feature encoder, a trained video feature dimension transformation model, a trained text feature encoder and a trained text feature dimension transformation model.

[0035] In an optional embodiment, at least one classification network includes a target classification network corresponding to a multi-label task, and the target classification network is used to classify the first sample video features to obtain sample label information, including:

[0036] Classifying the first sample video features through the target classification network to obtain sample label information of whether the sample video is relevant for each label information;

[0037] Based on the sample label information and the annotation label information, the preset video feature encoder and target classification network are trained, including:

[0038] Based on the sample label information of whether each label information is relevant for the sample video, and the annotated label information corresponding to multiple label information related to the sample video, a preset video feature encoder and target classification network are trained.

[0039] In an optional embodiment, based on the sample label features of the predetermined dimension, the second sample video features of the predetermined dimension and the annotated label information, a pre-trained video feature encoder, a preset video feature dimension transformation model, a preset text feature encoder and a preset text feature dimension transformation model are trained, including:

[0040] Calculating the similarity between the sample label feature of the predetermined dimension and the second sample video feature of the predetermined dimension;

[0041] Based on the similarity and the labeled label information of whether the label information corresponding to the sample label feature is related to the sample video, the pre-trained video feature encoder, the preset video feature dimension transformation model, the preset text feature encoder and the preset text feature dimension transformation model are trained.

[0042] In an optional embodiment, based on the similarity and the label information of whether the label information corresponding to the sample label feature is related to the sample video, a pre-trained video feature encoder, a preset video feature dimension transformation model, a preset text feature encoder, and a preset text feature dimension transformation model are trained, including:

[0043] The pre-trained video feature encoder, the preset video feature dimensional transformation model, the preset text feature encoder and the preset text feature dimensional transformation model are trained so that the similarity is maximized for the case where the label information corresponding to the sample label feature is related to the sample video, and the similarity is less than a predetermined value for the case where the label information corresponding to the sample label feature is not related to the sample video.

[0044] According to another aspect of the embodiments of the present disclosure, there is provided a video processing apparatus, the apparatus comprising:

[0045] Video acquisition module, used to acquire the video to be processed;

[0046] The video feature extraction module is used to extract the target video features of the video to be processed through the trained video feature encoder, and convert the dimensions of the target video features into a predetermined dimension through the trained video feature dimension transformation model;

[0047] A retrieval module is used to retrieve a target tag feature similar to a target video feature of a predetermined dimension from a preset tag feature library, wherein the dimension of each tag feature in the tag feature library is a predetermined dimension;

[0048] The association module is used to establish an association relationship between the label information corresponding to the target label feature and the video to be processed.

[0049] In an optional embodiment, the device further comprises:

[0050] A tag acquisition module, used to obtain description text corresponding to at least one tag information;

[0051] The label feature extraction module is used to extract the text features of each descriptive text through the trained text feature encoder, and convert each text feature into a predetermined dimension through the trained text feature dimension transformation model to obtain each label feature;

[0052] A storage module, used for storing each tag feature in a tag feature library;

[0053] Among them, the video feature encoder, video feature dimension transformation model, text feature encoder and text feature dimension transformation model are jointly trained.

[0054] In an optional embodiment, when the video feature extraction module is used to extract target video features of the video to be processed using a trained video feature encoder, it is specifically used to:

[0055] Obtain multiple modal information corresponding to the video to be processed;

[0056] For each modal information, the modal features of the modal information are extracted through the modal feature encoder corresponding to the modal information;

[0057] Multiple modal features are fused to obtain the target video features of the video to be processed.

[0058] In an optional embodiment, the multiple modal features include text features, at least one video frame feature, and at least one audio feature, and the text features include tag words. When the video feature extraction module is used to fuse the multiple modal features to obtain target video features of the video to be processed, it is specifically used to:

[0059] For each word feature in the text feature, the word feature is fused with the corresponding position feature and the corresponding type feature to obtain the first feature;

[0060] For each video frame feature, the video frame feature is fused with the corresponding position feature and the corresponding type feature to obtain a second feature;

[0061] For each audio feature, the audio-visual feature is fused with the corresponding position feature and the corresponding type feature to obtain a third feature;

[0062] Each first feature, each second feature, and each third feature are fused to obtain a fused feature corresponding to the marked word, which is used as the target video feature of the video to be processed.

[0063] In an optional implementation, when the tag acquisition module is used to obtain the description text corresponding to at least one tag information, it is specifically used to:

[0064] Obtain the name, at least one level classification, and description information corresponding to at least one tag information;

[0065] For each tag information, the name, at least one level classification, and description information corresponding to the tag information are combined according to a predetermined format to obtain a description text corresponding to the tag information.

[0066] In an optional embodiment, the device further includes at least one of the following modules:

[0067] A deletion module is used to obtain the tag information to be deleted, determine the first tag feature corresponding to the tag information to be deleted, and delete the first tag feature in the tag feature library;

[0068] An adding module is used to obtain the target description text corresponding to the label information to be added, extract the target text features of the target description text through the trained text feature encoder, and convert the target text features into a predetermined dimension through the trained text feature dimension transformation model to obtain a second label feature, and add the second label feature to the label feature library.

[0069] According to another aspect of the embodiments of the present disclosure, a model training device is provided, the device comprising:

[0070] A sample acquisition module is used to acquire training samples, where the training samples include a sample video and at least one annotated label information corresponding to the sample video;

[0071] A first training module is configured to extract a first sample video feature of a sample video using a preset video feature encoder, and classify the first sample video feature using at least one preset classification network to obtain at least one sample label information; based on the sample label information and the annotated label information, the preset video feature encoder and the preset at least one classification network are trained until a first preset condition is met, thereby obtaining a pre-trained video feature encoder;

[0072] The second training module is used to extract the second sample video features of the sample video through a pre-trained video feature encoder, and convert the dimensions of the second sample video features into a predetermined dimension through a preset video feature dimension transformation model; extract the sample label features of each preset label information description text through a preset text feature encoder, and convert each sample label feature into a predetermined dimension through a preset text feature dimension transformation model; based on the sample label features of the predetermined dimension, the second sample video features of the predetermined dimension and the annotated label information, train the pre-trained video feature encoder, the preset video feature dimension transformation model, the preset text feature encoder and the preset text feature dimension transformation model until the first preset condition is met, thereby obtaining the trained video feature encoder, the trained video feature dimension transformation model, the trained text feature encoder and the trained text feature dimension transformation model.

[0073] In an optional embodiment, the at least one classification network includes a target classification network corresponding to the multi-label task, and the first training module, when used to classify the first sample video features through the target classification network to obtain sample label information, is specifically used to:

[0074] Classifying the first sample video features through the target classification network to obtain sample label information of whether the sample video is relevant for each label information;

[0075] The first training module is specifically used to train a preset video feature encoder and target classification network based on sample label information and annotation label information:

[0076] Based on the sample label information of whether each label information is relevant for the sample video, and the annotated label information corresponding to multiple label information related to the sample video, a preset video feature encoder and target classification network are trained.

[0077] In an optional embodiment, the second training module, when used to train a pre-trained video feature encoder, a preset video feature dimension transformation model, a preset text feature encoder, and a preset text feature dimension transformation model based on a sample label feature of a predetermined dimension, a second sample video feature of a predetermined dimension, and annotated label information, is specifically used to:

[0078] Calculating the similarity between the sample label feature of the predetermined dimension and the second sample video feature of the predetermined dimension;

[0079] Based on the similarity and the labeled label information of whether the label information corresponding to the sample label feature is related to the sample video, the pre-trained video feature encoder, the preset video feature dimension transformation model, the preset text feature encoder and the preset text feature dimension transformation model are trained.

[0080] In an optional embodiment, the second training module is used to train the pre-trained video feature encoder, the preset video feature dimensional transformation model, the preset text feature encoder, and the preset text feature dimensional transformation model based on the similarity and the label information of whether the label information corresponding to the sample label feature is related to the sample video, specifically for:

[0081] The pre-trained video feature encoder, the preset video feature dimensional transformation model, the preset text feature encoder and the preset text feature dimensional transformation model are trained so that the similarity is maximized for the case where the label information corresponding to the sample label feature is related to the sample video, and the similarity is less than a predetermined value for the case where the label information corresponding to the sample label feature is not related to the sample video.

[0082] According to another aspect of an embodiment of the present disclosure, an electronic device is provided, which includes a memory, a processor, and a computer program stored in the memory, and the processor executes the computer program to implement the video processing method and / or model training method provided by the embodiment of the present disclosure.

[0083] According to another aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the video processing method and / or model training method provided by the embodiment of the present disclosure is implemented.

[0084] According to another aspect of an embodiment of the present disclosure, a computer program product is provided, including a computer program, which, when executed by a processor, implements the video processing method and / or model training method provided by the embodiment of the present disclosure.

[0085] The video processing method, model training method, corresponding device and electronic device provided by the embodiments of the present disclosure extract the target video features of the video to be processed through a trained video feature encoder, and convert the dimensions of the target video features into predetermined dimensions through a trained video feature dimension transformation model; retrieve target label features similar to the target video features of the predetermined dimensions in a preset label feature library, wherein the dimensions of each label feature in the label feature library are all predetermined dimensions; establish an association relationship between the label information corresponding to the target label features and the video to be processed, that is, in the embodiments of the present disclosure, by aligning the target video features and the label features in the label feature library to the same space, the video label generation task is converted into a simpler label feature retrieval task, and the corresponding label information can be directly obtained through the cross-domain feature retrieval method, which effectively improves the inference speed of video labels and significantly improves the efficiency of automatic generation of video labels by machines. BRIEF DESCRIPTION OF THE DRAWINGS

[0086] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following briefly introduces the drawings required for describing the embodiments of the present disclosure.

[0087] Figure 1 Schematic diagram of classification-based approach in related video tag generation technology;

[0088] Figure 2 A schematic diagram of a method based on similar videos in the related video tag generation technology;

[0089] Figure 3 A flowchart of a video processing method provided by an embodiment of the present disclosure;

[0090] Figure 4 A flowchart of a method for generating video tags provided by an embodiment of the present disclosure;

[0091] Figure 5 A schematic diagram of extracting video features using a multimodal model provided in an embodiment of the present disclosure;

[0092] Figure 6 A schematic diagram of text modality feature processing provided by an embodiment of the present disclosure;

[0093] Figure 7 A schematic diagram of multimodal video feature fusion provided by an embodiment of the present disclosure;

[0094] Figure 8 A flowchart of a model training method provided in an embodiment of the present disclosure;

[0095] Figure 9 A schematic diagram of an application scenario of the video tag generation method provided by an embodiment of the present disclosure;

[0096] Figure 10 A schematic structural diagram of a video processing device provided in an embodiment of the present disclosure;

[0097] Figure 11 A schematic diagram of the structure of a model training device provided in an embodiment of the present disclosure;

[0098] Figure 12 A schematic structural diagram of an electronic device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0099] The following describes embodiments of the present disclosure in conjunction with the accompanying drawings. It should be understood that the embodiments described below in conjunction with the accompanying drawings are exemplary descriptions for explaining the technical solutions of the embodiments of the present disclosure and do not constitute a limitation on the technical solutions of the embodiments of the present disclosure.

[0100] Those skilled in the art will understand that, unless otherwise stated, the singular forms "a", "an", "" and "the" used herein may also include plural forms. It should be further understood that the terms "include" and "comprise" used in the embodiments of the present disclosure mean that the corresponding features can be implemented as the presented features, information, data, steps, operations, elements and / or components, but do not exclude implementation as other features, information, data, steps, operations, elements, components and / or combinations thereof supported by the technical field. It should be understood that when we refer to an element as being "connected" or "coupled" to another element, the element may be directly connected or coupled to the other element, or it may refer to a connection relationship between the element and the other element through an intermediate element. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The term "and / or" used herein indicates at least one of the items defined by the term, for example, "A and / or B" or "A, B" indicates implementation as "A", or implementation as "B", or implementation as "A and B".

[0101] In order to make the objectives, technical solutions and advantages of the present disclosure more clear, the embodiments of the present disclosure will be further described in detail below with reference to the accompanying drawings.

[0102] First, the technical terms involved in this disclosure are introduced and explained:

[0103] (1) NLP (Natural Language Processing): It is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can enable effective communication between humans and computers using natural language. Natural language processing involves natural language, that is, the language people use in daily life, and is closely related to linguistic research; it also involves computer science and mathematics. The pre-training model, an important technology for model training in the field of artificial intelligence, is developed from the large language model (Large Language Model) in the field of NLP. After fine-tuning, the large language model can be widely used in downstream tasks. Natural language processing technology usually includes text processing, semantic understanding, machine translation, robot question answering, knowledge graph and other technologies.

[0104] (2) Transformer model: A network structure model suitable for processing sequence data. It uses a full attention structure instead of the traditional recurrent neural network (RNN) or convolutional neural network (CNN). Its core idea is to model the relationship between each element in the input sequence and other elements as an attention weight matrix, and calculate the representation of each element through the self-attention mechanism. This structure enables the Transformer model to have a strong ability to capture long-distance dependencies in the sequence, thus achieving good performance in NLP tasks. The Transformer model is widely used in the field of NLP, including tasks such as machine translation, text summarization, and question-answering systems. It introduced the self-attention mechanism for the first time and surpassed traditional sequence-to-sequence (Seq2Seq) models such as LSTM (Long Short-Term Memory) and GRU (Gated Recurrent Unit) in many aspects. The success of Transformer has also spawned many variant models based on the self-attention mechanism, such as BERT (Bidirectional Encoder Representations from Transformer) and GPT (Generative Pre-Trained Transformer). These models have made significant progress in NLP tasks.

[0105] (3) Token: A token in Transformer refers to a word or subword in a text and is the basic unit used by the model when processing text.

[0106] (4) Optical Character Recognition (OCR): It is a technology that uses optical input methods such as scanning to identify and extract text content from various images, photos, bills, newspapers, books, manuscripts and other printed materials.

[0107] (5) Automatic Speech Recognition (ASR) technology, whose goal is to convert the vocabulary content in human speech into computer-readable input, such as keystrokes, binary codes, or character sequences.

[0108] In existing related technologies, methods for generating video tags can be roughly divided into the following two categories from the perspective of model implementation:

[0109] The first category: classification-based methods.

[0110] like Figure 1 As shown, this type of method usually manually labels a batch of videos with their corresponding content labels, constructs training set data to train the classification model, and then uses the classification model to label the new videos. This type of method relies heavily on manually labeled content. From the perspective of labels, it can be considered that each video labeled with the label is a positive sample of the label, and all other videos not labeled with the label are negative samples of the label. Negative samples are generally sufficient for a label, but usually for each label, sufficient positive samples (such as more than 50) are required to train the classification model so that the classification model can accurately predict the label. Since it is generally difficult to obtain good model results for labels with insufficient positive samples, it is generally necessary to filter out labels with insufficient positive samples and not train them, so that the classification model only supports labels with sufficient positive samples. On the other hand, when new labels need to be added, it is generally necessary to supplement the training data in a targeted manner and retrain the model to support them.

[0111] The second category: methods based on similar videos.

[0112] like Figure 2 As shown in the figure, this type of method typically extracts features from videos and processes the manually annotated data into a format of <video features, label list>, which is then stored in a video feature library. For videos that need to be labeled, the same method is first used to extract features from the video. Video features with similar content are searched within the video features. The label lists corresponding to similar video features are used as candidate labels. The final labeling result is obtained through a label ranking module (e.g., voting based on label frequency).

[0113] The relevant technologies mainly have the following shortcomings:

[0114] 1) The label list supported by the classification-based method is a closed set, making it difficult to quickly adjust the range of supported labels. When the label list changes, the model needs to be retrained, resulting in low iteration efficiency.

[0115] 2) The method based on similar videos also has the following two shortcomings:

[0116] This approach assumes that videos with similar content have similar tags. It uses similar videos as a bridge to obtain the tags of similar videos as a candidate tag set. However, it also requires designing a ranking module (such as voting), making it an inefficient end-to-end fast solution.

[0117] b. New tags still need to be labeled with a certain number of sample videos, which must be processed and then flushed into the video feature library before they can be supported.

[0118] In response to at least one of the above-mentioned technical problems or areas that need improvement in the relevant technology, the present disclosure proposes a video processing method, a model training method, a corresponding device and an electronic device. This solution can be understood as a video label generation method based on cross-domain feature alignment. By retrieving cross-domain alignment features, this solution can quickly obtain the label information corresponding to the video to be processed, thereby effectively improving the inference speed of video labels.

[0119] The following describes several exemplary embodiments to illustrate the technical solutions of the embodiments of the present disclosure and the technical effects produced by the technical solutions of the present disclosure. It should be noted that the following embodiments can refer to, draw on, or combine with each other, and the same terms, similar features, and similar implementation steps in different embodiments will not be repeated.

[0120] The present disclosure provides a video processing method, such as Figure 3 As shown, the method includes:

[0121] Step S101: Obtain the video to be processed;

[0122] In the embodiments of the present disclosure, a video to be processed refers to a video for which label information needs to be generated. For example, the video to be processed can be a video published on a certain platform, a video stored in a database or blockchain, or a video shot and / or edited in real time. Specifically, the video to be processed can be various types of videos, such as film and television videos and short videos.

[0123] Among them, the label information of the video to be processed refers to the information that classifies, labels or describes the entire video or a certain frame or certain pictures of the video with keywords. Specifically, the label information may include one or more classification methods. As an example, for a video about "meteor shower", you can select <technology> for labeling in the label information classified by dimensions such as <life> and <technology>, and you can select <astronomy> for labeling in the label information classified by dimensions such as <astronomy>, <food>, and <music>. In actual applications, those skilled in the art can set the classification method and content of the label information according to actual conditions, and the embodiments of the present disclosure are not limited here.

[0124] Step S102: extracting target video features of the video to be processed by using a trained video feature encoder, and converting the dimensions of the target video features into predetermined dimensions by using a trained video feature dimension transformation model;

[0125] That is, in the disclosed embodiment, by inputting a to-be-processed video into a trained video feature encoder, features of the to-be-processed video can be extracted based on the trained video feature encoder to obtain target video features. Furthermore, by inputting the extracted target video features into a trained video feature dimensionality transformation model, the trained video feature dimensionality transformation model can be used to transform (specifically, scale) the dimensions of the target video features into predetermined dimensions, thereby obtaining target video features of the predetermined dimensions.

[0126] Among them, those skilled in the art can set the network structure and training method adopted by the video feature encoder and the video feature dimensional transformation model according to actual conditions, and the embodiments of the present disclosure are not limited here. As an example, the video feature encoder can adopt a transformer-like network structure, the video feature dimensional transformation model can adopt at least one multilayer perceptron (MLP), or the video feature encoder and the video feature dimensional transformation model can also adopt other neural network models. In addition, the video feature encoder and the video feature dimensional transformation model can be jointly trained, that is, the video feature encoder and the video feature dimensional transformation model can be trained simultaneously in an end-to-end training manner, which can reduce the complexity of training and obtain better model performance. Alternatively, the video feature encoder and the video feature dimensional transformation model can also be trained separately.

[0127] In the embodiment of the present disclosure, the dimension of the extracted target video feature is converted into a predetermined dimension, with the purpose of aligning the target video feature and the tag feature to be retrieved in step S103 into one space.

[0128] Step S103: searching a preset tag feature library for target tag features that are similar to target video features of a predetermined dimension, wherein the dimensions of each tag feature in the tag feature library are all predetermined dimensions;

[0129] In the disclosed embodiment, tag information is explicitly encoded into tag features and stored in a tag feature library. The dimensions of each tag feature in the tag feature library are predetermined, meaning they are aligned to the target video features in a single space, facilitating comparison with the target video features.

[0130] In the embodiment of the present disclosure, the label feature library can also be understood as a retrieval library, which is used to obtain target label features similar to target video features of a predetermined dimension by feature retrieval, so as to obtain the corresponding label information of the video to be processed.

[0131] In the embodiments of the present disclosure, each video to be processed can be labeled with one or more tag information. That is, the number of target tag features retrieved that are similar to the target video features of a predetermined dimension can be one or more. Those skilled in the art can configure the method for determining the number of tags based on actual circumstances. For example, the number of tags can be a fixed number, or each target tag feature whose similarity to the target video features of the predetermined dimension is greater than a set value can be retrieved, and the embodiments of the present disclosure do not limit this.

[0132] Step S104: establishing an association relationship between the tag information corresponding to the target tag feature and the video to be processed.

[0133] In the disclosed embodiment, since the target tag feature is similar to the target video feature, the tag information corresponding to the target tag feature can be used as the tag recognition result of the video to be processed, so an association relationship between the two can be directly established.

[0134] The video processing method provided by the embodiment of the present disclosure aligns the target video features with the label features in the label feature library into the same space, and can directly obtain the corresponding label information through the cross-domain feature retrieval method, thereby effectively improving the inference speed of video labels.

[0135] In the embodiment of the present disclosure, a feasible implementation method is provided for the construction of a label feature library. Specifically, each label feature in the label feature library is obtained by extracting the text features of the label information description text through a trained text feature encoder, and converting each text feature into a predetermined dimension through a trained text feature dimension transformation model.

[0136] Specifically, the processing may include:

[0137] Step S201: Obtain description text corresponding to at least one tag information;

[0138] In the embodiment of the present disclosure, each tag information can be specifically represented in the form of plain text. The specific content of the description text of the tag information and the number of tag information can be set by those skilled in the art according to actual conditions, and the embodiment of the present disclosure does not limit this.

[0139] Step S202: extracting text features of each descriptive text using a trained text feature encoder, and converting each text feature into a predetermined dimension using a trained text feature dimension conversion model to obtain each label feature;

[0140] That is, in the disclosed embodiment, each descriptive text is input into a trained text feature encoder, and features of the descriptive text are extracted based on the trained text feature encoder to obtain corresponding text features. Furthermore, the extracted text features are input into a trained text feature dimensionality transformation model, and the dimensions of the text features are converted (specifically, they can be scaled) to a predetermined dimension using the trained text feature dimensionality transformation model, thereby obtaining text features of the predetermined dimension as label features.

[0141] Among them, those skilled in the art can set the network structure and training method adopted by the text feature encoder and the text feature dimensional transformation model according to actual conditions, and the embodiments of the present disclosure are not limited here. As an example, the text feature encoder can adopt the BERT network structure, the text feature dimensional transformation model can also adopt MLP, or the text feature encoder and the text feature dimensional transformation model can also adopt other neural network models. Among them, the structure of the text feature dimensional transformation model and the video feature dimensional transformation model can be different, and the parameters are not shared. In addition, the video feature encoder, the video feature dimensional transformation model, the text feature encoder and the text feature dimensional transformation model can be jointly trained, that is, the video feature encoder, the video feature dimensional transformation model, the text feature encoder and the text feature dimensional transformation model can be trained simultaneously in an end-to-end training manner, which can reduce the complexity of training and obtain better model performance. Alternatively, these models can also be trained separately.

[0142] Step S203: storing each tag feature in a tag feature library;

[0143] In the disclosed embodiments, the trained video feature encoder, video feature dimensionality transformation model, text feature encoder, and text feature dimensionality transformation model can align the video features of the video to be processed and the label features of the tag information into the same space, facilitating cross-domain comparison and retrieval. Furthermore, because the extracted label features can be reused for different videos to be processed, the label features are pre-stored in a label feature library and can be directly accessed and used when labeling the video to be processed, thereby improving the efficiency of label recognition.

[0144] In general, the process of the video tag generation method provided by the embodiment of the present disclosure can be as follows: Figure 4 As shown, this can specifically include: processing each tag information description text into tag features of predetermined dimensions using a text feature encoder and a text feature dimensionality transformation model, and storing the tag features in a tag feature library. Whenever a video to be processed needs to be labeled, the video to be processed is processed into video features of predetermined dimensions using a video feature encoder and a video feature dimensionality transformation model. Similar tag features are searched in the tag feature library to obtain the corresponding label results.

[0145] In the embodiment of the present disclosure, an optional implementation is provided for "extracting target video features of the video to be processed by using a trained video feature encoder" in step S102. Specifically, the implementation may include:

[0146] Step S1021: Acquire multiple modal information corresponding to the video to be processed;

[0147] Step S1022: for each type of modal information, extracting modal features of the modal information using a modal feature encoder corresponding to the modal information;

[0148] Step S1023: Fusing multiple modal features to obtain target video features of the video to be processed.

[0149] In the embodiments of the present disclosure, it is taken into account that videos usually naturally contain multiple modalities, including but not limited to video frames, audio, text, etc., wherein the text of the video includes but is not limited to titles, subtitles, text in video frames (for example, recognized by OCR technology), text converted from speech (for example, recognized by ASR technology), video types filled in by video creators, related topics, etc.

[0150] In the embodiment of the present disclosure, the video feature extraction can train a video feature encoder that integrates multimodal features. That is, the video feature encoder can be understood as a multimodal model. Figure 5 As shown in the figure, after obtaining the multi-modal information corresponding to the video to be processed, the video features can be extracted through the multi-modal model for label recognition.

[0151] The structure of the video feature encoder (multimodal model) can include modal feature encoders for different modalities. That is, in the multimodal model, each modality has its own encoder. After extracting modal features, they are directly fused or further fused through a fusion model in the video feature encoder to obtain the fused video features.

[0152] In an embodiment of the present disclosure, an optional implementation scheme is provided for a video feature encoder (multimodal model). Specifically, the video feature encoder includes a text feature encoder, an audio feature encoder, and a visual feature encoder. The text feature encoder is used to extract text features of the video, the audio feature encoder is used to extract audio features of the video, and the visual feature encoder is used to extract video frame features of the video. The extracted multiple modal features include text features, at least one video frame feature, and at least one audio feature.

[0153] Optionally, the visual feature encoder may adopt a ResNet101 model, but is not limited thereto; the audio feature encoder may adopt a Vggish model, but is not limited thereto; and the text feature encoder may adopt a BERT model, but is not limited thereto.

[0154] In the embodiment of the present disclosure, the video feature encoder may further include a fusion model. Optionally, the fusion model may adopt a single-stream transformer structure to fuse and predict the tokens extracted from the video frame by the visual feature encoder, the tokens extracted from the audio frame by the audio feature encoder, and the video text tokens through the fusion model, but is not limited to this. The fusion model may also adopt other network structures.

[0155] In the embodiment of the present disclosure, the extracted text features may include a CLASS token (CLStoken), which is usually located at the first position token [0] among all tokens and is used to represent the entire feature and can be used to perform category prediction.

[0156] In the disclosed embodiment, a CLS token is used to represent the entire multimodal feature, that is, it is used to represent the features after the entire video is fused.

[0157] That is, step S1023 may specifically include:

[0158] For each word feature in the text feature, the word feature is fused with the corresponding position feature and the corresponding type feature to obtain the first feature;

[0159] For each video frame feature, the video frame feature is fused with the corresponding position feature and the corresponding type feature to obtain a second feature;

[0160] For each audio feature, the audio-visual feature is fused with the corresponding position feature and the corresponding type feature to obtain a third feature;

[0161] Each first feature, each second feature, and each third feature are fused to obtain a fused feature corresponding to the marked word, which is used as the target video feature of the video to be processed.

[0162] As an example, the multimodal information input of the video to be processed is tokenized, and the self-attention mechanism of the transformer is used to align and fuse the multimodal features. For the text part, the way to extract the embedding of the token is to add the embedding (word element feature) of the word itself extracted by the text feature encoder to the encoding of the position (corresponding position feature) and the encoding of the type (corresponding type feature) to obtain the first feature, such as Figure 6 As shown in the figure; for the video frame (which can be processed for each video frame or sampled at equal intervals), a method similar to word-unit features is used, using the features extracted by the visual feature encoder as visual token embedding (video frame features) plus the position encoding (corresponding position features) and type encoding (corresponding type features) to obtain the second feature; similarly, after the audio feature is extracted by the audio feature encoder, the position encoding (corresponding position features) and type encoding (corresponding type features) are added to obtain the third feature. See Figure 7 In this way, visual features (video frame features) and audio features, like text features (including CLStoken and SEP token (separated words, used for sentence segmentation)), can be input into the transformer in the form of tokens, where Figure 7 Each shadow module in Figure 6 The structure shown in FIG, and then the transformer can learn the fusion feature for prediction. In the embodiment of the present disclosure, the embedding output by the CLS token (the fusion feature corresponding to the tag word) is used as the target video feature of the video to be processed.

[0163] In the embodiment of the present disclosure, an optional implementation is provided for step S201, which may specifically include:

[0164] Step S2011: Obtain the name, at least one level classification, and description information corresponding to at least one tag information;

[0165] Step S2012: for each tag information, the name, at least one level classification, and description information corresponding to the tag information are combined according to a predetermined format to obtain a description text corresponding to the tag information.

[0166] For example, if the preset multiple tag information is as shown in the following table:

[0167]

[0168]

[0169]

[0170] The tag information includes a name, a first-level tag classification, a second-level tag classification, and a description information. For each tag information, these information can be combined according to a predetermined format to obtain a description text corresponding to the tag information.

[0171] Optionally, the description text of the tag information can be constructed using the following predefined format:

[0172] [Label Chinese name]: Belongs to the [Label Secondary Category] type under the [Label First Category] category, meaning: [Label description information].

[0173] For example, the description text corresponding to the "Interesting Experiments" label in the table above can be combined as follows:

[0174] Interesting experiments: belong to the "Highlights - Fun" type under the "Life" category, which means: interesting experiments and challenging content without scientific experimental content, as long as they are interesting (this label does not apply to normal scientific experiments). If the video experiment project is interesting and has a big imagination and a strong curiosity, it can be included in the "Bizarre Experiments".

[0175] Similarly, each tag information can generate a description text similar to this one, which is then fed into a text feature encoder to extract text features. For example, if the text feature encoder uses the BERT model, the description text in a predetermined format can be tokenized and fed into BERT. The embeddings corresponding to the CLS tokens are extracted as text features, and then processed into label features of the predetermined dimension using a text feature dimensionality transformation model.

[0176] Those skilled in the art will appreciate that the above-described predetermined format is merely an illustrative description and does not constitute a limitation on the embodiments of the present disclosure. Appropriate variations based on this example are also applicable to the present disclosure and should therefore be included within the scope of protection of the present disclosure. Furthermore, the label information in the table above is for illustrative purposes only. The specific label content and classification methods are subject to actual implementation. That is, the specific content and classification methods of the label information in the table above should not be construed as limiting the present disclosure.

[0177] In the embodiments of the present disclosure, feasible implementation methods are provided for adding and deleting tag information, specifically, the methods may include:

[0178] Step SA: obtaining the tag information to be deleted, determining the first tag feature corresponding to the tag information to be deleted, and deleting the first tag feature in the tag feature library;

[0179] In the embodiment of the present disclosure, since the label information is explicitly encoded into label features and stored in the label feature library, for the labels that need to be withdrawn from the library (that is, the label information to be deleted), the corresponding label features can be directly deleted, that is, the first label feature corresponding to the label information to be deleted is determined, and the first label feature in the label feature library is deleted.

[0180] Existing classification-based methods, however, do not support direct deletion of label information and require retraining the classification model. Existing similar video-based methods, for labels to be delisted, require searching the video label library for multiple video features corresponding to the label information and deleting them. This shows that the technical solution provided by the disclosed embodiments can better support open set formats compared to existing related technologies, greatly improving the efficiency and flexibility of label information maintenance.

[0181] Step SB: Obtain the target description text corresponding to the label information to be added, extract the target text features of the target description text through the trained text feature encoder, and convert the target text features into a predetermined dimension through the trained text feature dimension transformation model to obtain a second label feature, and add the second label feature to the label feature library.

[0182] In the embodiments of the present disclosure, for labels to be newly added to the library (label information to be added), the label information description text can be directly generated according to the method described in at least one of the above embodiments, and then the label features can be extracted and added to the label feature library to take effect. There is no need to annotate the data and re-update the model, and the demand for new labels can be flexibly and efficiently supported.

[0183] This is because the trained video feature encoder, video feature dimension transformation model, text feature encoder and text feature dimension transformation model can align the video features of the video to be processed and the label features of the label information into the same space, have strong generalization capabilities, and can directly generalize similar relationships to new labels.

[0184] Existing classification-based methods, however, do not support the direct addition of new label information and require retraining the classification model. Existing similar video-based methods, for labels that need to be added to the database, require the collection of a sufficient number of sample videos containing the newly added labels, processing them, and storing them in the database. This shows that the technical solution provided by the disclosed embodiments can better support open-set formats compared to existing related technologies, significantly improving the efficiency and flexibility of label information maintenance.

[0185] The video processing method provided by the embodiments of the present disclosure has at least the following two advantages over existing related technologies:

[0186] 1) Abandoning the two-stage approach of using similar videos as a bridge, first obtaining the video tags of similar videos as a subset of candidate tags and then performing tag screening, this approach directly aligns the multimodal features of the video and the text features describing the tags into the same space through the video and tag encoders. This converts the tag generation problem into a simple feature search problem, and directly obtains the tag results in one step in an end-to-end manner.

[0187] 2) Zero-sample support for new labels. This solution converts each label information into a label feature and stores it in a label feature library. When the system is adjusted, the label feature library (equivalent to the label information library) can be directly added or deleted. There is no need to refresh the library or retrain the model for each new label, as in existing technologies.

[0188] The present disclosure also provides a model training method, the training goal of which is to align (multimodal) video features and label features into a space, and support direct cross-domain search to generate relevant labels. Therefore, the model training method can also be called a cross-domain feature alignment training method. Figure 8 As shown, the method includes:

[0189] Step S301: Acquire a training sample, where the training sample includes a sample video and at least one annotated label information corresponding to the sample video;

[0190] In the embodiment of the present disclosure, each sample video is supported to be annotated with one or more annotated label information.

[0191] Step S302: extracting a first sample video feature of a sample video using a preset video feature encoder, and classifying the first sample video feature using at least one preset classification network to obtain at least one sample label information; training the preset video feature encoder and the at least one preset classification network based on the sample label information and the annotated label information until a first preset condition is met, thereby obtaining a pre-trained video feature encoder;

[0192] This step can be understood as the first stage of the training process, the main purpose of which is to train the video feature encoder (parameter is denoted as V) to have stronger video feature extraction and / or multimodal encoding capabilities.

[0193] Optionally, after the video feature encoder V (e.g., Figure 7 The transformer output (shown as CLS) is directly connected to at least one fully connected (FC) layer (i.e., a classification network) to predict the probability distribution of video categories (i.e., train classification capabilities). Optionally, multiple classification networks can be used to handle various classification tasks. For example, for the pre-formatted label information in the example above, two FC layers can be used to simultaneously train the video secondary classification and label names.

[0194] Optionally, for classification problems, the cross-entropy loss function (CELoss) can be used for training. This function measures the difference between the sample label information and the annotation label information. The closer the two are, the smaller the cross-entropy loss, indicating that the model prediction results are more accurate.

[0195] In addition, as can be seen from the above description, the embodiment of the present disclosure supports labeling multiple label information for each sample video, and the at least one classification network includes a target classification network corresponding to the multi-label task, and the target classification network is used to classify the first sample video feature to obtain sample label information, including: classifying the first sample video feature through the target classification network to obtain sample label information of whether the sample video is relevant for each label information;

[0196] Furthermore, based on the sample label information and the annotated label information, a preset video feature encoder and a target classification network are trained, including: based on the sample label information of whether each label information of the sample video is relevant, and the annotated label information corresponding to multiple label information related to the sample video, the preset video feature encoder and the target classification network are trained.

[0197] In other words, multi-label classification can be viewed as multiple binary classification problems. For each preset label, predict whether the sample video is associated with each label and measure the degree of difference between the predicted result and the actual annotated label information. Optionally, a binary classification focal loss function is used for training.

[0198] In the embodiment of the present disclosure, the first stage training termination conditions (first preset conditions) may include, but are not limited to, model convergence, loss value less than a preset value, and the number of training times reaching a predetermined number. Those skilled in the art may set these conditions based on actual conditions, and the embodiment of the present disclosure does not limit these conditions.

[0199] For the embodiment of the present disclosure, the first stage of training is to improve the effect of the (multimodal) video feature encoder before the second stage of training, rather than using this method to predict video label information. Therefore, the classification network (such as the FC layer) trained here will be discarded after the training of this stage is completed.

[0200] In addition, for the text feature encoder, a similar scheme can be used for pre-training, or BERT that has been pre-trained on MLM (Masked Language Model, a pre-training method commonly used in natural language processing) and NSP (next sentence predict, used to train the ability to understand between sentences) tasks can be directly used for initialization without the need for first-stage training.

[0201] Step S303: extract the second sample video features of the sample video through a pre-trained video feature encoder, and convert the dimension of the second sample video features into a predetermined dimension through a preset video feature dimension transformation model; extract the sample label features of each preset label information description text through a preset text feature dimension transformation model, and convert each sample label feature into a predetermined dimension through a preset text feature dimension transformation model; based on the sample label features of the predetermined dimension, the second sample video features of the predetermined dimension and the annotated label information, train the pre-trained video feature encoder, the preset video feature dimension transformation model, the preset text feature encoder and the preset text feature dimension transformation model until the first preset condition is met, thereby obtaining the trained video feature encoder, the trained video feature dimension transformation model, the trained text feature encoder and the trained text feature dimension transformation model.

[0202] This step can be understood as the second stage of the training process, the purpose of which is to train the ability to align video features and label features into a space.

[0203] Specifically, a small model is connected after the (multimodal) video feature encoder V (such as after the embedding corresponding to the output CLS token) as a video feature dimension transformation model (which can be a 2-layer MLP, with the parameter denoted as V'), which can scale the second sample video feature to d dimensions (for the convenience of description, the d-dimensional second sample video feature is denoted as fv).

[0204] Similarly, a small model is connected after the label feature encoder (the parameter is denoted as T) as the text feature dimension transformation model (the parameter is denoted as T', which can be different from the structure of V' above and the parameters are not shared), and the sample label features are also scaled to d dimensions (for the convenience of description, the d-dimensional sample label features are denoted as ft).

[0205] In short, in the second phase of training, V' and T' are trained after V and T, respectively. Optionally, the parameters of V' and T' are randomly initialized, while V and T are initialized with the parameters trained in the first phase. V, T, V', and T' are trained based on the two features fv and ft output during training, both of which have d dimensions, combined with the annotated label information.

[0206] In an optional embodiment, the similarity between the sample label feature of a predetermined dimension and the second sample video feature of a predetermined dimension can be calculated; based on the similarity, and the annotated label information of whether the label information corresponding to the sample label feature is related to the sample video, a pre-trained video feature encoder, a preset video feature dimension transformation model, a preset text feature encoder and a preset text feature dimension transformation model are trained.

[0207] Optionally, the similarity may be calculated using Euclidean distance, that is, for the two output features fv and ft whose feature dimensions are both d-dimensional, the Euclidean distance may be directly calculated, but the present invention is not limited thereto.

[0208] In this stage, the pre-trained video feature encoder, the preset video feature dimension transformation model, the preset text feature encoder and the preset text feature dimension transformation model can be trained so that the similarity is maximized when the label information corresponding to the sample label feature is related to the sample video, and the similarity is less than a predetermined value when the label information corresponding to the sample label feature is not related to the sample video.

[0209] Based on sampling a batch of sample videos and label annotation information pairs for each batch, if the sample video i should be marked with label annotation information j, then the feature fv of this video i (hereinafter referred to as i) and sample label feature ft j (hereinafter abbreviated as j) to form a positive sample pair, the two features i and j should be close (large similarity); otherwise they should be far apart (small similarity).

[0210] Optionally, a contrastive loss function can be used for training. The training goal is to pull related video and label features closer together, otherwise they are pulled further apart. It can be expressed as follows:

[0211] L(i,j)=I[y i =y j ]*dist(i,j) 2 +I[y i ≠y j ]*max(0,m–dist(i,j)) 2

[0212] Where dist is the Euclidean distance function, I[·] is the sign function, and the input is True (i.e., when the sample video i should be labeled with the label information j, y i =y j True, otherwise, y i ≠y jis True), otherwise it returns 1; m is a set hyperparameter (i.e., corresponding to the above predetermined value), indicating that the distance between negative sample pairs needs to exceed m (i.e., the similarity is less than the predetermined value).

[0213] For example, if sample video i should be labeled with label information j, I[y i =y j ]=1,I[y i ≠y j ]=0,L(i,j)=dist(i,j) 2 , we can train the two features i and j to be close. If the sample video i should be marked with the label information j, I[y i =y j ]=0,I[y i ≠y j ]=1,L(i,j)=max(0,m–dist(i,j)) 2 , at this time, if dist(i,j) is less than m, then L(i,j) = 0, when dist(i,j) is greater than m, L(i,j) = (m – dist(i,j)) 2 , the distance between the two features i and j can be trained to exceed m.

[0214] Alternatively, you can use a metric learning loss function for training, which can also achieve similar functions.

[0215] Similarly, the second stage training termination condition (second preset condition) can be set by those skilled in the art according to actual circumstances, and is not limited in this embodiment of the present disclosure. As an example, each pair can be set to calculate the loss according to the above formula, and the total loss can be calculated as the average of the losses of each pair in the batch. BP (Back Propagation) can be used to train the model until convergence.

[0216] In the embodiment of the present disclosure, after the model training is completed, referring to steps S201 to S203, it is only necessary to process the labels in the label library into description texts, extract d-dimensional label features using T and T' obtained from the second stage of training, and import them into the label feature library.

[0217] Similarly, referring to steps S101 to S104, for the video to be processed, V and V' obtained from the second stage of training are used to extract d-dimensional video features, and similar label features are searched in the label feature library in a feature retrieval manner to directly obtain the corresponding label information of the video.

[0218] The model training method provided by the embodiment of the present disclosure obtains the final result in one step from end to end, without the need for a two-step calculation like the prior art (first obtaining candidate tags by searching for similar videos and then screening the candidate tags) to obtain the final result.

[0219] The model training method provided by the embodiments of the present disclosure aligns (multimodal) video features and label features into the same space by training two feature encoders and two feature dimension transformation models, thus transforming the video label generation problem into a simple feature retrieval task. This method has at least the following two significant advantages:

[0220] 1) Fast inference speed: the corresponding label information can be directly obtained through cross-domain feature retrieval, without the need to use similar videos as a bridge to obtain a candidate label set and then sort and filter the labels;

[0221] 2) Zero-sample support for adding new tags is achieved by simply extracting the tag features of the new tags and adding them to the tag feature library. There is no need to store labeled data or train models, and tags can be easily added or deleted.

[0222] The video tag automatic generation method (machine tagging method) provided by the embodiment of the present disclosure can be applied to various business scenarios (such as short video platforms, news platforms, etc.). Figure 9 As shown in the figure, the video enters the content processing link from the content production link, obtains the corresponding content features through human-computer collaboration, and enters the downstream content distribution link.

[0223] The automatic video tag generation method provided by the embodiment of the present disclosure can provide video content information of different granularities for downstream content distribution links (such as recommendation systems and content operations), thereby improving the efficiency of content distribution and significantly reducing the cost of manual content review.

[0224] The method for automatically generating video tags provided by the embodiments of the present disclosure is simple to implement and is applicable to existing mainstream content understanding algorithm systems. The proposed method can directly obtain tag results in an end-to-end retrieval manner, and can achieve zero-sample support in scenarios where new tags are added, significantly improving the efficiency of machine review and saving the cost of manually annotating data and retraining models.

[0225] Optionally, the video tag generation method provided in the embodiments of the present disclosure can be applied to a terminal or a server. The server can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The terminal can be a smartphone, tablet computer, laptop computer, desktop computer, smart speaker, smart watch, smart home appliance, vehicle terminal, aircraft, etc., but is not limited thereto.

[0226] Optionally, the video tag generation method provided in the embodiments of the present disclosure may be collaboratively performed by multiple computing devices or components with computing capabilities. For example, different computing devices or components may each perform a portion of the steps of each method provided in the embodiments of the present disclosure. For example, the model training function may be performed by one computing device, while the video processing function may be performed by another computing device, etc., but the present invention is not limited thereto.

[0227] That is, the video tag generation method provided by the embodiment of the present disclosure can be implemented based on cloud technology. Among them, cloud technology is a general term for network technology, information technology, integration technology, management platform technology, application technology, etc. based on the application of cloud computing business model, which can form a resource pool and be used on demand, flexibly and conveniently. Cloud computing technology will become an important support. The background services of the technical network system require a large amount of computing and storage resources, such as video websites, picture websites and more portal websites. With the rapid development and application of the Internet industry, every item may have its own identification mark in the future, and all need to be transmitted to the background system for logical processing. Data of different levels will be processed separately, and all types of industry data require strong system backing support, which can only be achieved through cloud computing.

[0228] The video tag generation method provided by the embodiment of the present disclosure can be used for big data processing. Big data refers to a collection of data that cannot be captured, managed, and processed by conventional software tools within a certain time frame. It is a massive, high-growth, and diversified information asset that requires new processing models to have stronger decision-making power, insight discovery, and process optimization capabilities. With the advent of the cloud era, big data has also attracted more and more attention. Big data requires special technologies to effectively process large amounts of data within a tolerable time frame. Technologies suitable for big data include large-scale parallel processing databases, data mining, distributed file systems, distributed databases, cloud computing platforms, the Internet, and scalable storage systems.

[0229] The artificial intelligence (AI) technology involved in the embodiments of this disclosure refers to theories, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI also studies the design principles and implementation methods of various intelligent machines, enabling them to have the capabilities of perception, reasoning, and decision-making. AI technology is an interdisciplinary discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Basic AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, pre-trained model technology, operating / interaction systems, and mechatronics. Pre-trained models, also known as large models or basic models, can be fine-tuned and widely applied to downstream tasks across various AI domains. AI software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0230] Specifically, the embodiments of the present disclosure relate to computer vision technology (CV), speech technology (Speech Technology), natural language processing (NLP) and machine learning (ML) technology, etc.

[0231] Computer vision is the science of making machines "see." Specifically, it refers to machine vision, which uses cameras and computers to replace the human eye in identifying and measuring objects, and then further processes the images to create images more suitable for human observation or transmission to instruments. As a scientific discipline, computer vision studies related theories and technologies, aiming to build artificial intelligence systems that can extract information from images or multidimensional data. Large model technology has brought significant changes to the development of computer vision. Pre-trained models in the field of vision, such as the Swin Transformer, ViT, V-MOE, and MAE, can be fine-tuned to quickly and widely apply to specific downstream tasks. Computer vision technologies generally include image processing, image recognition, image semantic understanding, image retrieval, optical character recognition (OCR), video processing, video semantic understanding, video content / action recognition, 3D object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, and common biometric recognition technologies such as facial recognition and fingerprint recognition.

[0232] Machine learning is a multidisciplinary field that encompasses probability theory, statistics, approximation theory, convex analysis, and algorithmic complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is at the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and self-learning. Pretrained models are the latest development in deep learning, integrating these techniques.

[0233] With the research and advancement of artificial intelligence technology, artificial intelligence technology has been studied and applied in many fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, unmanned driving, autonomous driving, drones, digital twins, virtual humans, robots, artificial intelligence generated content (AIGC), conversational interaction, smart medical care, smart customer service, game AI, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.

[0234] The present disclosure provides a video processing device, such as Figure 10 As shown, the video processing device 100 may include: a video acquisition module 1001, a video feature extraction module 1002, a retrieval module 1003 and an association module 1004, wherein:

[0235] The video acquisition module 1001 is used to acquire the video to be processed;

[0236] The video feature extraction module 1002 is used to extract the target video features of the video to be processed through the trained video feature encoder, and convert the dimensions of the target video features into a predetermined dimension through the trained video feature dimension conversion model;

[0237] The retrieval module 1003 is used to retrieve target tag features similar to target video features of predetermined dimensions in a preset tag feature library, wherein the dimensions of each tag feature in the tag feature library are predetermined dimensions;

[0238] The association module 1004 is used to establish an association relationship between the tag information corresponding to the target tag feature and the video to be processed.

[0239] In an optional embodiment, the device further comprises:

[0240] The tag acquisition module 1005 is used to obtain the description text corresponding to at least one tag information;

[0241] The tag feature extraction module 1006 is used to extract the text features of each description text through the trained text feature encoder, and convert each text feature into a predetermined dimension through the trained text feature dimension conversion model to obtain each tag feature;

[0242] The storage module 1007 is used to store each tag feature in a tag feature library;

[0243] Among them, the video feature encoder, video feature dimension transformation model, text feature encoder and text feature dimension transformation model are jointly trained.

[0244] In an optional embodiment, when the video feature extraction module 1002 is used to extract the target video features of the video to be processed using the trained video feature encoder, it is specifically used to:

[0245] Obtain multiple modal information corresponding to the video to be processed;

[0246] For each modal information, the modal features of the modal information are extracted through the modal feature encoder corresponding to the modal information;

[0247] Multiple modal features are fused to obtain the target video features of the video to be processed.

[0248] In an optional embodiment, the multiple modal features include text features, at least one video frame feature, and at least one audio feature, and the text features include tag words. When the video feature extraction module 1002 is used to fuse the multiple modal features to obtain the target video features of the video to be processed, it is specifically used to:

[0249] For each word feature in the text feature, the word feature is fused with the corresponding position feature and the corresponding type feature to obtain the first feature;

[0250] For each video frame feature, the video frame feature is fused with the corresponding position feature and the corresponding type feature to obtain a second feature;

[0251] For each audio feature, the audio-visual feature is fused with the corresponding position feature and the corresponding type feature to obtain a third feature;

[0252] Each first feature, each second feature, and each third feature are fused to obtain a fused feature corresponding to the marked word, which is used as the target video feature of the video to be processed.

[0253] In an optional implementation, when the tag acquisition module 1005 is used to acquire the description text corresponding to at least one tag information, it is specifically used to:

[0254] Obtain the name, at least one level classification, and description information corresponding to at least one tag information;

[0255] For each tag information, the name, at least one level classification, and description information corresponding to the tag information are combined according to a predetermined format to obtain a description text corresponding to the tag information.

[0256] In an optional embodiment, the device further includes at least one of the following modules:

[0257] The deletion module 1008 is used to obtain the tag information to be deleted, determine the first tag feature corresponding to the tag information to be deleted, and delete the first tag feature in the tag feature library;

[0258] The adding module 1009 is used to obtain the target description text corresponding to the label information to be added, extract the target text features of the target description text through the trained text feature encoder, and convert the target text features into a predetermined dimension through the trained text feature dimension transformation model to obtain a second label feature, and add the second label feature to the label feature library.

[0259] The present disclosure provides a model training device, such as Figure 11 As shown, the model training device 110 may include: a sample acquisition module 1101, a first training module 1102 and a second training module 1103, wherein:

[0260] The sample acquisition module 1101 is used to acquire a training sample, where the training sample includes a sample video and at least one annotated label information corresponding to the sample video;

[0261] The first training module 1102 is configured to extract a first sample video feature of a sample video using a preset video feature encoder, and classify the first sample video feature using at least one preset classification network to obtain at least one sample label information; based on the sample label information and the annotated label information, the preset video feature encoder and the preset at least one classification network are trained until a first preset condition is met, thereby obtaining a pre-trained video feature encoder;

[0262] The second training module 1103 is used to extract the second sample video features of the sample video through a pre-trained video feature encoder, and convert the dimensions of the second sample video features into a predetermined dimension through a preset video feature dimension transformation model; extract the sample label features of each preset label information description text through a preset text feature encoder, and convert each sample label feature into a predetermined dimension through a preset text feature dimension transformation model; based on the sample label features of the predetermined dimension, the second sample video features of the predetermined dimension and the annotated label information, train the pre-trained video feature encoder, the preset video feature dimension transformation model, the preset text feature encoder and the preset text feature dimension transformation model until the first preset condition is met, thereby obtaining a trained video feature encoder, a trained video feature dimension transformation model, a trained text feature encoder and a trained text feature dimension transformation model.

[0263] In an optional embodiment, the at least one classification network includes a target classification network corresponding to the multi-label task. When the first training module 1102 is used to classify the first sample video features through the target classification network to obtain sample label information, it is specifically used to:

[0264] Classifying the first sample video features through the target classification network to obtain sample label information of whether the sample video is relevant for each label information;

[0265] The first training module is specifically used to train a preset video feature encoder and target classification network based on sample label information and annotation label information:

[0266] Based on the sample label information of whether each label information is relevant for the sample video, and the annotated label information corresponding to multiple label information related to the sample video, a preset video feature encoder and target classification network are trained.

[0267] In an optional embodiment, the second training module 1103, when used to train a pre-trained video feature encoder, a preset video feature dimension transformation model, a preset text feature encoder, and a preset text feature dimension transformation model based on the sample label feature of the predetermined dimension, the second sample video feature of the predetermined dimension, and the annotated label information, is specifically used to:

[0268] Calculating the similarity between the sample label feature of the predetermined dimension and the second sample video feature of the predetermined dimension;

[0269] Based on the similarity and the labeled label information of whether the label information corresponding to the sample label feature is related to the sample video, the pre-trained video feature encoder, the preset video feature dimension transformation model, the preset text feature encoder and the preset text feature dimension transformation model are trained.

[0270] In an optional embodiment, the second training module 1103 is used to train the pre-trained video feature encoder, the preset video feature dimension transformation model, the preset text feature encoder, and the preset text feature dimension transformation model based on the similarity and the label information of whether the label information corresponding to the sample label feature is related to the sample video, specifically for:

[0271] The pre-trained video feature encoder, the preset video feature dimensional transformation model, the preset text feature encoder and the preset text feature dimensional transformation model are trained so that the similarity is maximized for the case where the label information corresponding to the sample label feature is related to the sample video, and the similarity is less than a predetermined value for the case where the label information corresponding to the sample label feature is not related to the sample video.

[0272] The apparatus of the embodiments of the present disclosure can execute the methods provided by the embodiments of the present disclosure, and their implementation principles are similar and have corresponding technical effects. The actions performed by each module in the apparatus of each embodiment of the present disclosure correspond to the steps in the methods of each embodiment of the present disclosure. For detailed functional descriptions of each module of the apparatus, please refer to the descriptions of the corresponding methods shown above, and will not be repeated here.

[0273] An embodiment of the present disclosure provides an electronic device, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the method provided in any optional embodiment of the present disclosure.

[0274] In an alternative embodiment, an electronic device is provided, such as Figure 12 As shown, Figure 12 The electronic device 4000 shown includes: a processor 4001 and a memory 4003. The processor 4001 and the memory 4003 are connected, for example, via a bus 4002. Optionally, the electronic device 4000 may further include a transceiver 4004, which may be used for data exchange between the electronic device and other electronic devices, such as data transmission and / or data reception. It should be noted that in actual applications, the number of transceivers 4004 is not limited to one, and the structure of the electronic device 4000 does not constitute a limitation on the embodiments of the present disclosure.

[0275] Processor 4001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It may implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the present disclosure. Processor 4001 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, and the like.

[0276] Bus 4002 may include a path for transmitting information between the above components. Bus 4002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus. Bus 4002 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 12 Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.

[0277] The memory 4003 can be a ROM (Read Only Memory) or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory) or other types of dynamic storage devices that can store information and instructions, or an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory) or other optical disk storage, optical disk storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, other magnetic storage devices, or any other medium that can be used to carry or store computer programs and can be read by a computer, without limitation here.

[0278] The memory 4003 is used to store the computer program for executing the embodiments of the present disclosure, and the execution is controlled by the processor 4001. The processor 4001 is used to execute the computer program stored in the memory 4003 to implement the steps shown in the above method embodiments.

[0279] An embodiment of the present disclosure provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps and corresponding contents of the aforementioned method embodiment can be implemented.

[0280] The embodiments of the present disclosure further provide a computer program product, including a computer program, which can implement the steps and corresponding contents of the aforementioned method embodiments when executed by a processor.

[0281] It should be understood that, although the flowcharts of the embodiments of the present disclosure indicate the various operation steps by arrows, the order of implementation of these steps is not limited to the order indicated by the arrows. Unless otherwise clearly stated herein, in some implementation scenarios of the embodiments of the present disclosure, the implementation steps in each flowchart can be performed in other orders as required. In addition, some or all of the steps in each flowchart may include multiple sub-steps or multiple stages based on the actual implementation scenario. Some or all of these sub-steps or stages can be executed at the same time, and each sub-step or stage in these sub-steps or stages can also be executed at different times. In scenarios where the execution times are different, the order of execution of these sub-steps or stages can be flexibly configured as required, and the embodiments of the present disclosure do not limit this.

[0282] The above are only optional implementation methods for some implementation scenarios of the present disclosure. It should be pointed out that for ordinary technicians in this technical field, without departing from the technical concept of the solution of the present disclosure, other similar implementation methods based on the technical ideas of the present disclosure are also within the protection scope of the embodiments of the present disclosure.

Claims

1. A video processing method, characterized in that: include: Get the video to be processed; Extracting target video features of the video to be processed by a trained video feature encoder, and converting the dimensions of the target video features into predetermined dimensions by a trained video feature dimension transformation model; Retrieving a target tag feature similar to the target video feature of the predetermined dimension from a preset tag feature library, wherein the dimension of each tag feature in the tag feature library is the predetermined dimension; Establish an association relationship between the tag information corresponding to the target tag feature and the video to be processed.

2. The video processing method according to claim 1, wherein: Also includes: Obtain the description text corresponding to at least one tag information; Extracting the text features of each of the description texts through a trained text feature encoder, and converting each of the text features into the predetermined dimensions through a trained text feature dimension conversion model to obtain each label feature; Storing each of the tag features in the tag feature library; Wherein, the video feature encoder, the video feature dimensionality transformation model, the text feature encoder and the text feature dimensionality transformation model are jointly trained.

3. The video processing method according to claim 1, wherein: The step of extracting target video features of the video to be processed by using a trained video feature encoder includes: Obtaining multiple modal information corresponding to the video to be processed; For each modal information, the modal features of the modal information are extracted through the modal feature encoder corresponding to the modal information; Multiple modal features are fused to obtain target video features of the video to be processed.

4. The video processing method according to claim 1, wherein: The multiple modality features include text features, at least one video frame feature, and at least one audio feature, wherein the text features include tag words; The fusion of multiple modal features to obtain target video features of the video to be processed includes: For each word-unit feature in the text features, the word-unit feature is fused with the corresponding position feature and the corresponding type feature to obtain a first feature; For each video frame feature, the video frame feature is fused with the corresponding position feature and the corresponding type feature to obtain a second feature; For each audio feature, the audio-visual feature is fused with the corresponding position feature and the corresponding type feature to obtain a third feature; Each of the first features, each of the second features, and each of the third features are fused to obtain a fused feature corresponding to the marked word, which is used as a target video feature of the video to be processed.

5. The video processing method according to claim 2, wherein: The obtaining of the description text corresponding to at least one tag information includes: Obtain the name, at least one level classification, and description information corresponding to at least one tag information; For each tag information, the name, at least one level classification, and description information corresponding to the tag information are combined according to a predetermined format to obtain a description text corresponding to the tag information.

6. The video processing method according to claim 2, wherein: Also include at least one of the following: Acquire tag information to be deleted, determine a first tag feature corresponding to the tag information to be deleted, and delete the first tag feature in the tag feature library; Obtain the target description text corresponding to the label information to be added, extract the target text features of the target description text through the trained text feature encoder, and convert the target text features into the predetermined dimension through the trained text feature dimension transformation model to obtain a second label feature, and add the second label feature to the label feature library.

7. A model training method, characterized in that: include: Acquire a training sample, where the training sample includes a sample video and at least one annotated label information corresponding to the sample video; Extracting a first sample video feature of the sample video through a preset video feature encoder, and classifying the first sample video feature through at least one preset classification network to obtain at least one sample label information; training the preset video feature encoder and the at least one preset classification network based on the sample label information and the annotated label information until a first preset condition is satisfied, thereby obtaining a pre-trained video feature encoder; The second sample video feature of the sample video is extracted through the pre-trained video feature encoder, and the dimension of the second sample video feature is converted into a predetermined dimension through a preset video feature dimension transformation model; the sample label feature of each preset label information description text is extracted through a preset text feature dimension transformation model, and each sample label feature is converted into the predetermined dimension through a preset text feature dimension transformation model; based on the sample label feature of the predetermined dimension, the second sample video feature of the predetermined dimension and the annotated label information, the pre-trained video feature encoder, the preset video feature dimension transformation model, the preset text feature encoder and the preset text feature dimension transformation model are trained until the first preset condition is met, thereby obtaining a trained video feature encoder, a trained video feature dimension transformation model, a trained text feature encoder and a trained text feature dimension transformation model.

8. The video processing method according to claim 7, wherein: The at least one classification network includes a target classification network corresponding to the multi-label task, and classifies the first sample video features through the target classification network to obtain sample label information, including: Classifying the first sample video features through a target classification network to obtain sample label information of whether the sample video is relevant for each label information; Training the preset video feature encoder and the target classification network based on the sample label information and the annotated label information includes: The preset video feature encoder and the target classification network are trained based on the sample label information of whether each label information of the sample video is relevant, and the annotated label information corresponding to the multiple label information related to the sample video.

9. The video processing method according to claim 7, wherein: The training of the pre-trained video feature encoder, the preset video feature dimension transformation model, the preset text feature encoder, and the preset text feature dimension transformation model based on the sample label feature of the predetermined dimension, the second sample video feature of the predetermined dimension, and the annotated label information includes: Calculating the similarity between the sample label feature of the predetermined dimension and the second sample video feature of the predetermined dimension; Based on the similarity and the annotated label information of whether the label information corresponding to the sample label feature is related to the sample video, the pre-trained video feature encoder, the preset video feature dimension transformation model, the preset text feature encoder and the preset text feature dimension transformation model are trained.

10. The video processing method according to claim 9, characterized in that: The method includes training the pre-trained video feature encoder, the preset video feature dimension transformation model, the preset text feature encoder, and the preset text feature dimension transformation model based on the similarity and the annotated label information indicating whether the label information corresponding to the sample label feature is related to the sample video, including: The pre-trained video feature encoder, the preset video feature dimensional transformation model, the preset text feature encoder and the preset text feature dimensional transformation model are trained so that the similarity is maximized for the case where the label information corresponding to the sample label feature is related to the sample video, and the similarity is less than a predetermined value for the case where the label information corresponding to the sample label feature is not related to the sample video.

11. A video processing device, characterized in that: include: Video acquisition module, used to acquire the video to be processed; A video feature extraction module is used to extract target video features of the video to be processed using a trained video feature encoder, and convert the dimensions of the target video features into predetermined dimensions using a trained video feature dimension conversion model; A retrieval module is used to retrieve a target tag feature similar to the target video feature of the predetermined dimension from a preset tag feature library, wherein the dimension of each tag feature in the tag feature library is the predetermined dimension; The association module is used to establish an association relationship between the tag information corresponding to the target tag feature and the video to be processed.

12. A model training device, characterized in that: include: A sample acquisition module is used to acquire a training sample, wherein the training sample includes a sample video and at least one annotated label information corresponding to the sample video; A first training module is configured to extract a first sample video feature of the sample video using a preset video feature encoder, and classify the first sample video feature using at least one preset classification network to obtain at least one sample label information; based on the sample label information and the annotated label information, train the preset video feature encoder and the at least one preset classification network until a first preset condition is met, thereby obtaining a pre-trained video feature encoder; The second training module is used to extract the second sample video features of the sample video through the pre-trained video feature encoder, and convert the dimension of the second sample video features into a predetermined dimension through a preset video feature dimension transformation model; extract the sample label features of each preset label information description text through a preset text feature encoder, and convert each of the sample label features into the predetermined dimension through a preset text feature dimension transformation model; based on the sample label features of the predetermined dimension, the second sample video features of the predetermined dimension and the annotated label information, train the pre-trained video feature encoder, the preset video feature dimension transformation model, the preset text feature encoder and the preset text feature dimension transformation model until the first preset condition is met, thereby obtaining a trained video feature encoder, a trained video feature dimension transformation model, a trained text feature encoder and a trained text feature dimension transformation model.

13. An electronic device comprising a memory, a processor, and a computer program stored in the memory, wherein: The processor executes the computer program to implement the method according to any one of claims 1 to 10.

14. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 10 is implemented.

15. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 10 is implemented.