Video label identification model training method, video label identification method and device

In video tag recognition, the historical videos in the training sample set and the visual features of the newly added videos, combined with uniform sampling and feature compression methods, the problems of information loss and resource consumption in the prior art are solved, and the accuracy and efficiency of video tag recognition are improved.

CN120032198APending Publication Date: 2025-05-23TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202311583322.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-23
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

In the video tag recognition, sparse sampling or truncation methods lead to information loss and low accuracy; while intensive sampling increases the time and resource consumption of model training and inference.

Method used

By obtaining the training sample set, including the historical video set and the added video, its visual features are determined for each training sample. The visual features are obtained by compressing and splicing of the added video features and historical video features obtained by uniform sampling, and the video tag recognition model is trained based on these features.

Benefits of technology

It reduces information loss, improves the accuracy of video tag recognition, and reduces the time and resource consumption of model training and inference.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032198A_ABST
    Figure CN120032198A_ABST
Patent Text Reader

Abstract

Provided are a video tag identification model training method and device, and a video tag identification method and device, the method comprising: obtaining a training sample set, each training sample comprising a sample video set, text information of the sample video set, and a first tag and a second tag of the sample video set, the sample video set comprising a historical video set and a newly added video, determining a visual feature of a sample video set of the training sample, the visual feature of the sample video set being obtained by splicing a first visual feature and a second visual feature, the first visual feature being a visual feature obtained by uniformly sampling a newly added video, and the second visual feature being a visual feature obtained by uniformly sampling a newly added video; the second visual feature is obtained by performing feature compression on the visual feature of the historical video set and the first visual feature, and training a video tag identification model according to the visual feature of the sample video set, the text information of the sample video set, the first tag and the second tag of the sample video set and a first preset normal form instruction, and the training stopping condition is met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of artificial intelligence technology, and in particular to a video label recognition model training method, a video label recognition method and a device. Background Art

[0002] With the development of network technology, multimedia data such as video has become the main body of big data. Video tags are an important part of video content features. Video tags can provide different granularity video content features for downstream content distribution links, improve the efficiency of content distribution, and reduce manual review costs.

[0003] At present, when a video label is identified for a video collection including multiple short videos or medium videos, the video collection is first sampled to obtain a video frame sequence, and the video frame sequence and the text information of the video collection are input into a multimodal model to extract and fuse visual features to obtain a label for the video collection. In the prior art, in the process of sampling the video frames of a video collection to obtain a video frame sequence, when the content form of the video collection is different and the number of video frames of the video collection is large, a sparse sampling or truncation method is used to obtain the video frame sequence, or a dense sampling method is used to obtain the video frame sequence.

[0004] However, the video frame sequence obtained by sparse sampling or truncation cannot cover most of the content of the video collection, which will lead to information loss and low accuracy of video label recognition. The use of dense sampling will increase the time and memory / video memory consumption during training and inference of multimodal models. Summary of the invention

[0005] The embodiments of the present application provide a video tag recognition model training method, a video tag recognition method and a device, which can reduce information loss and improve the accuracy of video tag recognition.

[0006] In a first aspect, an embodiment of the present application provides a video tag recognition model training method, comprising:

[0007] Acquire a training sample set, each training sample includes a sample video set, text information of the sample video set, a first label and a second label of the sample video set, the sample video set includes a historical video set and a newly added video, the second label is a randomly transformed label of the label of the historical video set, the videos in the sample video set are sorted according to the chronological order of creation time, the last K videos in the sample video set are the newly added videos, and the videos in the sample video set other than the K videos constitute the historical video set, and K is a positive integer;

[0008] For each training sample, determining a visual feature of a sample video set of the training sample, wherein the visual feature of the sample video set is obtained by concatenating a first visual feature and a second visual feature, wherein the first visual feature is a visual feature obtained by uniformly sampling a newly added video in the sample video set, and the second visual feature is obtained by feature compression of a visual feature of a historical video set in the sample video set and the first visual feature;

[0009] According to the visual features of the sample video set, the text information of the sample video set, the first label and the second label of the sample video set and the first preset paradigm instruction, the video label recognition model is trained until the training stop condition is met to obtain a trained video label recognition model, wherein the video label recognition model includes a visual feature encoder, a pre-trained alignment module and a recognition model.

[0010] In a second aspect, an embodiment of the present application provides a video tag recognition method, including:

[0011] Acquire a video set to be identified and text information of the video set to be identified, wherein the video set to be identified includes a currently added video and a historical video set;

[0012] Obtaining predicted labels of the historical video set and visual features of the historical video set;

[0013] Determining visual features of the current newly added video;

[0014] Determining visual features of the to-be-identified video set according to the visual features of the currently newly added video and the visual features of the historical video set;

[0015] Determine the predicted labels of the video set to be identified based on the visual features of the video set to be identified, the text information of the video set to be identified, the predicted labels of the historical video set, a first preset paradigm instruction and a trained video label recognition model, wherein the video label recognition model is trained according to the method described in the first aspect, and the video label recognition model includes a visual feature encoder, an alignment module and a recognition model.

[0016] In a third aspect, an embodiment of the present application provides a video tag recognition model training device, comprising:

[0017] An acquisition module is used to acquire a training sample set, each training sample includes a sample video set, text information of the sample video set, a first label and a second label of the sample video set, the sample video set includes a historical video set and a newly added video, the second label is a randomly transformed label of the label of the historical video set, the videos in the sample video set are sorted according to the chronological order of creation time, the last K videos in the sample video set are the newly added videos, and the videos in the sample video set other than the K videos constitute the historical video set, and K is a positive integer;

[0018] A processing module, configured to determine, for each training sample, a visual feature of a sample video set of the training sample, wherein the visual feature of the sample video set is obtained by concatenating a first visual feature and a second visual feature, wherein the first visual feature is a visual feature obtained by uniformly sampling a newly added video in the sample video set, and the second visual feature is obtained by feature compression of a visual feature of a historical video set in the sample video set and the first visual feature;

[0019] A training module is used to train a video label recognition model according to the visual features of the sample video set, the text information of the sample video set, the first label and the second label of the sample video set, and a first preset paradigm instruction until a stop training condition is met to obtain a trained video label recognition model, wherein the video label recognition model includes a visual feature encoder, a pre-trained alignment module, and a recognition model.

[0020] In a fourth aspect, an embodiment of the present application provides a video tag recognition device, including:

[0021] An acquisition module, used to acquire a video set to be identified and text information of the video set to be identified, wherein the video set to be identified includes a currently added video and a historical video set;

[0022] The acquisition module is also used to: acquire the predicted labels of the historical video set and the visual features of the historical video set;

[0023] A processing module, used to: determine the visual features of the current newly added video;

[0024] Determining visual features of the to-be-identified video set according to the visual features of the currently newly added video and the visual features of the historical video set;

[0025] Determine the predicted labels of the video set to be identified based on the visual features of the video set to be identified, the text information of the video set to be identified, the predicted labels of the historical video set, a first preset paradigm instruction and a trained video label recognition model, wherein the video label recognition model is trained according to the method described in the first aspect, and the video label recognition model includes a visual feature encoder, an alignment module and a recognition model.

[0026] In a fifth aspect, an embodiment of the present application provides a computer device, comprising: a processor and a memory, the memory being used to store a computer program, the processor being used to call and run the computer program stored in the memory to execute the method of the first aspect or the second aspect.

[0027] In a sixth aspect, an embodiment of the present application provides a computer-readable storage medium, comprising instructions, which, when executed on a computer program, enable the computer to execute a method as in the first aspect or the second aspect.

[0028] In a seventh aspect, an embodiment of the present application provides a computer program product comprising instructions, which, when executed on a computer, causes the computer to execute the method of the first aspect or the second aspect.

[0029] In summary, in an embodiment of the present application, a training sample set is obtained, each training sample includes a sample video set, text information of the sample video set, a first label and a second label of the sample video set, the sample video set includes a historical video set and a newly added video, the second label is a randomly transformed label of the label of the historical video set, the videos in the sample video set are sorted in chronological order of creation time, the last K videos in the sample video set are newly added videos, and the videos other than K videos in the sample video set constitute a historical video set, then for each training sample, the visual features of the sample video set of the training sample are determined, the visual features of the sample video set are obtained by splicing the first visual features and the second visual features, the first visual features are visual features obtained by uniformly sampling the newly added videos in the sample video set, and the second visual features are obtained by feature compression of the visual features of the historical video set and the first visual features in the sample video set, then according to the visual features of the sample video set, the text information of the sample video set, the first label and the second label of the sample video set, and the first preset paradigm instruction, the video label recognition model is trained until the stop training condition is met to obtain a trained video label recognition model. Since the sample video set includes a historical video set and a newly added video when training the video label recognition model, when determining the visual features of the sample video set of the training sample, the visual features of the sample video set are obtained by splicing the first visual feature and the second visual feature, the first visual feature is the visual feature obtained by uniformly sampling the newly added video in the sample video set, and the second visual feature is obtained by feature compression of the visual features of the historical video set in the sample video set and the first visual feature. Therefore, when the trained video label recognition model obtains the visual features of the video set to be recognized during the process of video label recognition, it can also use uniform sampling for the newly added video and feature compression for the visual features of the historical video set, thereby reducing the information loss caused by sparse sampling or truncation, and can also reduce the reasoning time and resource consumption caused by dense sampling, thereby improving the accuracy of video label recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 A schematic diagram of a video tag recognition model training method and an implementation scenario of the video tag recognition method provided in an embodiment of the present application;

[0031] Figure 2 A schematic diagram of a flow chart of a video tag recognition method application provided in an embodiment of the present application;

[0032] Figure 3 A flowchart of a video tag recognition model training method provided in an embodiment of the present application;

[0033] Figure 4A schematic diagram of a method for obtaining text unit features provided in an embodiment of the present application;

[0034] Figure 5 A schematic diagram of a feature fusion process provided in an embodiment of the present application;

[0035] Figure 6 A flowchart of a video tag recognition model training method provided in an embodiment of the present application;

[0036] Figure 7 A schematic diagram of a video tag recognition model training method provided in an embodiment of the present application;

[0037] Figure 8 A flowchart of a video tag recognition method provided in an embodiment of the present application;

[0038] Fig. 9 A schematic diagram of a video tag recognition method provided in an embodiment of the present application;

[0039] Fig.10 A schematic diagram of the structure of a video tag recognition model training device provided in an embodiment of the present application;

[0040] Fig.11 A schematic diagram of the structure of a video tag recognition device provided in an embodiment of the present application;

[0041] Fig.12 It is a schematic block diagram of a computer device 300 provided in an embodiment of the present application. DETAILED DESCRIPTION

[0042] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the embodiments of the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the embodiments of the present application.

[0043] It should be noted that the terms "first", "second", etc. in the specification and claims of the embodiments of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product, or server comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products, or devices.

[0044] The embodiments of the present application may involve artificial intelligence technology.

[0045] Among them, artificial intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines so that machines have the functions of perception, reasoning and decision-making.

[0046] It should be understood that artificial intelligence technology is a comprehensive discipline that covers a wide range of fields, including both hardware-level and software-level technologies. Basic artificial intelligence technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operating / interactive systems, and mechatronics. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0047] With the research and advancement of artificial intelligence technology, artificial intelligence technology has been studied and applied in many fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, driverless cars, autonomous driving, drones, digital twins, virtual humans, robots, artificial intelligence generated content (AIGC), conversational interaction, smart medical care, smart customer service, game AI, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.

[0048] The embodiments of the present application may involve computer vision (CV) technology in artificial intelligence technology. Computer vision is a science that studies how to make machines "see". To put it more concretely, it refers to machine vision such as using cameras and computers to replace human eyes to identify, monitor and measure targets, and further perform graphic processing to make computer processing into images that are more suitable for human eye observation or transmission to instrument detection. As a scientific discipline, computer vision studies related theories and technologies, and attempts to establish an artificial intelligence system that can obtain information from images or multidimensional data. Computer vision technology generally includes image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous positioning and map construction, and other technologies, as well as common biometric recognition technologies such as face recognition and fingerprint recognition.

[0049] The embodiments of the present application may involve natural language processing (NLP) technology in artificial intelligence technology. Natural language processing is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can realize effective communication between people and computers using natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, research in this field will involve natural language, that is, the language used by people in daily life, so it is closely related to the study of linguistics. Natural language processing technology generally includes technologies such as text processing, semantic understanding, machine translation, robot question answering, and knowledge graphs.

[0050] The training method of the video tag prediction model provided in the embodiment of the present application and the video tag prediction method mainly relate to the computer vision technology (Computer Vision, CV) of artificial intelligence. Computer vision is a science that studies how to make machines "see". To put it more concretely, it refers to machine vision such as using cameras and computers to replace human eyes to identify and measure targets, and further do graphic processing to make computer processing become images more suitable for human eye observation or transmission to instrument detection. As a scientific discipline, computer vision studies related theories and technologies, and attempts to establish an artificial intelligence system that can obtain information from images or multidimensional data. Computer vision technology generally includes image processing, image segmentation, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, synchronous positioning and map construction, and other technologies, as well as common biometric recognition technologies such as face recognition and fingerprint recognition.

[0051] The video label prediction model training method and the video label prediction method provided in the embodiments of the present application mainly relate to the video semantic understanding technology (VSU) in the field of computer vision technology.

[0052] First, the relevant terms involved in the embodiments of the present application are introduced.

[0053] Pre-training model, also known as cornerstone model or big model, refers to a deep neural network (DNN) with large parameters. It is trained on massive unlabeled data. The function approximation ability of large-parameter DNN is used to enable PTM to extract common features from the data. After fine tuning, parameter efficient fine tuning (PEFT), prompt-tuning and other technologies, it is suitable for downstream tasks. Therefore, the pre-training model can achieve ideal results in few-shot or zero-shot scenarios. PTM can be divided into language models (ELMO, BERT, GPT), visual models (swin-transformer, ViT, V-MOE), speech models (VALL-E), multimodal models (ViBERT, CLIP, Flamingo, Gato), etc. according to the data modality processed. Among them, the multimodal model refers to a model that establishes two or more data modality feature representations. The pre-training model is an important tool for outputting artificial intelligence generated content (AIGC), and can also be used as a general interface to connect multiple specific task models.

[0054] Large Language Model (LLM): It learns the statistical laws and semantic information of language by training a large amount of text data, so as to predict the next word or sentence. LLM has a wide range of applications in the field of natural language processing (NLP), such as machine translation, speech recognition, text generation, etc. LLM can adopt different training strategies and model structures, such as pre-training and fine-tuning. Pre-training refers to large-scale training of the model under unsupervised or self-supervised conditions, so that it has a general understanding and expression ability of language data. Fine-tuning is to optimize the model for specific application scenarios or tasks through supervised learning based on the pre-trained model. LLM has achieved remarkable results in tasks such as language understanding, generation and translation, and its performance and application scope are also expanding with the continuous expansion of model scale and the increase of training data.

[0055] Multimodal large language model (MLLM): Based on LLM, it integrates media data of other modalities (such as images, videos, audio, etc.), so that the model can process information of different modalities at the same time, better understand and express semantics, and thus improve the effect and accuracy of the application.

[0056] Optical Character Recognition: (OCR) is a technology that detects and recognizes text content from images.

[0057] Automatic Speech Recognition (ASR) is a technology that converts human speech into text.

[0058] Token: In LLM, a token represents the smallest unit of meaning that the model can understand and generate, and is the basic unit used by the model when processing text. Depending on the specific tokenization scheme used, a token can represent a word, part of a word, or even just a character. Tokens are assigned numerical values ​​or identifiers and arranged in sequences or vectors, and are input or output from the model. They are the language building blocks of the model.

[0059] In the related art, when obtaining a video frame sequence of a video collection with a large number of videos, sparse sampling or truncation is used, which will lead to information loss and thus low accuracy in video label recognition. Dense sampling will increase the time and memory / video memory consumption during training and inference of the multimodal model.

[0060] To solve this problem, the embodiment of the present application divides the video set to be identified into a current newly added video and a historical video set when performing video label identification on the video set to be identified, uniformly samples the current newly added video and obtains the visual features of the current newly added video, compresses the visual features of the historical video set, and then obtains the visual features of the video set to be identified based on the visual features of the current newly added video and the visual features of the historical video set, and then determines the predicted labels of the video set to be identified based on the visual features of the video set to be identified, the text information of the video set to be identified, the predicted labels of the historical video set, the first preset paradigm instruction and the trained video label identification model. Since the number of newly added videos is less than the number of videos included in the entire video set to be identified when obtaining the visual features of the video set to be identified, uniform sampling can be used for the newly added videos, and feature compression can be performed on the visual features of the historical video set, thereby reducing the information loss caused by sparse sampling or truncation, and reducing the reasoning time and resource consumption caused by dense sampling, thereby improving the accuracy of video label identification.

[0061] The embodiments of the present application can be applied to various scenarios, including but not limited to video tag prediction scenarios, video classification scenarios, video recommendation scenarios, and the like. For example, for a large number of videos on a video website or video application, the method provided by the embodiments of the present application can be used to parse the video content offline or online to obtain the video tags corresponding to the videos. Based on the video tags corresponding to the videos, the large number of videos can be classified. For another example, based on the video tags corresponding to the videos, videos of interest can be recommended to users.

[0062] It should be noted that the application scenarios described above are only used to illustrate the embodiments of the present application and are not intended to be limiting. In specific implementation, the technical solutions provided in the embodiments of the present application can be flexibly applied according to actual needs.

[0063] For example, Figure 1 A video tag recognition model training method and a video tag recognition method implementation scenario diagram provided in the embodiment of the present application are as follows: Figure 1 As shown, the implementation scenario of the embodiment of the present application involves a server 1 and a terminal device 2, and the terminal device 2 can communicate data with the server 1 through a communication network. The communication network can be a wireless or wired network such as an intranet, the Internet, the Global System of Mobilecommunication (GSM), Wideband Code Division Multiple Access (WCDMA), 4G network, 5G network, Bluetooth, Wi-Fi, and a call network.

[0064] Among them, in some possible implementations, the terminal device 2 refers to a type of device that has rich human-computer interaction methods, has the ability to access the Internet, is usually equipped with various operating systems, and has strong processing capabilities. The terminal device can be a terminal device such as a smart phone, a tablet computer, a portable laptop, a desktop computer, or a phone watch, etc., but is not limited thereto. Optionally, in the embodiment of the present application, various applications are installed in the terminal device 2, such as video applications, news applications, etc.

[0065] Among them, in some possible implementations, the terminal device 2 includes but is not limited to mobile phones, computers, intelligent voice interaction devices, smart home appliances, car terminals, etc.

[0066] Figure 1The server 1 in the example can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. This embodiment of the application does not limit this. In this embodiment of the application, the server 1 can be a background server of an application installed in the terminal device 2.

[0067] In some possible implementations, Figure 1 One terminal device and one server are shown as an example, but other numbers of terminal devices and servers may actually be included, and the embodiments of the present application do not limit this.

[0068] In some embodiments, when video tag recognition is required, the server 1 may use the method provided in the embodiment of the present application to first train a video tag recognition model, which may be: obtaining a training sample set, each training sample including a sample video set, text information of the sample video set, a first label and a second label of the sample video set, the sample video set including a historical video set and a newly added video, the second label being a randomly transformed label of the label of the historical video set, the videos in the sample video set being sorted in chronological order of creation time, the last K videos in the sample video set being newly added videos, and the videos other than the K videos in the sample video set forming a historical video set; for each training sample, determining a sample video of the training sample; The visual features of the sample video set are obtained by concatenating the first visual feature and the second visual feature, the first visual feature is the visual feature obtained by uniformly sampling the newly added videos in the sample video set, and the second visual feature is obtained by feature compression of the visual features of the historical video set in the sample video set and the first visual feature; according to the visual features of the sample video set, the text information of the sample video set, the first label and the second label of the sample video set and the first preset paradigm instruction, the video label recognition model is trained until the stopping training condition is met to obtain the trained video label recognition model, the video label recognition model includes a visual feature encoder, a pre-trained alignment module and a recognition model.

[0069] After obtaining the trained video label recognition model, video label recognition can be performed, which can be specifically as follows: obtaining a video set to be recognized and text information of the video set to be recognized, the video set to be recognized including a currently added video and a historical video set, obtaining predicted labels of the historical video set and visual features of the historical video set, determining the visual features of the currently added video, determining the visual features of the video set to be recognized based on the visual features of the currently added video and the visual features of the historical video set, and determining the predicted labels of the video set to be recognized based on the visual features of the video set to be recognized, the text information of the video set to be recognized, the predicted labels of the historical video set, the first preset paradigm instruction and the trained video label recognition model.

[0070] It is understandable that in the specific implementation of the present application, related data such as user information (such as sample video collection) is involved. When the method of the embodiment of the present application is applied to a specific product or technology, it is necessary to obtain user permission or consent, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards of relevant countries and regions.

[0071] The embodiment of the present application can be applied to video collection label recognition scenarios. Figure 2 A schematic diagram of a process flow of a video tag recognition method application provided in an embodiment of the present application, the process flow of the application is as follows Figure 2 As shown, the video set is passed through the video tag recognition model to obtain the predicted tags of the video set, and the predicted tags of the video set are applied in downstream applications (such as recommendation systems or content operation systems, etc.). If a new video is added to a video set that has been processed (has obtained the predicted tags), an updated video set is obtained, which can trigger a tag prediction, that is, the updated video set is passed through the video tag recognition model to obtain the predicted tags of the updated video set, and the predicted tags of the updated video set are applied in downstream applications. The trigger here can also be that the system determines that a new video has been added to a video set that has been processed, which automatically triggers a tag prediction. The above tag prediction process does not require human participation, the efficiency of video tag prediction is high, and the labor cost is low.

[0072] The technical solution of the embodiment of the present application will be described in detail below:

[0073] Figure 3 A flowchart of a video tag recognition model training method provided in an embodiment of the present application. The execution subject of the embodiment of the present application is a device with a model training function, and the model training device can be, for example, a server, such as Figure 3 As shown, the method may include:

[0074] S101. Obtain a training sample set, where each training sample includes a sample video set, text information of the sample video set, a first label and a second label of the sample video set, the sample video set includes a historical video set and a newly added video, the second label is a randomly transformed label of the label of the historical video set, the videos in the sample video set are sorted in chronological order according to the creation time, the last K videos in the sample video set are newly added videos, and the videos other than K videos in the sample video set constitute the historical video set, where K is a positive integer.

[0075] Specifically, the video tag recognition model training may be an iterative training. In any iterative training process, the operations of S101 to S103 may be performed until a training stop condition is met to obtain a trained video tag recognition model.

[0076] Obtain a training sample set. During any iterative training process, multiple (e.g., I, where I is a positive integer) training samples may be randomly selected from the original training sample set to form a training sample set. Each training sample includes a sample video set, text information of the sample video set, a first label and a second label of the sample video set. The sample video set includes a historical video set and a newly added video. The second label is a randomly transformed label of the label of the historical video set. The first label is a real label of the sample video set (which may be obtained by manual review and annotation). The first label may include one or more labels. The label of the historical video set is also the real label of the historical video set. The label of the historical video set may include one or more labels. The label of the historical video set is randomly transformed, which may be a random addition of a label or a random deletion of a label, or a replacement of one or more labels therein to obtain a second label. The sample video set includes a historical video set and a newly added video. The videos in the sample video set are sorted according to the creation time. The last K videos in the sample video set are newly added videos. The videos other than the K videos in the sample video set constitute the historical video set. K is a preset positive integer. Taking K=1 as an example, for example, a sample video set includes video 1, video 2 and video 3 arranged in chronological order of creation time, wherein video 3 is a newly added video of the sample video set, and video 1 and video 2 constitute the historical video set of the sample video set. Taking K=2 as an example, for example, a sample video set includes video 1, video 2, video 3, video 4 and video 5, wherein video 4 and video 5 are newly added videos of the sample video set, and video 1, video 2 and video 3 constitute the historical video set of the sample video set.

[0077] Optionally, in this embodiment, an original training sample set is constructed in advance, and the training samples in the original training sample set can come from multiple video sets. For example, the original training sample set includes 5 training samples, and these 5 training samples come from two video sets. Video set 1 includes video vid1 at t1, and video vid1 constitutes training sample 1. Video set 1 includes video vid1 and video vid2 at t2, and video vid1 and video vid2 constitute training sample 2. Video set 1 includes video vid1, video vid2, and video vid3 at t3, and video vid1, video vid2, and video vid3 constitute training sample 3. Video set 2 includes video vd1 at t1', and video vd1 constitutes training sample 4. Video set 2 includes video vd1 and video vd2 at t2', and video vd1 and video vd2 constitute training sample 5. The label of each training sample can be manually reviewed and annotated, and the label of the historical video set in each training sample can also be manually reviewed and annotated.

[0078] S102. For each training sample, determine the visual features of the sample video set of the training sample, where the visual features of the sample video set are obtained by concatenating a first visual feature and a second visual feature. The first visual feature is a visual feature obtained by uniformly sampling a newly added video in the sample video set, and the second visual feature is obtained by feature compression of a visual feature of a historical video set in the sample video set and the first visual feature.

[0079] Specifically, after obtaining the training sample set, for each training sample, the visual features of the sample video set of the training sample are determined, and the visual features of the sample video set of the training sample are obtained by splicing the first visual feature and the second visual feature, wherein the first visual feature is the visual feature obtained by uniformly sampling the newly added video in the sample video set, and the second visual feature is obtained by feature compression of the visual features of the historical video set in the sample video set and the first visual feature, because when the historical video set in the sample video set of the training sample is empty, the visual features of the sample video set of the training sample are generated according to a preset feature compression generation method. Exemplarily, the generation process of the visual features of the sample video set of each training sample is shown below.

[0080] Specifically, in an implementable manner, when the historical video set in the sample video set of the training sample is empty, for example, the training sample 1 including only one video has an empty historical video set, or the video set at the initial moment itself has no historical video set, the visual features of the sample video set of the training sample are determined in S102, which may be:

[0081] S1021. Arrange the video frames in the sample video set of the training sample in time sequence and uniformly sample N frames, input the sampled N video frames into a visual feature encoder, and output N text unit features.

[0082] Wherein, N is a preset fixed value for uniform sampling. Specifically, Figure 4 A schematic diagram of a method for obtaining text unit features provided in an embodiment of the present application, such as Figure 4As shown, the video frames in the sample video set of the training sample can be arranged in chronological order, and then N frames are uniformly sampled therefrom to obtain N video frames. The input of the video feature encoder is N video frames, and the output is N text unit features (tokens). It should be noted that the name "text unit features" here is to align the input name of the text model. It does not specifically refer to text features, but can also be text-like features obtained after tokenization of other modal features. Optionally, the video feature encoder can be (VisionTransformer, ViT) or a convolutional neural network (CNN). The video feature encoder can also be a network of other structures that can extract features of video frame sequences.

[0083] S1022. Arrange the video frames in the sample video set of the training sample in time sequence and uniformly sample Q frames, input the sampled Q video frames into the visual feature encoder, output Q text unit features, and perform feature compression on the Q text unit features to generate M text unit features, where Q is greater than M, and N, Q, and M are all positive integers.

[0084] Optionally, in one embodiment, in S1022, the Q text unit features are compressed to generate M text unit features, which may be:

[0085] S10221, perform feature compression on the Q text unit features, and the feature compression is:

[0086] Calculate the similarity between two adjacent text unit features among the Q text unit features, fuse the two text unit features with the largest similarity to obtain Q-1 text unit features, and continue to compress the Q-1 text unit features until M text unit features are obtained.

[0087] Specifically, after obtaining Q text unit features, Q is greater than M, and the similarity of two adjacent text unit features among the Q text unit features is calculated, and the two text unit features with the greatest similarity are feature fused. For example, Q=4, M=2, and there are 4 text unit features V_0, V_1, V_2, and V_3. First, the similarity between V_0 and V_1, the similarity between V_1 and V_2, and the similarity between V_2 and V_3 are calculated respectively. For example, the similarity between V_1 and V_2 is the greatest among the three, and then V_1 and V_2 are feature compressed to obtain V_1', and the three text unit features of V_0, V_1' and V_3 are obtained. Then, the three text unit features are feature compressed again until M=2 text unit features are obtained. Among them, feature compression of V_1 and V_2 can be performed by directly merging the average features calculated for V_1 and V_2 to obtain V_1'.

[0088] S1023: Concatenate the N text unit features and the M text unit features to obtain visual features of a sample video set of training samples.

[0089] Specifically, the text unit feature is a vector, and N text unit features and M text unit features can be directly concatenated to obtain the visual features of the sample video set of the training sample.

[0090] Optionally, in an implementable manner, when the historical video set in the sample video set of the training sample is not empty, the visual features of the sample video set of the training sample are determined in S102, which may specifically be:

[0091] S1021', uniformly sample N frames of the newly added video in the sample video set, input the sampled N video frames into a visual feature encoder, and output a first visual feature, where the first visual feature includes N text unit features, where N is a positive integer.

[0092] Wherein, N is a preset fixed value for uniform sampling. Specifically, the newly added videos in the sample video set can be arranged in chronological order, and then N frames are uniformly sampled therefrom to obtain N video frames. The input of the video feature encoder is N video frames, and the output is N text unit features (tokens).

[0093] S1022', according to the identification of the video included in the historical video set in the sample video set, search the visual features of the historical video set in the sample video set from the pre-stored visual feature library, the visual feature library storing the identification of the video included in the video set and the visual features of the video set.

[0094] Specifically, when constructing the original training sample set, the visual features of the historical video set in the sample video set in each training sample are calculated and stored in the visual feature library, and the visual features are directly searched from the pre-stored visual feature library during training. Optionally, the visual features of the historical video set in the sample video set in each training sample can also be calculated online.

[0095] S1023', compressing the visual features of the historical video set in the sample video set and the first visual features to generate second visual features, where the second visual features include M text unit features, where M is a positive integer.

[0096] Specifically, the first visual feature includes N text unit features, the visual feature of the historical video set in the sample video set includes M text unit features, and the two are feature compressed to generate a second visual feature, and the generated second visual feature includes M text unit features.

[0097] Optionally, in an implementable manner, the visual features of the historical video set in the sample video set include M text unit features, and in S1023′, the visual features of the historical video set in the sample video set and the first visual features are compressed to generate the second visual features, which may be:

[0098] The N text unit features included in the first visual feature are fused with the M text unit features to obtain the second visual feature. The feature fusion is:

[0099] For each of the N text unit features, add the text unit feature to the M text unit features, calculate the similarity between two adjacent text unit features in the obtained M+1 text unit features, and fuse the two text unit features with the largest similarity to obtain the fused M text unit features, until the fusion of N text unit features is completed.

[0100] Specifically, taking the addition of a new text unit feature as an example, the following is combined with Figure 5 To explain the feature fusion process in detail, Figure 5 A schematic diagram of a feature fusion process provided in an embodiment of the present application is shown in FIG. Figure 5 As shown in the figure, feature compression always maintains M text unit features. For example, when no new text unit features are added, the visual features of the historical video set in the sample video set include M text unit features, namely: V_0, V_1, V_2, V_3...V_M-1. A new text unit feature (such as Figure 4 When a new text unit feature V_M shown in is obtained, M+1 text unit features are obtained, and the similarity between two adjacent text unit features in the M+1 text unit features is calculated. For example, the similarity between V_0 and V_1, the similarity between V_1 and V_2, ..., the similarity between V_M-1 and V_M are calculated, and the two text unit features with the greatest similarity are fused to obtain the fused M text unit features. Figure 4 Take a newly added text unit feature as an example. For N text unit features, the above steps are executed for each of the N text unit features. Figure 4 The process shown is continued until the N text unit features are fused and a second visual feature including M text unit features is obtained.

[0101] Optionally, the similarity can be cosine similarity, i.e. cos<v_{j},v_{j+1}> .

[0102] S1023', concatenate the N text unit features and the M text unit features to obtain visual features of the sample video set of the training sample.

[0103] Specifically, the text unit feature is a vector, and the N text unit features included in the first visual feature and the M text unit features included in the second visual feature can be directly concatenated to obtain the visual features of the sample video set of the training sample.

[0104] S103. Train a video label recognition model according to the visual features of the sample video set, the text information of the sample video set, the first label and the second label of the sample video set, and the first preset paradigm instruction until a stop training condition is met to obtain a trained video label recognition model, wherein the video label recognition model includes a visual feature encoder, a pre-trained alignment module, and a recognition model.

[0105] Specifically, after obtaining the visual features of the sample video set, the video label recognition model can be trained according to the visual features of the sample video set, the text information of the sample video set, the first label and the second label of the sample video set, and the first preset paradigm instruction. The specific training process is shown below.

[0106] Optionally, in an implementable manner, in S103, the video label recognition model is trained according to the visual features of the sample video set, the text information of the sample video set, the first label and the second label of the sample video set, and the first preset paradigm instruction, which may be:

[0107] S1031. Fix the parameters of the visual feature encoder, input the visual features of the sample video set into the trained alignment module, and output the aligned visual features.

[0108] Specifically, the visual feature of the sample video set is obtained by concatenating the first visual feature and the second visual feature. The first visual feature is the visual feature obtained by uniformly sampling the newly added video in the sample video set, and the second visual feature is obtained by feature compression of the visual feature of the historical video set in the sample video set and the first visual feature. The video tag recognition model can focus on the visual features of the newly added video and the visual features of the historical video set.

[0109] In this embodiment, the parameters of the visual feature encoder are first fixed, the visual features of the sample video set are input into the trained alignment module, and the aligned visual features are output. The alignment module is a trained alignment module, which includes a learnable parameter matrix. The alignment module is mainly used to map the visual features to the input dimension of the recognition model, and convert the visual features into text unit features that are compatible with the recognition model.

[0110] S1032: Input the aligned visual features, text information of the sample video set, the second label and the first preset paradigm instruction into the recognition model, and output the predicted label of the sample video set.

[0111] Specifically, the sample video set includes a historical video set and a newly added video, the second label is a randomly transformed label of the label of the historical video set, the second label is used to simulate the model recognition result to train the model, and the first label is the real label of the sample video set. Among them, the text information of the sample video set may include the title of the sample video set, the description text of the sample video set, and the key text extracted by optical character recognition (OCR) / automatic speech recognition (ASR) of the sample video set, etc., wherein OCR detects and recognizes text content from video frames, and ASR converts speech into text. It should be noted that for training samples with an empty historical video set, it has no second label, and the second label item entered in S1032 is empty.

[0112] The first preset paradigm instruction may be any one of the following three examples:

[0113] "What tags can be used to summarize the above video collection?";

[0114] "Please output some key content tags based on the content of the above video collection.";

[0115] "Based on the above video content, generate the main content tags."

[0116] Optionally, the aligned visual features, the text information of the sample video set, the second label and the first preset paradigm instruction may be combined into a prompt word, for example, the prompt word may be composed in the form of [aligned visual features][text information of the sample video set + the second label][first preset paradigm instruction], the prompt word may be input into the recognition model, and the predicted label of the sample video set may be output. Optionally, the recognition model may be a multimodal model. Specifically, for example, it may be a large language model (LLM) or a multimodal large language model (MLLM). For the introduction of LLM and MLLM, please refer to the introduction of the aforementioned terms, which will not be repeated here.

[0117] S1033. According to the first label of the sample video set and the predicted label of the sample video set, adjust the parameters of the recognition model and the parameters of the alignment module until the training stop condition is met.

[0118] Specifically, the first label of the sample video set is the true label of the sample video set. According to the true label of the sample video set and the predicted label of the sample video set, the parameters of the recognition model and the parameters of the alignment module can be adjusted until the training stop condition is met, thereby obtaining a trained video label recognition model.

[0119] Optionally, as an implementable manner, in S1033, according to the first label of the sample video set and the predicted label of the sample video set, the parameters of the recognition model and the parameters of the alignment module are adjusted until the training stop condition is met, which may be:

[0120] S10331. Construct a loss function according to the first label of the sample video set and the predicted label of the sample video set.

[0121] S10332. According to the loss function, back-propagation is used to adjust the parameters of the recognition model and the parameters of the alignment module until the training stop condition is met.

[0122] The training stop condition may be preset, such as reaching a preset number of training times, or other preset conditions, which are not limited in this embodiment.

[0123] Through the training of S101-S103 above, a trained video tag recognition model is obtained, and the trained video tag recognition model includes a visual feature encoder, an alignment module and a recognition model.

[0124] In this embodiment, the alignment module is pre-trained. The method of this implementation may further include a training process of the alignment module before S101, which is described in detail below.

[0125] Furthermore, before S101, the method of this embodiment may further include:

[0126] S104: Obtain a training sample set, where each training sample includes a sample video and text description information of the sample video.

[0127] Specifically, the training sample sets in the training sample set S101 here are different training sample sets. The training sample set here includes multiple training samples, each training sample includes a sample video and text description information of the sample video, and the text description information of the sample video is used to describe the video content of the sample video, for example, it can be the title of the sample video and / or key text extracted from the OCR / ASR recognition results of the sample video, etc.

[0128] S105 , fixing the parameters of the visual feature encoder and the parameters of the recognition model, and training the alignment module according to the training sample set to obtain a pre-trained alignment module.

[0129] Specifically, the video tag recognition model to be trained in this embodiment includes a visual feature encoder, an alignment module and a recognition model. In this embodiment, the parameters of the visual feature encoder and the parameters of the recognition model are fixed first, and the alignment module is first trained according to the training sample set to obtain a pre-trained alignment module. The main goal of the pre-trained alignment module is to align the visual features to the recognition model.

[0130] Optionally, in S105, the alignment module is trained according to the training sample set to obtain a pre-trained alignment module, which may specifically be:

[0131] S1051. For each training sample, input the sample video of the training sample into a visual feature encoder, and output the visual features of the sample video.

[0132] S1052: Input the visual features of the sample video into an alignment module, and output the aligned visual features.

[0133] Specifically, the alignment module includes a learnable parameter matrix for mapping visual features to the input dimension of the recognition model and converting them into text unit features adapted to the recognition model.

[0134] S1053, input the aligned visual features and the second preset paradigm instruction into the recognition module, and output the predicted description information of the sample video.

[0135] Specifically, the second preset paradigm instruction may be, for example, any one of the following examples:

[0136] "Briefly describe the following video.";

[0137] "Provides a short description of a given video clip.";

[0138] "Concisely explain the video clips provided.";

[0139] "An overview of the visual content of the following video.";

[0140] "Give a brief and clear explanation of the following video clip.";

[0141] "Briefly describe the meaning of the video provided.";

[0142] "Briefly describe the key features of the fragment.";

[0143] "Briefly describe the content of the video being shown.";

[0144] "Provide a clear and concise summary of the following video.";

[0145] "Write an informative summary of the following video clip.";

[0146] "Present the provided videos in a concise narrative.".

[0147] Optionally, the aligned visual features and the second preset paradigm instructions may be combined into a prompt word, for example, a prompt word may be combined into the form of [aligned visual features][second preset paradigm instructions], the prompt word may be input into the recognition model, and the predicted description information of the sample video may be output.

[0148] S1054. Adjust the parameters of the alignment module according to the text description information of the sample video and the predicted description information of the sample video until the training stop condition is met to obtain a pre-trained alignment module.

[0149] Optionally, the parameters of the alignment module are adjusted according to the text description information of the sample video and the predicted description information of the sample video. A loss function may be constructed according to the text description information of the sample video and the predicted description information of the sample video, and the parameters of the alignment module are adjusted by back propagation according to the loss function until the stop training condition is met to obtain a pre-trained alignment module. The stop training condition here may be a preset stop training condition, such as reaching a preset number of iterative training times or other conditions, which is not limited in this embodiment.

[0150] The video label recognition model training method provided in the present embodiment obtains a training sample set, each training sample includes a sample video set, text information of the sample video set, a first label and a second label of the sample video set, the sample video set includes a historical video set and a newly added video, the second label is a randomly transformed label of the label of the historical video set, the videos in the sample video set are sorted in chronological order of creation time, the last K videos in the sample video set are newly added videos, and the videos other than K videos in the sample video set constitute a historical video set, then for each training sample, the visual features of the sample video set of the training sample are determined, the visual features of the sample video set are obtained by splicing the first visual features and the second visual features, the first visual features are visual features obtained by uniformly sampling the newly added videos in the sample video set, and the second visual features are obtained by feature compression of the visual features of the historical video set and the first visual features in the sample video set, then according to the visual features of the sample video set, the text information of the sample video set, the first label and the second label of the sample video set, and the first preset paradigm instruction, the video label recognition model is trained until the training stop condition is met to obtain a trained video label recognition model. Since the sample video set includes a historical video set and a newly added video when training the video label recognition model, when determining the visual features of the sample video set of the training sample, the visual features of the sample video set are obtained by splicing the first visual feature and the second visual feature, the first visual feature is the visual feature obtained by uniformly sampling the newly added video in the sample video set, and the second visual feature is obtained by feature compression of the visual features of the historical video set in the sample video set and the first visual feature. Therefore, when the trained video label recognition model obtains the visual features of the video set to be recognized during the process of video label recognition, it can also use uniform sampling for the newly added video and feature compression for the visual features of the historical video set, thereby reducing the information loss caused by sparse sampling or truncation, and can also reduce the reasoning time and resource consumption caused by dense sampling, thereby improving the accuracy of video label recognition.

[0151] Combine the following Figure 6 and Figure 7 , a specific implementation example is used to describe in detail the process of video tag recognition model training.

[0152] Figure 6 A flowchart of a video tag recognition model training method provided in an embodiment of the present application. The execution subject of the embodiment of the present application is a device with a model training function, and the model training device can be, for example, a server, such as Figure 6 As shown, the method may include:

[0153] S201: Obtain a training sample set, where each training sample includes a sample video and text description information of the sample video.

[0154] Specifically, the video tag recognition model to be trained in this embodiment includes a visual feature encoder, an alignment module and a recognition model. The training of the video tag recognition model in this embodiment is divided into two stages, S201-S205 is stage one, first fix the parameters of the visual feature encoder and the parameters of the recognition model, first train the alignment module according to the training sample set, and obtain a pre-trained alignment module. The main goal of the pre-trained alignment module is to align the visual features to the recognition model. S206-S208 is stage two, training the recognition model and the alignment module, and finally obtaining a trained video tag recognition model.

[0155] Among them, the training samples can use the <short video, video title> pair in the business scenario as the sample video and the text description information of the sample video.

[0156] S202, fixing the parameters of the visual feature encoder and the parameters of the recognition model, for each training sample, inputting the sample video of the training sample into the visual feature encoder, and outputting the visual features of the sample video.

[0157] S203: Input the visual features of the sample video into an alignment module, and output the aligned visual features.

[0158] Specifically, the alignment module includes a learnable parameter matrix for mapping visual features to the input dimension of the recognition model and converting them into text unit features adapted to the recognition model.

[0159] S204: Input the aligned visual features and the second preset paradigm instruction into a recognition module, and output predicted description information of the sample video.

[0160] Optionally, the aligned visual features and the second preset paradigm instruction may be combined into a prompt word, for example, a prompt word may be combined into a form of [aligned visual features][second preset paradigm instruction], the prompt word may be input into the recognition model, and the predicted description information of the sample video may be output. Optionally, the second preset paradigm instruction may be, for example, any one of the above examples.

[0161] S205. Adjust the parameters of the alignment module according to the text description information of the sample video and the predicted description information of the sample video until the training stop condition is met to obtain a pre-trained alignment module.

[0162] Optionally, the parameters of the alignment module are adjusted according to the text description information of the sample video and the predicted description information of the sample video. A loss function may be constructed according to the text description information of the sample video and the predicted description information of the sample video, and the parameters of the alignment module are adjusted by back propagation according to the loss function until the stop training condition is met to obtain a pre-trained alignment module. The stop training condition here may be a preset stop training condition, such as reaching a preset number of iterative training times or other conditions, which is not limited in this embodiment.

[0163] After obtaining the pre-trained alignment module, the second stage of training is carried out.

[0164] S206. Obtain a training sample set, where each training sample includes a sample video set, text information of the sample video set, a first label and a second label of the sample video set. The sample video set includes a historical video set and a newly added video. The second label is a randomly transformed label of the label of the historical video set. The videos in the sample video set are sorted in chronological order according to the creation time. The last K videos in the sample video set are newly added videos. The videos other than K videos in the sample video set constitute the historical video set, and K is a positive integer.

[0165] Specifically, in an implementable manner, a training sample can be obtained based on a video set in an actual business scenario and a manual review label of the video set. The manual review label is the first label of the video set. The second label of the video set can be obtained by randomly transforming the manual review label of the video set. The text information of the video set can include the title of the video set, the description text of the video set, and the key text extracted by OCR / ASR of the video set. After obtaining the above information, the video set can be used as a training sample.

[0166] Exemplarily, for example, the training sample set includes Q training samples: <video set_{i,t}, first label_{i,t}, second label_{i,t}, text information of video set_{i,t}> (i=1,2,3,…,Q), where i is the number of the video set, and t represents the status of the video set at time t and the human review label result (because new videos can be added to the video set after time t, so that the content of the video set and the corresponding label will change).

[0167] S207. For each training sample, determine the visual features of the sample video set of the training sample, where the visual features of the sample video set are obtained by concatenating the first visual feature and the second visual feature. The first visual feature is a visual feature obtained by uniformly sampling the newly added video in the sample video set, and the second visual feature is obtained by feature compression of the visual features of the historical video set in the sample video set and the first visual feature.

[0168] Specifically, in an implementable manner, when the historical video set in the sample video set of the training sample is empty, the visual features of the sample video set of the training sample are determined in S207, which may be:

[0169] S2071. Arrange the video frames in the sample video set of the first training sample in time sequence and uniformly sample N frames, input the sampled N video frames into a visual feature encoder, and output N text unit features.

[0170] S2072. Arrange the video frames in the sample video set of the first training sample in time sequence and uniformly sample Q frames, input the sampled Q video frames into a visual feature encoder, output Q text unit features, and perform feature compression on the Q text unit features to generate M text unit features, where Q is greater than M, and N, Q, and M are all positive integers.

[0171] Optionally, in one embodiment, in S2072, the Q text unit features are compressed to generate M text unit features, which may be:

[0172] S20721. Perform feature compression on the Q text unit features. The feature compression is:

[0173] Calculate the similarity between two adjacent text unit features among the Q text unit features, fuse the two text unit features with the largest similarity to obtain Q-1 text unit features, and continue to compress the Q-1 text unit features until M text unit features are obtained.

[0174] S2073. Concatenate the N text unit features and the M text unit features to obtain visual features of the sample video set of the first training sample.

[0175] Specifically, the text unit feature is a vector, and N text unit features and M text unit features can be directly concatenated to obtain the visual features of the sample video set of the training sample.

[0176] Optionally, in an implementable manner, when the historical video set in the sample video set of the training sample is not empty, the visual features of the sample video set of the training sample are determined in S207, which may specifically be:

[0177] S2071', uniformly sample N frames of the newly added video in the sample video set, input the sampled N video frames into a visual feature encoder, and output a first visual feature, where the first visual feature includes N text unit features, where N is a positive integer.

[0178] Wherein, N is a preset fixed value for uniform sampling. Specifically, the newly added videos in the sample video set can be arranged in chronological order, and then N frames are uniformly sampled therefrom to obtain N video frames. The input of the video feature encoder is N video frames, and the output is N text unit features (tokens).

[0179] S2072', according to the identification of the video included in the historical video set in the sample video set, search the visual features of the historical video set in the sample video set from the pre-stored visual feature library, the visual feature library stores the identification of the video included in the video set and the visual features of the video set.

[0180] S2073', compressing the visual features of the historical video set in the sample video set and the first visual features to generate second visual features, where the second visual features include M text unit features, where M is a positive integer.

[0181] Specifically, the first visual feature includes N text unit features, the visual feature of the historical video set in the sample video set includes M text unit features, and the two are feature compressed to generate a second visual feature, and the generated second visual feature includes M text unit features.

[0182] Optionally, in an implementable manner, the visual features of the historical video set in the sample video set include M text unit features, and in S2073′, the visual features of the historical video set in the sample video set and the first visual features are feature compressed to generate the second visual features, which may be:

[0183] The N text unit features included in the first visual feature are fused with the M text unit features to obtain the second visual feature. The feature fusion is:

[0184] For each of the N text unit features, add the text unit feature to the M text unit features, calculate the similarity between two adjacent text unit features in the obtained M+1 text unit features, and fuse the two text unit features with the largest similarity to obtain the fused M text unit features, until the fusion of N text unit features is completed.

[0185] Optionally, the similarity can be cosine similarity, i.e. cos<v_{j},v_{j+1}> .

[0186] S2073', concatenate the N text unit features and the M text unit features to obtain visual features of the sample video set of the training sample.

[0187] Specifically, the text unit feature is a vector, and the N text unit features included in the first visual feature and the M text unit features included in the second visual feature can be directly spliced ​​to obtain the visual features of the sample video set of the training sample.

[0188] S208. Train the video label recognition model according to the visual features of the sample video set, the text information of the sample video set, the first label and the second label of the sample video set, and the first preset paradigm instruction until the training stop condition is met to obtain a trained video label recognition model, wherein the video label recognition model includes a visual feature encoder, a pre-trained alignment module, and a recognition model.

[0189] Optionally, in an implementable manner, in S208, the video label recognition model is trained according to the visual features of the sample video set, the text information of the sample video set, the first label and the second label of the sample video set, and the first preset paradigm instruction, which may be:

[0190] S2081. Fix the parameters of the visual feature encoder, input the visual features of the sample video set into the trained alignment module, and output the aligned visual features.

[0191] In this embodiment, the parameters of the visual feature encoder are first fixed, the visual features of the sample video set are input into the trained alignment module, and the aligned visual features are output. The alignment module is a trained alignment module, which includes a learnable parameter matrix. The alignment module is mainly used to map the visual features to the input dimension of the recognition model, and convert the visual features into text unit features that are compatible with the recognition model.

[0192] S2082, input the aligned visual features, text information of the sample video set, the second label and the first preset paradigm instruction into the recognition model, and output the predicted label of the sample video set.

[0193] Specifically, the text information of the sample video set may include the title of the sample video set, the description text of the sample video set, key text extracted by performing optical character recognition (OCR) / automatic speech recognition (ASR) on the sample video set, etc., where OCR detects and recognizes text content from video frames and ASR converts speech into text.

[0194] The first preset paradigm instruction may be any one of the following three examples:

[0195] "What tags can be used to summarize the above video collection?";

[0196] "Please output some key content tags based on the content of the above video collection.";

[0197] "Based on the above video content, generate the main content tags."

[0198] Optional, Figure 7 A process diagram of a video tag recognition model training method provided in an embodiment of the present application is shown in FIG. Figure 7 As shown, for the training sample, the newly added video is first uniformly sampled to obtain the first visual feature, the visual features of the historical video set in the sample video set and the first visual feature are feature compressed to obtain the second visual feature, the first visual feature and the second visual feature are spliced ​​to obtain the visual features of the sample video set in the training sample, the visual features of the sample video set are input into the trained alignment module, and the aligned visual features are output. The aligned visual features, the text information of the sample video set, the second label and the first preset paradigm instruction may be used to form a prompt word, for example, the prompt word is formed in the form of [aligned visual features][text information of the sample video set + second label][first preset paradigm instruction], the prompt word is input into the recognition model, and the predicted label of the sample video set is output. Optionally, the recognition model may be a multimodal model. Specifically, for example, it may be a large language model (LLM) or a multimodal large language model (MLLM). For the introduction of LLM and MLLM, please refer to the introduction of the aforementioned terms, which will not be repeated here.

[0199] S2083. According to the first label of the sample video set and the predicted label of the sample video set, adjust the parameters of the recognition model and the parameters of the alignment module until the training stop condition is met.

[0200] Specifically, the first label of the sample video set is the true label of the sample video set. According to the true label of the sample video set and the predicted label of the sample video set, the parameters of the recognition model and the parameters of the alignment module can be adjusted until the training stop condition is met, thereby obtaining a trained video label recognition model.

[0201] Optionally, as an implementable manner, in S2083, the parameters of the recognition model and the parameters of the alignment module are adjusted according to the first label of the sample video set and the predicted label of the sample video set until the training stop condition is met, which may be:

[0202] S20831. Construct a loss function according to the first label of the sample video set and the predicted label of the sample video set.

[0203] S20832. According to the loss function, back propagation is used to adjust the parameters of the recognition model and the parameters of the alignment module until the training stop condition is met.

[0204] The training stop condition may be preset, such as reaching a preset number of training times, or other preset conditions, which are not limited in this embodiment.

[0205] Through the above training, a trained video label recognition model is obtained, and the trained video label recognition model includes a visual feature encoder, an alignment module and a recognition model.

[0206] The video label recognition model training method provided in this embodiment is that when training the video label recognition model, the sample video set includes a historical video set and a newly added video. When determining the visual features of the sample video set of the training sample, the visual features of the sample video set are obtained by splicing the first visual feature and the second visual feature. The first visual feature is the visual feature obtained by uniformly sampling the newly added video in the sample video set, and the second visual feature is obtained by feature compression of the visual features of the historical video set and the first visual feature in the sample video set. Therefore, when the trained video label recognition model obtains the visual features of the video set to be recognized during the process of video label recognition, it can also use uniform sampling for the newly added video and feature compression for the visual features of the historical video set, thereby reducing the information loss caused by sparse sampling or truncation, and can also reduce the reasoning time and resource consumption caused by dense sampling, thereby improving the accuracy of video label recognition.

[0207] Figure 8 A flowchart of a video tag recognition method provided in an embodiment of the present application, the execution subject of the method may be a server, such as Figure 8 As shown, the method may include:

[0208] S301: Obtain a video set to be identified and text information of the video set to be identified. The video set to be identified includes a currently added video and a historical video set.

[0209] S302: Obtain prediction labels of the historical video set and visual features of the historical video set.

[0210] S303: Determine the visual features of the currently added video.

[0211] Optionally, S303 may specifically be:

[0212] N frames are uniformly sampled for the current newly added video, the sampled N video frames are input into the visual feature encoder, and the visual features of the current newly added video are output, where N is a positive integer.

[0213] S304: Determine visual features of the video set to be identified based on the visual features of the current newly added video and the visual features of the historical video set.

[0214] Optionally, in an implementable manner, S304 may specifically be:

[0215] S3041. Compress the visual features of the current newly added video and the visual features of the historical video set to generate visual features of the video set to be identified.

[0216] Furthermore, the visual features of the current newly added video include N text unit features, and the visual features of the historical video set include M text unit features, where M is a positive integer. S3041 may specifically be:

[0217] S11, perform feature fusion on the N text unit features and the M text unit features to obtain fused M text unit features. The feature fusion is:

[0218] For each of the N text unit features, add the text unit feature to the M text unit features, calculate the similarity between two adjacent text unit features in the obtained M+1 text unit features, and fuse the two text unit features with the largest similarity to obtain the fused M text unit features, until the fusion of N text unit features is completed.

[0219] Specifically, the specific process of feature fusion can be found in Figure 3 The description in the illustrated embodiment will not be repeated here.

[0220] S12: Concatenate the N text unit features and the fused M text unit features to obtain visual features of the video set to be identified.

[0221] S305 , determining predicted labels of the video set to be identified according to visual features of the video set to be identified, text information of the video set to be identified, predicted labels of the historical video set, a first preset paradigm instruction and a trained video label recognition model.

[0222] Among them, the video tag recognition model is based on Figure 3 or Figure 6 The video tag recognition model obtained by training with the method of the illustrated embodiment includes a visual feature encoder, an alignment module and a recognition model.

[0223] Optionally, S305 may specifically be:

[0224] S3051. Input the visual features of the video set to be identified into an alignment module, and output the aligned visual features.

[0225] S3052, input the aligned visual features, the text information of the video set to be identified, the predicted labels of the historical video set and the first preset paradigm instruction into the recognition model, and output the predicted labels of the video set to be identified.

[0226] Specifically, the first preset paradigm instruction may be the same as the first preset paradigm instruction used in the video tag recognition model training. It should be noted that for a video to be recognized that has no predicted label of a historical video set, the predicted label item of the historical video set input in S3052 is empty.

[0227] Fig. 9 A process diagram of a video tag recognition method provided in an embodiment of the present application is shown as follows: Fig. 9 As shown, for the video to be identified including the current newly added video and the historical video set, N frames are uniformly sampled for the current newly added video, and the sampled N video frames are input into the visual feature encoder to output the visual features of the current newly added video. Next, the predicted labels of the historical video set and the visual features of the historical video set are obtained, and the visual features of the current newly added video and the visual features of the historical video set are feature compressed to generate the visual features of the video set to be identified, and the visual features of the video set to be identified are input into the alignment module, and the aligned visual features are output. Finally, the aligned visual features, the text information of the video set to be identified, the predicted labels of the historical video set and the first preset paradigm instruction are input into the recognition model, and the predicted labels of the video set to be identified are output. Optionally, the aligned visual features, the text information of the video set to be identified, the predicted labels of the historical video set and the first preset paradigm instruction can be used to form a prompt word, for example, the prompt word is formed in the form of [aligned visual features][text information of the video set to be identified + predicted labels of the historical video set][first preset paradigm instruction], and the prompt word is input into the recognition model to output the predicted label of the video set to be identified.

[0228] The video tag recognition method provided in this embodiment divides the video set to be recognized into a current newly added video and a historical video set when performing video tag recognition on the video set to be recognized, uniformly samples the current newly added video and obtains the visual features of the current newly added video, compresses the visual features of the historical video set, and then obtains the visual features of the video set to be recognized based on the visual features of the current newly added video and the visual features of the historical video set, and then determines the predicted labels of the video set to be recognized based on the visual features of the video set to be recognized, the text information of the video set to be recognized, the predicted labels of the historical video set, the first preset paradigm instruction and the trained video tag recognition model. Since uniform sampling is used for the newly added video and feature compression is performed on the visual features of the historical video set when obtaining the visual features of the video set to be recognized, the information loss caused by sparse sampling or truncation is reduced, and the reasoning time and resource consumption caused by dense sampling can also be reduced, thereby improving the accuracy of video tag recognition.

[0229] Fig.10A schematic diagram of the structure of a video tag recognition model training device provided in an embodiment of the present application is shown in FIG. Fig.10 As shown, the device may include: an acquisition module 11, a processing module 12 and a training module 13.

[0230] Wherein, the acquisition module 11 is used to acquire a training sample set, each training sample includes a sample video set, text information of the sample video set, a first label and a second label of the sample video set, the sample video set includes a historical video set and a newly added video, the second label is a randomly transformed label of the label of the historical video set, the videos in the sample video set are sorted in chronological order according to the creation time, the last K videos in the sample video set are newly added videos, and the videos other than K videos in the sample video set constitute a historical video set, and K is a positive integer;

[0231] The processing module 12 is used to determine the visual features of the sample video set of the training sample for each training sample, where the visual features of the sample video set are obtained by splicing the first visual feature and the second visual feature, where the first visual feature is the visual feature obtained by uniformly sampling the newly added video in the sample video set, and the second visual feature is obtained by feature compression of the visual features of the historical video set in the sample video set and the first visual feature;

[0232] The training module 13 is used to train the video label recognition model according to the visual features of the sample video set, the text information of the sample video set, the first label and the second label of the sample video set and the first preset paradigm instruction until the training stop condition is met to obtain a trained video label recognition model, which includes a visual feature encoder, a pre-trained alignment module and a recognition model.

[0233] In one embodiment, when the historical video set in the sample video set of the training sample is empty, the processing module 12 is used to:

[0234] Arrange the video frames in the sample video set of the training sample in time sequence and uniformly sample N frames, input the sampled N video frames into the visual feature encoder, and output N text unit features;

[0235] Arrange the video frames in the sample video set of the training sample in time sequence and evenly sample Q frames, input the sampled Q video frames into the visual feature encoder, output Q text unit features, and perform feature compression on the Q text unit features to generate M text unit features, where Q is greater than M, and N, Q, and M are all positive integers;

[0236] The N text unit features and the M text unit features are concatenated to obtain the visual features of the sample video set of the training samples.

[0237] In one embodiment, the processing module 12 is specifically used for:

[0238] The features of Q text units are compressed as follows:

[0239] Calculate the similarity between two adjacent text unit features among the Q text unit features;

[0240] The two text unit features with the greatest similarity are fused to obtain Q-1 text unit features;

[0241] Continue to perform feature compression on Q-1 text unit features until M text unit features are obtained.

[0242] In one embodiment, when the historical video set in the sample video set of the training sample is not empty, the processing module 12 is used to:

[0243] Uniformly sample N frames of the newly added video in the sample video set, input the sampled N video frames into a visual feature encoder, and output a first visual feature, where the first visual feature includes N text unit features, where N is a positive integer;

[0244] According to the identifiers of the videos included in the historical video set in the sample video set, searching for the visual features of the historical video set in the sample video set from a pre-stored visual feature library, wherein the visual feature library stores the identifiers of the videos included in the video set and the visual features of the video set;

[0245] Performing feature compression on the visual features of the historical video set in the sample video set and the first visual features to generate a second visual feature, where the second visual feature includes M text unit features, where M is a positive integer;

[0246] The N text unit features and the M text unit features are concatenated to obtain the visual features of the sample video set of the training samples.

[0247] In one embodiment, the visual features of the historical video set in the sample video set include M text unit features, and the processing module 12 is specifically used to:

[0248] The N text unit features included in the first visual feature are fused with the M text unit features to obtain the second visual feature. The feature fusion is:

[0249] For each of the N text unit features, add the text unit feature to the M text unit features, calculate the similarity between two adjacent text unit features in the obtained M+1 text unit features, and fuse the two text unit features with the largest similarity to obtain the fused M text unit features, until the fusion of N text unit features is completed.

[0250] In one embodiment, the processing module 12 is further configured to:

[0251] Obtain a training sample set, each training sample including a sample video and text description information of the sample video;

[0252] The parameters of the visual feature encoder and the parameters of the recognition model are fixed, and the alignment module is trained according to the training sample set to obtain a pre-trained alignment module.

[0253] In one embodiment, the processing module 12 is specifically used for:

[0254] For each training sample, input the sample video of the training sample into a visual feature encoder, and output the visual features of the sample video;

[0255] Input the visual features of the sample video into the alignment module, and output the aligned visual features;

[0256] Input the aligned visual features and the second preset paradigm instruction into the recognition module, and output the predicted description information of the sample video;

[0257] According to the text description information of the sample video and the predicted description information of the sample video, the parameters of the alignment module are adjusted until the training stop condition is met to obtain a pre-trained alignment module.

[0258] In one embodiment, the processing module 12 is specifically used for:

[0259] Fix the parameters of the visual feature encoder, input the visual features of the sample video set into the trained alignment module, and output the aligned visual features;

[0260] Inputting the aligned visual features, the text information of the sample video set, the second label and the first preset paradigm instruction into the recognition model, and outputting the predicted label of the sample video set;

[0261] According to the first label of the sample video set and the predicted label of the sample video set, the parameters of the recognition model and the parameters of the alignment module are adjusted until the training stop condition is met.

[0262] In one embodiment, the training module 13 is used to:

[0263] Constructing a loss function according to the first label of the sample video set and the predicted label of the sample video set;

[0264] According to the loss function, back propagation adjusts the parameters of the recognition model and the parameters of the alignment module until the training stop condition is met.

[0265] Fig.11 A schematic diagram of the structure of a video tag recognition device provided in an embodiment of the present application is shown in FIG. Fig.11As shown, the device may include: an acquisition module 21 and a processing module 22.

[0266] The acquisition module 21 is used to acquire the video set to be identified and the text information of the video set to be identified, and the video set to be identified includes the current newly added video and the historical video set;

[0267] The acquisition module is also used to: obtain the predicted labels of the historical video collection and the visual features of the historical video collection;

[0268] The processing module 22 is used to: determine the visual features of the current newly added video;

[0269] Determine the visual features of the video set to be identified based on the visual features of the current newly added video and the visual features of the historical video set;

[0270] Determine the predicted labels of the video set to be identified according to the visual features of the video set to be identified, the text information of the video set to be identified, the predicted labels of the historical video set, the first preset paradigm instruction and the trained video label recognition model, and the video label recognition model is based on Figure 2 The video tag recognition model obtained by training with the method of the illustrated embodiment includes a visual feature encoder, an alignment module and a recognition model.

[0271] In one embodiment, the processing module 22 is used to:

[0272] The current newly added video is uniformly sampled by N frames, the sampled N video frames are input into the visual feature encoder, and the visual features of the current newly added video are output, where N is a positive integer.

[0273] In one embodiment, the processing module 22 is used to:

[0274] The visual features of the current newly added video and the visual features of the historical video set are compressed to generate the visual features of the video set to be identified.

[0275] In one embodiment, the visual features of the current newly added video include N text unit features, and the visual features of the historical video set include M text unit features, where M is a positive integer. The processing module 22 is specifically configured to:

[0276] The N text unit features are fused with the M text unit features to obtain the fused M text unit features. The feature fusion is:

[0277] For each of the N text unit features, add the text unit features to the M text unit features, calculate the similarity between two adjacent text unit features in the obtained M+1 text unit features, fuse the two text unit features with the greatest similarity, and obtain the fused M text unit features until the fusion of N text unit features is completed;

[0278] The N text unit features and the fused M text unit features are concatenated to obtain the visual features of the video set to be identified.

[0279] In one embodiment, the processing module 22 is specifically used for:

[0280] Input the visual features of the video set to be identified into the alignment module, and output the aligned visual features;

[0281] The aligned visual features, text information of the video set to be identified, predicted labels of the historical video set and the first preset paradigm instruction are input into the recognition model, and the predicted labels of the video set to be identified are output.

[0282] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.

[0283] It should be understood that the device embodiment and the method embodiment may correspond to each other, and similar descriptions may refer to the method embodiment. To avoid repetition, they will not be described here. Specifically, Fig.10 The video label recognition model training device shown or Fig.11 The video tag identification device shown can execute the method embodiment corresponding to the computer device, and the aforementioned and other operations and / or functions of each module in the device are respectively for implementing the method embodiment corresponding to the computer device, which will not be repeated here for the sake of brevity.

[0284] The video tag recognition model training device and the video tag recognition device of the embodiment of the present application are described above from the perspective of the functional module in conjunction with the accompanying drawings. It should be understood that the functional module can be implemented in hardware form, can be implemented by instructions in software form, and can also be implemented by a combination of hardware and software modules. Specifically, the steps of the method embodiment in the embodiment of the present application can be completed by the hardware integrated logic circuit and / or software instructions in the processor, and the steps of the method disclosed in the embodiment of the present application can be directly embodied as a hardware decoding processor to execute, or a combination of hardware and software modules in the decoding processor to execute. Optionally, the software module can be located in a mature storage medium in the field such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an electrically erasable programmable memory, a register, etc. The storage medium is located in a memory, and the processor reads the information in the memory, and completes the steps in the above method embodiment in conjunction with its hardware.

[0285] Fig.12 It is a schematic block diagram of a computer device 300 provided in an embodiment of the present application.

[0286] like Fig.12 As shown, the computer device 300 may include:

[0287] The memory 310 and the processor 320, the memory 310 is used to store the computer program and transmit the program code to the processor 320. In other words, the processor 320 can call and run the computer program from the memory 310 to implement the method in the embodiment of the present application.

[0288] For example, the processor 320 may be configured to execute the above method embodiments according to instructions in the computer program.

[0289] In some embodiments of the present application, the processor 320 may include but is not limited to:

[0290] General-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic device, discrete hardware components, etc.

[0291] In some embodiments of the present application, the memory 310 includes but is not limited to:

[0292] Volatile memory and / or non-volatile memory. Among them, the non-volatile memory can be read-only memory (ROM), programmable ROM (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM) or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM), and direct RAM bus random access memory (DR RAM).

[0293] In some embodiments of the present application, the computer program may be divided into one or more modules, which are stored in the memory 310 and executed by the processor 320 to complete the method provided by the present application. The one or more modules may be a series of computer program instruction segments capable of completing specific functions, and the instruction segments are used to describe the execution process of the computer program in the electronic device.

[0294] like Fig.12 As shown, the computer device may also include:

[0295] The transceiver 330 may be connected to the processor 320 or the memory 310 .

[0296] The processor 320 may control the transceiver 330 to communicate with other devices, specifically, to send information or data to other devices, or to receive information or data sent by other devices. The transceiver 330 may include a transmitter and a receiver. The transceiver 330 may further include an antenna, and the number of antennas may be one or more.

[0297] It should be understood that the various components in the electronic device are connected via a bus system, wherein the bus system includes not only a data bus but also a power bus, a control bus and a status signal bus.

[0298] The embodiment of the present application also provides a computer storage medium on which a computer program is stored, and when the computer program is executed by a computer, the computer can perform the method of the above method embodiment. In other words, the embodiment of the present application also provides a computer program product containing instructions, and when the instructions are executed by a computer, the computer can perform the method of the above method embodiment.

[0299] When software is used for implementation, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function according to the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website site, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (digital subscriber line, DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, server or data center. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integration. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a digital video disc (digital video disc, DVD)), or a semiconductor medium (e.g., a solid state drive (solid state disk, SSD)), etc.

[0300] Those of ordinary skill in the art will appreciate that the modules and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the embodiments of the present application.

[0301] In the several embodiments provided in the embodiments of the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the module is only a logical function division. There may be other division methods in actual implementation, such as multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or modules, which can be electrical, mechanical or other forms.

[0302] The modules described as separate components may or may not be physically separated, and the components displayed as modules may or may not be physical modules, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. For example, each functional module in each embodiment of the present application may be integrated into a processing module, or each module may exist physically separately, or two or more modules may be integrated into one module.

[0303] The above contents are only specific implementation methods of the embodiments of the present application, but the protection scope of the embodiments of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or replacements within the technical scope disclosed in the embodiments of the present application, which should be included in the protection scope of the embodiments of the present application. Therefore, the protection scope of the embodiments of the present application should be based on the protection scope of the claims.

Claims

1. A video tag recognition model training method, It is characterized in that include: Acquire a training sample set, each training sample includes a sample video set, text information of the sample video set, a first label and a second label of the sample video set, the sample video set includes a historical video set and a newly added video, the second label is a randomly transformed label of the label of the historical video set, the videos in the sample video set are sorted according to the chronological order of creation time, the last K videos in the sample video set are the newly added videos, and the videos in the sample video set other than the K videos constitute the historical video set, and K is a positive integer; For each training sample, determining a visual feature of a sample video set of the training sample, wherein the visual feature of the sample video set is obtained by concatenating a first visual feature and a second visual feature, wherein the first visual feature is a visual feature obtained by uniformly sampling a newly added video in the sample video set, and the second visual feature is obtained by feature compression of a visual feature of a historical video set in the sample video set and the first visual feature; According to the visual features of the sample video set, the text information of the sample video set, the first label and the second label of the sample video set and the first preset paradigm instruction, the video label recognition model is trained until the training stop condition is met to obtain a trained video label recognition model, wherein the video label recognition model includes a visual feature encoder, a pre-trained alignment module and a recognition model.

2. The method according to claim 1, It is characterized in that When the historical video set in the sample video set of the training sample is empty, determining the visual features of the sample video set of the training sample includes: Arrange the video frames in the sample video set of the training sample in time sequence and uniformly sample N frames, input the sampled N video frames into the visual feature encoder, and output the N text unit features; Arrange the video frames in the sample video set of the training sample in time sequence and uniformly sample Q frames, input the sampled Q video frames into the visual feature encoder, output the Q text unit features, perform feature compression on the Q text unit features to generate M text unit features, where Q is greater than M, and N, Q and M are all positive integers; The N text unit features and the M text unit features are concatenated to obtain visual features of the sample video set of the training sample.

3. The method according to claim 2, It is characterized in that The step of compressing the Q text unit features to generate M text unit features includes: The Q text unit features are subjected to feature compression, and the feature compression is: Calculating the similarity between two adjacent text unit features among the Q text unit features; The two text unit features with the greatest similarity are fused to obtain Q-1 text unit features; The feature compression is continued for Q-1 text unit features until the M text unit features are obtained.

4. The method according to claim 1, It is characterized in that When the historical video set in the sample video set of the training sample is not empty, determining the visual features of the sample video set of the training sample includes: Uniformly sampling N frames of the newly added video in the sample video set, inputting the sampled N video frames into the visual feature encoder, and outputting the first visual feature, where the first visual feature includes the N text unit features, where N is a positive integer; According to the identifiers of the videos included in the historical video set in the sample video set, searching for the visual features of the historical video set in the sample video set from a pre-stored visual feature library, wherein the visual feature library stores the identifiers of the videos included in the video set and the visual features of the video set; Performing feature compression on the visual features of the historical video set in the sample video set and the first visual features to generate the second visual features, where the second visual features include M text unit features, where M is a positive integer; The N text unit features and the M text unit features are concatenated to obtain visual features of the sample video set of the training sample.

5. The method according to claim 4, It is characterized in that The visual features of the historical video set in the sample video set include M text unit features, and the step of performing feature compression on the visual features of the historical video set in the sample video set and the first visual features to generate the second visual features includes: The N text unit features included in the first visual feature are fused with the M text unit features to obtain the second visual feature, where the feature fusion is: For each of the N text unit features, add the text unit feature to the M text unit features in turn, calculate the similarity between two adjacent text unit features in the obtained M+1 text unit features, and fuse the two text unit features with the greatest similarity to obtain the fused M text unit features, until the fusion of the N text unit features is completed.

6. The method according to any one of claims 1 to 5, It is characterized in that The method further comprises: Obtaining a training sample set, each training sample including a sample video and text description information of the sample video; The parameters of the visual feature encoder and the parameters of the recognition model are fixed, and an alignment module is trained according to the training sample set to obtain the pre-trained alignment module.

7. The method according to claim 6, It is characterized in that The training of the alignment module according to the training sample set to obtain the pre-trained alignment module includes: For each training sample, inputting a sample video of the training sample into the visual feature encoder, and outputting visual features of the sample video; Inputting the visual features of the sample video into an alignment module, and outputting the aligned visual features; Inputting the aligned visual features and the second preset paradigm instruction into the recognition module, and outputting the predicted description information of the sample video; According to the text description information of the sample video and the predicted description information of the sample video, the parameters of the alignment module are adjusted until a training stop condition is met to obtain the pre-trained alignment module.

8. The method according to any one of claims 1 to 5, It is characterized in that The training of the video label recognition model according to the visual features of the sample video set, the text information of the sample video set, the first label and the second label of the sample video set, and the first preset paradigm instruction includes: Fixing the parameters of the visual feature encoder, inputting the visual features of the sample video set into the trained alignment module, and outputting the aligned visual features; Inputting the aligned visual features, the text information of the sample video set, the second label and the first preset paradigm instruction into the recognition model, and outputting the predicted label of the sample video set; According to the first label of the sample video set and the predicted label of the sample video set, the parameters of the recognition model and the parameters of the alignment module are adjusted until the training stop condition is met.

9. The method according to claim 8, It is characterized in that The adjusting the parameters of the recognition model and the parameters of the alignment module according to the first label of the sample video set and the predicted label of the sample video set until the training stop condition is met includes: Constructing a loss function according to the first label of the sample video set and the predicted label of the sample video set; According to the loss function, back propagation is used to adjust the parameters of the recognition model and the parameters of the alignment module until the training stop condition is met.

10. A video tag recognition method, It is characterized in that include: Acquire a video set to be identified and text information of the video set to be identified, wherein the video set to be identified includes a currently added video and a historical video set; Obtaining predicted labels of the historical video set and visual features of the historical video set; Determining visual features of the current newly added video; Determining visual features of the to-be-identified video set according to the visual features of the currently newly added video and the visual features of the historical video set; Determine the predicted labels of the video set to be identified based on the visual features of the video set to be identified, the text information of the video set to be identified, the predicted labels of the historical video set, a first preset paradigm instruction and a trained video label recognition model, wherein the video label recognition model is trained according to the method described in any one of claims 1-9, and the video label recognition model includes a visual feature encoder, an alignment module and a recognition model.

11. The method according to claim 10, It is characterized in that The determining of the visual features of the current newly added video includes: The current newly added video is uniformly sampled by N frames, the sampled N video frames are input into the visual feature encoder, and the visual features of the current newly added video are output, where N is a positive integer.

12. The method according to claim 10, It is characterized in that The determining the visual features of the to-be-identified video set according to the visual features of the current newly added video and the visual features of the historical video set includes: The visual features of the current newly added video and the visual features of the historical video set are compressed to generate the visual features of the video set to be identified.

13. The method according to claim 12, It is characterized in that The visual features of the current newly added video include the N text unit features, the visual features of the historical video set include M text unit features, where M is a positive integer, and the visual features of the current newly added video and the visual features of the historical video set are compressed to generate the visual features of the video set to be identified, including: The N text unit features are fused with the M text unit features to obtain fused M text unit features, wherein the fused features are: For each of the N text unit features, add the text unit feature to the M text unit features, calculate the similarity between two adjacent text unit features in the obtained M+1 text unit features, and perform feature fusion on the two text unit features with the greatest similarity to obtain fused M text unit features, until the fusion of the N text unit features is completed; The N text unit features and the fused M text unit features are concatenated to obtain visual features of the video set to be identified.

14. The method according to any one of claims 10 to 13, wherein the predicted labels of the video set to be identified are determined based on the visual features of the video set to be identified, the text information of the video set to be identified, the predicted labels of the historical video set, the first preset paradigm instruction and the trained video label recognition model, include: Inputting the visual features of the video set to be identified into the alignment module, and outputting the aligned visual features; The aligned visual features, the text information of the video set to be identified, the predicted labels of the historical video set and the first preset paradigm instruction are input into the recognition model, and the predicted labels of the video set to be identified are output.

15. A video tag recognition model training device, It is characterized in that include: An acquisition module is used to acquire a training sample set, each training sample includes a sample video set, text information of the sample video set, a first label and a second label of the sample video set, the sample video set includes a historical video set and a newly added video, the second label is a randomly transformed label of the label of the historical video set, the videos in the sample video set are sorted according to the chronological order of creation time, the last K videos in the sample video set are the newly added videos, and the videos in the sample video set other than the K videos constitute the historical video set, and K is a positive integer; A processing module, configured to determine, for each training sample, a visual feature of a sample video set of the training sample, wherein the visual feature of the sample video set is obtained by concatenating a first visual feature and a second visual feature, wherein the first visual feature is a visual feature obtained by uniformly sampling a newly added video in the sample video set, and the second visual feature is obtained by feature compression of a visual feature of a historical video set in the sample video set and the first visual feature; A training module is used to train a video label recognition model according to the visual features of the sample video set, the text information of the sample video set, the first label and the second label of the sample video set, and a first preset paradigm instruction until a stop training condition is met to obtain a trained video label recognition model, wherein the video label recognition model includes a visual feature encoder, a pre-trained alignment module, and a recognition model.

16. A video tag recognition device, It is characterized in that include: An acquisition module, used to acquire a video set to be identified and text information of the video set to be identified, wherein the video set to be identified includes a currently added video and a historical video set; The acquisition module is also used to: acquire the predicted labels of the historical video set and the visual features of the historical video set; A processing module, used to: determine the visual features of the current newly added video; Determining visual features of the to-be-identified video set according to the visual features of the currently newly added video and the visual features of the historical video set; Determine the predicted labels of the video set to be identified based on the visual features of the video set to be identified, the text information of the video set to be identified, the predicted labels of the historical video set, a first preset paradigm instruction and a trained video label recognition model, wherein the video label recognition model is trained according to the method described in any one of claims 1-9, and the video label recognition model includes a visual feature encoder, an alignment module and a recognition model.

17. A computer device, It is characterized in that include: A processor and a memory, the memory being used to store a computer program, the processor being used to call and run the computer program stored in the memory to execute the method according to any one of claims 1 to 9 or 10 to 14.

18. A computer-readable storage medium, It is characterized in that Comprising instructions which, when run on a computer program, cause the computer to perform the method as claimed in any one of claims 1 to 9 or 10 to 14.

19. A computer program product comprising instructions, It is characterized in that When the instructions are executed on a computer, the computer is caused to perform the method according to any one of claims 1 to 9 or 10 to 14.

Citation Information

Cited By

  • Prostate cancer diagnosis method based on multimodal large model prompt learning mechanism

    CN120954689A