Label generation method, label model training method and information distribution method
Automatically generate tags through the generative language model, the problem of inefficient tagging is solved, and efficient tag generation and information distribution is achieved.
Patent Information
- Application Number
- CN202410016191.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-04
- Publication Date
- 2025-07-04
AI Technical Summary
In the prior art, tag labeling is inefficient and cannot efficiently generate and distribute information.
By inputting the content information and tags of the second sample into the pre-trained generative language model, a target generative language model is generated, which is used to automatically generate the label of the first sample, and the target information is processed using the trained label model to generate the corresponding label.
It improves the efficiency and accuracy of label generation, simplifies the label marking process, reduces the demand for manual labeling, and improves the efficiency of information distribution.
Smart Images

Figure CN120257949A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology. Specifically, this application relates to a method for generating labels, a method for training a label model, and a method for information distribution. Background Art
[0002] Currently, in scenarios such as model training or information distribution, it is necessary to use label information for certain processing. For example, during model training, it is necessary to use samples and the corresponding labels of the samples to train the model. Another example is that during information distribution, it is necessary to use the labels of the information for information distribution.
[0003] However, currently, labels are mainly marked manually, and the efficiency of label marking is low. Summary of the Invention
[0004] Embodiments of this application provide a method for generating labels, a method for training a label model, and a method for information distribution, which are used to solve the technical problem of low efficiency of label marking and can achieve the technical effect of improving the efficiency of label generation.
[0005] On the one hand, an embodiment of this application provides a method for generating labels, including:
[0006] Obtain a first sample for which a label is to be generated, where the first sample includes at least one of a first text sample, a first image sample, or a first voice sample;
[0007] Identify the first content information of the first sample;
[0008] Input the first content information into a target generative language model to obtain a first label determined by the target generative language model based on the first content information, and the first label is used as the label corresponding to the first sample;
[0009] Among them, the target generative language model is obtained through the following method:
[0010] Input label generation prompt information into a pre-trained initial generative language model to obtain a target generative language model, where the label generation prompt information includes the second content information of a second sample related to the first sample and the second label corresponding to the second sample, and the second sample includes at least one of a second text sample, a second image sample, or a second voice sample.
[0011] On the other hand, an embodiment of this application also provides a method for training a label model, including:
[0012] Obtain a training sample set, where the training sample set includes multiple first samples, each first sample corresponds to first sample information, and the first sample information includes a first label. The first label corresponding to the first sample is obtained through the method of any embodiment of this application;
[0013] The preset label model is trained using a plurality of first samples and the first sample information corresponding to each first sample until the training of the preset label model is completed, obtaining a target label model. The target label model is used to process the target information input to the target label model to obtain the target label corresponding to the target information, where the target information includes at least one of target text, target image, or target voice.
[0014] On the other hand, an embodiment of the present application further provides an information distribution method, including:
[0015] Obtain the target information to be distributed, where the target information includes at least one of target text, target image, or target voice;
[0016] Process the target information through the trained target label model to obtain the target label corresponding to the target information, where the target tagging label is trained based on the method of any embodiment of the present application;
[0017] Distribute the target information based on the target label.
[0018] On the other hand, an embodiment of the present application further provides a label generation device, including:
[0019] A sample acquisition module for acquiring a first sample for which a label is to be generated, where the first sample includes at least one of a first text sample, a first image sample, or a first voice sample;
[0020] A content information recognition module for recognizing the first content information of the first sample;
[0021] A label generation module for inputting the first content information into a target generative language model to obtain a first label determined by the target generative language model based on the first content information, where the first label is used as the label corresponding to the first sample;
[0022] Wherein, the target generative language model is obtained by the label generation module in the following manner:
[0023] Input the label generation prompt information into a pre-trained initial generative language model to obtain the target generative language model, where the label generation prompt information includes the second content information of a second sample related to the first sample and the second label corresponding to the second sample, and the second sample includes at least one of a second text sample, a second image sample, or a second voice sample.
[0024] On the other hand, an embodiment of the present application further provides a training device for a label model, including:
[0025] A training sample set acquisition module for acquiring a training sample set, where the training sample set includes a plurality of first samples, each first sample corresponds to first sample information, and the first sample information includes a first label, and the first label corresponding to the first sample is obtained by the device according to any embodiment of the present application;
[0026] A training module for training a preset label model by using a plurality of first samples and the first sample information corresponding to each first sample until the training of the preset label model ends, to obtain a target label model, where the target label model is used to process target information input to the target label model to obtain a target label corresponding to the target information, and the target information includes at least one of target text, target image, or target voice.
[0027] On the other hand, an embodiment of the present application further provides an information distribution device, including:
[0028] An information acquisition module for acquiring target information to be distributed, where the target information includes at least one of target text, target image, or target voice;
[0029] A label determination module for processing the target information through the trained target label model to obtain a target label corresponding to the target information, where the target tagging label is trained based on the device according to any embodiment of the present application;
[0030] An information distribution module for distributing the target information based on the target label.
[0031] On the other hand, an embodiment of the present application further provides an electronic device, including a memory, a processor, and a computer program stored on the memory, and the processor executes the computer program to implement the steps of the method according to any embodiment of the present application.
[0032] On the other hand, an embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method according to any embodiment of the present application are implemented.
[0033] On the other hand, an embodiment of the present application further provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the method according to any embodiment of the present application are implemented.
[0034] In the embodiments of the present application, by inputting the second content information of the second sample and the second label of the second sample into a pre-trained initial generative language model, the obtained target generative language model can acquire the knowledge between the sample and the label. Then, by inputting the first content information of the first sample into the target generative language model, the label corresponding to the first sample can be determined by the target generative language model based on the first content information. Since labels can be automatically generated, and by inputting label generation prompt information into the pre-trained initial generative language model, the target generative language model can learn the relevant knowledge of the relationship between the sample and the label. Therefore, the efficiency of label generation can be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] To more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the accompanying drawings required for the description in the embodiments of the present application.
[0036] Figure 1 Schematic diagram of the implementation environment of a label generation method provided by an embodiment of the present application;
[0037] Figure 2 Schematic diagram of the flow of a label generation method provided by an embodiment of the present application;
[0038] Figure 3 Schematic diagram of a frame image in a video clip provided by an embodiment of the present application;
[0039] Figure 4 Schematic diagram of the flow of a training method for a label model provided by an embodiment of the present application;
[0040] Figure 5 Schematic diagram of the structure of an MLLM model provided by an embodiment of the present application;
[0041] Figure 6 Schematic diagram of the flow of an information distribution method provided by an embodiment of the present application;
[0042] Figure 7 Schematic diagram of the flow of video distribution provided by an embodiment of the present application;
[0043] Figure 8 Schematic diagram of the flow of label generation, model training, and information distribution provided by an embodiment of the present application;
[0044] Figure 9 Schematic diagram of the structure of a label generation device provided by an embodiment of the present application;
[0045] Figure 10 Schematic diagram of the structure of a training device for a label model provided by an embodiment of the present application;
[0046] Figure 11 The structural schematic diagram of an information distribution device provided by an embodiment of this application;
[0047] Figure 12 The structural schematic diagram of an electronic device provided by an embodiment of this application. Detailed implementation manners
[0048] The embodiments of this application will be described below with reference to the accompanying drawings in this application. It should be understood that the implementation manners described below with reference to the drawings are exemplary descriptions for explaining the technical solutions of the embodiments of this application, and do not constitute limitations on the technical solutions of the embodiments of this application.
[0049] Those skilled in the art of this technology can understand that, unless specifically stated otherwise, the singular forms "a", "an", and "the" used here may also include the plural forms. It should be further understood that the terms "including" and "comprising" used in the embodiments of this application mean that the corresponding features can be implemented as the presented features, information, data, steps, operations, elements, and / or components, but do not exclude the implementation of other features, information, data, steps, operations, elements, components, and / or their combinations, etc. supported by this technology field. It should be understood that when we say an element is "connected" or "coupled" to another element, this element can be directly connected or coupled to the other element, or it can mean that this element and the other element establish a connection relationship through an intermediate element. In addition, the "connection" or "coupling" used here can include wireless connection or wireless coupling. The term "and / or" used here indicates at least one of the items defined by this term. For example, "A and / or B" indicates the implementation as "A", or the implementation as "B", or the implementation as "A and B".
[0050] To make the purpose, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the drawings.
[0051] First, several terms related to this application will be introduced and explained:
[0052] UGC: The abbreviation of User-generated Content, which means user-generated content.
[0053] LLM (Large Language Model): By training on a large amount of data (such as text data), it learns the statistical laws and semantic information of language, enabling it to predict the next word or sentence. LLM has a wide range of applications in the field of natural language processing (NLP), such as machine translation, speech recognition, text generation, etc. LLM can adopt different training strategies and model structures, such as pre-training and fine-tuning. Pre-training refers to training the model on a large scale in an unsupervised or self-supervised manner, enabling it to have a general understanding and expression ability of language data. Fine-tuning is based on the pre-trained model and, through supervised learning, enables the model to be optimized for specific application scenarios or tasks. LLM has achieved remarkable results in tasks such as language understanding, generation, and translation, and with the continuous expansion of the model scale and the increase in training data, its performance and application scope are also continuously expanding.
[0054] MLLM (multimodal large language model): Based on LLM, it integrates media data of other modalities (such as images, videos, audio, etc.), enabling the model to process information of different modalities simultaneously, better understand and express semantics, thereby improving the effect and accuracy of applications.
[0055] VFM (Visual Foundation Model): Usually refers to a visual model obtained through pre-training and fine-tuning on a large amount of data, mainly used to solve various computer vision tasks, such as classification, detection, segmentation, etc.
[0056] In-context learning (ICL): Mainly through a small number of labeled samples, design a template for task-related instructions to guide test samples to generate corresponding results in an analogical manner.
[0057] Transformer: A network structure model suitable for processing sequential data, which adopts a full attention structure to replace traditional recurrent neural networks (RNNs) or convolutional neural networks (CNNs). Its core idea is to model the relationships between each element in the input sequence and other elements as an attention weight matrix, and calculate the representation of each element through the self-attention mechanism. This structure enables Transformer to have strong capabilities in capturing long-range dependencies in sequences, thus achieving good performance in NLP tasks. The Transformer model has been widely applied in the field of NLP, including tasks such as machine translation, text summarization, and question-answering systems. It first introduced the self-attention mechanism and surpassed traditional sequence-to-sequence (Seq2Seq) models, such as LSTMs and GRUs, in many aspects. The success of Transformer has also spawned many variant models based on the self-attention mechanism, which have made significant progress in NLP tasks.
[0058] Token: In Transformer, a token refers to a word or sub-word in the text, which is the basic unit used by the model when processing text.
[0059] In view of the above at least one technical problem or area for improvement in the background, the present application proposes a label generation method, a training method for a label model, and an information distribution method. By inputting the second content information of the second sample and the second label of the second sample into a pre-trained initial generative language model, the resulting target generative language model can acquire the knowledge between the sample and the label. Then, by inputting the first content information of the first sample into the target generative language model, the target generative language model can determine the label corresponding to the first sample based on the first content information. Since labels can be automatically generated, and by inputting label generation prompt information into the pre-trained initial generative language model, the target generative language model can learn the relevant knowledge of the relationship between the sample and the label, the efficiency of label generation can be improved.
[0060] The technical solutions of the embodiments of the present application and the technical effects produced by the technical solutions of the present application will be described below through the description of several exemplary embodiments. It should be noted that the following embodiments can refer to, draw on, or combine with each other. For the same terms, similar features, and similar implementation steps in different embodiments, they will not be described repeatedly.
[0061] Optionally, the present application may be related to artificial intelligence (AI) technology.
[0062] Artificial intelligence uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, including theories, methods, technologies, and application systems that can perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce an intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable machines to have the functions of perception, reasoning, and decision-making.
[0063] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include sensors, special artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-training model technology, operation / interaction systems, mechatronics, etc. Among them, the pre-training model, also known as the large model or the foundation model, can be widely applied to downstream tasks in various directions of artificial intelligence after fine-tuning. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0064] The pre-training model (PTM), also known as the foundation model or the large model, refers to a deep neural network (DNN) with a large number of parameters. It is trained on a large amount of unlabeled data, and the function approximation ability of the large-parameter DNN is used to extract common features from the data by the PTM. Through technologies such as fine-tuning, parameter-efficient fine-tuning (PEFT), and prompt-tuning, it is suitable for downstream tasks. Therefore, the pre-training model can achieve ideal results in few-shot or zero-shot scenarios. PTMs can be classified into language models (ELMO, BERT, GPT), vision models (swin-transformer, ViT, V-MOE), speech models (VALL-E), multi-modal models (ViBERT, CLIP, Flamingo, Gato), etc. according to the data modalities they process. Among them, the multi-modal model refers to a model that establishes feature representations of two or more data modalities. The pre-training model is an important tool for outputting artificial intelligence-generated content (AIGC) and can also be used as a general interface connecting multiple specific task models.
[0065] It should be noted that in the optional embodiments of the present application, for relevant data such as object information (account data), when the embodiments in the present application are applied to specific products or technologies, object permission or consent needs to be obtained, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions. That is to say, if the embodiments in the present application involve data related to the object, it needs to be obtained under the authorization and consent of the object, the authorization and consent of relevant departments, and compliance with the relevant laws, regulations, and standards of the country and region. In the embodiments, if personal information is involved, the consent of the individual needs to be obtained for all personal information. If sensitive information is involved, the separate consent of the information subject needs to be obtained, and the embodiments also need to be implemented under the authorization and consent of the object.
[0066] Please refer to Figure 1 , Figure 1 which is a schematic diagram of the implementation environment of a label generation method provided by an embodiment of the present application.
[0067] Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process, such as storing samples. The data storage system can be integrated on the server 104, or can be placed in the cloud or on other network servers 104.
[0068] In this embodiment, the second content information of the second sample and the second label corresponding to the second sample can be input into the initial generative language model through the server 104, and then the terminal 102 generates the first label corresponding to the first sample. Or it can be that the second content information of the second sample and the second label corresponding to the second sample are input into the initial generative language model through the server 104 and the first label corresponding to the first sample is generated. In addition, the second content information of the second sample and the second label corresponding to the second sample can also be input into the initial generative language model through the terminal 102 and the first label corresponding to the first sample is generated, which can be set according to needs and is not limited here.
[0069] Among them, the terminal 102 can be but is not limited to at least one of various personal computers, laptop computers, smartphones, tablet computers, Internet of Things devices, or portable wearable devices. The Internet of Things devices can be at least one of smart speakers, smart TVs, smart air conditioners, or smart in-vehicle devices, etc. The portable wearable devices can be at least one of smart watches, smart bracelets, or head-mounted devices, etc. The server 104 can be implemented by an independent server 104 or a server 104 cluster composed of multiple servers 104.
[0070] Please refer to Figure 2 , Figure 2A flowchart of a method for generating a tag provided by an embodiment of the present application. The method of this embodiment can be executed by an electronic device, and the electronic device can include at least one of a terminal or a server. As Figure 2 shown, the method may include:
[0071] S210. Obtain a first sample of the tag to be generated.
[0072] Among them, the first sample includes at least one of a first text sample, a first image sample, or a first voice sample. That is to say, the solution of this embodiment can generate at least one of a tag corresponding to text, a tag corresponding to an image, or a tag corresponding to voice. A tag may refer to a type of information characterizing the relevant characteristics of a sample. For example, a tag can characterize the category of a sample or the target audience of a sample, etc., and the meaning of the tag can be configured according to needs and is not limited herein.
[0073] It should be noted that the image in this embodiment may refer to a single frame of image or a series of consecutive frames of images (video), and the image sample may refer to a single frame of image sample or a video sample composed of a series of consecutive frames of image samples.
[0074] Exemplarily, the tag of the first text sample can characterize the classification of the first text sample, such as a narrative or an essay; in addition, the tag of the first text sample can characterize the target audience of the first text sample, such as a narrative enthusiast or an essay enthusiast, etc. The tag of the first image sample can characterize the classification of the first image sample, such as a game image or an anime image, etc.; in addition, the tag of the first image sample can characterize the target audience of the first image sample, such as a game enthusiast or an anime enthusiast, etc.
[0075] S220. Identify the first content information of the first sample.
[0076] Among them, the first content information may refer to content information related to the first sample. Exemplarily, if the first sample includes a first text sample, the first content information may include the text content of the first text sample; if the first sample includes a first image sample, the first content information may include the image content and text content of the first image sample; if the first sample includes a first voice sample, the first content information may include the voice content of the first language sample.
[0077] S230. Input the first content information into a target generative language model to obtain a first tag determined by the target generative language model based on the first content information, and the first tag serves as the tag corresponding to the first sample.
[0078] Among them, the target generative language model is obtained through the following method:
[0079] Input the label generation prompt information into a pre-trained initial generative language model to obtain a target generative language model. The label generation prompt information includes the second content information of a second sample related to the first sample and the second label corresponding to the second sample. The second sample includes at least one of a second text sample, a second image sample, or a second voice sample.
[0080] In this embodiment, if the first sample includes a first text sample, the second sample includes a second text sample; if the first sample includes a first image sample, the second sample includes a second image sample; if the first sample includes a first voice sample, the second sample includes a second voice sample. Specifically, the second sample related to the first sample may mean that the first sample and the second sample have the same attributes or the similarity is higher than a preset similarity threshold. For example, the first sample and the second sample have the same source, are from the same video, are from the same author, are from the same application program, etc., which are not limited herein. Or the content similarity between the first sample and the second sample is higher than a preset content similarity threshold, which is not limited herein.
[0081] Specifically, by inputting the label generation prompt information into a pre-trained initial generative language model, and the label generation prompt information includes the second content information of a second sample related to the first sample and the second label corresponding to the second sample, that is, at least inputting the second content information of the second sample related to the first sample and the second label corresponding to the second sample into the initial generative language model, the obtained target generative language model can learn knowledge such as the relationship between the sample and the label.
[0082] It should be noted that the label generation prompt information can be constructed manually based on the second content information of the second sample, or can be constructed by an electronic device based on the second content information of the second sample, which is not limited herein.
[0083] The technical solution of this embodiment inputs the second content information of the second sample and the second label of the second sample into a pre-trained initial generative language model, then the obtained target generative language model can obtain the knowledge between the sample and the label. Then, inputting the first content information of the first sample into the target generative language model, the label corresponding to the first sample can be determined by the target generative language model based on the first content information. Since the label can be automatically generated, and inputting the label generation prompt information into a pre-trained initial generative language model can enable the target generative language model to learn the relevant knowledge of the relationship between the sample and the label, the efficiency of label generation can be improved.
[0084] In a possible implementation manner, the first sample includes at least one frame of a first image sample. Identifying the first content information of the first sample includes:
[0085] Processing at least one first image sample through a pre-trained content extraction model to obtain first content information of the at least one first image sample, where the first content information includes at least one of first text information or second text information, the first text information is used to describe the image content of the at least one first image sample, and the second text information includes the text content in the at least one first image sample;
[0086] Wherein, the second sample includes at least one second image sample related to the at least one first image sample, and the second content information includes at least one of third text information or fourth text information, the third text information is used to describe the image content of the at least one second image sample, and the fourth text information includes the text content in the at least one second image sample.
[0087] In this embodiment, the first content information and the second content information exist in the form of text. Then the label generation prompt information may include label generation prompt text, and input the sample and the label corresponding to the sample into the initial generative language model in the form of text, so that the initial generative language model can learn the relationship between the sample and the label corresponding to the sample to obtain the target generative language model.
[0088] It should be noted that if there is no text content in the at least one first image sample, the second text information is empty. Similarly, if there is no text content in the at least one second image sample, the fourth text information is empty.
[0089] Optionally, the label generation prompt text can be input in the ICL manner. The second content information is obtained by processing at least one second image sample through a content extraction model. In addition, the second content information can be obtained through manual analysis, which is not limited herein.
[0090] The technical solution of this embodiment inputs the label generation prompt text in the ICL manner, designs a template for task-related instructions through a small number of labeled samples to guide the test samples to generate corresponding results in an analogical manner. Since a small number of samples can enable the initial generative language model to learn relevant knowledge, it can further improve the efficiency and simplicity of label generation and reduce the requirement for the number of samples.
[0091] In a possible implementation manner, the first text information includes at least one of first sub-text information for describing the image scene of the at least one first image sample or second sub-text information for describing the image object of the at least one first image sample, and the second text information includes at least one of the caption content in the at least one first image sample or the title content in the at least one first image sample;
[0092] Processing at least one frame of first image samples through a content extraction model to obtain first content information of the at least one frame of first image samples, including at least one of the following:
[0093] Performing image scene recognition processing on at least one frame of first image samples through a content extraction model to obtain first sub-text information;
[0094] Performing image object recognition processing on at least one frame of first image samples through a content extraction model to obtain second sub-text information;
[0095] Performing caption content extraction processing on at least one frame of first image samples through a content extraction model to obtain caption content of the at least one frame of first image samples;
[0096] Performing title content extraction processing on at least one frame of first image samples through a content extraction model to obtain title content of the at least one frame of first image samples;
[0097] Wherein, the third text information includes at least one of third sub-text information for describing the image scene of at least one frame of second image samples or fourth sub-text information for describing the image objects of at least one frame of second image samples, and the fourth text information includes at least one of caption content in at least one frame of second image samples or title content in at least one frame of second image samples.
[0098] Wherein, the content extraction model can be a VFM model. The image object can be a static object and a dynamic object; it can also be an object object or a human object, etc.
[0099] It can be understood that the more text sub-information included in the first text information and the second text information, the higher the accuracy of label generation can be.
[0100] The following takes the sample as a video sample as an example for illustration.
[0101] Please refer to Figure 3 , Figure 3 is a schematic diagram of a frame of an image in a video clip provided by an embodiment of the present application. Then the first text information or the second text information corresponding to the video clip can be as follows:
[0102] "Video clip content:
[0103] 00:00 - 00:10: [Video clip content description] (such as: two men and a woman, wearing outdoor windbreakers, carrying backpacks, walking together in the primeval forest, with thick trees and mist all around)
[0104] 00:10 - 00:20: [Video clip content description]
[0105] ……
[0106] Character dialogue content:
[0107] 00:00-00:03: [Subtitle content] (eg: As expected of a counsellor)
[0108] 00:03-00:09: [Subtitle content]
[0109] …
[0110] Objects in the picture:
[0111] 00:00-00:01: Person: [x1, x2, y1, y2], Person: [x1, x2, y1, y2], Gun: [x1, x2, y1, y2], Backpack: [x1, x2, y1, y2], Tree: [x1, x2, y1, y2], Tree: [x1, x2, y1, y2], Tree: [x1, x2, y1, y2], Mountain: [x1, x2, y1, y2], …
[0112] …
[0113] People in the video
[0114] Figure 1: 00:00 [x1, x2, y1, y2], 00:01 [x1, x2, y1, y2], …
[0115] Figure 2: 00:01 [x1, x2, y1, y2], …
[0116] ..."
[0117] Among them, x, y, z represent the position of the object in the image.
[0118] In a possible implementation manner, at least one frame of first image samples and at least one frame of second image samples belong to the same target video.
[0119] In this embodiment, taking the target video as a TV series as an example, at least one frame of the first image sample and at least one frame of the second image sample belong to the same target video. The at least one frame of the first image sample and the at least one frame of the second image sample may be different video segments of the same episode of the video, or the at least one frame of the first image sample and the at least one frame of the second image sample may be different video segments of different episodes of the video.
[0120] In a possible implementation, inputting the label generation prompt information into a pre-trained initial generative language model to obtain a target generative language model includes:
[0121] Acquire target video information of a target video, where the target video information includes at least one of a video type or a video time length of the target video;
[0122] Determine the target quantity of the second image samples based on the target video information;
[0123] Input the second content information of the target quantity of second image samples and the second labels corresponding to the target quantity of second image samples into a pre-trained initial generative language model to obtain a target generative language model.
[0124] In this embodiment, the target quantity may be related to at least one of the video type or the video time length of the target video. Specifically, if the video type of the target video is a science fiction video, the image scenes of the science fiction video are more complex, so more second image samples of the target quantity are required to improve the learning accuracy of the initial generative language model; if the video type of the target video is a comedy video, the image scenes of the comedy video are relatively simpler than those of the science fiction video, so fewer second image samples of the target quantity may be required to improve the learning efficiency of the initial generative language model. In addition, the longer the video time length, generally the more complex the plot, then more second image samples of the target quantity can be input into the initial generative language model to improve the learning effect of the initial generative language model; while the shorter the video time length, generally the simpler the plot, then fewer second image samples of the target quantity can be input into the initial generative language model to improve the learning efficiency of the generative language model.
[0125] The technical solution of this embodiment can balance the learning efficiency and learning effect of the initial generative language model by obtaining the target video information of the target video, where the target video information includes at least one of the video type or the video time length of the target video; determining the target quantity of the second image samples based on the target video information; and inputting the second content information of the target quantity of second image samples and the second labels corresponding to the target quantity of second image samples into a pre-trained initial generative language model to obtain a target generative language model.
[0126] In a possible implementation manner, the label generation prompt information includes a label generation prompt text, and the label generation prompt text is obtained by the following method:
[0127] Obtain a pre-configured label prompt text template, where the label prompt text template includes at least one content field, a content filling area corresponding to each content field, a label prompt field, and a label filling area corresponding to the label prompt field;
[0128] Determine the field that matches the second content information from at least one content field;
[0129] Fill the second content information into the content filling area corresponding to the field that matches the second content information, and fill the second label into the label filling area to obtain the label generation prompt text.
[0130] Among them, at least one content field may include, but is not limited to, title content and detailed content. The detailed content may include subtitle content, image frame content, image object content, etc.
[0131] Exemplarily, one of the label prompt text templates may be as follows:
[0132] Video content: Descriptions such as the title of the video are:
video_desc
textualizing_video
video_tags
[0133] Among them, "Descriptions such as the title of the video" and "detailed content" may be content fields.
video_desc
textualizing_video
video_tags
[0134] It can be understood that the specific label prompt text template can be set according to requirements and is not limited here. The number of label generation prompt texts can be set as needed and is not limited here.
[0135] In the technical solution of this embodiment, by obtaining a pre-configured label prompt text template, the label prompt text template includes at least one content field, the content filling area corresponding to each content field, a label prompt field, and the label filling area corresponding to the label prompt field; determining the field that matches the second content information from at least one content field; filling the second content information into the content filling area corresponding to the field that matches the second content information, and filling the second label into the label filling area to obtain a label generation prompt text. Then, when the first content information of the first sample is recognized, a label generation prompt text can be automatically constructed, which can improve the efficiency of constructing the label generation prompt text and further improve the efficiency of label generation.
[0136] Optionally, the label prompt text template may further include a general prompt text. Optionally, an example of one of the general prompt texts may be as follows:
[0137] "You are a chatbot based on video content. Please extract key content labels according to the detailed information of the video (including: description of segmented content, objects and position coordinates in each frame, trajectories of people in the frame, content text of people's conversations, etc.). The description gives specific time nodes for each piece of information. You can aggregate content from the time dimension to eliminate ambiguity, but you cannot fabricate non-existent content out of thin air."
[0138] In this embodiment, by constraining the learning of the initial generative language model through general prompt texts, the learning accuracy of the initial generative language model can be further improved.
[0139] In one possible implementation, inputting the first content information into the target generative language model includes:
[0140] Obtaining a pre-configured label request text template, where the label request text template includes at least one content field and a content filling area corresponding to each content field;
[0141] Determining a field that matches the first content information from the at least one content field;
[0142] Filling the first content information into the content filling area corresponding to the field that matches the first content information to obtain a label generation request text;
[0143] Inputting the label generation request text into the target generative language model.
[0144] In this embodiment, the layout of at least one content field and the content filling area corresponding to each content field in the label request text template and the label prompt text template may be the same.
[0145] The technical solution of this embodiment can improve the efficiency of the label generation request text and thus improve the efficiency of label generation by obtaining a pre-configured label request text template, where the label request text template includes at least one content field and a content filling area corresponding to each content field; determining a field that matches the first content information from the at least one content field; filling the first content information into the content filling area corresponding to the field that matches the first content information to obtain a label generation request text; and inputting the label generation request text into the target generative language model. In addition, the layout of at least one content field and the content filling area corresponding to each content field in the label request text template and the label prompt text template may be the same, which can improve the accuracy of label generation.
[0146] Please refer to Figure 4 , Figure 4 which is a schematic flowchart of a method for training a label model provided by an embodiment of this application. As Figure 4 shown, the method can be executed by an electronic device, and the electronic device may include at least one of a terminal or a server. As Figure 4 shown, the method may include:
[0147] S410. Obtaining a training sample set, where the training sample set includes a plurality of first samples, and each first sample corresponds to first sample information, and the first sample information includes a first label.
[0148] Among them, the first label in the embodiments of the present application can be obtained in the manner of the above embodiments, which will not be elaborated here.
[0149] S420. Use multiple first samples and the first sample information corresponding to each first sample to train a preset label model until the training of the preset label model ends, and obtain a target label model.
[0150] Among them, the target label model is used to process the target information input to the target label model to obtain a target label corresponding to the target information, and the target information includes at least one of target text, target image, or target voice.
[0151] In this embodiment, if the first sample includes a first text sample, the target information includes target text; if the first sample includes a first image sample, the target information includes target image; if the first sample includes a first voice sample, the target information includes target voice.
[0152] The technical solution of this embodiment is to obtain a training sample set, where the training sample set includes multiple first samples, each first sample corresponds to first sample information, and the first sample information includes a first label. Use multiple first samples and the first sample information corresponding to each first sample to train a preset label model until the training of the preset label model ends, and obtain a target label model. The first label can be obtained in the manner of the above embodiments, which can improve the generation efficiency of the first label and thus improve the training efficiency of the model.
[0153] In a possible implementation manner, the preset label model includes a feature extraction network and a generative language network. Using multiple first samples and the first sample information corresponding to each first sample to train the preset label model includes:
[0154] Extract the features of the first sample through the feature extraction network to obtain first sample features;
[0155] Perform label prediction based on the first sample features through the generative language network to obtain a predicted label corresponding to the first sample;
[0156] If it is determined that the preset label model does not meet the preset first training end condition based on the predicted label and the first label, update the parameters of the generative language network and continue to train the preset label model;
[0157] If it is determined that the preset label model meets the first training end condition based on the predicted label and the first label, determine the preset label model that meets the first training end condition as the target label model.
[0158] In this embodiment, if the first sample includes a first text sample, the feature extraction network includes a text feature extraction network; if the first sample includes a first image sample, the feature extraction network includes an image feature extraction network; if the first sample includes a first voice sample, the feature extraction network includes a voice feature extraction network.
[0159] In a possible implementation, the first sample information further includes first content information, and the preset label model further includes an alignment network. Training the preset label model using multiple first samples and the first sample information corresponding to each first sample further includes:
[0160] Extracting the features of the first sample through the feature extraction network to obtain second sample features;
[0161] Mapping the second sample features to the input dimension of the generative language network through the alignment network to obtain predicted content information;
[0162] If it is determined that the preset label model does not meet the second training end condition based on the predicted content information and the first content information, update the parameters of the alignment network;
[0163] If it is determined that the preset label model meets the second training end condition based on the predicted content information and the first content information, extract the features of the first sample through the feature extraction network in the preset label model that meets the second training end condition to obtain first sample features, and perform label prediction on the first sample based on the first sample features through the generative language network in the preset label model that meets the second training end condition to obtain the predicted label corresponding to the first sample;
[0164] If it is determined that the preset label model does not meet the preset first training end condition based on the predicted label and the first label, update the parameters of the generative language network, including:
[0165] If it is determined that the preset label model does not meet the preset first training end condition based on the predicted label and the first label, update the parameters of the generative language network and the parameters of the alignment network.
[0166] The technical solution of this embodiment can make the matching degree between the feature extraction network and the generative extraction network by setting an alignment network to map the second sample features to the input dimension of the generative language network, thereby improving the model training effect of the preset label model.
[0167] Among them, the preset label model of this embodiment can be an MLLM model. Please refer to Figure 5 , Figure 5 which is a schematic structural diagram of an MLLM model provided by an embodiment of this application. This embodiment takes the sample as a video for illustration. For example Figure 5The MLLM model shown can include a visual feature extraction network, an alignment network, and a generative language network.
[0168] As Figure 5 shown, the MLLM structure is as shown in the figure. There are two types of inputs received by the model: 1) Visual feature tokens: The visual feature encoder can be either ViT or CNN, and its function is to convert an input video frame image into visual features (column vectors). To align with the input of the LLM, a visual token is followed by an alignment module P, which contains a learnable parameter matrix. Its main function is to map the visual features to the input dimension of the LLM and at the same time convert them into tokens adapted to the LLM, denoted as [video_emb]; 2) Text: The text information includes the title of the video, the descriptive text, etc., denoted as [video_txt]. The LLM can choose open-source Chinese LLMs such as chinese-llama, BLOOM, etc. The input prompt template is:
[0169] "Human: [video_emb][video_txt][instruct]Assistant:"
[0170] The LLM generates the inferred label results based on this input.
[0171] The training of the MLLM is divided into two stages
[0172] Stage 1: Pretraining, feature alignment
[0173] In this stage, the parameters of the visual feature encoder and the LLM are all fixed, and only the alignment module P is trained. The main goal is to align the visual tokens to the LLM. The training data for this stage uses <video, video descriptive text> pairs. Here, <short video, video title> pairs in the business scenario can be used to train the model with the goal of generating a brief description of the video content. The following prompt is used:
[0174] "Human:[video_emb][alignment instruction]Assistant:"
[0175] Among them, the alignment instructions can be diversified. For example, one of the following statements can be used:
[0176] "Briefly describe the following video."
[0177] "Provide a short description of the given video clip."
[0178] "Briefly explain the provided video clip."
[0179] "Outline the visual content of the following video."
[0180] "Provide a short and clear explanation of the following video clip."
[0181] "Briefly elaborate on the meaning of the provided video."
[0182] "Briefly describe the key features of the clip."
[0183] "Briefly describe the content of the video shown."
[0184] "Provide a clear and concise summary of the following video."
[0185] "Write an informative summary of the following video clip."
[0186] "Present the provided video in a concise narrative manner.
[0187] Phase II: Instruction Tuning
[0188] Fix the parameters of the visual feature encoder and release the parameters of the alignment module and the LLM for training."
[0189] The training objective of this phase is: Based on the <video, label> data in the actual business scenario, construct corresponding instructions and perform instruction tuning with the generated label list text as the target:
[0190] “Human:[video_emb][video_txt][instruct]Assistant:”
[0191] Similarly, similar to the alignment instructions in Phase I, synonymous instructions can also be designed for instruct, which can improve the training effect. For example:
[0192] "What labels can be used to summarize the above video?"
[0193] "Please output some key content labels based on the content of the above video."
[0194] "Generate the main content labels based on the above video content."
[0195] In this embodiment, the visual feature extraction network may include a visual feature encoder. The alignment network may include an alignment module P.
[0196] It should be noted that the visual encoder in this embodiment can be replaced with a network of other structures as long as it can extract the features of the video frame sequence; the alignment module is used to align the visual features to the input of the LLM and can also be replaced with other structures, such as q-former; for the main process of this embodiment, the framework of the LLM can also be replaced with other multi-modal models, such as visual-bert, uniter, etc.
[0197] Please refer to Figure 6 , Figure 6 , which is a schematic flowchart of an information distribution method provided by an embodiment of the present application. The method of this embodiment can be executed by an electronic device, and the electronic device can include at least one of a terminal or a server. As Figure 6 shown, the method may include:
[0198] S610. Obtain the target information to be distributed.
[0199] Among them, the information in this embodiment may include multimedia information. In this embodiment, the target information includes at least one of target text, target image, or target voice.
[0200] S620. Process the target information through a trained target label model to obtain a target label corresponding to the target information.
[0201] Among them, the target tagging label is trained based on the method of any of the above embodiments, and will not be elaborated here.
[0202] S630. Distribute the target information based on the target label.
[0203] In this embodiment, distributing the target information may refer to distributing the target information corresponding to the same or similar labels to the same destination. For example, distributing the same or similar labels to the same object or the same recommendation algorithm system, etc., which is not limited here.
[0204] The technical solution of this embodiment can improve the efficiency of information distribution by obtaining the target information to be distributed, where the target information includes at least one of target text, target image, or target voice, processing the target information through a trained target label model to obtain a target label corresponding to the target information, and distributing the target information based on the target label.
[0205] In a possible implementation manner, distributing the target information based on the target label includes:
[0206] If the confidence level corresponding to the target label is higher than or equal to a preset confidence level threshold, then distribute the target information based on the target label;
[0207] The method further includes:
[0208] If the confidence level corresponding to the target label is lower than the confidence level threshold, then send the target information and the target label corresponding to the target information to the target object;
[0209] Receive a new target label corresponding to the target information from the target object;
[0210] Distribute the target information based on the new target label.
[0211] Optionally, when the target label model of this embodiment outputs the target label corresponding to the target information, it can synchronously output the confidence level corresponding to the target label. The confidence level can characterize the credibility of the target label. The greater the confidence level, the higher the credibility. The new target label may be different from the original label or the same, and is determined according to the review result of the target object.
[0212] In this embodiment, if the confidence level corresponding to the target label is higher than or equal to the preset confidence level threshold, the target information is distributed based on the target label; if the confidence level corresponding to the target label is lower than the confidence level threshold, the target information and the target label corresponding to the target information are sent to the target object so that the target object can perform manual review, and then a new target label is returned, and then the target information is distributed based on the new target label.
[0213] The technical solution of this embodiment is as follows: if the confidence level corresponding to the target label is higher than or equal to the preset confidence level threshold, the target information is distributed based on the target label; if the confidence level corresponding to the target label is lower than the confidence level threshold, the target information and the target label corresponding to the target information are sent to the target object; a new target label corresponding to the target information from the target object is received; the target information is distributed based on the new target label. By combining the human-machine method, the accuracy of information distribution can be improved.
[0214] For the sake of easy understanding, in the following embodiments, based on any of the above embodiments, the generation of video labels, the training of video label models, and the distribution of videos are described.
[0215] Video label recognition is an important part of video content features. Automatically generating labels for a large number of UGC videos by machines can provide video content features at different granularities for downstream content distribution links (such as recommendation systems and content operations), improve the efficiency of content distribution, and at the same time greatly reduce the cost of manual content review. Due to the diversity of UGC video content, the number of labels in the label library used in the general business scenario can reach hundreds of thousands or even millions or more, making it difficult to assign corresponding content labels to each video.
[0216] Generally speaking, there are methods for generating human-machine collaborative video tags, model training, and information distribution based on large models. In the startup phase, an off-the-shelf VFM is used to "translate" video content into plain text. By using examples of a small number of human-reviewed results and leveraging the context learning ability of large models, training data is directly generated with the help of ChatGPT, thereby training the target model. This process is highly efficient and only requires a small number of annotation results, avoiding the problem of manual large-scale data annotation without the assistance of AI. In the evolution phase, through a human-machine collaborative approach, high-confidence tag prediction results directly bypass human review and enter the content distribution link. Human review only examines low-confidence tags, overall improving the link efficiency and labor costs. At the same time, the human-reviewed results are fed back into the training and test sets, and the model can achieve automatic and high-frequency iteration. This process does not require manual intervention and can more efficiently utilize newly added human-reviewed data for model updates to ensure the model's performance.
[0217] Please refer to Figure 7 , Figure 7 which is a schematic diagram of the video distribution process provided by an embodiment of this application. As Figure 7 shown, the video enters the content processing link from the content production link, and corresponding content features are obtained through a human-machine collaborative method, and then enter the downstream content distribution link.
[0218] Please refer to Figure 8 , Figure 8 which is a schematic diagram of the processes for tag generation, model training, and information distribution provided by an embodiment of this application.
[0219] (1) Startup phase:
[0220] The goal of this phase is to use tools such as off-the-shelf VFM and ChatGPT (which can also be replaced by other dialogue LLMs, such as Hunyuan Assistant, ChatGLM, etc.) to generate <video, pseudo-label> samples based on a very small amount of annotated data for training the tag model.
[0221] a. Video content textification:
[0222] This step uses a variety of perception models (content extraction models) to explicitly encode all video content into text description information. For example, the video classification model obtains the behavior category, the image description model obtains the spatial detail information of different frames, and speech recognition generates subtitles, etc.
[0223] Specifically, the video is encoded to obtain text information (textualizing_video). The textification result can refer to Figure 3 the relevant textification results and will not be elaborated here.
[0224] b. Context learning to generate tags:
[0225] After multiple VFM inferences, the main visual information of the video is encoded as textualizing_video. The video also naturally contains some text information, such as video titles, descriptions, and CP accounts, denoted as video_desc. Next, we can utilize the context learning ability of large models. First, a small number of videos are annotated by human annotators, and the labeled tag results are denoted as video_tags. Then, a large number of <video, pseudo-tag result> data are generated in batches. The specific approach is as follows:
[0226] Design a prompt template. Present the video information in plain text, and combine the human-reviewed tag results as examples for question and answer. Feed N (N = 3 in the example) examples along with the question to chatGPT to directly obtain the results.
[0227] # Prompt Template
[0228] You are a chatbot based on video content. Please extract key content tags according to the detailed information of the video (including: description of segmented content, objects and their position coordinates in each frame, trajectories of people in the frame, text content of people's conversations, etc.). The description provides specific time nodes for each piece of information. You can aggregate the content from the time dimension to eliminate ambiguity, but you cannot fabricate content that does not exist.
[0229] Video content: The title and other descriptions of the video are:
video_desc
textualizing_video
video_tags
[0230] Video content: The title and other descriptions of the video are:
video_desc
textualizing_video
video_tags
[0231] Video content: The title and other descriptions of the video are:
video_desc
textualizing_video
video_tags
[0232] Video content: The title and other descriptions of the video are:
video_desc
textualizing_video
[0233] This method only requires a small number of samples to be labeled, avoiding the difficulty of large-scale labeling by human review without AI assistance. On the other hand, different from supervised learning in the training stage that requires using reverse gradients to update model parameters, ICL does not require parameter updates and directly makes predictions on a pre-trained language model. It is hoped that the model will learn the patterns hidden in the demonstrations and make correct predictions accordingly.
[0234] c. Model training method:
[0235] According to the method above, a batch of initial training data: {<video, pseudo-label>} is generated, and at this time, a multi-modal model can be trained to predict labels. To illustrate the process of this embodiment in detail, MLLM is used as an example here. It should be noted that the multi-modal model here can be trained using other algorithms.
[0236] This embodiment can refer to Figure 5 for relevant descriptions and will not be elaborated here.
[0237] (2) Evolution stage:
[0238] A. Human-machine collaborative label generation:
[0239] After the model is deployed online, it can predict the label results of UGC videos. Based on the confidence level as the division criterion, videos with high confidence levels do not need to be sent for review and directly enter the content distribution link, which can save a large amount of human review labor costs and shorten the time-consuming of the content processing link; other videos are sent for human review to correct the label results and then continue to be distributed. At this time, the prediction results of the model are also provided to the annotators as auxiliary bottom results for reference and modification to improve the review efficiency.
[0240] B. Model automatic iteration:
[0241] Online video content will inevitably change continuously over time. The model needs to maintain a sufficient update frequency to maintain an ideal effect. The data labeled by human review can flow back into a new batch of training data, and after preprocessing, it is fed to the model for automatic iteration and update.
[0242] The content understanding algorithm system of the embodiment of this application. The proposed human-machine collaborative label generation method can significantly improve the efficiency of machine review and save the cost of manual review.
[0243] Please refer to Figure 9 , Figure 9 which is a schematic structural diagram of a label generation device provided by an embodiment of this application. As Figure 9 shown, the device 900 may include a sample acquisition module 910, a content information recognition module 920, and a label generation module 930, where:
[0244] A sample acquisition module 910, configured to acquire a first sample for which a label is to be generated, where the first sample includes at least one of a first text sample, a first image sample, or a first voice sample;
[0245] A content information recognition module 920, configured to recognize first content information of the first sample;
[0246] A label generation module 930, configured to input the first content information into a target generative language model to obtain a first label determined by the target generative language model based on the first content information, where the first label serves as the label corresponding to the first sample;
[0247] Wherein, the target generative language model is obtained by the label generation module 930 in the following manner:
[0248] Inputting label generation prompt information into a pre-trained initial generative language model to obtain the target generative language model, where the label generation prompt information includes second content information of a second sample related to the first sample and a second label corresponding to the second sample, and the second sample includes at least one of a second text sample, a second image sample, or a second voice sample.
[0249] Optionally, the first sample includes at least one frame of first image samples. When the content information recognition module 920 recognizes the first content information of the first sample, it can be used to:
[0250] Processing at least one frame of first image samples through a pre-trained content extraction model to obtain first content information of the at least one frame of first image samples, where the first content information includes at least one of first text information or second text information, the first text information is used to describe the image content of the at least one frame of first image samples, and the second text information includes the text content in the at least one frame of first image samples;
[0251] Wherein, the second sample includes at least one frame of second image samples related to the at least one frame of first image samples, and the second content information includes at least one of third text information or fourth text information, the third text information is used to describe the image content of the at least one frame of second image samples, and the fourth text information includes the text content in the at least one frame of second image samples.
[0252] Optionally, the first text information includes at least one of first sub-text information for describing the image scene of the at least one frame of first image samples or second sub-text information for describing the image objects of the at least one frame of first image samples, and the second text information includes at least one of the caption content in the at least one frame of first image samples or the title content in the at least one frame of first image samples;
[0253] The content information recognition module 920 processes at least one frame of the first image sample through a content extraction model to obtain the first content information of at least one frame of the first image sample, which can be used for at least one of the following:
[0254] Perform image frame recognition processing on at least one frame of the first image sample through a content extraction model to obtain first sub-text information;
[0255] Perform image object recognition processing on at least one frame of the first image sample through a content extraction model to obtain second sub-text information;
[0256] Perform subtitle content extraction processing on at least one frame of the first image sample through a content extraction model to obtain the subtitle content of at least one frame of the first image sample;
[0257] Perform title content extraction processing on at least one frame of the first image sample through a content extraction model to obtain the title content of at least one frame of the first image sample;
[0258] Among them, the third text information includes at least one of the third sub-text information for describing the image frame of at least one frame of the second image sample or the fourth sub-text information for describing the image object of at least one frame of the second image sample, and the fourth text information includes at least one of the subtitle content in at least one frame of the second image sample or the title content in at least one frame of the second image sample.
[0259] Optionally, at least one frame of the first image sample and at least one frame of the second image sample belong to the same target video. When the label generation module 930 inputs label generation prompt information into a pre-trained initial generative language model to obtain a target generative language model, it can be used for:
[0260] Obtain the target video information of the target video, where the target video information includes at least one of the video type or video time length of the target video;
[0261] Determine the target number of the second image samples based on the target video information;
[0262] Input the second content information of the target number of second image samples and the second labels corresponding to the target number of second image samples into a pre-trained initial generative language model to obtain a target generative language model.
[0263] Optionally, the label generation prompt information includes label generation prompt text, and the label generation prompt text is obtained by the label generation module 930 in the following way:
[0264] Obtain a pre-configured label prompt text template, which includes at least one content field, a content filling area corresponding to each content field, a label prompt field, and a label filling area corresponding to the label prompt field;
[0265] Determine the field that matches the second content information from at least one content field;
[0266] Fill the second content information into the content filling area corresponding to the field that matches the second content information, and fill the second label into the label filling area to obtain a label generation prompt text.
[0267] Optionally, when the label generation module 930 inputs the first content information into the target generative language model, it can be used for:
[0268] Obtain a pre-configured label request text template, which includes at least one content field and a content filling area corresponding to each content field;
[0269] Determine the field that matches the first content information from at least one content field;
[0270] Fill the first content information into the content filling area corresponding to the field that matches the first content information to obtain a label generation request text;
[0271] Input the label generation request text into the target generative language model.
[0272] Please refer to Figure 10 , Figure 10 which is a schematic structural diagram of a training device for a label model provided by an embodiment of this application. As Figure 10 shown, the device 1000 may include a training sample set acquisition module 1010 and a training module 1020, where:
[0273] The training sample set acquisition module 1010 is used to obtain a training sample set, which includes a plurality of first samples, and each first sample corresponds to first sample information, and the first sample information includes a first label;
[0274] The training module 1020 is used to train a preset label model by using a plurality of first samples and the first sample information corresponding to each first sample until the training of the preset label model ends, to obtain a target label model, and the target label model is used to process the target information input to the target label model to obtain a target label corresponding to the target information, and the target information includes at least one of target text, target image, or target voice.
[0275] Optionally, the preset label model includes a feature extraction network and a generative language network. When the training module 1020 trains the preset label model using multiple first samples and the first sample information corresponding to each first sample, it can be used to:
[0276] Extract the features of the first sample through the feature extraction network to obtain the first sample features;
[0277] Perform label prediction based on the first sample features through the generative language network to obtain the predicted labels corresponding to the first samples;
[0278] If it is determined that the preset label model does not meet the preset first training end condition based on the predicted labels and the first labels, update the parameters of the generative language network and continue to train the preset label model;
[0279] If it is determined that the preset label model meets the first training end condition based on the predicted labels and the first labels, determine the preset label model that meets the first training end condition as the target label model.
[0280] Optionally, the first sample information further includes first content information, and the preset label model further includes an alignment network. When the training module 1020 trains the preset label model using multiple first samples and the first sample information corresponding to each first sample, it is further used to:
[0281] Extract the features of the first sample through the feature extraction network to obtain the second sample features;
[0282] Map the second sample features to the input dimension of the generative language network through the alignment network to obtain the predicted content information;
[0283] If it is determined that the preset label model does not meet the second training end condition based on the predicted content information and the first content information, update the parameters of the alignment network;
[0284] If it is determined that the preset label model meets the second training end condition based on the predicted content information and the first content information, extract the features of the first sample through the feature extraction network in the preset label model that meets the second training end condition to obtain the first sample features, and perform label prediction based on the first sample features through the generative language network in the preset label model that meets the second training end condition to obtain the predicted labels corresponding to the first samples;
[0285] If it is determined that the preset label model does not meet the preset first training end condition based on the predicted labels and the first labels, update the parameters of the generative language network, including:
[0286] If it is determined that the preset label model does not meet the preset first training end condition based on the predicted label and the first label, the parameters of the generative language network and the parameters of the alignment network are updated.
[0287] Please refer to Figure 11 , Figure 11 which is a schematic structural diagram of an information distribution device provided by an embodiment of the present application. As Figure 11 shown, the device 1100 may include an information acquisition module 1110, a label determination module 1120, and an information distribution module 1130, where:
[0288] The information acquisition module 1110 is configured to acquire target information to be distributed, and the target information includes at least one of target text, target image, or target voice;
[0289] The label determination module 1120 is configured to process the target information through a trained target label model to obtain a target label corresponding to the target information;
[0290] The information distribution module 1130 is configured to distribute the target information based on the target label.
[0291] Optionally, when the distribution module distributes the target information based on the target label, it may be configured to:
[0292] If the confidence corresponding to the target label is higher than or equal to a preset confidence threshold, the target information is distributed based on the target label;
[0293] The information distribution module 1130 is further configured to:
[0294] If the confidence corresponding to the target label is lower than the confidence threshold, the target information and the target label corresponding to the target information are sent to the target object;
[0295] Receive a new target label corresponding to the target information from the target object;
[0296] Distribute the target information based on the new target label.
[0297] The device according to the embodiment of the present application can execute the method provided by the embodiment of the present application, and its implementation principle is similar. The actions performed by each module in the device according to the embodiments of the present application correspond to the steps in the method according to the embodiments of the present application. For the detailed function description of each module of the device, reference may specifically be made to the description in the corresponding method shown above, and details are not described herein again.
[0298] An electronic device is provided in an embodiment of the present application, including a memory, a processor, and a computer program stored on the memory, and the processor executes the above computer program to implement the steps of the method according to any embodiment of the present application.
[0299] In an alternative embodiment, an electronic device is provided, such as Figure 12 shown, Figure 12 the electronic device 1200 shown includes: a processor 1201 and a memory 1203. Among them, the processor 1201 and the memory 1203 are connected, such as connected through a bus 1202. Optionally, the electronic device 1200 may further include a transceiver 1204, and the transceiver 1204 may be used for data interaction between the electronic device and other electronic devices, such as data sending and / or data receiving, etc. It should be noted that in practical applications, the transceiver 1204 is not limited to one, and the structure of the electronic device 1200 does not constitute a limitation to the embodiments of the present application.
[0300] The processor 1201 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logical blocks, modules, and circuits described in connection with the disclosure of the present application. The processor 1201 may also be a combination that implements a computing function, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.
[0301] The bus 1202 may include a path for transmitting information between the above components. The bus 1202 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. The bus 1202 may be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 12 only a thick line is used to represent it in [the figure], but it does not mean that there is only one bus or one type of bus.
[0302] The memory 1203 can be a ROM (Read Only Memory), or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory), or other types of dynamic storage devices that can store information and instructions. It can also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, other magnetic storage devices, or any other medium that can be used to carry or store computer programs and can be read by a computer, which is not limited herein.
[0303] The memory 1203 is used to store the computer program for implementing the embodiments of the present application and is controlled by the processor 1201 for execution. The processor 1201 is used to execute the computer program stored in the memory 1203 to implement the steps shown in the foregoing method embodiments.
[0304] The embodiments of the present application provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps and corresponding contents of the foregoing method embodiments can be implemented.
[0305] The embodiments of the present application further provide a computer program product, including a computer program. When the computer program is executed by a processor, the steps and corresponding contents of the foregoing method embodiments can be implemented.
[0306] It should be understood that although the flowcharts of the embodiments of the present application indicate various operation steps by arrows, the execution order of these steps is not limited to the order indicated by the arrows. Unless there is a clear description in this article, in some implementation scenarios of the embodiments of the present application, the implementation steps in each flowchart can be executed in other orders according to requirements. In addition, some or all of the steps in each flowchart may include multiple sub-steps or multiple stages based on the actual implementation scenarios. Some or all of these sub-steps or stages can be executed at the same time, and each sub-step or stage of these sub-steps or stages can also be executed at different times respectively. In the scenario where the execution times are different, the execution order of these sub-steps or stages can be flexibly configured according to requirements, and the embodiments of the present application do not limit this.
[0307] The above are only optional implementation manners of some implementation scenarios of this application. It should be noted that for those of ordinary skill in the art, without departing from the technical concept of the solution of this application, adopting other similar implementation means based on the technical idea of this application also belongs to the protection scope of the embodiments of this application.
Claims
1. A label generation method, characterized in that, Including: Obtain a first sample for which a tag is to be generated, where the first sample includes at least one of a first text sample, a first image sample, or a first voice sample; Identify first content information of the first sample; Input the first content information into a target generative language model to obtain a first tag determined by the target generative language model based on the first content information, and the first tag serves as the tag corresponding to the first sample; Wherein, the target generative language model is obtained through the following method: Input tag generation prompt information into a pre-trained initial generative language model to obtain the target generative language model, where the tag generation prompt information includes second content information of a second sample related to the first sample and a second tag corresponding to the second sample, and the second sample includes at least one of a second text sample, a second image sample, or a second voice sample.
2. The method according to claim 1, characterized in that The first sample includes at least one frame of a first image sample, and the identifying first content information of the first sample includes: Process the at least one frame of the first image sample through a pre-trained content extraction model to obtain first content information of the at least one frame of the first image sample, where the first content information includes at least one of first text information or second text information, the first text information is used to describe the image content of the at least one frame of the first image sample, and the second text information includes the text content in the at least one frame of the first image sample; Wherein, the second sample includes at least one frame of a second image sample related to the at least one frame of the first image sample, and the second content information includes at least one of third text information or fourth text information, the third text information is used to describe the image content of the at least one frame of the second image sample, and the fourth text information includes the text content in the at least one frame of the second image sample.
3. The method according to claim 2, wherein The first text information includes at least one of first sub-text information for describing the image scene of the at least one frame of the first image sample or second sub-text information for describing the image object of the at least one frame of the first image sample, and the second text information includes at least one of the caption content in the at least one frame of the first image sample or the title content in the at least one frame of the first image sample; The processing the at least one frame of the first image sample through the content extraction model to obtain first content information of the at least one frame of the first image sample includes at least one of the following: Perform image scene recognition processing on the at least one frame of the first image sample through the content extraction model to obtain the first sub-text information; Perform image object recognition processing on the at least one frame of the first image sample through the content extraction model to obtain the second sub-text information; Perform caption content extraction processing on the at least one frame of the first image sample through the content extraction model to obtain the caption content of the at least one frame of the first image sample; Perform title content extraction processing on the at least one frame of the first image sample through the content extraction model to obtain the title content of the at least one frame of the first image sample; Wherein, the third text information includes at least one of third sub - text information for describing the image frame of the at least one frame of second image samples or fourth sub - text information for describing the image objects of the at least one frame of second image samples, and the fourth text information includes at least one of the caption content in the at least one frame of second image samples or the title content in the at least one frame of second image samples.
4. The method according to claim 2, characterized in that, The at least one frame of first image samples and the at least one frame of second image samples belong to the same target video. The step of inputting the tag generation prompt information into a pre - trained initial generative language model to obtain the target generative language model includes: Obtaining target video information of the target video, where the target video information includes at least one of the video type or video time length of the target video; Determining the target number of the second image samples based on the target video information; Inputting the second content information of the target number of second image samples and the second tags corresponding to the target number of second image samples into a pre - trained initial generative language model to obtain the target generative language model.
5. The method according to any one of claims 1-4, characterized in that The tag generation prompt information includes a tag generation prompt text, which is obtained by the following method: Obtaining a pre - configured tag prompt text template, which includes at least one content field, a content filling area corresponding to each content field, a tag prompt field, and a tag filling area corresponding to the tag prompt field; Determining the field that matches the second content information from the at least one content field; Filling the second content information into the content filling area corresponding to the field that matches the second content information, and filling the second tag into the tag filling area to obtain the tag generation prompt text.
6. The method according to claim 5, wherein The step of inputting the first content information into the target generative language model includes: Obtaining a pre - configured tag request text template, which includes the at least one content field and a content filling area corresponding to each content field; Determining the field that matches the first content information from the at least one content field; Filling the first content information into the content filling area corresponding to the field that matches the first content information to obtain a tag generation request text; Inputting the tag generation request text into the target generative language model.
7. A training method for a label model, characterized in that, Including: Obtaining a training sample set, where the training sample set includes a plurality of first samples, each first sample corresponds to first sample information, and the first sample information includes a first tag. The first tag corresponding to the first sample is obtained by the method according to any one of claims 1 - 6; Training a preset tag model using the plurality of first samples and the first sample information corresponding to each first sample until the training of the preset tag model ends, to obtain a target tag model. The target tag model is used to process the target information input to the target tag model to obtain the target tag corresponding to the target information. The target information includes at least one of target text, target image, or target voice.
8. The method according to claim 7, characterized in that, The preset label model includes a feature extraction network and a generative language network. Training the preset label model by using the multiple first samples and the first sample information corresponding to each first sample includes: Extracting the features of the first sample through the feature extraction network to obtain first sample features; Performing label prediction based on the first sample features through the generative language network to obtain the predicted labels corresponding to the first sample; If it is determined that the preset label model does not meet the preset first training end condition based on the predicted labels and the first labels, then update the parameters of the generative language network and continue to train the preset label model; If it is determined that the preset label model meets the first training end condition based on the predicted labels and the first labels, then determine the preset label model that meets the first training end condition as the target label model.
9. The method according to claim 8, wherein The first sample information further includes the first content information, and the preset label model further includes an alignment network. Training the preset label model by using the multiple first samples and the first sample information corresponding to each first sample further includes: Extracting the features of the first sample through the feature extraction network to obtain second sample features; Mapping the second sample features to the input dimension of the generative language network through the alignment network to obtain predicted content information; If it is determined that the preset label model does not meet the second training end condition based on the predicted content information and the first content information, then update the parameters of the alignment network; If it is determined that the preset label model meets the second training end condition based on the predicted content information and the first content information, then extract the features of the first sample through the feature extraction network in the preset label model that meets the second training end condition to obtain first sample features, and perform label prediction based on the first sample features through the generative language network in the preset label model that meets the second training end condition to obtain the predicted labels corresponding to the first sample; The step of, if it is determined that the preset label model does not meet the preset first training end condition based on the predicted labels and the first labels, then update the parameters of the generative language network, includes: If it is determined that the preset label model does not meet the preset first training end condition based on the predicted labels and the first labels, then update the parameters of the generative language network and the parameters of the alignment network.
10. An information distribution method, characterized in that, Including: Obtaining target information to be distributed, where the target information includes at least one of target text, target image, or target voice; Processing the target information through the trained target label model to obtain the target labels corresponding to the target information, where the target labeling labels are trained by the method according to any one of claims 7-9; Distributing the target information based on the target labels.
11. The method according to claim 10, wherein The distributing the target information based on the target labels includes: If the confidence level corresponding to the target label is higher than or equal to a preset confidence level threshold, then distribute the target information based on the target label; The method further includes: If the confidence level corresponding to the target label is lower than the confidence level threshold, sending the target information and the target label corresponding to the target information to a target object; Receiving a new target label corresponding to the target information from the target object; Distributing the target information based on the new target label.
12. A label generating device, characterized in that, It includes: A sample acquisition module, configured to acquire a first sample for which a label is to be generated, where the first sample includes at least one of a first text sample, a first image sample, or a first voice sample; A content information recognition module, configured to recognize first content information of the first sample; A label generation module, configured to input the first content information into a target generative language model to obtain a first label determined by the target generative language model based on the first content information, and the first label serves as the label corresponding to the first sample; Wherein, the target generative language model is obtained by the label generation module in the following manner: Inputting label generation prompt information into a pre-trained initial generative language model to obtain the target generative language model, where the label generation prompt information includes second content information of a second sample related to the first sample and a second label corresponding to the second sample, and the second sample includes at least one of a second text sample, a second image sample, or a second voice sample.
13. A training device for a label model, characterized in that, It includes: A training sample set acquisition module, configured to acquire a training sample set, where the training sample set includes a plurality of first samples, and each first sample corresponds to first sample information, and the first sample information includes a first label, and the first label corresponding to the first sample is obtained by the device according to claim 12; A training module, configured to use the plurality of first samples and the first sample information corresponding to each first sample to train a preset label model until the preset label model finishes training, to obtain a target label model, and the target label model is configured to process target information input to the target label model to obtain a target label corresponding to the target information, where the target information includes at least one of target text, target image, or target voice.
14. An information distribution device, characterized in that, It includes: An information acquisition module, configured to acquire target information to be distributed, where the target information includes at least one of target text, target image, or target voice; A label determination module, configured to process the target information through a trained target label model to obtain a target label corresponding to the target information, where the target tagging label is trained based on the device according to claim 13; An information distribution module, configured to distribute the target information based on the target label.
15. An electronic device, comprising a memory, a processor, and a computer program stored on the memory, characterized in that, The processor executes the computer program to implement the steps of the method according to any one of claims 1-11.
16. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1-11.
17. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1-11.