Text label prediction method and device and storage medium

Through joint training of the comprehensive label scoring model, the problems of multimedia resource label prediction accuracy and score prediction are solved, and high-accurate multimedia resource recommendation is achieved.

CN120030217APending Publication Date: 2025-05-23TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311564569.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-21
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

In the application scenarios of searching and recommendation of multimedia resources, it is difficult to accurately predict the labels and scores of multimedia resources, which affects the recommendation effect.

Method used

A comprehensive label scoring model is adopted. This model can predict both text labels and tag scores to be trained, and improve prediction accuracy by jointly training the label generation model to be trained.

Benefits of technology

Accurate label prediction and score prediction of multimedia resources are achieved, improving the accuracy and user experience of the recommendation system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030217A_ABST
    Figure CN120030217A_ABST
Patent Text Reader

Abstract

The invention discloses a text label prediction method and device and a storage medium, which can be applied to various scenes such as cloud technology, artificial intelligence, smart traffic, Internet of Vehicles and the like, and the method comprises the following steps: inputting a target multimedia resource into a comprehensive label scoring model to obtain a target text label and a target score of the target text label; the comprehensive label scoring model is obtained by training a to-be-trained label scoring model based on the difference between the first sample score result and the first sample score label; the first sample score result is obtained by inputting the first sample prediction placeholder into a to-be-trained scoring model in the to-be-trained label scoring model for score prediction training; the first sample prediction placeholder and the first sample prediction result are obtained by inputting the first sample multimedia resource marked with the first sample text label into a to-be-trained label generation model in a to-be-trained label scoring model and performing label prediction processing. According to the invention, the prediction accuracy of the comprehensive label scoring model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a text label prediction method, device and storage medium. Background Art

[0002] In the related art, in application scenarios such as searching and recommending multimedia resources, it is necessary to accurately determine the labels of multimedia resources. In one example, after determining the labels of multimedia resources, multimedia resources can be recommended to users based on the labels of multimedia resources. For example, in a search scenario, based on the labels of multimedia resources, corresponding multimedia resources are recommended to users as search results; for another example, based on the labels of media content, corresponding multimedia resources are actively recommended to users. Therefore, it is particularly important to accurately determine the labels of multimedia resources. Summary of the invention

[0003] The present application provides a text tag prediction method, device and storage medium, which can accurately predict text tags and a comprehensive tag scoring model for accurately predicting the scores of text tags.

[0004] On the one hand, the present application provides a text label prediction method, the method comprising:

[0005] Acquire target multimedia resources;

[0006] Inputting the target multimedia resource into a comprehensive label scoring model for label prediction processing and predicted label scoring processing to obtain a target text label and a target score for the target text label; the target score represents the prediction accuracy of the target text label;

[0007] Among them, the comprehensive label scoring model is obtained by training the label scoring model to be trained based on the difference between the first sample score result and the first sample score label; the first sample score result is obtained by inputting the first sample prediction placeholder into the to-be-trained scoring model in the to-be-trained label scoring model for score prediction training; the first sample score label is constructed based on the difference between the first sample prediction result and the first sample text label; the first sample prediction placeholder represents the score of the first sample prediction result; the first sample prediction placeholder and the first sample prediction result are obtained by inputting the first sample multimedia resource marked with the first sample text label into the to-be-trained label generation model in the to-be-trained label scoring model for label prediction processing.

[0008] On the other hand, a text label prediction device is provided, the device comprising:

[0009] A target resource acquisition module is used to acquire target multimedia resources;

[0010] A prediction module is used to input the target multimedia resource into a comprehensive label scoring model for label prediction processing and predicted label scoring processing to obtain a target text label and a target score for the target text label; the target score represents the prediction accuracy of the target text label;

[0011] Among them, the comprehensive label scoring model is obtained by training the label scoring model to be trained based on the difference between the first sample score result and the first sample score label; the first sample score result is obtained by inputting the first sample prediction placeholder into the to-be-trained scoring model in the to-be-trained label scoring model for score prediction training; the first sample score label is constructed based on the difference between the first sample prediction result and the first sample text label; the first sample prediction placeholder represents the score of the first sample prediction result; the first sample prediction placeholder and the first sample prediction result are obtained by inputting the first sample multimedia resource marked with the first sample text label into the to-be-trained label generation model in the to-be-trained label scoring model for label prediction processing.

[0012] On the other hand, an electronic device is provided, comprising a processor and a memory, wherein the memory stores at least one instruction or at least one program, and the at least one instruction or the at least one program is loaded and executed by the processor to implement the text label prediction method as described above.

[0013] On the other hand, a computer storage medium is provided, wherein the computer storage medium stores at least one instruction or at least one program, and the at least one instruction or at least one program is loaded and executed by a processor to implement the text tag prediction method as described above.

[0014] On the other hand, a computer program product or a computer program is provided, the computer program product or the computer program includes computer instructions, the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes to implement the text label prediction method as described above.

[0015] The text label prediction method, device and storage medium provided by this application have the following technical effects:

[0016] The present application obtains a target multimedia resource; inputs the target multimedia resource into a comprehensive label scoring model for label prediction processing and predicted label scoring processing to obtain a target text label and a target score of the target text label; the target score represents the prediction accuracy of the target text label; the comprehensive label scoring model of the present application can not only predict the target text label of the target multimedia resource, but also predict the target score of the target text label, so that the user can know the prediction accuracy of the target text label; wherein the comprehensive label scoring model is obtained by training a label scoring model to be trained based on the difference between the first sample score result and the first sample score label; the first sample score result is obtained by inputting a first sample prediction placeholder into the label scoring model to be trained for score prediction training; the first sample score label is constructed based on the difference between the first sample prediction result and the first sample text label; the first sample prediction placeholder represents the score of the first sample prediction result; the first sample prediction placeholder and the first sample prediction result are obtained by inputting a first sample multimedia resource marked with the first sample text label into a label generation model to be trained in the label scoring model to be trained, and performing label prediction processing. The present application jointly trains a label generation model to be trained and a label scoring model to be trained, thereby obtaining a comprehensive label scoring model that can predict both text labels and scores of text labels. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions and advantages of the embodiments of this specification or the prior art, the drawings required for use in the embodiments or the prior art descriptions are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0018] Figure 1 is a schematic diagram of a text label prediction system provided by an embodiment of this specification;

[0019] Figure 2 It is a flowchart of a text label prediction method provided by an embodiment of this specification;

[0020] Figure 3 It is a flowchart of a training method for a comprehensive label scoring model provided in an embodiment of this specification;

[0021] Figure 4 It is a flowchart of a method provided in an embodiment of this specification for training the above-mentioned label scoring model to be trained based on the difference between the above-mentioned first sample score result and the above-mentioned first sample score label to obtain the above-mentioned comprehensive label scoring model;

[0022] Figure 5 It is a flowchart of a method for training the above-mentioned label generation model to be trained based on the second sample multimedia resource and the above-mentioned scoring model to obtain a label generation model provided by an embodiment of this specification;

[0023] Figure 6 It is a schematic diagram of a training process of a comprehensive label scoring model provided in an embodiment of this specification;

[0024] Figure 7 It is a flowchart of a method for training the above-mentioned large language model based on the third loss data to obtain the above-mentioned label generation model to be trained, provided by an embodiment of this specification;

[0025] Figure 8 It is a schematic diagram of a training process of a label generation model to be trained provided in an embodiment of this specification;

[0026] Fig. 9 This is a schematic diagram of an application scenario of a text label prediction method provided in an embodiment of this specification;

[0027] Fig.10 is a structural schematic diagram of a text label prediction device provided in an embodiment of this specification;

[0028] Fig.11 It is a structural diagram of a server provided in an embodiment of this specification. DETAILED DESCRIPTION

[0029] The following will be combined with the drawings in the embodiments of this specification to clearly and completely describe the technical solutions in the embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0030] First, some nouns or terms that appear in the description of the embodiments of this specification are explained as follows:

[0031] Large language model: A large language model is a natural language processing model built based on deep learning technology. It can generate, understand, and process natural language text by learning the statistical laws and semantic features of language. It is widely used in tasks such as language modeling, machine translation, text summarization, and sentiment analysis. The large language model generates coherent text by predicting the probability of the next word or character. Because of the large number of parameters, the large language model is more effective than the general language model.

[0032] Short video tag system: The short video tag system is a system that inputs short video content and outputs tags related to the video. The output tags can better manage and organize content. Tags can help users find videos of interest more easily and increase video exposure and playback volume. Operators can improve the recommendation effect of videos by reasonably allocating tags and accurately describing video content and themes. The tagging system helps short video platforms achieve refined operations and improve user experience.

[0033] Tag generation: Tag generation is a natural language processing technology that generates appropriate tags for text content by analyzing it. These tags can reflect information such as the subject, keywords, and entities of the text, which helps to better manage and organize content. Tag generation methods include rule-based, statistical model, and large language model. Through tag generation, the accuracy and efficiency of information retrieval can be improved, and refined operations and personalized recommendations can be achieved.

[0034] Label sorting: Re-score and sort the recalled and generated labels in the labeling system to obtain better accuracy and recall indicators.

[0035] Artificial Intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines so that machines have the functions of perception, reasoning and decision-making.

[0036] Artificial intelligence technology is a comprehensive discipline that covers a wide range of fields, including both hardware-level and software-level technologies. Basic artificial intelligence technologies generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, mechatronics, etc. Among them, pre-trained models are also called large models and basic models. After fine-tuning, they can be widely used in downstream tasks in various major directions of artificial intelligence. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0037] Machine Learning (ML) is a multi-disciplinary interdisciplinary subject involving probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory and other disciplines. It specializes in studying how computers simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications are spread across all areas of artificial intelligence. Machine learning and deep learning usually include artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning by teaching. The pre-trained model is the latest development of deep learning, which integrates the above technologies.

[0038] With the research and advancement of artificial intelligence technology, artificial intelligence technology has been studied and applied in many fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, driverless cars, autonomous driving, drones, digital twins, virtual humans, robots, artificial intelligence generated content (AIGC), conversational interaction, smart medical care, smart customer service, game AI, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.

[0039] It is understandable that in the specific implementation of the present application, related data such as user information and user browsing history are involved. When the above embodiments of the present application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of relevant data need to comply with relevant laws, regulations and standards of relevant countries and regions.

[0040] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or server that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0041] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.

[0042] In the short video application scenario, there are usually two types of schemes for labeling short videos: discriminative schemes and generative schemes. In the discriminative scheme, the labeling system needs an additional recall subsystem, such as a face label recall subsystem that is responsible for recalling the labels corresponding to the characters that appear in the video. The labeling system of the discriminative scheme also requires a sorting model to fuse and re-score the recalled labels from each sub-recall source. In the generative scheme, the text news related to the short video is directly input into a generative language model to generate relevant labels.

[0043] The disadvantage of the discriminative approach to short video tagging is that an additional recall subsystem is required to provide recall for the sorting model, which adds additional costs. In addition, the recall rate of the tag system is determined by the recall subsystem. When the recall subsystems cannot be updated quickly, they cannot quickly support the addition of new and popular tags. In addition, since the discriminative approach requires more models, multiple services need to be deployed at the system level, which increases the update and maintenance costs.

[0044] The generative solution utilizes the ability of the generative model and can generate labels using the autoregressive method, so there is no need for an additional recall subsystem to provide recall labels. However, the generative solution can only generate text and cannot obtain the probability of the generated labels, that is, it has no ability to output the generated label scores. The labeling system of the generative solution cannot filter the generated labels by scores, and the model obtained by its training solution is also difficult to achieve high accuracy and high recall.

[0045] This embodiment combines the advantages of the discriminative scheme and the generative scheme, and uses two-stage training. The first stage focuses on the generative training of the model to enable it to have the ability to recall labels. The second stage, based on the first stage, scores the latent variables corresponding to the flag bits of the generated labels to complete the discriminative supervised learning, so that the generative model also has the ability to output scores. This embodiment jointly trains the label generation model to be trained and the label scoring model to be trained, so as to obtain a comprehensive label scoring model that can predict both text labels and scores of text labels. This model has both high recall rate and high accuracy.

[0046] See also Figure 1 , Figure 1 is a schematic diagram of a text label prediction system provided by an embodiment of this specification, such as Figure 1 As shown, the text tag prediction system may include at least a server 01 and a client 02 .

[0047] Specifically, in the embodiments of this specification, the server 01 may include an independently operated server, or a distributed server, or a server cluster composed of multiple servers, and may also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. Server 01 may include a network communication unit, a processor, a memory, and the like. Specifically, the server 01 can be used to train a comprehensive label scoring model, and predict the target text label of the target multimedia resource and the target score of the target text label according to the comprehensive label scoring model.

[0048] Specifically, in the embodiments of this specification, the client 02 may include physical devices such as smart phones, desktop computers, tablet computers, laptop computers, digital assistants, smart wearable devices, smart speakers, vehicle terminals, smart TVs, etc., and may also include software running in physical devices, such as web pages provided to users by some service providers, or applications provided to users by these service providers. Specifically, the client 02 may be used to display the target text label of the target multimedia resource predicted by the comprehensive label scoring model and the target score of the target text label.

[0049] The following introduces a text label prediction method of this application. Figure 2 It is a flowchart of a text label prediction method provided in an embodiment of this specification. This specification provides method operation steps as described in the embodiment or flowchart, but may include more or fewer operation steps based on conventional or non-creative labor. The order of steps listed in the embodiment is only one way of executing the steps among many orders, and does not represent the only order of execution. When the actual system or server product is executed, it can be executed in the order shown in the embodiment or the accompanying drawings or in parallel (for example, in a parallel processor or multi-threaded processing environment). Specifically, Figure 2 As shown, the above method may include:

[0050] S201: Acquire target multimedia resources.

[0051] In the embodiments of the present specification, the target multimedia resources may include but are not limited to video, audio, image and other information.

[0052] S203: Inputting the target multimedia resource into the comprehensive label scoring model for label prediction processing and predicted label scoring processing to obtain a target text label and a target score for the target text label; the target score represents the prediction accuracy of the target text label;

[0053] Among them, the above-mentioned comprehensive label scoring model is obtained by training the label scoring model to be trained based on the difference between the first sample score result and the first sample score label; the above-mentioned first sample score result is obtained by inputting the first sample prediction placeholder into the above-mentioned label scoring model to be trained for score prediction training; the above-mentioned first sample score label is constructed based on the difference between the first sample prediction result and the first sample text label; the above-mentioned first sample prediction placeholder represents the score of the above-mentioned first sample prediction result; the above-mentioned first sample prediction placeholder and the above-mentioned first sample prediction result are obtained by inputting the first sample multimedia resource marked with the above-mentioned first sample text label into the to-be-trained label generation model in the above-mentioned label scoring model to be trained, and performing label prediction processing.

[0054] In an embodiment of the present specification, the target multimedia resource can be input into a comprehensive label scoring model for label prediction processing and predicted label scoring processing, and the text label features can be extracted through the comprehensive label scoring model to obtain the target text label, and the label score features can be extracted through the comprehensive label scoring model to obtain the target score of the above-mentioned target text label.

[0055] In the embodiment of the present specification, the target multimedia resource is input into the comprehensive label scoring model for label prediction processing and predicted label scoring processing to obtain the target text label and the target score of the target text label, including:

[0056] Input the target multimedia resource into the label generation model in the comprehensive label scoring model for label prediction processing to obtain the target text label and the label score placeholder corresponding to the target text label;

[0057] The above-mentioned label score placeholder is input into the scoring model in the above-mentioned comprehensive label scoring model for label scoring processing to obtain the target score corresponding to the above-mentioned target text label.

[0058] In an embodiment of the present specification, the comprehensive label scoring model may include a label generation model and a scoring model. The target multimedia resource is input into the label generation model in the comprehensive label scoring model for label prediction processing to obtain the target text label and the label score placeholder corresponding to the target text label; the label score placeholder is then input into the scoring model in the comprehensive label scoring model for label scoring processing to obtain the target score corresponding to the target text label; wherein the target score can represent the pre-stored accuracy of the target text label.

[0059] In the embodiment of the present specification, the target multimedia resource is input into the comprehensive label scoring model for label prediction processing and predicted label scoring processing to obtain the target text label and the target score of the target text label, including:

[0060] Input the target multimedia resource into the comprehensive label scoring model for label prediction and predicted label scoring to obtain at least two target text labels and a target score corresponding to each target text label;

[0061] Exemplarily, the above method further includes:

[0062] According to the target score corresponding to each target text label, a label to be recommended is screened out from the at least two target text labels; the target score of the label to be recommended is greater than the target score of the non-recommended label; the non-recommended label is a label other than the label to be recommended among the at least two target text labels;

[0063] The tag to be recommended is determined as the text tag of the target multimedia resource.

[0064] In the embodiment of the present specification, the comprehensive tag scoring model can predict at least two target text tags and the target score of each target text tag; thereby, the at least two target text tags can be sorted according to their corresponding target scores, and a preset number of tags to be recommended can be screened out according to the sorting results;

[0065] Exemplarily, at least two target text tags can be sorted from large to small according to their corresponding target scores, and a preset number of tags with a higher order can be used as tags to be recommended; or at least two target text tags can be sorted from small to large according to their corresponding target scores, and a preset number of tags with a lower order can be used as tags to be recommended; then the tag to be recommended is determined as the text tag of the above-mentioned target multimedia resource, and the target multimedia resource and its corresponding tag to be recommended can be stored in a database, thereby improving the accuracy of the text tag of the target multimedia resource.

[0066] In the embodiment of the present specification, the number of the target multimedia resources is at least two, and the method further includes:

[0067] Acquire historical multimedia resources of the target object;

[0068] Based on the above historical multimedia resources, determine the historical text label corresponding to the above target object;

[0069] Determine the target multimedia resource that matches the to-be-recommended tag with the historical text tag as the to-be-recommended multimedia resource;

[0070] The recommended multimedia resources are recommended to the target object.

[0071] In an embodiment of the present specification, in a multimedia resource recommendation scenario, historical multimedia resources of a target object can be obtained, and at least two corresponding candidate text tags and the tag scores corresponding to each candidate text tag can be predicted through a comprehensive tag scoring model; and based on the tag scores, tags with higher scores are screened out from at least two candidate text tags as historical text tags; and then, based on the historical text tags, tags to be recommended that match the above historical text tags are screened out from multiple tags to be recommended, thereby screening out multimedia resources to be recommended from at least two target multimedia resources; and the above recommended multimedia resources are recommended to the target object, thereby improving the accuracy of resource recommendations and the conversion rate of recommended multimedia resources.

[0072] In the embodiment of the present specification, the number of the target multimedia resources is at least two, and the method further includes:

[0073] Get the search tag of the target object;

[0074] Determine the to-be-recommended tags that match the search tags, and obtain the matching recommended tags;

[0075] Obtain the target multimedia resource corresponding to the matching recommendation tag to obtain the recommended multimedia resource;

[0076] The recommended multimedia resources are recommended to the target object.

[0077] In an embodiment of the present specification, in a search application scenario, a search tag of a target object can be obtained; a tag to be recommended that matches the search tag is determined to obtain a matching recommended tag; then the target multimedia resource corresponding to the matching recommended tag is obtained to obtain a recommended multimedia resource; thereby recommending the recommended multimedia resource to the target object, which can improve the accuracy of resource recommendation.

[0078] In the embodiments of this specification, Figure 3 As shown, the training method of the above comprehensive label scoring model includes:

[0079] S301: Acquire a first sample multimedia resource, where the first sample multimedia resource is marked with a first sample text tag;

[0080] In the embodiment of the present specification, the first sample multimedia resource may include but is not limited to video, audio, image and other information. The first sample text label may be a label representing the content of the first sample multimedia resource. During the training process, the first sample text label may be one or at least two.

[0081] S303: Inputting the first sample multimedia resource into the to-be-trained label generation model in the to-be-trained label scoring model, performing label prediction processing, and obtaining a first sample prediction result and a first sample prediction placeholder corresponding to the first sample prediction result; the first sample prediction placeholder represents the score of the first sample prediction result;

[0082] In the embodiments of this specification, the above-mentioned label generation model to be trained is obtained based on the training of a large language model; the label generation model to be trained has the function of recalling labels using its own generation ability, and through training, each label generated by it is followed by a placeholder. The training goal of this stage is to use these placeholders to perform scoring layer training for each candidate label, so that the model can obtain the scoring ability that the generative model does not have.

[0083] In an embodiment of the present specification, a video title, speech recognition text, and character recognition text can be extracted based on a first sample multimedia resource, and then the video title, speech recognition text, and character recognition text can be input into a label generation model to be trained for label prediction processing to obtain a first sample prediction result and a first sample prediction placeholder corresponding to the first sample prediction result; the first sample prediction placeholder represents the score of the first sample prediction result.

[0084] S305: constructing a first sample score label based on the difference between the first sample prediction result and the first sample text label;

[0085] In the embodiments of the present specification, the model parameters of the label generation model to be trained can be fixed first, and the scoring model to be trained can be trained; according to the output results of the label generation model to be trained, the first sample score label can be constructed, and the scoring model to be trained can be trained according to the first sample score label.

[0086] S307: Inputting the first sample prediction placeholder into the to-be-trained scoring model in the to-be-trained label scoring model to perform score prediction training to obtain a first sample score result;

[0087] In the embodiments of the present specification, the above-mentioned first sample score result represents the accuracy of the above-mentioned first sample prediction result; the scoring model to be trained can be composed of a fully connected layer, a relu (activation) layer and a sigmoid function, and can also be a model of other structures; the sigmoid function is a mathematical function in a neural network, which maps the input value to an output between 0 and 1. It is usually used as an activation function in artificial neural networks. The input of the scoring model is the feature of the last_hidden_states (last hidden layer) of the large language model. The high-dimensional features are mapped to a 1-dimensional score through a fully connected layer, and the score is mapped to a value range of 0-1 using a sigmoid function.

[0088] S309: Based on the difference between the first sample score result and the first sample score label, train the to-be-trained label scoring model to obtain the comprehensive label scoring model.

[0089] In the embodiments of the present specification, the model parameters of the label generation model to be trained can be fixed (frozen) first, and loss data can be constructed based on the difference between the first sample score result and the above-mentioned first sample score label. The model parameters of the label scoring model to be trained can be adjusted according to the loss data until the training end conditions are met to obtain a comprehensive label scoring model.

[0090] In the embodiments of this specification, Figure 4 As shown, based on the difference between the first sample score result and the first sample score label, the above-mentioned label scoring model to be trained is trained to obtain the above-mentioned comprehensive label scoring model, including:

[0091] S3091: constructing first loss data based on the difference between the first sample score result and the first sample score label;

[0092] S3093: Perform score prediction training on the to-be-trained scoring model based on the first loss data to obtain a scoring model;

[0093] S3095: Training the to-be-trained label generation model based on the second sample multimedia resource and the scoring model to obtain a label generation model;

[0094] In the embodiments of the present specification, after the scoring model training is completed, the model parameters of the scoring model can be fixed, and the label generation model to be trained can be trained to obtain a label generation model. By alternately training the two models to be trained, a comprehensive label scoring model that can predict both text labels and scores of text labels can be quickly trained.

[0095] In the embodiments of this specification, Figure 5As shown, the above-mentioned label generation model to be trained is trained based on the second sample multimedia resource and the above-mentioned scoring model to obtain the label generation model, including:

[0096] S30951: Acquire a second sample multimedia resource, where the second sample multimedia resource is marked with a second sample score label;

[0097] In the embodiment of this specification, the second sample multimedia resource and the first sample multimedia resource are resources of the same type, and the two can be the same resource or different resources. The second sample multimedia resource is also marked with a second sample text label, and the second sample score label represents the accuracy of the second sample text label.

[0098] S30953: Input the second sample multimedia resource into the label generation model to be trained, perform label prediction processing, and obtain a second sample prediction result and a second sample prediction placeholder corresponding to the second sample prediction result;

[0099] In this embodiment of the specification, the second sample prediction placeholder and the first sample prediction placeholder are placeholders of the same type.

[0100] S30955: Input the second sample prediction placeholder into the scoring model to perform score prediction processing to obtain a second sample score result;

[0101] S30957: constructing second loss data based on the difference between the second sample score result and the second sample score label;

[0102] S30959: Train the label generation model to be trained based on the second loss data to obtain the label generation model.

[0103] In the embodiment of the present specification, the model parameters of the scoring model can be fixed, and according to the second loss data, the model parameters of the above-mentioned label generation model to be trained are adjusted until the training end condition is met, and the label generation model to be trained at the end of the training is used as the label generation model. The training end condition can be determined by the number of training iterations and / or the loss threshold corresponding to the second loss data.

[0104] S3097: Based on the above-mentioned scoring model and the above-mentioned label generation model, determine the above-mentioned comprehensive label scoring model.

[0105] In the embodiments of the present specification, the scoring model can be combined with the label generation model to obtain a comprehensive label scoring model, thereby realizing the prediction of text labels and the scores corresponding to the text labels.

[0106] In the embodiments of this specification, Figure 6 As shown, Figure 6The figure is a flowchart of a training method for a comprehensive label scoring model, wherein the second sample multimedia resource is a video. For each video, the corresponding title, ASR (Automatic Speech Recognition) text, OCR (Optical Character Recognition) text, and machine classification text of the video are spliced ​​together to form a sentence. "The basic information of the video is:" is added in front of this sentence, and "Please generate the key tags corresponding to the video" is added after this sentence. This complete text is processed through the encode function of the tokenizer (vectorized text class) to become multiple tokens, and these tokens are input into the large language model (the label generation model to be trained) to generate the corresponding answer. Unlike the training goal of the first stage, the second stage mainly trains the ability of the scoring model, so the parameters of the large language model are fixed and only the parameters of the scoring model are trained. Since the answers generated by the large language model are random, the label of each training data needs to be adjusted dynamically. Assume that the answer text generated by the large language model is: "Humanities <s>Character Stories <s>Historical Archives <s>Historical Documentary <s>Tongji University <s>Domestic documentary <s>CCTV Records <s>Documentary <s>", these labels are used as candidate labels, and they are calculated in order to see if they are in the training labels. The training gt (label) can be obtained as [0, 1, 0, 1, 0, 1, 0, 1], where 1 represents the generated candidate label is correct and 0 represents the generated candidate label is wrong. During the training process, the model's last_hidden_states (last hidden layer) is replaced by <s>The corresponding vector is input into the scoring model to obtain the sample score result. Finally, the loss data is calculated based on the sample score result and the training GT, and the large language model (the label generation model to be trained) is trained through back propagation of the loss data.

[0107] In the embodiments of this specification, Figure 7 As shown, the above method also includes:

[0108] S701: Acquire a third sample multimedia resource, where the third sample multimedia resource is marked with a third sample text tag;

[0109] S703: Input the third sample multimedia resource into the large language model for label prediction processing to obtain a third sample prediction result and a third sample prediction placeholder corresponding to the third sample prediction result;

[0110] S705: constructing third loss data based on the difference between the third sample prediction result and the third sample text label;

[0111] S707: Based on the third loss data, the large language model is trained to obtain the label generation model to be trained.

[0112] In the embodiments of this specification, the training goal at this stage is to make full use of the generative large language model's generation capabilities, so that the large language model learns to extract, summarize, and generalize keywords from the text to form key tags, so the training data must be in the data format of training the generative model. Exemplarily, first collect business data, the required number is 60,000 videos; perform ASR speech recognition on the voices in these 60,000 videos to obtain the ASR text of the video, perform OCR optical character recognition on the video frames of these 60,000 videos to obtain the OCR text of the 60,000 videos. In addition, these 60,000 videos can be submitted to manual review, marked with corresponding labels, and obtain the training goal of 60,000 videos. Since the content classification of multimedia resources has a great improvement effect on this task, the machine review classification function can be used to mark the multimedia resources with machine review classification labels.

[0113] The large language model in this embodiment may include but is not limited to BERT, T5, and GLM; BERT is a large language model developed by Google, which is pre-trained using the Transformer architecture and has achieved state-of-the-art results in various natural language processing tasks. T5 is a pre-trained language model developed by Google with 1.1 billion parameters that can perform a variety of natural language processing tasks, such as text generation, question answering, summarization, etc. It is based on the Transformer structure and is pre-trained in an autoregressive manner, and can be fine-tuned to adapt to various natural language processing tasks, such as question answering, summarization, translation, etc. Unlike other language models, T5 uses natural language prompts to guide the model to generate text, and can be used to process natural language text in different languages. GLM (General Language Model) is a general language model that is pre-trained using an autoregressive fill-in-the-blank objective and can be fine-tuned for various natural language understanding and generation tasks.

[0114] Exemplarily, the large language model used in this embodiment can be a bloomz-3b model, which has 3 billion parameters. Bloom is a Transformer language model with only a decoder. It is trained on the ROOTS corpus, which includes hundreds of sources in 46 natural languages ​​and 13 programming languages. This model is a generative model. When a text is input, the model has the ability to continue writing the text, that is, the ability to answer questions.

[0115] Exemplarily, the third sample multimedia resource may be a video, such as Figure 8 As shown, Figure 8 The following is a diagram of the training process of a label generation model to be trained. For each video, the corresponding title, ASR (Automatic Speech Recognition) text, OCR (Optical Character Recognition) text, and machine classification text are spliced ​​together to form a sentence. "The basic information of the video is:" is added in front of this sentence, and "Please generate the key tags corresponding to this video" is added after this sentence. This complete text is processed through the encode function of the tokenizer (vectorized text class) and turned into multiple tokens. These tokens are input into the large language model (label generation model to be trained) to generate the corresponding answers. The loss data is calculated based on the difference between the answer generated by the large language model and the training target, and the large language model is trained by backpropagation of the loss data. After this stage of training, the model has learned the ability to generate relevant labels through input content, and a placeholder will be generated after each label. <s>In this stage of training, the placeholder has no practical meaning, and the model only needs to generate the placeholder according to the training data.

[0116] In the embodiment of the present specification, the first sample multimedia resource is marked with a current sample score label, and the above-mentioned label scoring model to be trained is trained based on the difference between the first sample score result and the first sample score label to obtain the above-mentioned comprehensive label scoring model, including:

[0117] Based on the difference between the first sample score result and the first sample score label, score prediction training is performed on the to-be-trained scoring model to obtain a current scoring model;

[0118] The above-mentioned label generation model to be trained is used as the current label generation model, and the above-mentioned first sample prediction placeholder is input into the above-mentioned current scoring model for score prediction processing to obtain the current sample score result;

[0119] Based on the difference between the current sample score result and the current sample score label, the current label generation model is trained, and the trained model is used as the current label generation model again;

[0120] Input the first sample multimedia resource into the current label generation model, perform label prediction processing, and obtain a current sample prediction result and a current sample prediction placeholder corresponding to the current sample prediction result;

[0121] Based on the current sample prediction placeholder, the current scoring model and the current label generation model are iteratively trained to obtain the comprehensive label scoring model.

[0122] In the embodiments of the present specification, the parameters of the current scoring model can be fixed, the current label generation model can be trained, and the current label generation model obtained by this training can be used to update the current label generation model trained last time, and then the parameters of the updated current label generation model can be fixed to train the current scoring model; the training method for the current scoring model is similar to the training method for the current scoring model, and alternating iterative training of the current scoring model and the current label generation model is implemented, thereby further improving the accuracy of the comprehensive label scoring model.

[0123] In the embodiment of the present specification, the current scoring model and the current label generation model are iteratively trained based on the current sample prediction placeholder to obtain the comprehensive label scoring model, including:

[0124] Input the current sample prediction placeholder into the current scoring model to perform score prediction training to obtain a scoring result, and use the scoring result as the current sample score result again;

[0125] Repeat the above steps of training the current label generation model based on the difference between the current sample score result and the current sample score label, and reusing the trained model as the current label generation model; and inputting the current sample prediction placeholder into the current scoring model for score prediction training to obtain a scoring result, and reusing the scoring result as the current sample score result, until the training end condition is met;

[0126] The current scoring model at the end of training is determined as the scoring model, and the current scoring model at the end of training is determined as the label generation model;

[0127] Based on the above-mentioned scoring model and the above-mentioned label generation model, the above-mentioned comprehensive label scoring model is determined.

[0128] In the embodiments of the present specification, after the current label generation model is trained, the model parameters of the current label generation model can be fixed and the current scoring model can be trained; when training the current label generation model, the model parameters of the current scoring model are fixed, thereby realizing alternating iterative training of the two models, and determining the training end conditions based on the number of iterations, thereby further improving the accuracy of the comprehensive label scoring model.

[0129] In the embodiments of this specification, Fig. 9 As shown, Fig. 9 The figure is a schematic diagram of application scenarios of a text label prediction method. The application scenarios may include three scenarios: recommendation scenario, search scenario and storage scenario.

[0130] (1) In a tag storage scenario, at least two target text tags can be sorted according to their corresponding target scores, and a preset number of tags to be recommended can be screened out based on the sorting results; the target multimedia resources and their corresponding tags to be recommended are stored in a database, thereby improving the accuracy of the text tags of the target multimedia resources.

[0131] (2) In the multimedia resource recommendation scenario, the historical multimedia resources of the target object can be obtained, and at least two candidate text tags corresponding to the target object and the tag scores corresponding to each candidate text tag can be predicted through a comprehensive tag scoring model; and based on the tag scores, tags with higher scores are screened out from at least two candidate text tags as historical text tags; and then, based on the historical text tags, tags to be recommended that match the above historical text tags are screened out from multiple tags to be recommended, thereby screening out multimedia resources to be recommended from at least two target multimedia resources; and the above recommended multimedia resources are recommended to the target object, thereby improving the accuracy of resource recommendation and the conversion rate of recommended multimedia resources.

[0132] (3) In a search application scenario, a search tag of a target object can be obtained; a tag to be recommended that matches the search tag is determined to obtain a matching recommended tag; then a target multimedia resource corresponding to the matching recommended tag is obtained to obtain a recommended multimedia resource; thereby recommending the recommended multimedia resource to the target object to improve the accuracy of resource recommendation.

[0133] The technical solution of the present invention combines the advantages of the discriminative method and the generative solution, and uses two-stage training. The first stage focuses on the generative training of the model to enable it to have the ability to recall labels. The second stage, based on the first stage, uses the hidden variables corresponding to the flag bits of the generated labels to score, completes the supervised learning of the discriminative method, and enables the generative model to also have the ability to output scores, with both high recall rate and high accuracy.

[0134] It can be seen from the technical solutions provided by the above embodiments of this specification that the embodiments of this specification obtain a target multimedia resource; input the target multimedia resource into a comprehensive label scoring model for label prediction processing and predicted label scoring processing to obtain a target text label and a target score of the target text label; the target score represents the prediction accuracy of the target text label; the comprehensive label scoring model of the present application can not only predict the target text label of the target multimedia resource, but also predict the target score of the target text label, so that the user can know the prediction accuracy of the target text label; wherein, the comprehensive label scoring model is obtained by training the label scoring model to be trained based on the difference between the first sample score result and the first sample score label; the first sample score result is obtained by inputting the first sample prediction placeholder into the label scoring model to be trained for score prediction training; the first sample score label is constructed based on the difference between the first sample prediction result and the first sample text label; the first sample prediction placeholder represents the score of the first sample prediction result; the first sample prediction placeholder and the first sample prediction result are obtained by inputting the first sample multimedia resource marked with the first sample text label into the label generation model to be trained in the label scoring model to be trained, and performing label prediction processing. The present application jointly trains a label generation model to be trained and a label scoring model to be trained, thereby obtaining a comprehensive label scoring model that can predict both text labels and scores of text labels.

[0135] The present specification also provides a text label prediction device, such as Fig.10 As shown, the above device comprises:

[0136] The target resource acquisition module 1010 is used to acquire the target multimedia resource;

[0137] Prediction module 1020, used for inputting the target multimedia resource into the comprehensive label scoring model for label prediction processing and predicted label scoring processing, and obtaining a target text label and a target score of the target text label; the target score represents the prediction accuracy of the target text label;

[0138] Among them, the comprehensive label scoring model is obtained by training the label scoring model to be trained based on the difference between the first sample score result and the first sample score label; the first sample score result is obtained by inputting the first sample prediction placeholder into the to-be-trained scoring model in the to-be-trained label scoring model for score prediction training; the first sample score label is constructed based on the difference between the first sample prediction result and the first sample text label; the first sample prediction placeholder represents the score of the first sample prediction result; the first sample prediction placeholder and the first sample prediction result are obtained by inputting the first sample multimedia resource marked with the first sample text label into the to-be-trained label generation model in the to-be-trained label scoring model for label prediction processing.

[0139] In some embodiments, the above apparatus further comprises:

[0140] Model training module, used to train comprehensive label scoring model;

[0141] Exemplarily, the above model training module includes:

[0142] A first sample acquisition submodule, configured to acquire a first sample multimedia resource, wherein the first sample multimedia resource is annotated with a first sample text label;

[0143] A first input submodule is used to input the first sample multimedia resource into the to-be-trained label generation model in the to-be-trained label scoring model, perform label prediction processing, and obtain a first sample prediction result and a first sample prediction placeholder corresponding to the first sample prediction result; the first sample prediction placeholder represents the score of the first sample prediction result;

[0144] A first label construction submodule, used to construct a first sample score label based on the difference between the first sample prediction result and the first sample text label;

[0145] A first sample result determination submodule, configured to input the first sample prediction placeholder into the to-be-trained scoring model in the to-be-trained label scoring model to perform score prediction training to obtain a first sample score result;

[0146] The model training submodule is used to train the above-mentioned label scoring model to be trained based on the difference between the above-mentioned first sample score result and the above-mentioned first sample score label to obtain the above-mentioned comprehensive label scoring model.

[0147] In some embodiments, the above-mentioned model training submodule includes:

[0148] A first construction unit, configured to construct first loss data based on a difference between the first sample score result and the first sample score label;

[0149] A scoring model determination unit, configured to perform score prediction training on the scoring model to be trained based on the first loss data to obtain a scoring model;

[0150] A label generation model determination unit, used to train the label generation model to be trained based on the second sample multimedia resource and the scoring model to obtain a label generation model;

[0151] The comprehensive model determination unit is used to determine the comprehensive label scoring model based on the scoring model and the label generation model.

[0152] In some embodiments, the label generation model determination unit includes:

[0153] A second resource acquisition subunit is used to acquire a second sample multimedia resource, where the second sample multimedia resource is marked with a second sample score label;

[0154] A second label prediction subunit is used to input the second sample multimedia resource into the label generation model to be trained, perform label prediction processing, and obtain a second sample prediction result and a second sample prediction placeholder corresponding to the second sample prediction result;

[0155] A second score prediction subunit is used to input the second sample prediction placeholder into the scoring model to perform score prediction processing to obtain a second sample score result;

[0156] A second loss construction subunit, configured to construct second loss data based on a difference between the second sample score result and the second sample score label;

[0157] The model subunit is used to train the above-mentioned label generation model to be trained based on the above-mentioned second loss data to obtain the above-mentioned label generation model.

[0158] In some embodiments, the model training submodule includes:

[0159] a current scoring model determining unit, configured to perform score prediction training on the to-be-trained scoring model based on a difference between the first sample score result and the first sample score label, to obtain a current scoring model;

[0160] A current sample score result determination unit, configured to use the to-be-trained label generation model as the current label generation model, input the first sample prediction placeholder into the current scoring model for score prediction processing, and obtain a current sample score result;

[0161] A current training unit, used to train the current label generation model based on the difference between the current sample score result and the current sample score label, and use the trained model as the current label generation model again;

[0162] A label prediction unit, used to input the first sample multimedia resource into the current label generation model, perform label prediction processing, and obtain a current sample prediction result and a current sample prediction placeholder corresponding to the current sample prediction result;

[0163] The iterative training unit is used to iteratively train the current scoring model and the current label generation model based on the current sample prediction placeholder to obtain the comprehensive label scoring model.

[0164] In some embodiments, the iterative training unit includes:

[0165] A first updating subunit is used to input the current sample prediction placeholder into the current scoring model to perform score prediction training, obtain a scoring result, and use the scoring result as the current sample score result again;

[0166] A repeating subunit is used to repeat the above steps of training the current label generation model based on the difference between the current sample score result and the current sample score label, and reusing the trained model as the current label generation model; inputting the current sample prediction placeholder into the current scoring model for score prediction training, obtaining a scoring result, and reusing the scoring result as the current sample score result, until the training end condition is met;

[0167] Two model determination subunits, used to determine the current scoring model at the end of training as the scoring model, and determine the current scoring model at the end of training as the label generation model;

[0168] The comprehensive model determination subunit is used to determine the comprehensive label scoring model based on the scoring model and the label generation model.

[0169] In some embodiments, the above apparatus further comprises:

[0170] A third resource acquisition module, used to acquire a third sample multimedia resource, where the third sample multimedia resource is marked with a third sample text tag;

[0171] A third result determination module is used to input the third sample multimedia resource into the large language model for label prediction processing to obtain a third sample prediction result and a third sample prediction placeholder corresponding to the third sample prediction result;

[0172] A third loss determination module, configured to construct third loss data based on a difference between the third sample prediction result and the third sample text label;

[0173] The module for determining the model to be trained is used to train the large language model based on the third loss data to obtain the label generation model to be trained.

[0174] In some embodiments, the prediction module includes:

[0175] A first prediction submodule is used to input the target multimedia resource into the label generation model in the comprehensive label scoring model to perform label prediction processing, so as to obtain the target text label and the label score placeholder corresponding to the target text label;

[0176] The second prediction submodule is used to input the above-mentioned label score placeholder into the scoring model in the above-mentioned comprehensive label scoring model for label scoring processing to obtain the target score corresponding to the above-mentioned target text label.

[0177] In some embodiments, the prediction module includes:

[0178] A score prediction submodule is used to input the above-mentioned target multimedia resources into the comprehensive label scoring model for label prediction processing and predicted label scoring processing, and obtain at least two target text labels and a target score corresponding to each target text label;

[0179] In some embodiments, the above apparatus further comprises:

[0180] A tag screening module, used to screen out a tag to be recommended from the at least two target text tags according to the target score corresponding to each target text tag; the target score of the tag to be recommended is greater than the target score of the non-recommended tag; the non-recommended tag is a tag other than the tag to be recommended among the at least two target text tags;

[0181] The text tag determination module is used to determine the above-mentioned tag to be recommended as the text tag of the above-mentioned target multimedia resource.

[0182] In some embodiments, the number of the target multimedia resources is at least two, and the apparatus further includes:

[0183] A historical resource acquisition module is used to acquire historical multimedia resources of a target object;

[0184] A historical tag determination module, used to determine the historical text tag corresponding to the target object based on the historical multimedia resources;

[0185] A first resource-to-be-recommended determining module, configured to determine a target multimedia resource whose to-be-recommended tag matches the historical text tag as a multimedia resource to be recommended;

[0186] The first recommendation module is used to recommend the recommended multimedia resources to the target object.

[0187] In some embodiments, the number of the target multimedia resources is at least two, and the apparatus further includes:

[0188] A search tag acquisition module, used to obtain the search tag of the target object;

[0189] A matching module is used to determine the recommended tags that match the search tags and obtain the matching recommended tags;

[0190] A second to-be-recommended resource determination module is used to obtain the target multimedia resource corresponding to the matching recommendation tag to obtain the recommended multimedia resource;

[0191] The second recommendation module is used to recommend the recommended multimedia resources to the target object.

[0192] The device and method embodiments in the above-mentioned device embodiments are based on the same inventive concept.

[0193] An embodiment of the present specification provides an electronic device, which includes a processor and a memory, wherein the memory stores at least one instruction or at least one program, and the at least one instruction or at least one program is loaded and executed by the processor to implement a text label prediction method as provided in the above method embodiment.

[0194] An embodiment of the present application also provides a computer storage medium, which can be set in a terminal to store at least one instruction or at least one program related to a text label prediction method in a method embodiment. The at least one instruction or at least one program is loaded and executed by the processor to implement the text label prediction method provided by the above method embodiment.

[0195] The embodiment of the present application also provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes to implement the text label prediction method provided by the above method embodiment.

[0196] Optionally, in the embodiments of this specification, the storage medium may be located in at least one of the multiple network servers of the computer network. Optionally, in this embodiment, the storage medium may include, but is not limited to: a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and other media that can store program codes.

[0197] The above-mentioned memory in the embodiment of this specification can be used to store software programs and modules, and the processor executes various functional applications and data processing by running the software programs and modules stored in the memory. The memory may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, application programs required for functions, etc.; the data storage area may store data created according to the use of the above-mentioned devices, etc. In addition, the memory may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage device. Accordingly, the memory may also include a memory controller to provide the processor with access to the memory.

[0198] The text tag prediction method embodiment provided in the embodiments of this specification can be executed in a mobile terminal, a computer terminal, a server or a similar computing device. Taking running on a server as an example, Fig.11 is a hardware structure block diagram of a server of a text tag prediction method provided in an embodiment of this specification. Fig.11 As shown, the server 1100 may have relatively large differences due to different configurations or performances, and may include one or more central processing units (CPU) 1111 (the central processing unit 1111 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 1130 for storing data, and one or more storage media 1120 (such as one or more mass storage devices) for storing application programs 1123 or data 1122. Among them, the memory 1130 and the storage medium 1120 can be short-term storage or permanent storage. The program stored in the storage medium 1120 may include one or more modules, each of which may include a series of instruction operations on the server. Furthermore, the central processing unit 1111 can be configured to communicate with the storage medium 1120 and execute a series of instruction operations in the storage medium 1120 on the server 1100. The server 1100 may also include one or more power supplies 1160, one or more wired or wireless network interfaces 1150, one or more input and output interfaces 1140, and / or one or more operating systems 1121, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.

[0199] The input / output interface 1140 may be used to receive or send data via a network. The specific example of the network may include a wireless network provided by a communication provider of the server 1100. In one example, the input / output interface 1140 includes a network adapter (Network Interface Controller, NIC), which may be connected to other network devices via a base station so as to communicate with the Internet. In one example, the input / output interface 1140 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0200] It can be understood by those skilled in the art that Fig.11 The structure shown is only for illustration and does not limit the structure of the above electronic device. Fig.11 More or fewer components as shown, or with Fig.11 Different configurations are shown.

[0201] It can be seen from the embodiments of the text tag prediction method, device, equipment or storage medium provided by the present application that the present application obtains a target multimedia resource; inputs the target multimedia resource into a comprehensive tag scoring model for tag prediction processing and predicted tag scoring processing to obtain a target text tag and a target score for the target text tag; the target score represents the prediction accuracy of the target text tag; the comprehensive tag scoring model of the present application can not only predict the target text tag of the target multimedia resource, but also predict the target score of the target text tag, so that the user can know the prediction accuracy of the target text tag; wherein the comprehensive tag scoring model is based on the first sample score result The difference between the first sample score label and the first sample score label is obtained by training the to-be-trained label scoring model; the first sample score result is obtained by inputting the first sample prediction placeholder into the to-be-trained label scoring model for score prediction training; the first sample score label is constructed based on the difference between the first sample prediction result and the first sample text label; the first sample prediction placeholder represents the score of the first sample prediction result; the first sample prediction placeholder and the first sample prediction result are obtained by inputting the first sample multimedia resource marked with the first sample text label into the to-be-trained label generation model in the to-be-trained label scoring model for label prediction processing. The present application jointly trains the to-be-trained label generation model and the to-be-trained label scoring model, so as to obtain a comprehensive label scoring model that can predict both text labels and scores of text labels.

[0202] It should be noted that the above sequence of the embodiments of this specification is for description only and does not represent the advantages and disadvantages of the embodiments. The above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0203] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device, equipment, and storage medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0204] A person skilled in the art will appreciate that all or part of the steps to implement the above embodiments may be accomplished by hardware or by instructing related hardware through a program, and the program may be stored in a computer storage medium, and the above-mentioned storage medium may be a read-only memory, a disk or an optical disk, etc.

[0205] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application should be included in the protection scope of the present application.< / s> < / s> < / s> < / s> < / s> < / s> < / s> < / s> < / s> < / s>

Claims

1. A text label prediction method, It is characterized in that The method comprises: Acquire target multimedia resources; Inputting the target multimedia resource into a comprehensive label scoring model for label prediction processing and predicted label scoring processing to obtain a target text label and a target score for the target text label; the target score represents the prediction accuracy of the target text label; Among them, the comprehensive label scoring model is obtained by training the label scoring model to be trained based on the difference between the first sample score result and the first sample score label; the first sample score result is obtained by inputting the first sample prediction placeholder into the to-be-trained scoring model in the to-be-trained label scoring model for score prediction training; the first sample score label is constructed based on the difference between the first sample prediction result and the first sample text label; the first sample prediction placeholder represents the score of the first sample prediction result; the first sample prediction placeholder and the first sample prediction result are obtained by inputting the first sample multimedia resource marked with the first sample text label into the to-be-trained label generation model in the to-be-trained label scoring model for label prediction processing.

2. The method according to claim 1, It is characterized in that The training method of the comprehensive label scoring model includes: Acquire the first sample multimedia resource, where the first sample multimedia resource is marked with a first sample text tag; Inputting the first sample multimedia resource into a to-be-trained label generation model in a to-be-trained label scoring model, performing label prediction processing, and obtaining the first sample prediction result and a first sample prediction placeholder corresponding to the first sample prediction result; the first sample prediction placeholder represents a score of the first sample prediction result; Constructing the first sample score label based on the difference between the first sample prediction result and the first sample text label; Inputting the first sample prediction placeholder into the to-be-trained scoring model in the to-be-trained label scoring model to perform score prediction training to obtain a first sample score result; Based on the difference between the first sample score result and the first sample score label, the to-be-trained label scoring model is trained to obtain the comprehensive label scoring model.

3. The method according to claim 2, It is characterized in that The step of training the to-be-trained label scoring model based on the difference between the first sample score result and the first sample score label to obtain the comprehensive label scoring model includes: constructing first loss data based on a difference between the first sample score result and the first sample score label; Performing score prediction training on the to-be-trained scoring model based on the first loss data to obtain a scoring model; Training the to-be-trained label generation model based on the second sample multimedia resource and the scoring model to obtain a label generation model; Based on the scoring model and the label generation model, the comprehensive label scoring model is determined.

4. The method according to claim 3, It is characterized in that The step of training the to-be-trained label generation model based on the second sample multimedia resource and the scoring model to obtain the label generation model includes: Acquire a second sample multimedia resource, where the second sample multimedia resource is marked with a second sample score label; Inputting the second sample multimedia resource into the to-be-trained label generation model, performing label prediction processing, and obtaining a second sample prediction result and a second sample prediction placeholder corresponding to the second sample prediction result; Inputting the second sample prediction placeholder into the scoring model to perform score prediction processing to obtain a second sample score result; constructing second loss data based on a difference between the second sample score result and the second sample score label; The label generation model to be trained is trained based on the second loss data to obtain the label generation model.

5. The method according to any one of claims 1 to 4, It is characterized in that The first sample multimedia resource is annotated with a current sample score label, and the training of the to-be-trained label scoring model based on the difference between the first sample score result and the first sample score label to obtain the comprehensive label scoring model includes: Based on the difference between the first sample score result and the first sample score label, performing score prediction training on the to-be-trained scoring model to obtain a current scoring model; The to-be-trained label generation model is used as the current label generation model, and the first sample prediction placeholder is input into the current scoring model for score prediction processing to obtain a current sample score result; Based on the difference between the current sample score result and the current sample score label, training the current label generation model, and using the trained model as the current label generation model again; Inputting the first sample multimedia resource into the current label generation model, performing label prediction processing, and obtaining a current sample prediction result and a current sample prediction placeholder corresponding to the current sample prediction result; Based on the current sample prediction placeholder, the current scoring model and the current label generation model are iteratively trained to obtain the comprehensive label scoring model.

6. The method according to claim 5, It is characterized in that The iterative training of the current scoring model and the current label generation model based on the current sample prediction placeholder to obtain the comprehensive label scoring model includes: Input the current sample prediction placeholder into the current scoring model for score prediction training to obtain a scoring result, and use the scoring result as the current sample score result again; Repeat the steps of training the current label generation model based on the difference between the current sample score result and the current sample score label, and reusing the trained model as the current label generation model; to the step of inputting the current sample prediction placeholder into the current scoring model for score prediction training, obtaining a scoring result, and reusing the scoring result as the current sample score result, until the training end condition is met; The current scoring model at the end of training is determined as the scoring model, and the current scoring model at the end of training is determined as the label generation model; Based on the scoring model and the label generation model, the comprehensive label scoring model is determined.

7. The method according to any one of claims 1 to 4, It is characterized in that The method further comprises: Acquire a third sample multimedia resource, where the third sample multimedia resource is marked with a third sample text tag; Inputting the third sample multimedia resource into the large language model for label prediction processing to obtain a third sample prediction result and a third sample prediction placeholder corresponding to the third sample prediction result; Constructing third loss data based on the difference between the third sample prediction result and the third sample text label; Based on the third loss data, the large language model is trained to obtain the label generation model to be trained.

8. The method according to any one of claims 1 to 4, It is characterized in that The step of inputting the target multimedia resource into a comprehensive label scoring model for label prediction processing and predicted label scoring processing to obtain a target text label and a target score of the target text label includes: Inputting the target multimedia resource into the label generation model in the comprehensive label scoring model for label prediction processing to obtain the target text label and the label score placeholder corresponding to the target text label; The label score placeholder is input into the scoring model in the comprehensive label scoring model for label scoring processing to obtain the target score corresponding to the target text label.

9. The method according to claim 8, It is characterized in that The step of inputting the target multimedia resource into a comprehensive label scoring model for label prediction processing and predicted label scoring processing to obtain a target text label and a target score of the target text label includes: Inputting the target multimedia resource into a comprehensive label scoring model for label prediction processing and predicted label scoring processing to obtain at least two target text labels and a target score corresponding to each target text label; The method further comprises: According to the target score corresponding to each target text label, a label to be recommended is screened out from the at least two target text labels; the target score of the label to be recommended is greater than the target score of the non-recommended label; the non-recommended label is a label other than the label to be recommended among the at least two target text labels; The tag to be recommended is determined as a text tag of the target multimedia resource.

10. The method according to claim 9, It is characterized in that The number of target multimedia resources is at least two, and the method further includes: Acquire historical multimedia resources of the target object; Based on the historical multimedia resources, determining the historical text label corresponding to the target object; Determine the target multimedia resource that matches the to-be-recommended tag with the historical text tag as the to-be-recommended multimedia resource; The recommended multimedia resource is recommended to the target object.

11. The method according to claim 9, It is characterized in that The number of target multimedia resources is at least two, and the method further includes: Get the search tag of the target object; Determine a tag to be recommended that matches the search tag, and obtain a matching recommended tag; Acquire the target multimedia resource corresponding to the matching recommendation tag to obtain the recommended multimedia resource; The recommended multimedia resource is recommended to the target object.

12. A text label prediction device, It is characterized in that The device comprises: A target resource acquisition module is used to acquire target multimedia resources; A prediction module is used to input the target multimedia resource into a comprehensive label scoring model for label prediction processing and predicted label scoring processing to obtain a target text label and a target score for the target text label; the target score represents the prediction accuracy of the target text label; Among them, the comprehensive label scoring model is obtained by training the label scoring model to be trained based on the difference between the first sample score result and the first sample score label; the first sample score result is obtained by inputting the first sample prediction placeholder into the to-be-trained scoring model in the to-be-trained label scoring model for score prediction training; the first sample score label is constructed based on the difference between the first sample prediction result and the first sample text label; the first sample prediction placeholder represents the score of the first sample prediction result; the first sample prediction placeholder and the first sample prediction result are obtained by inputting the first sample multimedia resource marked with the first sample text label into the to-be-trained label generation model in the to-be-trained label scoring model for label prediction processing.

13. An electronic device, It is characterized in that The device includes: a processor and a memory, wherein the memory stores at least one instruction or at least one program, and the at least one instruction or at least one program is loaded and executed by the processor to implement the text label prediction method as described in any one of claims 1-11.

14. A computer storage medium, It is characterized in that The computer storage medium stores at least one instruction or at least one program, and the at least one instruction or at least one program is loaded and executed by the processor to implement the text tag prediction method according to any one of claims 1-11.