Label generation method and device, equipment, storage medium and program product
By predicting and screening the target items, combined with the method of querying and updating the tag library, the problems of poor label generation ability and low consistency in the existing technology are solved, and more accurate and consistent label generation is achieved.
Patent Information
- Application Number
- CN202510099992.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-27
AI Technical Summary
In the prior art, the tag generation method has poor ability to generate new tags, high training cost, and low consistency of generated tags. Especially when using a generative large language model, the tag generation is too divergent.
By making label predictions on the target items, the first candidate tag with high confidence is selected, and the tag library is queried based on the tag to determine the target tag. If there is no synonym tag with the same semantics in the tag library, the first candidate tag is added to the tag library.
Improve the accuracy and consistency of label generation, reduce training costs, ensure consistency of labels used by different target projects, and avoid confusion caused by inconsistent label naming or definition.
Smart Images

Figure CN120045938A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a label generation method, device, equipment, storage medium and program product. Background Art
[0002] In the label generation method of the related technology, a large number of commonly used labels are collected and the discriminant model is trained to fit the labeled labels. The ability to generate new labels is poor, and the label generation ability can only be obtained by adding new labeled labels and corresponding data to retrain the model. The training cost is high. The labels generated by the generative large language model in the related technology also face the problems of too divergent label generation (the same meaning is easy to generate words with different expressions) and low label consistency. Summary of the invention
[0003] The embodiments of the present application provide a label generation method, apparatus, device, storage medium and program product, which can improve the accuracy and consistency of label generation.
[0004] The technical solution of the embodiment of the present application is implemented as follows:
[0005] The present application provides a label generation method, the method comprising:
[0006] Perform label prediction on the target item to obtain multiple predicted labels and the confidence level of each predicted label;
[0007] Filtering at least one first candidate label from the multiple predicted labels according to the confidence of each predicted label;
[0008] Querying a tag library based on each of the first candidate tags to obtain a tag query result for each of the first candidate tags;
[0009] According to the tag query result of each of the first candidate tags, a target tag corresponding to each of the first candidate tags is determined, wherein the target tag is used to characterize the target item.
[0010] In the above solution, after determining the target tag corresponding to each first candidate tag according to the tag query result of each first candidate tag, the method further includes:
[0011] In response to the tag query result indicating that there is no synonymous tag with the same semantics as the first candidate tag in the tag library, the first candidate tag is added to the tag library.
[0012] The present application embodiment provides a method for training a second language model, the method comprising:
[0013] Acquire a plurality of project samples, and divide the plurality of project samples into a first project sample set and a second project sample set according to a preset ratio;
[0014] A first training task is performed on the second language model to be trained using the first project sample set, and a second training task is performed on the second language model to be trained using the second project sample set, to obtain a trained second language model, wherein the first training task is used to perform the label prediction using the project samples, and the second training task is used to perform the label prediction using the project samples and extended information samples of the project samples.
[0015] In the above solution, performing the first training task on the second language model to be trained using the first project sample set includes:
[0016] The following processing is performed on each item sample in the first item sample set:
[0017] Obtaining a first pre-labeled label for each of the project samples, wherein the first pre-labeled label includes at least one first pre-labeled positive label and at least one first pre-labeled negative label;
[0018] Performing a second feature encoding process on the project sample to obtain a first project feature sample;
[0019] Performing a second feature mapping process on the first project feature sample to obtain at least one first predicted positive label and at least one first predicted negative label;
[0020] determining a first loss value based on the at least one first pre-labeled positive label, the at least one first pre-labeled negative label, the at least one first predicted positive label, and the at least one first predicted negative label;
[0021] The parameters of the second language model to be trained are updated based on the first loss value to obtain a second language model that completes the first training task.
[0022] In the above solution, performing the second training task on the second language model to be trained by using the second project sample set includes:
[0023] The following processing is performed on each item sample in the second item sample set:
[0024] Acquire a second pre-labeled label for each of the project samples, wherein the second pre-labeled label includes at least one second pre-labeled positive label and at least one second pre-labeled negative label;
[0025] generating an extended information sample of the project sample, wherein the extended information sample includes content related to the project sample;
[0026] adding the extended information sample to the project sample to obtain a new project sample, and performing a third feature encoding process based on the new project sample to obtain a second project feature sample;
[0027] Performing a third feature mapping process on the second project feature sample to obtain at least one second predicted positive label and at least one second predicted negative label;
[0028] determining a second loss value based on the at least one second pre-labeled positive label, the at least one second pre-labeled negative label, the at least one second predicted positive label, and the at least one second predicted negative label;
[0029] The parameters of the second language model to be trained are updated based on the second loss value to obtain a second language model that completes the second training task.
[0030] The present application provides a label generation device, including:
[0031] A label prediction module is used to predict labels for target items and obtain multiple predicted labels and the confidence level of each predicted label;
[0032] A label screening module, used for screening at least one first candidate label from the multiple predicted labels according to the confidence of each predicted label;
[0033] A tag detection module, configured to perform a tag query based on a tag library for each of the first candidate tags to obtain a tag query result for each of the first candidate tags;
[0034] A tag generation module is used to determine a target tag corresponding to each of the first candidate tags according to the tag query result of each of the first candidate tags, wherein the target tag is used to characterize the target project.
[0035] An embodiment of the present application provides an electronic device, the electronic device comprising:
[0036] A memory for storing computer executable instructions or computer programs;
[0037] The processor is used to implement the label generation method provided in the embodiment of the present application when executing the computer executable instructions or computer program stored in the memory.
[0038] An embodiment of the present application provides a computer-readable storage medium storing a computer program or computer-executable instructions for implementing the label generation method provided in the embodiment of the present application when executed by a processor.
[0039] An embodiment of the present application provides a computer program product, including a computer program or computer executable instructions. When the computer program or computer executable instructions are executed by a processor, the label generation method provided in the embodiment of the present application is implemented.
[0040] The embodiments of the present application have the following beneficial effects:
[0041] By obtaining multiple predicted labels of the target project and the confidence of each predicted label, the correlation between the predicted label and the target project is quantified, which provides data support for subsequent screening and makes the processing process more explainable. By screening out the first candidate label according to the confidence, the low-confidence predicted labels are filtered out to obtain a more accurate set of first candidate labels. At the same time, when there are a large number of predicted labels, screening out high-confidence first candidate labels can reduce the amount of calculation for subsequent processing. By querying the label library based on each first candidate label, the label query result is obtained, and the target label corresponding to each first candidate label is determined according to the label query result, the first candidate label is mapped to the target label. This mapping can ensure the consistency of the labels used by different target projects and avoid confusion caused by inconsistent label naming or definition, thereby improving the standardization and consistency of label generation. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 It is a structural diagram of the label generation system architecture provided in an embodiment of the present application;
[0043] Figure 2 is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application;
[0044] Figure 3A This is a first flow chart of the label generation method provided in an embodiment of the present application;
[0045] Figure 3B It is a second flow chart of the label generation method provided in an embodiment of the present application;
[0046] Figure 3C It is a third flow chart of the label generation method provided in the embodiment of the present application;
[0047] Figure 3D 4 is a schematic diagram of a fourth process of the label generation method provided in an embodiment of the present application;
[0048] Figure 3E It is a fifth flow chart of the label generation method provided in the embodiment of the present application;
[0049] Figure 3F 6 is a schematic diagram of a sixth flow chart of a label generation method provided in an embodiment of the present application;
[0050] Figure 3G 7 is a schematic diagram of a seventh flow chart of a label generation method provided in an embodiment of the present application;
[0051] Figure 3H This is an eighth flow chart of the label generation method provided in the embodiment of the present application;
[0052] Fig. 3I This is a ninth flow chart of the label generation method provided in an embodiment of the present application;
[0053] Figure 3J is a tenth flow chart of the label generation method provided in an embodiment of the present application;
[0054] Figure 3K This is a schematic diagram of the eleventh process of the label generation method provided in the embodiment of the present application;
[0055] Figure 4 It is a schematic diagram of the application principle of the tag generation method provided in the embodiment of the present application in a content management system;
[0056] Figure 5 This is a schematic diagram of the application process of the tag generation method provided in the embodiment of the present application in the news recommendation scenario;
[0057] Fig. 6A It is a first schematic diagram of the principle of the label generation method provided in the embodiment of the present application;
[0058] Figure 6B is a second schematic diagram of the principle of the label generation method provided in an embodiment of the present application;
[0059] Figure 6C is a third schematic diagram of the principle of the label generation method provided in an embodiment of the present application;
[0060] Fig.6D is a fourth schematic diagram of the principle of the label generation method provided in an embodiment of the present application;
[0061] Fig. 6E is a fifth schematic diagram of the principle of the label generation method provided in an embodiment of the present application;
[0062] Fig. 6F This is the sixth schematic diagram of the principle of the label generation method provided in the embodiment of the present application.
[0063] It should be pointed out that the above-mentioned "first" and "second" are only used to distinguish different solutions, and do not represent the degree of superiority or inferiority of the solutions or the priority in the implementation process. DETAILED DESCRIPTION
[0064] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the accompanying drawings. The described embodiments should not be regarded as limiting the present application. All other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of this application.
[0065] In the following description, reference is made to “some embodiments”, which describe a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0066] In the following description, the terms "first\second\third" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0067] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.
[0068] Unless otherwise defined, all technical and scientific terms used in the embodiments of the present application have the same meanings as those commonly understood by those skilled in the art. The terms used in the embodiments of the present application are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.
[0069] The relevant data collection and processing in the embodiments of this application should be strictly in accordance with the requirements of relevant laws and regulations when applied in examples, obtain the informed consent or separate consent of the personal information subject, and carry out subsequent data use and processing within the scope of authorization of laws and regulations and the personal information subject.
[0070] Before further describing the embodiments of the present application in detail, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations.
[0071] 1) In response, it is used to indicate the conditions or states on which the executed operations depend. When the dependent conditions or states are met, one or more operations executed may be in real time or have a set delay. Unless otherwise specified, there is no restriction on the order in which the multiple operations executed are executed.
[0072] 2) Human-computer interaction interface, which is used to provide an interface for human-computer interaction functions / an interface for displaying label generation information.
[0073] For example, graphical user interface (GUI) display, such as augmented reality (AR) interface, virtual reality (VR) interface, voice user interface (VUI), interactive projection interface (using projection technology to display information on a plane), eye movement detection interface (interface controlled by detecting the user's line of sight), holographic interface (three-dimensional hologram formed by projecting images through holographic projection technology, and stereoscopic images can be seen without wearing special glasses), multimodal interface (interface that combines multiple interaction methods, such as touch, vision, hearing, etc.), brain-machine interface (Brain-Machine Interface, BMI) interface, etc.
[0074] 3) Positive Labels are labels related to the target item and can accurately describe the attributes, characteristics or categories of the target item.
[0075] 4) Negative Labels are labels that are irrelevant to the target item and cannot accurately describe the attributes, characteristics or categories of the target item.
[0076] 5) Confidence Score is an important concept in machine learning, statistics or data analysis. It is used to measure the credibility of the label predicted by the label. It can be expressed in the form of a probability value (between 0 and 1) or a percentage (0% to 100%). The higher the value, the higher the credibility of the label.
[0077] 6) Tag Library is a system or database that stores and manages tags, which is used to organize, classify and retrieve tags related to specific projects. Tag library plays an important role in content management, recommendation system, search engine, data analysis and knowledge graph.
[0078] 7) The first candidate labels are obtained by screening the confidence of the predicted labels. These labels are considered to be highly relevant and credible to the target project and can be used for further analysis, recommendation or decision-making.
[0079] 8) Large Language Model (LLM) is a large-scale language model designed to understand and generate human language. They are trained on large amounts of text data and can perform a wide range of tasks, including text summarization, translation, sentiment analysis, etc. Large language models are characterized by their large scale and billions of parameters that help them learn complex patterns in language data. They are usually based on deep learning architectures. Large language models refer to deep learning models trained with large amounts of text data, containing billions or even larger parameters. They can be used to generate natural language text and understand the meaning of natural language text. Through training, the model can learn the statistical laws and semantic associations of language to build a huge language knowledge base, thereby simulating human language understanding and generation capabilities. Large language models have the following characteristics:
[0080] Learning ability: Through training on massive amounts of text data, large language models can learn a wealth of language knowledge and expressions, including grammar, semantics, and common expression habits.
[0081] Pattern recognition: Large language models can recognize common text patterns and semantic associations, such as the co-occurrence relationship between words, the logical structure and semantic roles of sentences, etc.
[0082] Contextual understanding: Large language models can capture contextual information in text, understand the impact of previous text on subsequent text, and generate appropriate responses based on the context.
[0083] Generation capability: The large language model can generate relevant natural language text based on the input information, including answering questions, generating articles, and conducting conversations.
[0084] Resolving ambiguity: Despite the polysemy and ambiguity of language, large language models resolve ambiguity through contextual information and language regularities, providing more accurate and appropriate text generation or understanding.
[0085] The application scenarios of large language models are very broad. They can be applied to intelligent customer service, intelligent question and answer, natural language generation, advertising recommendation, games and other fields. They can improve the efficiency and accuracy of human-computer interaction and enhance user experience.
[0086] 9) Transformer refers to a time series model based on the self-attention mechanism. The encoder part can effectively encode time series information. Its processing ability of time series information is much better than that of the long short-term memory network (LSTM), and its speed is fast. It is widely used in natural language processing, computer vision, machine translation, speech recognition and other fields.
[0087] In the label generation method of the related technology, a large number of commonly used labels are collected and the discriminant model is trained to fit the labeled labels, but the ability to generate new labels is poor. For long-tail labels, such as names of people and drama titles, a special model is trained to generate names and drama titles corresponding to the generated items, but this method can only solve the problem of names and drama titles in a targeted manner, and cannot solve the long-tail recognition problem of other entities. There are also some label generation methods that use the base library annotation method to annotate a batch of data as a retrieval base library. For example, for text items (articles), by retrieving base library articles similar to the text items, the annotated labels of several similar base library articles are combined as the label of the text item, but this method requires tens of thousands of large-scale base library annotations, and the annotation cost is high. The label generation method based on the large language model has the ability to adaptively generate new labels, but due to the probabilistic randomness of the generated labels of the large language model, the large language model cannot generate reliable and stable confidence. At the same time, the generated labels also face the problems of too divergent label generation (the same meaning is easy to generate words with different expressions) and low label consistency.
[0088] The embodiments of the present application provide a label generation method, apparatus, device, computer-readable storage medium and computer program product, which can improve the accuracy and consistency of label generation. The exemplary application of the electronic device provided by the embodiments of the present application is described below. The electronic device provided by the embodiments of the present application can be implemented as various types of terminals such as laptop computers, tablet computers, desktop computers, set-top boxes, smart phones, smart speakers, smart watches, smart TVs, and vehicle-mounted terminals, and can also be implemented as servers.
[0089] See also Figure 1 , Figure 1 is a structural diagram of the label generation system architecture provided in an embodiment of the present application, Figure 1 The server 100, the terminal device 200 and the network 300 are involved. The terminal device 200 is connected to the server 100 via the network 300, wherein the network 300 can be a wide area network or a local area network, or a combination of the two.
[0090] In some embodiments, the embodiments of the present application can be implemented by a server and a terminal device in collaboration. For example, the terminal device 200 sends the target item to the server 100, the server 100 receives the target item, generates a target tag representing the target item through the tag generation method provided in the embodiments of the present application, and sends the target tag to the terminal device 200.
[0091] The label generation system provided in the embodiment of the present application can be applied to various scenarios where label generation is required, such as news recommendation, video recommendation, etc., which are illustrated by examples below.
[0092] 1) News recommendation system. For example, the server generates tags for news articles through a tag generation system, such as "economy", "technology", "entertainment", etc. Terminal devices (such as users' mobile phones or tablets) can use these tags to push personalized news content based on users' reading history and interest preferences, thereby improving user experience.
[0093] 2) E-commerce platforms, for example, the server generates labels for products through a label generation system, such as "hot sale", "new product", etc. Terminal devices (such as the user's computer or mobile phone) can quickly filter and recommend products based on these labels, helping users to find the required products efficiently.
[0094] 3) Video recommendation system. For example, the server generates tags for video content through a tag generation system, such as "movie", "variety show", "education", "funny", etc. The terminal device (such as the user's smart TV or tablet) can use these tags to recommend related videos based on the user's viewing history and preferences, thereby increasing the user's viewing time.
[0095] 4) Social media platforms, for example, servers dynamically generate tags for users’ posts through a tag generation system, such as “travel,” “food,” and “fitness.” Terminal devices (such as users’ smartphones) can classify content and make interest recommendations based on these tags, thereby enhancing user interaction and stickiness.
[0096] 5) Content management system. For example, the server generates tags for articles, pictures, videos and other content through a tag generation system, such as "technology", "entertainment", "education", etc. Terminal devices (such as the editor's computer or the user's tablet) can classify, retrieve and manage content based on these tags to improve content operation efficiency.
[0097] 6) Search engine optimization. For example, the server generates tags for web page content through a tag generation system, such as "technical tutorials" and "industry news". Terminal devices (such as users' computers or mobile phones) can quickly match user search intent based on these tags to improve the relevance and accuracy of search results.
[0098] In other embodiments, the present invention can be implemented by a terminal device alone. The terminal device 200 generates a target tag for a target item using the tag method provided in the present invention.
[0099] See also Figure 2 , Figure 2 is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application, Figure 2 The electronic device 400 shown may be the server 100 or the terminal device 200. Figure 2 The electronic device 400 shown includes: at least one processor 410, a memory 430 and at least one network interface 420. The various components in the electronic device 400 are coupled together via a bus system 440. It is understood that the bus system 440 is used to achieve connection and communication between these components. In addition to the data bus, the bus system 440 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, the bus system 440 is not described in detail. Figure 2 Various buses are labeled as bus system 440 .
[0100] The processor 410 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor, etc.
[0101] The memory 430 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard disk drives, optical disk drives, etc. The memory 430 may optionally include one or more storage devices that are physically remote from the processor 410.
[0102] The memory 430 includes a volatile memory or a non-volatile memory, and may also include both volatile and non-volatile memories. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 430 described in the embodiments of the present application is intended to include any suitable type of memory.
[0103] In some embodiments, memory 430 can store data to support various operations, examples of which include programs, modules, and data structures, or a subset or superset thereof, as exemplarily described below.
[0104] Operating system 431, including system programs for processing various basic system services and performing hardware-related tasks, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks;
[0105] A network communication module 432, used to reach other electronic devices via one or more (wired or wireless) network interfaces 420, exemplary network interfaces 420 include: Bluetooth, Wireless Compatibility Authentication (WiFi), and Universal Serial Bus (USB), etc.;
[0106] In some embodiments, the device provided in the embodiments of the present application can be implemented in software. Figure 2 The label generation device 433 stored in the memory 430 is shown, which can be software in the form of a program and a plug-in, etc., including the following software modules: a label prediction module 4331, a label screening module 4332, a label detection module 4333 and a label generation module 4334. These modules are logical, so they can be arbitrarily combined or further split according to the functions implemented. The functions of each module will be explained below.
[0107] In some embodiments, the terminal device or server can implement the label generation method provided by the embodiment of the present application by running various computer executable instructions or computer programs. For example, computer executable instructions can be commands, machine instructions or software instructions at the microprogram level. The computer program can be a native program or software module in the operating system; it can be a local (Native) application (APPlication, APP), that is, a program that needs to be installed in the operating system to run; it can also be a small program that can be embedded in any A PP (such as an e-commerce platform client, a news media client, etc.), that is, a program that can be run only by downloading it to a browser environment. In short, the above-mentioned computer executable instructions can be instructions in any form, and the above-mentioned computer program can be an application, module or plug-in in any form.
[0108] In other embodiments, the device provided in the embodiments of the present application can be implemented in hardware. As an example, the device provided in the embodiments of the present application can be a processor in the form of a hardware decoding processor, which is programmed to execute the label generation method provided in the embodiments of the present application. For example, the processor in the form of a hardware decoding processor can adopt one or more application-specific integrated circuits (Application Specific Integrated Circuit, ASIC), digital signal processors (Digital Signal Processor, DSP), programmable logic devices (Programmable Logic Device, PLD), complex programmable logic devices (Complex Programmable Logic Device, CPLD), field programmable gate arrays (Field-Programmable Gate Array, FPGA) or other electronic components.
[0109] The tag generation method provided in the embodiment of the present application will be described below in combination with the exemplary application and implementation of the server provided in the embodiment of the present application, with the server as the execution subject. Fig. 6A , Fig. 6A 1 is a first schematic diagram of the principle of the label generation method provided by the embodiment of the present application. First, the project content of the target project is input into the pre-trained second language model (see the description of steps 201 to 202 below for the training steps), so as to perform label prediction processing on the target project (corresponding to step 101 below) to obtain multiple predicted labels. Next, the label confidence of each predicted label is obtained. When the confidence of each predicted label is greater than the confidence threshold (for example, 0.5, here, if the predicted label is a positive label, the corresponding confidence is directly compared with the confidence threshold. If the predicted label is a negative label, the complementary confidence pre-confidence of the negative label is used. The confidence threshold is used for comparison. For the implementation method, please refer to the description of step 1021 below), at least one first candidate tag is screened out from multiple predicted tags (corresponding to step 102 below), and the tag query result of each first candidate tag is obtained (corresponding to step 103 below). According to the tag query result, the target tag corresponding to each first candidate tag is output (corresponding to step 104 below). When the confidence of each tag is less than or equal to the confidence threshold, the extended information of the target project is generated (the specific implementation method is described in the embodiment below), the extended information is integrated into the target project, and the tag prediction process and subsequent processing steps are performed again.
[0110] See also Figure 3A , Figure 3A This is a first flow chart of the label generation method provided in the embodiment of the present application, which will be combined with Figure 3AThe steps shown are explained.
[0111] In step 101, a label prediction is performed on the target item to obtain a plurality of predicted labels and the confidence of each predicted label.
[0112] In some embodiments, the target item may be various types of content, such as text, images (including videos consisting of multiple consecutive image frames, which may or may not carry audio), audio, etc. The embodiments of the present application do not limit the data modality of the target item.
[0113] In some embodiments, see Figure 3B , Figure 3A The step 101 shown can be implemented by following the steps 1011 to 1012, which are described in detail below.
[0114] In step 1011, a first feature encoding process is performed on the target project to obtain project features.
[0115] In some embodiments, when the target project is text, the project content of the target project is obtained. For example, if the target project is an article, the project content may be the title and content of the article. The target project is subjected to a first feature encoding process, which may be achieved in the following manner: performing word segmentation on the project content to obtain a plurality of input units (tokens); performing word embedding on the plurality of input units to obtain word embedding features of each input unit; performing position encoding on the plurality of input units to obtain a position encoding feature of each input unit; fusing the word embedding feature of each input unit with the position encoding feature to obtain a fused feature of each input unit; performing attention encoding on the fused feature of each input unit (e.g., self-attention mechanism, multi-head attention mechanism, etc.) to obtain an attention encoding feature of each input unit; and combining the attention encoding features of each input unit to obtain a project feature of the target project.
[0116] For example, punctuation marks such as spaces, periods, and commas may be used as tokenization marks to perform tokenization processing on the project content, that is, to divide the project content into multiple input units.
[0117] For example, the project content can be segmented by a tokenizer based on Byte-Pair Encoding (BPE), a tokenizer based on Byte-Level Byte Pair Encoding, etc. The embodiments of the present application do not limit the specific tokenizer (Tokenizer) or tokenization method used to segment the project content.
[0118] For example, word embedding processing is performed on multiple input units to obtain the word embedding features of each input unit. This can be achieved in the following way: first, a vocabulary is constructed. The vocabulary contains all the words or symbols (tokens) that appear in the text. For each word or symbol in the vocabulary, it is converted into a vector of a fixed size. This process is called word embedding. The dimension of word embedding is a hyperparameter, and values such as 50, 100, and 300 can be selected. The higher the embedding dimension, the richer the information that the model can capture, but the higher the computational cost. The embedding matrix of words or symbols can be learned through training models such as Word2Vec, GloVe, BERT, etc., so that each input unit can be converted into the corresponding word embedding features through the word embedding model.
[0119] For example, position encoding (Position Embeddings) processing is performed on multiple input units to obtain the position encoding features of each input unit. The position encoding process can be generated by a fixed algorithm, for example, it can be implemented using a combination of sine and cosine functions. For the dimension index i of the position encoding feature of each position (the position of each input unit in the project content) (for example, the dimension number of the position encoding feature of each token is d, then the range value of the dimension index i is from 0 to (d-1)), the position encoding of the dimension where i is an even number uses a sine function, and the position encoding of the dimension where i is an odd number uses a cosine function. The embodiments of the present application do not limit the specific implementation method of the position encoding process. Here, the position encoding process is performed because when the project content is converted into a numerical representation (word embedding feature), the order information in the text will be lost, and the purpose of the position encoding is to encode the order information of the token.
[0120] For example, the word embedding feature of each input unit is added to the position encoding feature to obtain the fused feature of each input unit.
[0121] For example, before attention encoding, the fused features of each input unit can also be normalized (for example, layer normalization) to obtain the normalized features of each input unit. For the normalized features of each input unit, the following attention encoding process is performed: First, a linear transformation is performed to generate three matrices: query vector (Query, Q), key vector (Key, K) and value vector (Value, V). The linear transformation is composed of a learnable weight matrix W Q , W K and W VImplementation, next, calculate the dot product of Q and K to obtain the attention score (here, the dot product operation is performed on the query vector of the currently processed input unit and the key vectors of all input units (including the currently processed input unit) to obtain the attention scores corresponding to all input units). Next, normalize the attention score, for example, apply a normalization function (such as a softmax function) to convert the attention score into a probability distribution, and use the attention probability distribution to weight V to obtain a rich representation of the correlation of input units at different positions, that is, the attention encoding features of the input units, and concatenate the attention encoding features of the input units into project features.
[0122] The encoder for feature encoding can be constructed based on a Transformer network. For example, the fusion feature of each input unit can be encoded by a Transformer structure to obtain project features. Here, the Transformer structure can be stacked with multiple layers for encoding. The embodiment of the present application does not limit the specific encoder structure.
[0123] In some embodiments, when the target item is a single image, the item content of the target item is obtained. For example, when the target item is a product image, the item content may include the product title, product introduction, product evaluation and product image. The target item is subjected to a first feature encoding process. This may be achieved in the following ways: performing image feature extraction on the image in the item content (such as the product image) to obtain image features; performing text feature extraction on the text in the item content (such as the product title, product introduction and product evaluation, etc.) to obtain text features; and fusing the image features and text features to obtain the item features of the target item.
[0124] For example, a pre-configured deep learning model (such as a vision transformer (ViT)) can be used to perform image feature extraction processing on images in the project content to obtain image features.
[0125] Here, text features are extracted from the text in the project content. Please refer to the above description when the target project is text, which will not be repeated here.
[0126] For example, the fusion of image features and text features can be achieved in the following way: through the Transformer architecture, the image features and text features are fused to obtain project features. For example, the extracted image features and text features are encoded separately and converted into a format that can be processed by the Transformer, such as converting one-dimensional sequence features into a two-dimensional matrix form, and ensuring that features of different modalities are aligned in length through padding and axis alignment. Then, the features of different modalities are spliced, and the spliced features are further extracted through a multi-head attention mechanism to capture the high-order relationship between features of different modalities, thereby obtaining project features.
[0127] In some embodiments, when the target item is audio, the project content of the target item is obtained. For example, when the target item is song audio, the project content may include the song title, song author, song type (such as pop, classical, etc.), song evaluation and song audio, and the target item is subjected to a first feature encoding process. This can be achieved in the following ways: extracting audio features from the audio in the project content (such as the song audio) to obtain audio features; extracting text features from the text in the project content (such as the song title, song author, song type and song evaluation) to obtain text features; and fusing the audio features and text features to obtain project features of the target item.
[0128] For example, audio features can be extracted from the audio in the project content using a pre-configured deep learning model (such as an Audio Spectrogram Transformer (AST) etc.) to obtain audio features.
[0129] Here, text features are extracted from the text in the project content. Please refer to the above description when the target project is text, which will not be repeated here.
[0130] For example, the audio text corresponding to the audio in the project content can also be obtained through speech recognition technology (Automatic Speech Recognition, ASR), and the audio text and the text in the project content are concatenated and then transferred to perform the above-mentioned text feature extraction to obtain text features.
[0131] For example, the fusion of audio features and text features can be achieved in the following way: through the Transformer architecture, the audio features and text features are fused to obtain project features. For example, the extracted audio features and text features are encoded separately and converted into a format that can be processed by the Transformer, such as converting one-dimensional sequence features into a two-dimensional matrix form, and ensuring that features of different modalities are aligned in length through padding and axis alignment. Then, the features of different modalities are spliced, and the spliced features are further extracted through a multi-head attention mechanism to capture the high-order relationship between features of different modalities, thereby obtaining project features.
[0132] In some embodiments, when the target project is a video composed of multiple image frames and carries audio, the project content of the target project is obtained. For example, when the target project is a short video, the project content may include the short video title, multiple video frames of the short video, and audio data corresponding to the short video. The target project is subjected to a first feature encoding process. This can be achieved in the following ways: performing image feature extraction on multiple video frames in the project content to obtain image features; performing text feature extraction on text in the project content (such as the short video title) to obtain text features; performing audio feature extraction on audio data in the project content to obtain audio features; and fusing image features, text features, and audio features to obtain project features of the target project.
[0133] For example, the audio data corresponding to the short video can be extracted by using a multimedia processing library (such as FFmpeg, etc.).
[0134] Here, for text feature extraction of the text in the project content, please refer to the instructions when the target project is text above. For image feature extraction of multiple video frames in the project content, please refer to the instructions when the target project is a single image above. For audio feature extraction of audio data in the project content, please refer to the processing when the target project is audio above. No further details will be given here.
[0135] For example, the fusion of image features, text features and audio features can be achieved in the following way: through the Transformer architecture, the image features, text features and audio features are fused to obtain project features. For example, the extracted image features, text features and audio features are encoded separately and converted into a format that can be processed by Transformer, such as converting one-dimensional sequence features into a two-dimensional matrix form, and ensuring that features of different modalities are aligned in length through padding and axis alignment. Then, the features of different modalities are spliced, and the spliced features are further extracted through a multi-head attention mechanism to capture the high-order relationship between features of different modalities, thereby obtaining project features.
[0136] In step 1012, a first feature mapping process is performed on the project features to obtain at least one positive label, at least one negative label, and the confidence of each positive label and the confidence of each negative label, wherein the positive label is related to the target project and the negative label is not related to the target project.
[0137] In some embodiments, the project features can be linearly mapped through a feed-forward neural network (FFN) layer to obtain linear features; nonlinear mapping is performed based on the linear features to obtain at least one positive label, at least one negative label, and the confidence of each positive label and the confidence of each negative label (which can be represented by a probability value).
[0138] For example, the feedforward neural network layer can adopt a multilayer perceptron (MLP) structure. First, the project features are linearly transformed through a linear layer to obtain linear features. The linear layer can be, for example, a weight matrix W 1 , the bias vector is b 1 Next, the linear features are nonlinearly transformed through an activation function (such as softmax) to obtain at least one positive label, at least one negative label, and the confidence of each positive label and the confidence of each negative label.
[0139] For example, see Figure 6B , Figure 6B is a second schematic diagram of the principle of the label generation method provided in the embodiment of the present application. After the target item is input into the second language model, the model outputs multiple predicted labels (corresponding to Figure 6B Label 1, Label 2, and Label 3) shown in the figure and a “positive / negative label” for each predicted label, where the “positive / negative label” is used to indicate whether the predicted label is a positive label or a negative label.
[0140] For example, the "positive / negative tag" of the positive label can be represented by "[UNUSED_TOKEN_131|", and the "positive / negative tag" of the negative label can be represented by "[UNUSED_TOKEN_132]". The output of the second language model is "label 1 [UNUSED_TOKEN_131] label 2 [UNUSED_TOKEN_131] label 3 [UNUSED_TOKEN_132]", which means that label 1 and label 2 are positive labels, and label 3 is a negative label. At the same time, the second language model outputs the confidence of each predicted label (corresponding to Figure 6B Label 1 confidence, label 2 confidence, and label 3 confidence shown in ).
[0141] In some embodiments, label prediction is achieved by a pre-trained second language model, see Figure 3C The second language model is obtained by training from step 201 to step 202, which is described in detail below.
[0142] In step 201, a plurality of project samples are obtained, and the plurality of project samples are divided into a first project sample set and a second project sample set according to a preset ratio.
[0143] In some embodiments, the project sample can be various types of content, such as text, image, audio, etc. The present application embodiment does not limit the data modality of the project sample. For ease of description, the project sample is described as text, and no further details are given below.
[0144] For example, the plurality of project samples are divided into a first project sample set and a second project sample set according to a preset ratio (eg, 50%), that is, the plurality of project samples are equally divided into two sample sets.
[0145] For example, the first project sample set and the second project sample set both include project samples and pre-labeled labels for each project sample, the pre-labeled labels include pre-labeled positive labels and pre-labeled negative labels (as well as the confidence of the labels), and the embodiment of the present application does not limit the number of labels (pre-labeled positive labels and pre-labeled negative labels) included in the pre-labeled labels.
[0146] In step 202, a first training task is performed on the second language model to be trained using the first project sample set, and a second training task is performed on the second language model to be trained using the second project sample set, to obtain a trained second language model, wherein the first training task is used to perform label prediction using project samples, and the second training task is used to perform label prediction using project samples and extended information samples of project samples.
[0147] It should be noted that the first training task and the second training task are performed based on the same second language model to be trained. The first training task and the second training task can be performed independently. For example, the second language model is first trained for the first training task, and then the second training task is trained on this basis (without restricting the execution order of the training tasks). They can also be combined. For example, the loss value generated by executing the first training task (corresponding to the first loss value below) and the loss value generated by executing the second training task (corresponding to the second loss value below) are added to obtain a combined loss value, and then the parameters of the second language model to be trained are updated based on the combined loss value, thereby obtaining a second language model that completes the first training task and the second training task.
[0148] For example, see Figure 6C , when the item sample is text (such as a news article), the input of the first training task corresponds to Figure 6C The title and article content shown in (the project content of the project sample includes the title and article content), the input of the second training task corresponds to Figure 6C Sample title, article content and extended information shown in .
[0149] In some embodiments, see Figure 3D , Figure 3C The step 202 in which the first training task is performed on the second language model to be trained through the first project sample set can be implemented by performing the following steps 301 to 305 on each project sample in the first project sample set, which are described in detail below.
[0150] In step 301 , a first pre-labeled label of each project sample is obtained, wherein the first pre-labeled label includes at least one first pre-labeled positive label and at least one first pre-labeled negative label.
[0151] For example, assuming that the project content (news title and news text content) of a project sample (news sample) in the first project sample set is: title "Parks carry out water-saving irrigation projects during drought", text "Faced with the challenges of drought weather, the park management department of a certain city has launched a series of water-saving irrigation projects to improve the efficiency of water resource utilization and ensure the survival of green plants...", then the first pre-labeled positive label of the project sample may include: water saving (confidence: 0.8), environmental protection (confidence: 0.7), urban greening (confidence: 0.75), public facilities (confidence: 0.6); the first pre-labeled negative label of the project sample may include: sports (confidence: 0.95), finance (confidence: 0.9), entertainment (confidence: 0.9), technology (confidence: 0.85), health care (confidence: 0.8).
[0152] Here, a higher confidence of the first pre-labeled positive label indicates a higher correlation between the first pre-labeled positive label and the project sample, and a higher confidence of the first pre-labeled negative label indicates a lower correlation between the first pre-labeled negative label and the project sample.
[0153] In step 302, a second feature encoding process is performed on the project sample to obtain a first project feature sample.
[0154] In some embodiments, obtaining the project content of the project sample and performing a second feature encoding process on the project sample can be achieved in the following ways: performing word segmentation on the project content of the project sample to obtain multiple input units; performing word embedding on the multiple input units to obtain word embedding features of each input unit; performing position encoding on the multiple input units to obtain position encoding features of each input unit; fusing the word embedding features of each input unit with the position encoding features to obtain fused features of each input unit; performing attention encoding on the fused features of each input unit (such as self-attention mechanism, multi-head attention mechanism, etc.) to obtain attention encoding features of each input unit; combining the attention encoding features of each input unit to obtain a first project feature sample of the project sample.
[0155] Here, the implementation of the second feature encoding process can refer to the description of the first feature encoding process in step 1011 above, and will not be repeated here.
[0156] In step 303, a second feature mapping process is performed on the first project feature sample to obtain at least one first predicted positive label and at least one first predicted negative label.
[0157] In some embodiments, a linear mapping can be performed on the first project feature sample through a feed-forward neural network (FFN) layer to obtain a linear feature; a nonlinear mapping is performed based on the linear feature to obtain at least one first predicted positive label (corresponding to Figure 6C predicted positive label in ), at least one first predicted negative label (corresponding to Figure 6C The predicted negative labels shown in ) as well as the confidence of each first predicted positive label and the confidence of each first predicted negative label (which can be represented by a probability value).
[0158] Here, the implementation of the second feature mapping process can refer to the description of the first feature mapping process in step 1012 above, and will not be repeated here.
[0159] In step 304 , a first loss value is determined based on at least one first pre-labeled positive label, at least one first pre-labeled negative label, at least one first predicted positive label, and at least one first predicted negative label.
[0160] In some embodiments, a first loss value is determined based on at least one first pre-labeled positive label, at least one first pre-labeled negative label, at least one first predicted positive label, and at least one first predicted negative label through a pre-set loss function (e.g., cross entropy loss, mean square error, mean absolute error, etc.).
[0161] In step 305, the parameters of the second language model to be trained are updated based on the first loss value to obtain a second language model that completes the first training task.
[0162] In some embodiments, gradient information is obtained through the first loss value, and parameters of the second language model to be trained are updated according to the gradient information to obtain the second language model that completes the first training task.
[0163] For example, the gradient information of the first loss value for each parameter of the second language model to be trained is obtained through the back-propagation algorithm, and the parameters of the second language model to be trained are updated using the obtained gradient information according to a gradient descent optimization algorithm (such as batch gradient descent, stochastic gradient descent, etc.). The above process is repeated until a certain number of iterations is reached or the second language model to be trained converges, thereby obtaining the second language model that completes the first training task.
[0164] In some embodiments, see Figure 3E , Figure 3C The step 202 in which the second training task is performed on the second language model to be trained through the second project sample set can be implemented by performing the following steps 401 to 406 on each project sample in the second project sample set, which are described in detail below.
[0165] In step 401 , a second pre-labeled label of each project sample is obtained, wherein the second pre-labeled label includes at least one second pre-labeled positive label and at least one second pre-labeled negative label.
[0166] Here, examples of the second pre-labeled labels can refer to the description of the first pre-labeled labels in step 301 above, which will not be repeated here.
[0167] In step 402, an extended information sample of the project sample is generated, wherein the extended information sample includes content related to the project sample.
[0168] In some embodiments, generating an extended information sample of a project sample can be achieved by executing any of the following methods: performing at least one of the following processing: querying from a preset project library at least one first project sample whose similarity with the project sample is greater than or equal to a preset first sample similarity threshold, and using the description text of the at least one first project sample (such as an abstract, introduction, etc. of the first project sample) as the extended information sample; querying from the project library at least one second project sample whose similarity with the project sample is greater than or equal to a second sample similarity threshold, and determining a candidate label sample from the project labels of the at least one second project sample as the extended information sample, wherein the second sample similarity threshold is greater than the first sample similarity threshold.
[0169] For example, see Fig.6D , by extracting search term samples from the project content of the project sample, obtaining the similarity between the search term feature samples of the search term samples and the project sample features in the project library, obtaining the project samples in the project library whose similarity is higher than a preset threshold (such as the second sample similarity threshold or the first sample similarity threshold mentioned above) (corresponding to Fig.6D ItemSample-1 to ItemSample-n).
[0170] For example, the project library includes project sample features for each project sample, and querying at least one first project sample from a preset project library whose similarity with the project sample is greater than or equal to a preset first sample similarity threshold can be achieved in the following way: extracting a search term sample from the project sample; performing word embedding encoding on the search term sample to obtain a search term feature sample; obtaining the similarity between the search term feature sample and each project sample feature in the project library; querying at least one first project sample from the project library whose similarity is greater than or equal to a first sample similarity threshold (for example, the similarity value range is [0, 1], 0 means completely dissimilar, and the first sample similarity threshold is set to 0.8).
[0171] For example, see Fig. 6E ,The project sample features of the project samples in the project library can be ,achieved in the following ways: preprocessing the text data of the ,external data source (e.g., removing irrelevant characters, removing stop words, etc.) to obtain the project samples, embedding encoding the ,project samples to obtain the project sample features, and associating the ,project samples and their project sample features to be stored in the project library.
[0172] For example, embedding encoding is performed on the project sample to obtain the project sample features of the project sample. This can be achieved in the following way: obtaining the project content (such as title and article content) of the project sample, performing word segmentation processing on the project content to obtain multiple input units; performing word embedding processing on the multiple input units to obtain word embedding features of each input unit; performing position encoding processing on the multiple input units to obtain position encoding features of each input unit; fusing the word embedding features of each input unit with the position encoding features to obtain the fused features of each input unit; performing attention encoding processing (such as self-attention mechanism, multi-head attention mechanism, etc.) on the fused features of each input unit to obtain the attention encoding features of each input unit; combining the attention encoding features of each input unit to obtain the project sample features of the project sample (for the specific implementation method, please refer to the description of step 1011 above).
[0173] For example, embedding encoding is performed on the project sample to obtain the project sample features of the project sample. This can also be achieved in the following ways: obtaining at least one project label of the project sample (such as extracting keywords in the project content of the project sample as project labels), performing word embedding processing on each project label to obtain the label features of each project label, and combining (concatenating) the label features of each project label to obtain the project sample features of the project sample.
[0174] Here, the text data of the external data source may include news website data, academic paper data, book or e-book data, etc.
[0175] For example, extract the search term sample from the project sample (corresponding to Fig.6D Extracting search terms in the sample) can be achieved by: obtaining the project content of the project sample (corresponding to Fig.6D The keyword extraction algorithm (such as the term frequency-inverse document frequency statistical method (TF-IDF)) is used to extract keywords from the project content of the project sample. These keywords are usually words that tend to express the title or main content of the text. The frequency of the extracted keywords is counted, and the number of occurrences or frequency of each keyword in the project content of the project sample is calculated. For example, a counter object in Python or other statistical methods are used to implement keyword frequency statistics, so as to obtain multiple keyword frequency values, and the keywords with the top 3 keyword frequency values are used as search term samples, or the keywords with a keyword frequency greater than a preset keyword frequency are used as search term samples.
[0176] For example, the search term samples are encoded with word embedding (corresponding to Fig.6DThe word embedding encoding in the text) is used to obtain the search term feature sample, which can be achieved in the following way: first, a vocabulary is constructed. The vocabulary contains all the words or symbols (t oken) that appear in the text. For each word or symbol in the vocabulary, it is converted into a vector of a fixed size. This process is called word embedding. The dimension of word embedding is a hyperparameter, and values such as 50, 100, and 300 can be selected. The higher the embedding dimension, the richer the information that the model can capture, but the higher the computational cost. The embedding matrix of words or symbols can be learned through training models such as Word2Vec, GloVe, and BERT, so that each search term sample can be converted into the corresponding search term feature sample through the word embedding model.
[0177] For example, the cosine similarity, Euclidean distance, etc. can be calculated to obtain the sample of search term features and the sample features of each item in the item library (through Fig.6D The project sample shown in the figure is embedded in the coding, and the coding object can be the project content of the project sample (see step 1011 above for the coding method) or the similarity between the project labels of the project sample (see the word embedding coding above for the coding method).
[0178] For example, the implementation method of querying at least one second project sample from the project library whose similarity with the project sample is greater than or equal to the second sample similarity threshold (for example, the second sample similarity threshold is set to 0.9) can be referred to the above description of obtaining at least one first project sample, which will not be repeated here.
[0179] For example, each second item sample (corresponding to Fig.6D Any one of the project samples -1 to -n shown in the figure) corresponds to at least one project label sample (which can be understood as the annotated label of the second project sample), and a candidate label sample is determined from the project label samples of at least one second project sample as an extended information sample, which can be achieved in the following manner: obtaining the cumulative number of uses of each project label sample, wherein the cumulative number of uses is the number of times the project label sample is used as the target label of any project; taking the project label sample whose cumulative number of uses is greater than or equal to a preset number of uses threshold (for example, 5 times) as the candidate label sample (corresponding to Fig.6D The candidate label samples shown in FIG); the candidate label samples are used as extended information samples.
[0180] Here, the cumulative number of uses refers to the total number of times a tag is applied to any project, which reflects the frequency with which the tag is used.
[0181] Here, the project library contains related documents, articles, or other types of information resources. The size and diversity of the project library depends on the specific application scenario. It can be a huge database containing millions of articles or a relatively small collection of documents dedicated to a certain title. Although the project library is pre-set, it can be dynamically updated as needed to include the latest information or remove outdated content.
[0182] In step 403, the extended information sample is added to the project sample to obtain a new project sample, and a third feature encoding process is performed based on the new project sample to obtain a second project feature sample.
[0183] In some embodiments, the project content of the project sample is obtained, and the extended information sample is added to the project content to obtain a new project sample. For example, the original project content of the project sample includes an article title and article content, and the extended information sample (corresponding to the description text of at least one first project sample above, or a candidate label sample determined from the project label of at least one second project sample) is added to the project content of the project sample to obtain a new project sample.
[0184] In some embodiments, performing a third feature encoding process based on a new project sample can be achieved in the following manner: obtaining the project content of the new project sample, performing word segmentation on the new project content, and obtaining multiple input units; performing word embedding processing on the multiple input units to obtain word embedding features of each input unit; performing position encoding processing on the multiple input units to obtain position encoding features of each input unit; fusing the word embedding features of each input unit with the position encoding features to obtain the fused features of each input unit; performing attention encoding processing (such as self-attention mechanism, multi-head attention mechanism, etc.) on the fused features of each input unit to obtain the attention encoding features of each input unit; combining the attention encoding features of each input unit to obtain a second project feature sample of the new project sample.
[0185] Here, the implementation of the third feature encoding process can refer to the description of the first feature encoding process in step 1011 above, which will not be repeated here.
[0186] In step 404, a third feature mapping process is performed on the second project feature sample to obtain at least one second predicted positive label and at least one second predicted negative label.
[0187] In some embodiments, the second project feature samples can be linearly mapped through a feed-forward neural network (FFN) layer to obtain linear features; nonlinear mapping is performed based on the linear features to obtain at least one second predicted positive label, at least one second predicted negative label, and the confidence of each second predicted positive label and the confidence of each second predicted negative label (which can be represented by a probability value).
[0188] Here, the implementation of the third feature mapping process can refer to the description of the first feature mapping process in step 1012 above, and will not be repeated here.
[0189] In step 405 , a second loss value is determined based on at least one second pre-labeled positive label, at least one second pre-labeled negative label, at least one second predicted positive label, and at least one second predicted negative label.
[0190] In some embodiments, a second loss value is determined based on at least one second pre-labeled positive label, at least one second pre-labeled negative label, at least one second predicted positive label, and at least one second predicted negative label through a pre-set loss function (e.g., cross entropy loss, mean square error, mean absolute error, etc.).
[0191] In step 406, the parameters of the second language model to be trained are updated based on the second loss value to obtain a second language model that completes the second training task.
[0192] In some embodiments, gradient information is obtained through the second loss value, and parameters of the second language model to be trained are updated according to the gradient information to obtain a second language model that completes the second training task.
[0193] For example, the gradient information of the second loss value for each parameter of the second language model to be trained is obtained through the back-propagation algorithm, and the parameters of the second language model to be trained are updated using the obtained gradient information according to a gradient descent optimization algorithm (such as batch gradient descent, stochastic gradient descent, etc.). The above process is repeated until a certain number of iterations is reached or the second language model to be trained converges, thereby obtaining a second language model that completes the second training task.
[0194] Through step 201 to step 202, the first training task and the second training task for the second language model are implemented, wherein the first training task enables the second language model to have the ability to predict labels based only on project samples, and the second training task enables the second language model to have the ability to predict labels based on project samples and extended information samples of project samples. Through this training method, it is ensured that the second language model can learn effectively with or without extended information input, and the balance of training is improved, thereby achieving the beneficial effect of reducing dependence on external knowledge sources such as extended information and improving the independence and robustness of the second language model.
[0195] Continue to see Figure 3A In step 102, at least one first candidate tag is selected from multiple predicted tags according to the confidence of each predicted tag.
[0196] In some embodiments, the plurality of predicted labels includes at least one negative label and at least one positive label, see Figure 3F , Figure 3A The step 102 shown can be implemented by following the steps 1021 to 1022, which are described in detail below.
[0197] In step 1021, the confidence of each negative label is converted into a complementary probability to obtain a complementary confidence of each negative label.
[0198] For example, assuming that the confidence of the negative label is 0.6, the confidence of the negative label is subjected to complementary probability transformation, and the complementary confidence of the negative label can be expressed as: 1-0.6=0.4. Here, the role of the complementary probability transformation is to convert the confidence of the negative label that is irrelevant to the target project into a complementary confidence that is relevant to the target project, thereby facilitating subsequent label screening.
[0199] In step 1022, at least one first candidate label is determined based on the confidence of each positive label and the complementary confidence of each negative label.
[0200] In some embodiments, the confidence of each positive label and the complementary confidence of each negative label are screened by a pre-set first candidate label confidence threshold, and labels (including positive labels and negative labels) whose confidence of the positive label or the complementary confidence of the negative label is greater than or equal to the first candidate label confidence threshold are obtained as first candidate labels.
[0201] For example, assuming that 3 positive labels are predicted, with corresponding confidences of 0.8, 0.7 and 0.6 respectively, and assuming that 1 negative label is predicted, with corresponding confidence of 0.6, then the complementary confidence of the negative label is 0.4, and the confidence threshold of the first candidate label is set to 0.7. Then, 2 positive labels are selected from the 3 positive labels as the first candidate labels, and the first candidate label is not selected from the negative labels.
[0202] Through steps 1021 to 1022, the first candidate labels are screened out according to the confidence level, and low-confidence labels (whether positive or negative labels) are filtered out, thereby obtaining a more accurate set of first candidate labels. At the same time, when there are a large number of labels, screening out high-confidence first candidate labels can reduce the amount of computation for subsequent processing.
[0203] Continue to see Figure 3A In step 103, the tag library is queried based on each first candidate tag to obtain a tag query result for each first candidate tag.
[0204] For example, see Fig. 6F , query whether there is a tag with the same name as the first candidate tag in the tag library (that is, whether the first candidate tag is already in the tag library, corresponding to step 1031 below), if there is a tag with the same name, output the tag with the same name as the tag query result, if there is no tag with the same name, extract the tag embedding vector of the first candidate tag (corresponding to step 10331 below), query similar tags similar to the first candidate tag from the tag library (corresponding to step 10332 below), perform semantic relationship classification on the first candidate tag and each similar tag, and obtain a semantic relationship classification result, wherein the semantic relationship classification result indicates whether the similar tag is a synonymous tag of the first candidate tag (corresponding to step 10333 below), if the similar tag is a synonymous tag of the first candidate tag, output the synonymous tag as the tag query result (one or more synonymous tags can be output here), if all similar tags are not synonymous tags of the first candidate tag, output the first candidate tag as the tag query result, and add the first candidate tag to the tag library.
[0205] In some embodiments, see Figure 3G , Figure 3A The step 103 shown can be implemented by following the steps 1031 to 1033, which are described in detail below.
[0206] In step 1031, a tag with the same name as the first candidate tag is searched from the tag library.
[0207] In some embodiments, a tag that is completely consistent with the first candidate tag is queried from a tag library (a predefined set containing a series of tags) as a tag with the same name. For example, if the first candidate tag is "technological innovation", the tag with the same name in the tag library that is consistent with the first candidate tag is also "technological innovation".
[0208] Continue to see Figure 3G In step 1032, in response to finding a tag with the same name as the first candidate tag, the tag with the same name is used as the tag query result of the first candidate tag.
[0209] In some embodiments, when there is a tag in the tag library that is identical to the first candidate tag, the tag with the same name is used as the tag query result.
[0210] In step 1033, in response to not finding a tag with the same name as the first candidate tag in the tag library, a synonymous tag of the first candidate tag is queried from the tag library. If a synonymous tag of the first candidate tag is found, the synonymous tag is used as the tag query result of the first candidate tag. If a synonymous tag of the first candidate tag is not found in the tag library, a tag query result is generated to indicate that a synonymous tag of the first candidate tag does not exist in the tag library.
[0211] In some embodiments, the tag library includes multiple tags and an embedding vector for each tag, see Figure 3H , Figure 3G In step 1033 shown, querying the tag library for synonymous tags of the first candidate tag can be implemented by following steps 10331 to 10333, which are described in detail below.
[0212] In step 10331, embedding encoding is performed on the first candidate tag to obtain an embedding vector of the first candidate tag.
[0213] In some embodiments, embedding encoding is performed on the first candidate tag to obtain an embedding vector of the first candidate tag, which can be achieved in the following way: first, a vocabulary is constructed, and the vocabulary contains all words or symbols (tokens) that appear in the text. For each word or symbol in the vocabulary, it is converted into a vector of a fixed size. This process is called word embedding. The dimension of word embedding is a hyperparameter, and values such as 50, 100, and 300 can be selected. The higher the embedding dimension, the richer the information that the model can capture, but the higher the computational cost. The embedding matrix of words or symbols can be learned through training models such as Word2Vec, GloVe, BERT, etc., so that the first candidate tag can be converted into the corresponding embedding vector through the word embedding model.
[0214] In step 10332, a first similarity between the embedding vector of the first candidate tag and the embedding vector of each tag is obtained, and at least one tag whose first similarity is greater than or equal to a first similarity threshold is taken as a similar tag.
[0215] In some embodiments, the first similarity between the embedding vector of the first candidate tag and the embedding vector of each tag can be obtained by calculating cosine similarity, Euclidean distance, etc., and at least one tag in the tag library whose first similarity is greater than or equal to a first similarity threshold (for example, 0.8) is taken as a similar tag.
[0216] In step 10333, semantic relationship classification is performed on the first candidate tag and each similar tag to obtain a semantic relationship classification result, wherein the semantic relationship classification result indicates whether the similar tag is a synonymous tag of the first candidate tag.
[0217] In some embodiments, see Fig. 3I , Figure 3H The illustrated step 10333 can be implemented by following the steps 501 to 504, which are described in detail below.
[0218] In step 501, a prompt word is generated, wherein the prompt word is used to instruct a pre-trained first language model to perform semantic understanding processing on a first candidate tag and similar tags.
[0219] In some embodiments, a preset prompt word template is obtained, and the first candidate tag and similar tags are filled into the prompt word template to obtain the prompt word.
[0220] For example, the preset prompt word template may be expressed as "Please determine whether (fill in the first candidate tag here) and (fill in the similar tag here) have the same semantics, and give a probability value that the two have the same semantics."
[0221] Here, the purpose of generating prompt words is to guide the pre-trained first language model to understand and process specific tasks, such as distinguishing the first candidate tag from its similar tags. These prompt words should be able to clearly indicate which aspects of the text the model should focus on in order to perform correct semantic understanding.
[0222] In step 502, the first language model is called through the prompt word to perform semantic understanding processing on the first candidate tag and similar tags to obtain a probability value that determines that the first candidate tag and the similar tags have the same semantics.
[0223] In some embodiments, after receiving the input prompt word, the first language model uses its internal mechanism (such as attention mechanism, word embedding, etc.) to understand the semantics of the first candidate tag and similar tags. The first language model evaluates the semantic similarity between the first candidate tag and similar tags based on pre-learned knowledge, and the model outputs a probability value between 0 and 1, indicating the possibility that the first candidate tag has the same semantics as the similar tag.
[0224] For example, suppose the input prompt word is: "Please determine whether the following two label words are semantically the same, and give the probability value of the two semantics being the same: (music) and (art)", the output of the first language model can be expressed as: "The semantic similarity between music and art is 0.65". This probability value of 0.65 means that the first language model believes that "music" and "art" are 65% similar in semantics. Considering that music is a form of art, this evaluation is reasonable.
[0225] In step 503, in response to the probability value being greater than or equal to a preset probability threshold, the semantic similarity between the first candidate tag and the similar tag is taken as a semantic relationship classification result.
[0226] In some embodiments, when the probability value is greater than or equal to a preset probability threshold (eg, 0.85), a semantic relationship classification result is obtained in which the first candidate tag has the same semantics as the similar tag.
[0227] In step 504, in response to the probability value being less than a preset probability threshold, the semantic difference between the first candidate tag and the similar tag is taken as a semantic relationship classification result.
[0228] In some embodiments, when the probability value is less than a preset probability threshold (eg, 0.85), a semantic relationship classification result is obtained in which the semantics of the first candidate tag and the similar tag are different.
[0229] Through steps 501 to 504, it is achieved that the more complex semantic relationship between the first candidate label and similar labels is captured through the pre-trained first language model, thereby providing more accurate semantic relationship judgment. The fully trained first language model can better generalize to unseen data, thereby achieving the beneficial effect of improving robustness and providing more refined semantic relationship classification results for downstream tasks.
[0230] In some embodiments, see Figure 3J , Figure 3H The illustrated step 10333 can be implemented by following the steps 601 to 605, which are described in detail below.
[0231] In step 601 , a first sentence is generated based on a first candidate tag, and a second sentence is generated based on a similar tag.
[0232] In some embodiments, context generation processing is performed based on the first candidate tag to obtain the first sentence, and context generation processing is performed based on the similar tag to obtain the second sentence.
[0233] Taking the context generation processing based on the first candidate label as an example, it can be achieved in the following way: converting the first candidate label into a vector representation; using the knowledge base of the pre-trained first language model, obtaining the association with other related words according to the vector representation converted from the first candidate label, and generating a coherent first sentence.
[0234] For example, the first candidate label (word) is converted into a vector representation using the embedding layer inside the first language model, and this vector representation is provided as input to the decoder part of the first language model to initialize the context. During the generation process, the first language model uses the attention mechanism to identify the words most relevant to the current context. The first language model predicts the probability distribution of the next word at each time step, selects the word with the highest probability or uses a sampling strategy to increase diversity. The first language model updates its internal state based on the predicted next word, and continues to generate the next word until the specified output length is reached or a specific stop condition (such as generating a terminator) is met. After the text is generated, post-processing can be performed to correct grammatical errors, eliminate repetitions, or perform other necessary text optimizations to obtain a smooth first sentence.
[0235] For example, assuming that the first candidate tag is "music" and the similar tag is "music", after context generation processing, the first sentence obtained can be expressed as "Music is a common language of mankind. It can transcend national boundaries and touch the depths of people's hearts. Whether it is a cheerful melody or a low rhythm, music has its unique charm", and the second sentence can be expressed as "This music is full of affection, and every note is like a confession of the performer's inner world. The audience can feel the strong emotional fluctuations and delicate emotional expression from it."
[0236] In step 602, the first candidate tag in the first sentence is replaced with a similar tag to obtain a third sentence, and the similar tag in the second sentence is replaced with the first candidate tag to obtain a fourth sentence.
[0237] Continuing with the above example, replacing the first candidate label in the first sentence with a similar label, the third sentence obtained can be expressed as "Music is a common language of mankind. It can transcend national boundaries and touch the depths of people's hearts. Whether it is a cheerful melody or a low rhythm, music has its unique charm." Replace the similar label in the second sentence with the first candidate label, and the fourth sentence obtained can be expressed as "This music is full of affection, and every note is like a confession of the performer's inner world. The audience can feel the strong emotional fluctuations and delicate emotional expressions from it."
[0238] In step 603, a second similarity between the first sentence and the third sentence is obtained, and a third similarity between the second sentence and the fourth sentence is obtained.
[0239] In some embodiments, obtaining the second similarity between the first sentence and the third sentence can be achieved by: performing feature extraction processing on the first sentence to obtain a first text feature, and performing feature extraction processing on the second sentence to obtain a second text feature; based on the first text feature and the second text feature, determining the second similarity.
[0240] Taking the feature extraction processing of the first sentence as an example, the feature extraction processing can be implemented in the following ways: perform word segmentation processing on the first sentence to obtain multiple input units; perform word embedding processing on the multiple input units to obtain the word embedding features of each input unit; perform position encoding processing on the multiple input units to obtain the position encoding features of each input unit; fuse the word embedding features of each input unit with the position encoding features to obtain the fused features of each input unit; perform attention encoding processing (such as self-attention mechanism, multi-head attention mechanism, etc.) on the fused features of each input unit to obtain the attention encoding features of each input unit; combine the attention encoding features of each input unit to obtain the first text feature of the first sentence. The specific implementation method can be found in the description of step 1011 above, which will not be repeated here.
[0241] For example, the second similarity between the first text feature and the second text feature may be obtained by calculating cosine similarity, Euclidean distance, etc.
[0242] In some embodiments, obtaining the second similarity between the second sentence and the fourth sentence can be achieved by: performing feature extraction processing on the second sentence to obtain a third text feature, and performing feature extraction processing on the fourth sentence to obtain a fourth text feature; and determining the third similarity based on the third text feature and the fourth text feature. The implementation method can refer to the description of obtaining the second similarity between the first sentence and the third sentence above, which will not be repeated here.
[0243] In step 604, in response to the second similarity and the third similarity being both greater than or equal to the second similarity threshold, a semantic relationship classification result indicating that the similar tag is a synonymous tag of the first candidate tag is generated.
[0244] In some embodiments, when the second similarity and the third similarity are both greater than or equal to the second similarity threshold, the similar tag has the same semantics as the first candidate tag, and the similar tag is regarded as a synonym of the first candidate tag as a semantic relationship classification result.
[0245] In step 605, in response to at least one of the second similarity and the third similarity being less than a second similarity threshold, a semantic relationship classification result indicating that the similar tag is not a synonymous tag of the first candidate tag is generated.
[0246] In some embodiments, when at least one of the second similarity and the third similarity is less than the second similarity threshold, the similar tag has different semantics from the first candidate tag, and the similar tag is not a synonym of the first candidate tag as a semantic relationship classification result.
[0247] Through steps 601 to 605, the interchangeability of similar labels and the first candidate label is tested in a specific context, which can more accurately capture the semantic nuances of words, thereby effectively handling the different semantics of polysemous words and synonyms in different contexts. Through this multi-angle context verification method, the misjudgment caused by a single similarity indicator is reduced, thereby achieving the beneficial effect of improving the accuracy of the semantic relationship classification results.
[0248] Continue to see Figure 3A In step 104, a target tag corresponding to each first candidate tag is determined according to the tag query result of each first candidate tag, wherein the target tag is used to characterize the target item.
[0249] In some embodiments, the target tag corresponding to each first candidate tag is determined based on the tag query result of each first candidate tag, which can be achieved by performing the following processing for each first candidate tag: in response to the tag query result indicating that there is a tag with the same name as the first candidate tag in the tag library, the tag with the same name is used as the target tag corresponding to the first candidate tag; in response to the tag query result indicating that there is a synonymous tag with the same semantics as the first candidate tag in the tag library, the synonymous tag is used as the target tag corresponding to the first candidate tag; in response to the tag query result indicating that there is no synonymous tag with the same semantics as the first candidate tag in the tag library, the first candidate tag is used as the target tag corresponding to the first candidate tag.
[0250] Through steps 103 to 104, it is achieved to query the tag library based on each first candidate tag to obtain the tag query result, and determine the target tag corresponding to each first candidate tag based on the tag query result, and map the first candidate tag to the target tag. This mapping can ensure the consistency of the tags used in different target projects, avoid confusion caused by inconsistent tag naming or definition, and thus improve the standardization and consistency of tag generation.
[0251] In some other embodiments, after the first candidate tag is used as the target tag, the following processing may be performed: in response to the tag query result indicating that there is no synonymous tag with the same semantics as the first candidate tag in the tag library, the first candidate tag is added to the tag library.
[0252] By adding the first candidate tag to the tag library, the tag library is enriched so that it can cover a wider range of semantics and concepts, which helps to reduce ambiguity and confusion caused by the lack of accurate tags. The dynamic update of the tag library enables the tag system to adapt to new data and user needs. Timely addition of missing tags helps to keep the tag library relevant to actual application scenarios.
[0253] In some other embodiments, before selecting at least one first candidate label from multiple predicted labels (including positive labels and negative labels) according to the confidence of each predicted label (corresponding to step 102 above), the following processing may also be performed: in response to the confidence of each positive label and the complementary confidence of each negative label being less than or equal to the confidence threshold, generating extended information of the target project, wherein the extended information includes content related to the target project; fusing the extended information into the target project, and performing a first feature encoding process on the target project based on the fused target project to obtain project features (corresponding to step 1011 above), thereby performing a first feature mapping process on the project features to recover at least one positive label, at least one negative label, and the confidence of each positive label and the confidence of each negative label (corresponding to step 1012 above).
[0254] For example, when the confidence of each positive label and the confidence of each negative label are less than or equal to a preset confidence threshold (for example, 0.5), extended information of the target project is generated, and the extended information is integrated into the target project. The first feature encoding processing is performed based on the integrated target project, and the first feature mapping is performed again based on the project features obtained by the first feature encoding processing to obtain predicted positive labels and negative labels, as well as the confidence of each positive label and each negative label.
[0255] This is done because when the confidence of the positive and negative labels is low (that is, less than or equal to the confidence threshold), it means that the information is incomplete or ambiguous. By adding extended information, the missing details can be supplemented to make the target project more complete. The extended information provides more contextual background, which helps to more fully understand the meaning and relevance of the target project. The fused target project contains more information and details, so that the subsequent first feature encoding can capture features of more dimensions, thereby improving the richness and accuracy of the feature representation, and further improving the accuracy of the target label generated to represent the target project. That is, in actual reasoning, the step of adding extended information is omitted to speed up reasoning and reduce computing costs. However, when the confidence of the label generated by the language model is low, the details of the target project are supplemented by adding extended information.
[0256] In some embodiments, generating extended information of the target project can be achieved by performing at least one of the following processes: querying from the project library at least one project sample whose similarity with the target project is greater than or equal to a third similarity threshold, and using the description text in the at least one project sample as the extended information; querying from the project library at least one project sample whose similarity with the target project is greater than or equal to a fourth similarity threshold, and determining a second candidate tag from the project tags of the at least one project sample as the extended information, wherein the fourth similarity threshold is greater than the third similarity threshold.
[0257] In some embodiments, see Figure 3K The project library includes project sample features of each project sample. Querying at least one project sample whose similarity with the target project is greater than or equal to a fourth similarity threshold from the project library can be achieved by following steps 701 to 704, which are described in detail below.
[0258] In step 701, search terms are extracted from the target item.
[0259] In some embodiments, extracting search terms from the target item can be achieved in the following ways: in response to the target item being text, extracting at least one first keyword from the text, and combining the at least one first keyword into a search term; in response to the target item being an image, performing target recognition on the image to obtain the type of object in the image, and using the type of object as a search term; in response to the target item being audio, identifying at least one second keyword from the audio, and combining the at least one second keyword as a search term.
[0260] For example, when the target project is a text (such as an article, etc.), extracting at least one first keyword from the text can be achieved in the following way: obtaining the project content of the target project (the title of the text and the content of the text), using a keyword extraction algorithm (such as the term frequency-inverse file frequency statistics method (TF-IDF)), etc.) to extract candidate keywords from the project content of the target project, these candidate keywords are words that express the title or main content of the text, performing frequency statistics on the extracted candidate keywords, and calculating the number of occurrences or frequency of each candidate keyword in the project content of the target project, for example, using a counter object in Python or other statistical methods to implement keyword frequency statistics, thereby obtaining multiple keyword frequency values, and using the candidate keyword with the top 3 keyword frequency values as the first keyword, or using the candidate keyword with a keyword frequency greater than a preset keyword frequency as the first keyword.
[0261] For example, when the target item is an image (such as an image of a product), target recognition is performed on the image to obtain the type of the object in the image. This can be achieved in the following way: through a pre-trained target recognition model (such as a Single Shot MultiBox Detector (SSD), a You Only Look Once (YOLO) model, etc.), target recognition processing is performed on the image to obtain the type of the object in the image.
[0262] For example, target recognition processing of an image can be performed using a pre-trained target recognition model in the following manner: performing convolution processing on the image to obtain a feature representation of the image; and performing classification processing on the feature representation to obtain the type of the object in the image.
[0263] For example, taking the target recognition model YOLO as an example, the training of the target recognition model can be achieved in the following way: obtain multiple image samples and the labeled category of the object in each image sample; perform convolution processing on each image sample through the initialized target recognition model to obtain the feature representation of each image sample; perform classification processing on the feature representation of each image sample (here refers to the position of the output bounding box and the category of the object in the bounding box) to obtain the predicted category of the image sample; determine the loss value of the target recognition model based on the predicted category and the labeled category; update the initialized target recognition model according to the loss value to obtain the trained target recognition model.
[0264] For example, updating the initialized target recognition model according to the loss value can be implemented in the following way: obtaining gradient information through the loss value, updating the parameters of the initialized target recognition model according to the gradient information, and obtaining the trained target recognition model. For example, the gradient information of the loss value for each parameter of the target recognition model is obtained through the back propagation algorithm, and the parameters of the initialized target recognition model are updated using the obtained gradient information according to the gradient descent optimization algorithm (such as batch gradient descent, stochastic gradient descent, etc.), and the above process is repeated until a certain number of iterations is reached or the target recognition model converges, thereby obtaining the trained target recognition model.
[0265] For example, when the target item is audio (such as audio data of a song), identifying at least one second keyword from the audio and combining the at least one second keyword as a search term can be achieved in the following ways: performing speech recognition processing on the audio to obtain speech text; using a keyword extraction algorithm (such as the term frequency-inverse document frequency statistical method (TF-IDF) etc.) to extract at least one second keyword from the speech text (see the instructions for extracting at least one first keyword above).
[0266] For example, speech recognition processing of audio can be achieved in the following ways: use a speech recognition algorithm (such as Deep Speech, a speech recognition algorithm based on a deep neural network, etc.) to convert the audio into speech text. After obtaining the preliminary speech text, some post-processing can be performed, such as deleting blank parts in the speech text, etc., to improve the accuracy of the speech text.
[0267] Taking speech recognition processing by the Deep Speech method as an example, the audio frame of the audio can be converted to the frequency domain by using the Short-Time Fourier Transform (STFT) algorithm, that is, the time-continuous audio is converted into a set of time-discrete spectrograms, each spectrogram represents the amplitude of the audio at a certain frequency component, and the Fourier transform is applied to each audio frame to obtain its spectrum. Then, for each spectrogram, its Mel-Frequency Cepstral Coef ficients (MFCC) can be calculated to obtain the spectral features of the audio, or feature extraction can be performed by a one-dimensional convolutional neural network or other methods to obtain the spectral features of the audio. Next, the relationship between the spectral features and the corresponding text is further learned through a neural network composed of multiple hidden layers to obtain a high-dimensional feature vector. Finally, the feature vector is converted into a text sequence as the corresponding speech text through a decoding algorithm (such as the Connectionist Temporal Classification (CTC) decoding algorithm).
[0268] In step 702, the search term is embedded and encoded to obtain the search term features.
[0269] In some embodiments, word embedding encoding is performed on search terms to obtain search term features, which can be achieved in the following way: first, a vocabulary is constructed, which contains all words or symbols (tokens) that appear in the text. For each word or symbol in the vocabulary, it is converted into a vector of a fixed size. This process is called word embedding. The dimension of word embedding is a hyperparameter, and values such as 50, 100, and 300 can be selected. The higher the embedding dimension, the richer the information that the model can capture, but the higher the computational cost. The embedding matrix of words or symbols can be learned through training models such as Word2Vec, GloVe, BERT, etc., so that each search term can be converted into a corresponding search term feature through a word embedding model.
[0270] In step 703, the fourth similarity between the search term feature and each item sample feature is obtained.
[0271] In some embodiments, the fourth similarity between the search term feature and each item sample feature may be obtained by calculating cosine similarity, Euclidean distance, etc.
[0272] In step 704, at least one project sample whose fourth similarity is greater than or equal to a fourth similarity threshold is searched from the project library.
[0273] For example, at least one project sample whose similarity is greater than or equal to a fourth similarity threshold (for example, the fourth similarity threshold is set to 0.8) is searched from the project library.
[0274] Through steps 701 to 704, it is possible to query project samples similar to the target project (project samples with a similarity greater than or equal to a third similarity threshold) from the project library. By using the descriptive text (such as introduction, abstract, title, etc.) in the project samples as extended information, it is possible to supplement the lack of details of the target project, provide more background information and relevant data, and help the language model to have a deeper understanding of the semantics and connotation of the target project.
[0275] In some embodiments, determining a second candidate tag from the project tags of at least one project sample as extended information can be achieved in the following manner: obtaining the cumulative number of uses of each project tag, wherein the cumulative number of uses is the number of times the project tag is used as the target tag of any project; using the project tag whose cumulative number of uses is greater than or equal to a preset usage threshold (for example, 5 times) as the second candidate tag; and using the second candidate tag as extended information.
[0276] Here, the cumulative number of uses refers to the total number of times a project tag is applied to any project, which reflects the frequency with which the project tag is used.
[0277] A project label with a high cumulative usage count indicates that it is used in multiple projects, which increases the reliability and universality of the project label. Frequently used project labels can more accurately describe the characteristics of the project and improve the semantic relevance of the extended information to the target project.
[0278] It should be noted that the embodiments of the present application can be applied to various scenarios that require tag generation, such as news recommendation, video recommendation, etc., which are illustrated below with examples.
[0279] 1) News recommendation: For example, when a user browses articles on a news platform, the platform server generates tags for each article, such as "technology" and "entertainment". The platform server builds a user interest model based on the article tags and the user's browsing data, and filters and recommends relevant news articles from the news library based on the user interest model. The terminal device receives the personalized news recommendations pushed by the platform server and displays them to the user.
[0280] 2) E-commerce platform: For example, when a user browses products on an e-commerce platform, the platform server generates labels for the products through a label generation method, such as "dress", "electronic equipment", and "household items". The platform server builds a user interest model based on the product labels and the user's browsing, clicking and other historical data. The user enters search keywords or selects labels through the terminal device. The platform server filters and recommends related products based on this information, and the terminal device displays the recommendation results.
[0281] 3) Video recommendation: For example, when a user watches a video on a video platform, the platform server adds relevant topic tags to each video, such as "travel", "food", and "fitness" through a tag generation method. The platform server detects the user's viewing data and tag preferences, and builds a user interest graph. The user selects the tags of interest through the terminal device, and the server recommends related videos based on these tags. The terminal device displays a list of recommended videos.
[0282] 4) Search engine optimization, for example, website administrators add keyword tags to web page content using tag generation methods, the server receives and stores these tags, optimizes the search engine index of the web page, users search for web pages through terminal devices, the search engine improves the web page ranking based on the tags, and the terminal device displays the search results to the user.
[0283] Next, an exemplary application of the present application embodiment in a news recommendation scenario will be described. Figure 4 , Figure 4 It is a schematic diagram of the application principle of the label generation method provided in the embodiment of the present application in the content management system. The content management system (CMS) serves as a central platform and provides a structured and organized content library for the recommendation system, so that the recommendation algorithm can process and recommend content more effectively. The content management system may include three modules: content production, content processing and content distribution. Among them, content production is responsible for producing article content, content processing is responsible for generating labels for articles, and content distribution is responsible for deciding which users to distribute the articles to based on the article labels.
[0284] Among them, the content processing module generates corresponding tags for articles (such as names of people, places, product names, general descriptive words in titles, etc.) to facilitate the recommendation system to better divide user interests, recommend other reading articles with the same tag interests to users based on the distribution of tags, or reduce the recommendation of other articles with the same tags for the tags of articles with negative user feedback. The tags of articles need to comprehensively cover points of interest (such as objects or types of articles that may arouse user interest in the article), and respond to the generation of new tags as quickly as possible (such as generating new tags for emergencies). At the same time, it is also necessary to maintain the consistency and stability of the tags of points of interest, and avoid too scattered synonyms or near-synonymous tags corresponding to the same point of interest, because this will cause confusion and dispersion of tags, which is not conducive to the organization and retrieval of the recommendation system.
[0285] See also Figure 5 , Figure 5 This is a schematic diagram of the application flow of the tag generation method provided in the embodiment of the present application in the news recommendation scenario, which takes the server as the execution subject and combines Figure 5 The steps shown are explained.
[0286] Step 801: Receive a news article for which a tag is to be generated.
[0287] In some embodiments, the server receives news articles for which tags are to be generated (corresponding to the target items in step 101 ) from a content management system or an editorial department.
[0288] Step 802: pre-process the news article.
[0289] In some embodiments, first, the news article is cleaned, for example, irrelevant characters and stop words are removed, and the cleaned news article is preprocessed, for example, word segmentation and part-of-speech tagging are performed.
[0290] In step 803, label prediction is performed on the preprocessed news article to obtain at least one positive label, at least one negative label, and the confidence of each positive label and the confidence of each negative label, wherein the positive label is related to the news article and the negative label is not related to the news article.
[0291] In some embodiments, a first feature encoding process is performed on the news article to obtain article features (corresponding to the project features above, see the description of step 1011 above); a first feature mapping process is performed on the article features to obtain at least one positive label, at least one negative label, and the confidence of each positive label and the confidence of each negative label (see the description of step 1012 above).
[0292] Here, the label prediction process can be implemented by a pre-trained second language model. The training of the second language model can refer to the description of step 201 to step 202 above, which will not be repeated here.
[0293] In step 804, at least one first candidate label is screened out from at least one positive label and at least one negative label according to the confidence of each positive label and the confidence of each negative label.
[0294] In some embodiments, the confidence of each negative label is converted into a complementary probability to obtain the complementary confidence of each negative label (see the description of step 1021 above); based on the confidence of each positive label and the complementary confidence of each negative label, at least one first candidate label is determined (see the description of step 1022 above).
[0295] In step 805, the tag library is queried based on each first candidate tag to obtain a tag query result for each first candidate tag.
[0296] In some embodiments, a tag with the same name as the first candidate tag is queried from the tag library (see the description of step 1031 above); in response to finding a tag with the same name as the first candidate tag, the tag with the same name is used as the tag query result of the first candidate tag (see the description of step 1032 above); in response to not finding a tag with the same name as the first candidate tag from the tag library, a synonymous tag of the first candidate tag is queried from the tag library, and if a synonymous tag of the first candidate tag is found, the synonymous tag is used as the tag query result of the first candidate tag; if a synonymous tag of the first candidate tag is not found from the tag library, a tag query result is generated that represents that a synonymous tag of the first candidate tag does not exist in the tag library (see the description of step 1033 above).
[0297] For example, querying synonymous tags of the first candidate tag from the tag library can be achieved in the following way: embedding the first candidate tag to obtain the embedding vector of the first candidate tag (see the description in step 10331 above); obtaining the first similarity between the embedding vector of the first candidate tag and the embedding vector of each tag, and taking at least one tag whose first similarity is greater than or equal to the first similarity threshold as a similar tag (see the description in step 10332 above); performing semantic relationship classification on the first candidate tag and each similar tag to obtain a semantic relationship classification result, wherein the semantic relationship classification result indicates whether the similar tag is a synonymous tag of the first candidate tag (see the description in step 10333 above).
[0298] In step 806, the target tag corresponding to each first candidate tag is determined according to the tag query result of each first candidate tag.
[0299] In some embodiments, the target tag corresponding to each first candidate tag is determined based on the tag query result of each first candidate tag, which can be achieved by performing the following processing for each first candidate tag: in response to the tag query result indicating that there is a tag with the same name as the first candidate tag in the tag library, the tag with the same name is used as the target tag corresponding to the first candidate tag; in response to the tag query result indicating that there is a synonymous tag with the same semantics as the first candidate tag in the tag library, the synonymous tag is used as the target tag corresponding to the first candidate tag; in response to the tag query result indicating that there is no synonymous tag with the same semantics as the first candidate tag in the tag library, the first candidate tag is used as the target tag corresponding to the first candidate tag.
[0300] In step 807, the news article and the corresponding target tag are stored in a news database.
[0301] In some embodiments, news articles are stored with associated target tags in a news database.
[0302] In step 808, a recommendation algorithm is executed according to the user's tag preference to determine the news articles to be recommended from the news database.
[0303] In some embodiments, tag preferences are extracted from user data, and the similarity between target tags (or news tags) in the news database and user preference tags is calculated using cosine similarity or other similarity calculation methods. The target tags are sorted according to the similarity scores, and news articles (which can be one or more) associated with the target tags ranked first (for example, the top 3 target tags) are selected as the news articles to be recommended.
[0304] Through step 801 to step 807, personalized news recommendations are generated for users, and the news recommendation platform can better meet user needs, thereby achieving the beneficial effect of improving user satisfaction. Specifically, by obtaining the positive label and negative label of the news article, as well as the confidence of each positive label and the confidence of each negative label, the correlation between the label and the news article is quantified, providing data support for subsequent screening, making the processing process more explainable, and by screening out the first candidate label according to the confidence, it is achieved to filter out low-confidence labels (whether positive labels or negative labels), thereby obtaining a more reliable set of first candidate labels. At the same time, when there are a large number of labels, screening out high-confidence first candidate labels can reduce the amount of calculation for subsequent processing. By querying the label library based on each first candidate label, obtaining the label query result, and determining the target label corresponding to each first candidate label according to the label query result, it is achieved to map the first candidate label to the target label. This mapping can ensure that the labels used in different news articles are consistent, avoid confusion caused by inconsistent label naming or definition, and thus improve the standardization and consistency of label generation.
[0305] The following is a description of an exemplary structure of the label generation device 433 provided in the embodiment of the present application implemented as a software module. In some embodiments, Figure 2 As shown, the software modules stored in the label generating device 433 of the memory 430 may include:
[0306] The label prediction module 4331 is used to perform label prediction on the target item to obtain multiple predicted labels and the confidence of each predicted label.
[0307] The label screening module 4332 is used to screen out at least one first candidate label from the multiple predicted labels according to the confidence of each predicted label.
[0308] The tag detection module 4333 is used to perform a tag query based on the tag library for each of the first candidate tags to obtain a tag query result for each of the first candidate tags.
[0309] The tag generation module 4334 is used to determine the target tag corresponding to each of the first candidate tags according to the tag query result of each of the first candidate tags, wherein the target tag is used to characterize the target project.
[0310] In some embodiments, the tag detection module 4333 is also used to query a tag with the same name that is consistent with the first candidate tag from the tag library; in response to querying a tag with the same name that is consistent with the first candidate tag, using the tag with the same name as the tag query result of the first candidate tag; in response to not querying a tag with the same name that is consistent with the first candidate tag from the tag library, querying a synonymous tag of the first candidate tag from the tag library, if a synonymous tag of the first candidate tag is queried, using the synonymous tag as the tag query result of the first candidate tag, and if a synonymous tag of the first candidate tag is not queried from the tag library, generating a tag query result indicating that a synonymous tag of the first candidate tag does not exist in the tag library.
[0311] In some embodiments, the tag library includes multiple tags and an embedding vector for each of the tags, and the tag detection module 4333 is further used to perform embedding encoding on the first candidate tag to obtain the embedding vector of the first candidate tag; obtain the first similarity between the embedding vector of the first candidate tag and the embedding vector of each of the tags, and take at least one of the tags whose first similarity is greater than or equal to a first similarity threshold as a similar tag; perform semantic relationship classification on the first candidate tag and each of the similar tags to obtain a semantic relationship classification result, wherein the semantic relationship classification result indicates whether the similar tag is a synonymous tag of the first candidate tag.
[0312] In some embodiments, the tag detection module 4333 is also used to generate a prompt word, wherein the prompt word is used to instruct the pre-trained first language model to perform semantic understanding processing on the first candidate tag and the similar tag; the first language model is called through the prompt word to perform semantic understanding processing on the first candidate tag and the similar tag to obtain a probability value for determining that the first candidate tag and the similar tag are semantically identical; in response to the probability value being greater than or equal to a preset probability threshold, the semantic identity between the first candidate tag and the similar tag is taken as the semantic relationship classification result; in response to the probability value being less than the preset probability threshold, the semantic dissimilarity between the first candidate tag and the similar tag is taken as the semantic relationship classification result.
[0313] In some embodiments, the label detection module 4333 is also used to generate a first sentence based on the first candidate label, and to generate a second sentence based on the similar label; replace the first candidate label in the first sentence with the similar label to obtain a third sentence, and replace the similar label in the second sentence with the first candidate label to obtain a fourth sentence; obtain a second similarity between the first sentence and the third sentence, and obtain a third similarity between the second sentence and the fourth sentence; in response to the second similarity and the third similarity being greater than or equal to a second similarity threshold, generate a semantic relationship classification result indicating that the similar label is a synonymous label of the first candidate label; in response to at least one of the second similarity and the third similarity being less than the second similarity threshold, generate a semantic relationship classification result indicating that the similar label is not a synonymous label of the first candidate label.
[0314] In some embodiments, the tag detection module 4333 is further configured to add the first candidate tag to the tag library in response to the tag query result indicating that there is no synonymous tag with the same semantics as the first candidate tag in the tag library.
[0315] In some embodiments, the tag generation module 4334 is also used to perform the following processing for each of the first candidate tags: in response to the tag query result indicating that there is a tag with the same name as the first candidate tag in the tag library, the tag with the same name is used as the target tag corresponding to the first candidate tag; in response to the tag query result indicating that there is a synonymous tag with the same semantics as the first candidate tag in the tag library, the synonymous tag is used as the target tag corresponding to the first candidate tag; in response to the tag query result indicating that there is no synonymous tag with the same semantics as the first candidate tag in the tag library, the first candidate tag is used as the target tag corresponding to the first candidate tag.
[0316] In some embodiments, the label prediction module 4331 is also used to perform a first feature encoding process on the target project to obtain project features; perform a first feature mapping process on the project features to obtain at least one positive label, at least one negative label, and a confidence level of each positive label and a confidence level of each negative label, wherein the positive label is related to the target project and the negative label is not related to the target project.
[0317] In some embodiments, the label prediction module 4331 is also used to generate extended information of the target project in response to the confidence of each of the positive labels and the complementary confidence of each of the negative labels being less than or equal to a confidence threshold, wherein the extended information includes content related to the target project; fuse the extended information into the target project, and perform the first feature encoding processing on the target project based on the fused target project to obtain project features.
[0318] In some embodiments, the label prediction module 4331 is also used to perform at least one of the following processing: querying at least one project sample from the project library whose similarity with the target project is greater than or equal to a third similarity threshold, and using the description text of the at least one project sample as the extended information; querying at least one project sample from the project library whose similarity with the target project is greater than or equal to a fourth similarity threshold, and determining a second candidate label from the project labels of the at least one project sample as the extended information, wherein the fourth similarity threshold is greater than the third similarity threshold.
[0319] In some embodiments, the label prediction module 4331 is also used to extract search terms from the target project; perform word embedding encoding on the search terms to obtain search term features; obtain the fourth similarity between the search term features and each of the project sample features; and query at least one project sample from the project library whose fourth similarity is greater than or equal to the fourth similarity threshold.
[0320] In some embodiments, the label prediction module 4331 is also used to, in response to the target item being text, extract at least one first keyword from the text, and combine the at least one first keyword into the search term; in response to the target item being an image, perform target recognition on the image to obtain the type of object in the image, and use the type of object as the search term; in response to the target item being audio, identify at least one second keyword from the audio, and combine the at least one second keyword as the search term.
[0321] In some embodiments, the label prediction module 4331 is also used to obtain the cumulative number of uses of each of the project labels, wherein the cumulative number of uses is the number of times the project label is used as the target label of any project; the project label whose cumulative number of uses is greater than or equal to a preset number of uses threshold is used as the second candidate label; and the second candidate label is used as the extended information.
[0322] In some embodiments, the label prediction module 4331 is further used to obtain multiple project samples, and divide the multiple project samples into a first project sample set and a second project sample set according to a preset ratio; perform a first training task on the second language model to be trained through the first project sample set, and perform a second training task on the second language model to be trained through the second project sample set to obtain a trained second language model, wherein the first training task is used to perform the label prediction using the project samples, and the second training task is used to perform the label prediction using the project samples and the extended information samples of the project samples.
[0323] In some embodiments, the label prediction module 4331 is further used to perform the following processing on each project sample in the first project sample set: obtain a first pre-labeled label for each of the project samples, wherein the first pre-labeled label includes at least one first pre-labeled positive label and at least one first pre-labeled negative label; perform a second feature encoding process on the project samples to obtain a first project feature sample; perform a second feature mapping process on the first project feature sample to obtain at least one first predicted positive label and at least one first predicted negative label; determine a first loss value based on the at least one first pre-labeled positive label, the at least one first pre-labeled negative label, the at least one first predicted positive label and the at least one first predicted negative label; and update the parameters of the second language model to be trained based on the first loss value to obtain a second language model that completes the first training task.
[0324] In some embodiments, the label prediction module 4331 is further used to perform the following processing on each project sample in the second project sample set: obtain a second pre-labeled label for each of the project samples, wherein the second pre-labeled label includes at least one second pre-labeled positive label and at least one second pre-labeled negative label; generate an extended information sample for the project sample, wherein the extended information sample includes content related to the project sample; add the extended information sample to the project sample to obtain a new project sample, and perform a third feature encoding process based on the new project sample to obtain a second project feature sample; perform a third feature mapping process on the second project feature sample to obtain at least one second predicted positive label and at least one second predicted negative label; determine a second loss value based on the at least one second pre-labeled positive label, the at least one second pre-labeled negative label, the at least one second predicted positive label and the at least one second predicted negative label; update the parameters of the second language model to be trained based on the second loss value to obtain a second language model that completes the second training task.
[0325] In some embodiments, the label screening module 4332 is also used to perform complementary probability conversion on the confidence of each of the negative labels to obtain complementary confidence of each of the negative labels; and determine at least one of the first candidate labels based on the confidence of each of the positive labels and the complementary confidence of each of the negative labels.
[0326] The embodiment of the present application provides a computer program product, which includes a computer program or a computer executable instruction, and the computer program or the computer executable instruction is stored in a computer-readable storage medium. The processor of the electronic device reads the computer executable instruction from the computer-readable storage medium, and the processor executes the computer executable instruction, so that the electronic device executes the label generation method described in the embodiment of the present application.
[0327] The embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions or computer programs are stored. When the computer-executable instructions or computer programs are executed by a processor, the processor will execute the label generation method provided in the embodiment of the present application, for example, Figure 3A The label generation method shown.
[0328] In some embodiments, the computer-readable storage medium may be a memory such as RAM, ROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or may be various devices including one or any combination of the above memories.
[0329] In some embodiments, computer executable instructions may be in the form of a program, software, software module, script or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine or other unit suitable for use in a computing environment.
[0330] As an example, computer-executable instructions may, but need not, correspond to a file in a file system, may be stored as part of a file that stores other programs or data, such as in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files storing one or more modules, subroutines, or code portions).
[0331] As an example, computer executable instructions may be deployed to be executed on one electronic device, or on multiple electronic devices located at one site, or on multiple electronic devices distributed at multiple sites and interconnected by a communication network.
[0332] In summary, through the embodiments of the present application, multiple predicted labels of the target project and the confidence of each predicted label are obtained, the correlation degree between the predicted label and the target project is quantified, data support is provided for subsequent screening, and the processing process is more explainable. By screening out the first candidate label according to the confidence, it is achieved to filter out the low-confidence predicted labels, thereby obtaining a more accurate set of first candidate labels. At the same time, when there are a large number of predicted labels, screening out high-confidence first candidate labels can reduce the amount of calculation for subsequent processing. By querying the label library based on each first candidate label, the label query result is obtained, and the target label corresponding to each first candidate label is determined according to the label query result, the first candidate label is mapped to the target label. This mapping can ensure that the labels used for different target projects are consistent, avoid confusion caused by inconsistent label naming or definition, and thus improve the standardization and consistency of label generation.
[0333] The above is only an embodiment of the present application and is not intended to limit the protection scope of the present application. Any modifications, equivalent replacements and improvements made within the spirit and scope of the present application are included in the protection scope of the present application.
Claims
1. A label generation method, characterized in that: The method comprises: Perform label prediction on the target item to obtain multiple predicted labels and the confidence level of each predicted label; Filtering at least one first candidate label from the multiple predicted labels according to the confidence of each predicted label; Querying a tag library based on each of the first candidate tags to obtain a tag query result for each of the first candidate tags; According to the tag query result of each of the first candidate tags, a target tag corresponding to each of the first candidate tags is determined, wherein the target tag is used to characterize the target item.
2. The method according to claim 1, characterized in that The querying of the tag library based on each of the first candidate tags to obtain a tag query result for each of the first candidate tags includes: Querying a tag with the same name as the first candidate tag from a tag library; In response to finding a tag with the same name as the first candidate tag, taking the tag with the same name as the tag query result of the first candidate tag; In response to not finding a tag with the same name as the first candidate tag from the tag library, a synonymous tag of the first candidate tag is queried from the tag library. If a synonymous tag of the first candidate tag is found, the synonymous tag is used as a tag query result of the first candidate tag. If a synonymous tag of the first candidate tag is not found from the tag library, a tag query result is generated indicating that a synonymous tag of the first candidate tag does not exist in the tag library.
3. The method according to claim 2, characterized in that The tag library includes a plurality of tags and an embedding vector of each of the tags, and querying a synonymous tag of the first candidate tag from the tag library includes: Performing embedding encoding on the first candidate tag to obtain an embedding vector of the first candidate tag; Obtaining a first similarity between the embedding vector of the first candidate tag and the embedding vector of each of the tags, and taking at least one of the tags whose first similarity is greater than or equal to a first similarity threshold as a similar tag; Semantic relationship classification is performed on the first candidate tag and each of the similar tags to obtain a semantic relationship classification result, wherein the semantic relationship classification result indicates whether the similar tag is a synonymous tag of the first candidate tag.
4. The method according to claim 3, characterized in that The performing semantic relationship classification on the first candidate tag and each of the similar tags to obtain a semantic relationship classification result includes: Generate a prompt word, wherein the prompt word is used to instruct the pre-trained first language model to perform semantic understanding processing on the first candidate tag and the similar tag; The first language model is called by the prompt word to perform semantic understanding processing on the first candidate tag and the similar tag to obtain a probability value for determining that the first candidate tag and the similar tag have the same semantics; In response to the probability value being greater than or equal to a preset probability threshold, taking the semantic similarity between the first candidate tag and the similar tag as the semantic relationship classification result; In response to the probability value being less than the preset probability threshold, the semantic difference between the first candidate tag and the similar tag is taken as the semantic relationship classification result.
5. The method according to claim 3, characterized in that: The performing semantic relationship classification on the first candidate tag and each of the similar tags to obtain a semantic relationship classification result includes: generating a first sentence based on the first candidate tag, and generating a second sentence based on the similar tag; Replacing the first candidate tag in the first sentence with the similar tag to obtain a third sentence, and replacing the similar tag in the second sentence with the first candidate tag to obtain a fourth sentence; Obtaining a second similarity between the first sentence and the third sentence, and obtaining a third similarity between the second sentence and the fourth sentence; In response to the second similarity and the third similarity being both greater than or equal to a second similarity threshold, generating a semantic relationship classification result indicating that the similar tag is a synonymous tag of the first candidate tag; In response to at least one of the second similarity and the third similarity being smaller than the second similarity threshold, a semantic relationship classification result indicating that the similar tag is not a synonymous tag of the first candidate tag is generated.
6. The method according to any one of claims 1 to 3, characterized in that: The determining, according to the tag query result of each of the first candidate tags, a target tag corresponding to each of the first candidate tags includes: For each of the first candidate tags, perform the following processing: In response to the tag query result indicating that there is a tag with the same name as the first candidate tag in the tag library, taking the tag with the same name as the target tag corresponding to the first candidate tag; In response to the tag query result indicating that there is a synonymous tag with the same semantics as the first candidate tag in the tag library, taking the synonymous tag as the target tag corresponding to the first candidate tag; In response to the tag query result indicating that there is no synonymous tag with the same semantics as the first candidate tag in the tag library, the first candidate tag is used as a target tag corresponding to the first candidate tag.
7. The method according to any one of claims 1 to 3, characterized in that: The step of performing label prediction on the target item to obtain a plurality of predicted labels and the confidence level of each predicted label includes: Performing a first feature encoding process on the target project to obtain project features; Perform a first feature mapping process on the project features to obtain at least one positive label, at least one negative label, and a confidence of each positive label and a confidence of each negative label, wherein the positive label is related to the target project, and the negative label is not related to the target project.
8. The method according to claim 7, characterized in that The step of selecting at least one first candidate tag from the plurality of predicted tags according to the confidence of each predicted tag comprises: Performing complementary probability conversion on the confidence of each of the negative labels to obtain complementary confidence of each of the negative labels; At least one of the first candidate labels is determined based on the confidence of each of the positive labels and the complementary confidence of each of the negative labels.
9. The method according to claim 8, characterized in that Before selecting at least one first candidate tag from the plurality of predicted tags according to the confidence of each predicted tag, the method further includes: In response to the confidence of each of the positive labels and the complementary confidence of each of the negative labels being less than or equal to a confidence threshold, generating extended information of the target item, wherein the extended information includes content related to the target item; The extended information is integrated into the target project, and based on the integrated target project, the first feature encoding process is performed on the target project to obtain project features.
10. The method according to claim 9, characterized in that The generating of the extended information of the target project includes: Perform at least one of the following actions: Searching for at least one project sample whose similarity with the target project is greater than or equal to a third similarity threshold from the project library, and using a description text of the at least one project sample as the extended information; At least one project sample whose similarity with the target project is greater than or equal to a fourth similarity threshold is queried from the project library, and a second candidate label is determined from the project labels of the at least one project sample as the extended information, wherein the fourth similarity threshold is greater than the third similarity threshold.
11. The method according to claim 10, characterized in that The project library includes project sample features of each of the project samples, and querying the project library for at least one project sample whose similarity with the target project is greater than or equal to a fourth similarity threshold comprises: extracting search terms from the target item; Performing word embedding encoding on the search term to obtain search term features; Obtaining a fourth similarity between the search term feature and each of the project sample features; At least one project sample whose fourth similarity is greater than or equal to the fourth similarity threshold is searched from the project library.
12. The method according to claim 11, characterized in that The extracting search terms from the target item includes: In response to the target item being a text, extracting at least one first keyword from the text, and combining the at least one first keyword into the search term; In response to the target item being an image, performing target recognition on the image to obtain a type of an object in the image, and using the type of the object as the search term; In response to the target item being audio, at least one second keyword is identified from the audio, and the at least one second keyword is combined as a search term.
13. The method according to claim 10, characterized in that The determining of a second candidate label from the project label of the at least one project sample as the extended information includes: Obtaining the cumulative number of times each of the project tags is used, wherein the cumulative number of times the project tag is used as a target tag for any project; The project tag whose cumulative usage times is greater than or equal to a preset usage times threshold is used as the second candidate tag.
14. The method according to any one of claims 1 to 3, characterized in that: The label prediction is achieved by a pre-trained second language model, and the second language model is trained in the following way: Acquire a plurality of project samples, and divide the plurality of project samples into a first project sample set and a second project sample set according to a preset ratio; A first training task is performed on the second language model to be trained using the first project sample set, and a second training task is performed on the second language model to be trained using the second project sample set, to obtain a trained second language model, wherein the first training task is used to perform the label prediction using the project samples, and the second training task is used to perform the label prediction using the project samples and extended information samples of the project samples.
15. The method according to claim 14, characterized in that The performing a first training task on the second language model to be trained by using the first project sample set includes: The following processing is performed on each item sample in the first item sample set: Acquire a first pre-labeled label for each of the project samples, wherein the first pre-labeled label includes at least one first pre-labeled positive label and at least one first pre-labeled negative label; Performing a second feature encoding process on the project sample to obtain a first project feature sample; Performing a second feature mapping process on the first project feature sample to obtain at least one first predicted positive label and at least one first predicted negative label; determining a first loss value based on the at least one first pre-labeled positive label, the at least one first pre-labeled negative label, the at least one first predicted positive label, and the at least one first predicted negative label; The parameters of the second language model to be trained are updated based on the first loss value to obtain a second language model that completes the first training task.
16. The method according to claim 14, characterized in that The performing a second training task on the second language model to be trained by using the second project sample set includes: The following processing is performed on each item sample in the second item sample set: Acquire a second pre-labeled label for each of the project samples, wherein the second pre-labeled label includes at least one second pre-labeled positive label and at least one second pre-labeled negative label; generating an extended information sample of the project sample, wherein the extended information sample includes content related to the project sample; adding the extended information sample to the project sample to obtain a new project sample, and performing a third feature encoding process based on the new project sample to obtain a second project feature sample; Performing a third feature mapping process on the second project feature sample to obtain at least one second predicted positive label and at least one second predicted negative label; determining a second loss value based on the at least one second pre-labeled positive label, the at least one second pre-labeled negative label, the at least one second predicted positive label, and the at least one second predicted negative label; The parameters of the second language model to be trained are updated based on the second loss value to obtain a second language model that completes the second training task.
17. A label generating device, characterized in that: The device comprises: A label prediction module is used to predict labels for target items and obtain multiple predicted labels and the confidence level of each predicted label; A label screening module, used for screening at least one first candidate label from the multiple predicted labels according to the confidence of each predicted label; A tag detection module, configured to perform a tag query based on a tag library for each of the first candidate tags to obtain a tag query result for each of the first candidate tags; A tag generation module is used to determine a target tag corresponding to each of the first candidate tags according to the tag query result of each of the first candidate tags, wherein the target tag is used to characterize the target project.
18. An electronic device, characterized in that: The electronic device comprises: A memory for storing computer executable instructions or computer programs; A processor, configured to implement the label generation method according to any one of claims 1 to 16 when executing the computer executable instructions or computer program stored in the memory.
19. A computer-readable storage medium storing computer-executable instructions or a computer program, characterized in that: When the computer executable instructions or computer program are executed by a processor, the label generation method according to any one of claims 1 to 16 is implemented.
20. A computer program product comprising computer executable instructions or a computer program, characterized in that When the computer executable instructions or computer program are executed by a processor, the label generation method according to any one of claims 1 to 16 is implemented.
Citation Information
Cited By
Intelligent customer service real-time intention analysis and response system based on pre-training large model
CN121071153A