Text encoder training method, related method, device, equipment and medium

By constructing a target loss function through a contrastive learning method and training the text encoder, the problem of insufficient text representation quality in multi-level, multi-label classification is solved, and the classification accuracy is improved.

CN121009947APending Publication Date: 2025-11-25TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410624320.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-05-20
Publication Date
2025-11-25

AI Technical Summary

Technical Problem

In existing technologies, the vector representation quality of text information is poor in multi-level, multi-label classification tasks, resulting in inaccurate classification results.

Method used

By acquiring a classification training set, determining the baseline samples and their positive and negative samples, constructing a target loss function using a contrastive learning method, and training the text encoder, the quality of text representation is improved.

Benefits of technology

It improves the quality of text representation output by the text encoder and enhances the accuracy of multi-level, multi-label classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121009947A_ABST
    Figure CN121009947A_ABST
Patent Text Reader

Abstract

The invention discloses a text encoder training method, a related method, a related device, equipment and a medium. Comprising the following steps: acquiring a classification training set, comparing sample level classification labels of a plurality of training samples with reference level classification labels of reference samples determined from the plurality of training samples of the classification training set, and determining positive samples and negative samples corresponding to the reference samples; inputting the text information corresponding to the reference sample, the positive sample and the negative sample into a text encoder to obtain a reference sample representation, a positive sample representation and a negative sample representation; and constructing a target loss function according to the first comparison similarity between the reference sample representation and the positive sample representation and the second comparison similarity between the reference sample representation and the negative sample representation, and training the text encoder through the target loss function to obtain a trained text encoder. The method can improve the sampling accuracy of positive and negative samples, so that the encoding performance of the text encoder is improved. The method can be used for various artificial intelligence scenes such as classification tasks.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence, and more particularly, to a text encoder training method and related method, device, equipment and medium. BACKGROUND

[0002] Multi-level multi-label classification is a complex machine learning or data mining task that aims to assign instances to multiple classes, and there can be a hierarchical structure among these classes. This means that each instance can belong to one or more labels, and these labels can be organized in a tree structure, where more specific labels are at lower levels of the tree, and more general labels are at higher levels of the tree.

[0003] The widespread application of natural language processing technology has enabled multi-level multi-label classification tasks based on text information for instances, which can assign instances to multiple classes based on text information related to the instances, and there can be a hierarchical structure among these classes. The challenge of this classification task is to effectively capture the semantic representation of the text information in order to accurately predict the labels of multiple levels.

[0004] When the related art technology performs multi-level multi-label classification on instances using the text information of the instances, it first encodes the text information to obtain a corresponding vector representation, and then inputs the vector representation into a classification model to perform multi-level multi-label classification to obtain a corresponding classification result. However, the classification result output by the above classification method is not accurate enough, which is due to the poor quality of the vector representation of the text information. SUMMARY

[0005] The embodiments of the present application provide a text encoder training method and related method, device, equipment and medium, which can improve the quality of the text representation output by the trained text encoder.

[0006] According to an aspect of the present application, a training method of a text encoder is provided, which comprises: obtaining a classification training set, the classification training set comprising at least a plurality of training samples, and determining a reference sample from the plurality of training samples of the classification training set; comparing a reference level classification label of the reference sample with sample level classification labels of the plurality of training samples, and determining a positive sample and a negative sample corresponding to the reference sample from the plurality of training samples according to a comparison result; inputting reference text information corresponding to the reference sample, positive sample text information corresponding to the positive sample, and negative sample text information corresponding to the negative sample into a text encoder respectively for text encoding to obtain a reference sample representation, a positive sample representation, and a negative sample representation; performing first similarity calculation based on the reference sample representation and the positive sample representation to obtain a first comparison similarity, and performing second similarity calculation based on the reference sample representation and the negative sample representation to obtain a second comparison similarity; constructing a target loss function, the target loss function being a decreasing function of the first comparison similarity and an increasing function of the second comparison similarity, and training the text encoder through the target loss function to obtain a trained text encoder.

[0007] According to an aspect of the present application, a training method of a classification model is provided, which comprises: obtaining a classification training set; inputting text description information of each training sample in the classification training set into a trained text encoder for text encoding to obtain a text description representation corresponding to each training sample; wherein the trained text encoder is obtained according to the training method of the text encoder; performing sample classification on each text description representation through a classification model to obtain a predicted classification result corresponding to each training sample; determining a classification loss function based on a sample level classification label and a predicted classification result corresponding to each training sample, and training the trained text encoder and the classification model according to the classification loss function to obtain a trained target text encoder and a trained classification model.

[0008] According to an aspect of the present application, a training device of a text encoder is provided, which comprises: a first obtaining module, a sample determining module, a first encoding module, a similarity calculating module, and a first training module.

[0009] The first obtaining module is configured to obtain a classification training set, the classification training set comprising at least a plurality of training samples, and determine a reference sample from the plurality of training samples of the classification training set;

[0010] The sample determining module is configured to compare a reference level classification label of the reference sample with sample level classification labels of the plurality of training samples, and determine a positive sample and a negative sample corresponding to the reference sample from the plurality of training samples according to a comparison result.

[0011] The first encoding module is configured to input the reference text information corresponding to the reference sample, the positive text information corresponding to the positive sample, and the negative text information corresponding to the negative sample into a text encoder for text encoding, to obtain a reference sample representation, a positive sample representation, and a negative sample representation;

[0012] The similarity calculation module is configured to perform first similarity calculation based on the reference sample representation and the positive sample representation to obtain a first contrast similarity, and perform second similarity calculation based on the reference sample representation and the negative sample representation to obtain a second contrast similarity.

[0013] The first training module is configured to construct a target loss function, the target loss function being a decreasing function of the first contrast similarity and an increasing function of the second contrast similarity, and train the text encoder through the target loss function to obtain a trained text encoder.

[0014] Optionally, the sample determination module can include a first label determination unit, a first sample acquisition unit, a second label determination unit, a second sample acquisition unit, and a sample comparison unit. The first label determination unit is configured to take the reference hierarchical classification label of the reference sample as the positive sample classification label corresponding to the reference sample. The first sample acquisition unit is configured to acquire a first mapping relationship corresponding to the positive sample classification label, and determine a first sample corresponding to the positive sample classification label based on the first mapping relationship. The second label determination unit is configured to determine a negative sample classification label corresponding to the reference sample based on a label structure tree corresponding to the classification training set. The second sample acquisition unit is configured to acquire a second mapping relationship corresponding to the negative sample classification label, and determine a second sample corresponding to the negative sample classification label based on the second mapping relationship. The sample comparison unit is configured to compare the plurality of training samples with the first sample and the second sample respectively, and determine the positive sample and the negative sample corresponding to the reference sample from the plurality of training samples according to the comparison result.

[0015] Optionally, the second label determination unit can be specifically configured to: find a low-level classification label corresponding to the reference hierarchical classification label in the label structure tree; acquire other hierarchical classification labels in the label structure tree except the reference hierarchical classification label and the low-level classification label; and take the other hierarchical classification labels as the negative sample classification label corresponding to the reference sample.

[0016] Optionally, the comparison result includes a first comparison result and a second comparison result, and the sample determination unit can be specifically configured to: perform comparison between the plurality of training samples and the first sample respectively to obtain the first comparison result; based on the first comparison result, take the training sample matched with the first sample in the plurality of training samples as the positive sample corresponding to the reference application; perform comparison between the plurality of training samples and the second sample respectively to obtain the second comparison result; and based on the second comparison result, take the training sample matched with the second sample in the plurality of training samples as the negative sample corresponding to the reference application.

[0017] Optionally, the sample determination module can further include a label information acquisition unit, a positive sample determination unit and a negative sample determination unit. The label information acquisition unit is configured to acquire first label information of a reference level classification label of the reference sample and second label information of a sample level classification label of the plurality of training samples; the positive sample determination unit is configured to determine the second label information same as the first label information as positive sample label information, and determine the training sample corresponding to the positive sample label information as the positive sample corresponding to the reference sample; and the negative sample determination unit is configured to determine the second label information different from the first label information as negative sample label information, and determine the training sample corresponding to the negative sample label information as the negative sample corresponding to the reference sample.

[0018] Optionally, the negative sample determination unit can be specifically configured to: determine the second label information different from the first label information as intermediate sample label information; determine the intermediate sample label information not covering the first label information as negative sample label information; and determine the training sample corresponding to the negative sample label information from the plurality of training samples as the negative sample corresponding to the reference sample.

[0019] Optionally, the first training module can include a first probability prediction unit, a second probability prediction unit and a loss construction unit. The first probability prediction unit is configured to perform probability prediction according to the first comparison similarity to obtain a positive example probability distribution corresponding to the positive sample; the second probability prediction unit is configured to perform probability prediction according to the second comparison similarity to obtain a negative example probability distribution corresponding to the negative sample; and the loss construction unit is configured to perform accumulation calculation based on the positive example probability distribution and the negative example probability distribution to construct a target loss function.

[0020] Optionally, the first training module can further include a training judgment unit, a first training unit and a second training unit. The training judgment unit is configured to iteratively train the text encoder by using the target loss function, and determine whether the target loss function meets a training end condition after each iteration; the first training unit is configured to adjust parameters of the text encoder when the target loss function does not meet the training end condition, and return to perform the iteration training of the text encoder by using the target loss function until the target loss function meets the training end condition; and the second training unit is configured to stop training the text encoder when the target loss function meets the training end condition, and obtain the trained text encoder.

[0021] Optionally, the first encoding module can include a text acquisition unit, a first encoding unit, a second encoding unit and a third encoding unit. The text acquisition unit is configured to acquire reference text information corresponding to the reference sample, positive sample text information corresponding to the positive sample and negative sample text information corresponding to the negative sample; the first encoding unit is configured to input the reference text information into the text encoder for text encoding to obtain a reference sample representation; the second encoding unit is configured to input the positive sample text information into the text encoder for text encoding to obtain a positive sample representation; and the third encoding unit is configured to input the negative sample text information into the text encoder for text encoding to obtain a negative sample representation.

[0022] According to an aspect of the present application, a training device of a classification model is provided, which includes a second acquisition module, a second encoding module, a sample classification module and a second training module.

[0023] The second acquisition module is configured to acquire a classification training set.

[0024] The second encoding module is configured to input text description information of each training sample in the classification training set into the trained text encoder for text encoding to obtain a text description representation corresponding to each training sample.

[0025] The trained text encoder is obtained according to the training method of the text encoder.

[0026] The sample classification module is configured to perform sample classification on each text description representation by using the classification model to obtain a predicted classification result corresponding to each training sample.

[0027] The second training module is configured to determine a classification loss function based on the sample hierarchical classification label and the predicted classification result corresponding to each training sample, and train the trained text encoder and the classification model according to the classification loss function to obtain a trained target text encoder and a trained classification model.

[0028] Optionally, the second training module can be specifically configured to: obtain a sample hierarchical classification label corresponding to each training sample; perform difference calculation based on the predicted classification result corresponding to each training sample and the sample hierarchical classification label corresponding to each training sample to construct a classification loss function; and iteratively train the trained text encoder and the classification model according to the classification loss function until a training end condition is met, to obtain a trained target text encoder and a trained classification model.

[0029] According to an aspect of the present application, a computer readable storage medium is provided, which stores a computer program, wherein the computer program, when executed by a processor, performs the training method of the text encoder and the related method.

[0030] According to an aspect of the present application, a computer device is provided, which comprises a processor and a memory, and the memory stores a computer program, which is called by the processor to perform the training method of the text encoder and the related method.

[0031] According to an aspect of the present application, a computer program product is provided, which comprises a computer program stored in a storage medium; a processor of a computer device reads the computer program from the storage medium, and the processor executes the computer program to enable the computer device to perform the training method of the text encoder and the related method.

[0032] The training method of the text encoder of the present application can obtain a classification training set, which comprises at least a plurality of training samples, determine a reference sample from the plurality of training samples of the classification training set, and compare the reference hierarchical classification label of the reference sample with the sample hierarchical classification labels of the plurality of training samples, to determine the positive sample belonging to the same category as the reference sample and the negative sample not belonging to the same category as the reference sample for the reference sample. In this way, the positive sample and the negative sample are determined for the reference sample based on the more strict classification standard of sample category, thereby improving the sampling accuracy of the positive and negative samples.

[0033] Further, the reference text information corresponding to the reference sample, the positive sample text information corresponding to the positive sample and the negative sample text information corresponding to the negative sample are respectively input into the text encoder for text encoding to obtain corresponding reference sample representation, positive sample representation and negative sample representation. Further, the first similarity calculation is performed based on the reference sample representation and the positive sample representation to obtain the first contrast similarity, and the second similarity calculation is performed based on the reference sample representation and the negative sample representation to obtain the second contrast similarity; the target loss function is constructed, wherein the target loss function is a decreasing function of the first contrast similarity and an increasing function of the second contrast similarity, and the text encoder is trained through the target loss function to obtain the trained text encoder. Therefore, by encoding the text information of the reference sample and the text information of the positive and negative samples of the reference sample obtained by accurate sampling through the text encoder, the semantic information in the text can be more accurately captured, the accuracy of learning the similarity between the reference sample and its positive sample and the accuracy of learning the difference between the reference sample and its negative sample by the text encoder are improved, thereby improving the text encoding performance of the text encoder and generating high-quality feature representation.

[0034] Other features and advantages of the present application will be set forth in the following description, and in part will be apparent from the description, or can be learned by practice of the present application. The objects and other advantages of the present application will be realized and attained by the structure particularly pointed out in the description and claims. BRIEF DESCRIPTION OF DRAWINGS

[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor on the basis of these drawings.

[0036] Figure 1 A system architecture diagram provided by an embodiment of the present application is shown.

[0037] Figure 2 A system deployment diagram provided by an embodiment of the present application is shown.

[0038] Figure 3 An application scenario diagram provided by an embodiment of the present application is shown.

[0039] Figure 4 Another application scenario diagram provided by an embodiment of the present application is shown.

[0040] Figure 5 A flowchart of a training method of a text encoder provided by an embodiment of the present application is shown.

[0041] Figure 6 A schematic diagram of a hierarchical classification label is shown.

[0042] Figure 7 A schematic diagram of a similarity matrix is shown.

[0043] Figure 8 A schematic diagram of another similarity matrix is shown.

[0044] Figure 9 An example diagram of a mapping relationship is shown.

[0045] Figure 10 A schematic diagram of a label structure tree is shown.

[0046] Figure 11 A label information schematic diagram of a drug classification is shown.

[0047] Figure 12 A flowchart of a training method of a classification model is shown.

[0048] Figure 13 A training flowchart of a classification model is shown.

[0049] Figure 14 A test flowchart is shown.

[0050] Figure 15 A cosine similarity heat map is shown.

[0051] Figure 16 Another cosine similarity heat map is shown.

[0052] Figure 17 A module block diagram of a text encoder training device is shown.

[0053] Figure 18 A module block diagram of a classification model training device is shown.

[0054] Figure 19 A module block diagram of a computer device is shown.

[0055] Figure 20 A module block diagram of a computer readable storage medium is shown. DETAILED DESCRIPTION

[0056] In order for those skilled in the art to better understand the solutions of the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all the embodiments. The embodiments described by reference to the accompanying drawings are exemplary and are only used to explain the present application, and cannot be understood as a limitation on the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present application.

[0057] In some processes described in the specification, claims and the above drawings, a plurality of steps are included, which appear in a specific order, but it should be clearly understood that these steps can be executed or executed in parallel without the order in which they appear in this text, and the step number is only used to distinguish different steps, and the number itself does not represent any execution order. In addition, the description of "first", "second" or "target" in this paper is used to distinguish similar objects, and does not necessarily describe a specific order, sequence or quantity.

[0058] It is worth noting that in the specific embodiments of the present application, data related to training samples and the like are involved, and when the above embodiments of the present application are applied to specific products or technologies, the permission or consent of the subject is required, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards. For example, when the embodiments of the present application need to obtain training samples, a pop-up window or a jump to a confirmation page can be used to obtain separate permission or separate consent of the training samples, and after obtaining separate permission or separate consent, the necessary related data for the normal operation of the embodiments of the present application is obtained.

[0059] Before further detailing the embodiments of the present application, the terms and terms involved in the embodiments of the present application are explained, and the terms and terms involved in the embodiments of the present application are applicable to the following explanations:

[0060] The training method of the text encoder proposed in the present application involves artificial intelligence (Artificial Intelligence, AI) technology. Artificial intelligence technology is to use digital computers or digital computer-controlled machines to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology of computer science, which tries to understand the essence of intelligence and produce a new intelligent machine that can react in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines, so that machines have the functions of perception, reasoning and decision-making.

[0061] Artificial intelligence technology is a comprehensive discipline, involving a wide range of fields, both hardware and software technologies. Artificial intelligence basic technologies generally include, such as sensors, special artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-training model technology, operation / interaction system, mechatronics, etc. Among them, the pre-training model is also called large model, basic model, which can be widely applied to downstream tasks in various directions of artificial intelligence after fine-tuning. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning, etc.

[0062] Deep learning is a new research direction in the field of machine learning. Deep learning can learn the internal rules and representation levels of sample data, and the information obtained in this learning process can greatly help the interpretation of data such as text, images and sound. The ultimate goal of deep learning is to enable machines to have analysis and learning capabilities like humans, and to be able to recognize text, images and sound data. Deep learning is a complex machine learning algorithm, and has achieved many results in other related fields such as search technology, data mining, machine translation and recommendation system. For example, in the embodiment of the present application, representation learning is performed by using a text encoder to generate a text feature representation for multi-level multi-label classification.

[0063] Among them, the training process of the text encoder adopts contrastive learning, which is a machine learning method that aims to learn a model by comparing rather than directly predicting labels or values. It learns the similarity or difference between instances in the input data by comparing them. The key to this method is to define a comparison metric to measure the similarity or difference between instances. Contrastive learning can be applied to a variety of tasks, including classification, clustering, ranking and recommendation, etc. Its advantage is that it can learn from a large amount of unlabeled data, and in some cases it can produce more robust and generalizable models.

[0064] At present, multi-level multi-label classification based on text information is a complex classification task. By using the text information of the instance to be classified, the instance can be assigned to multiple categories, and these categories may also have a hierarchical structure. This classification task has a wide range of applications in many fields, including news classification: classifying news articles by theme or category, such as economy, sports, etc., while considering more specific subcategories; product review analysis: analyzing product reviews on e-commerce websites, labeling product features, user experience, purchase inclination, etc.; In addition, it also includes legal document classification and application software (APP) classification, etc.

[0065] For example, software classification based on description text about software is a task of classifying software according to text data such as description of software functions. This classification method can help organize and manage a large number of software resources, making it easier for users to find software that meets their needs. For example, software can be classified into different categories such as office software, entertainment software, development tools, security software, etc., and under different categories, there can also be multi-level classification, such as entertainment software can include game software and social software, etc., and game software can include puzzle game software and strategy game software, etc.

[0066] In multi-level multi-label classification based on text information, first, the text information needs to be encoded using a text encoder, and then the encoded text representation is used for the classification task. Therefore, the encoding performance of the text encoder plays a crucial role in the accuracy of the classification task. However, most related technologies use open-source pre-trained text encoders, which have poor vector representation capabilities, resulting in poor quality text representations used by classification models, which cannot accurately classify instances into target categories. Especially in multi-level multi-label classification tasks, the quality of classification representation is higher than that of single-label classification, and related technologies cannot obtain accurate classification results.

[0067] To solve the above problems, the inventors have proposed a training method for a text encoder and related methods provided in the present application. First, the system architecture of the above-mentioned methods related to the present application and related application scenarios will be described.

[0068] Please refer to Figure 1 , Figure 1 A system architecture diagram is shown. As Figure 1 indicated, the above-mentioned methods provided in the present application can be applied in system 100. Data acquisition device 110 is used to acquire training data, which is a training sample used for network training, and the training sample includes corresponding text information and hierarchical classification labels. For example, the training sample is an application software, the text information is a text describing the function, attribute, etc. of the application software, and the hierarchical classification label marks the classification of the application software. The hierarchical classification label can be manually labeled. After obtaining the training data, the data acquisition device 110 can store the training data in the database 120, and the training device 130 can train the preset neural network based on the training data in the database 120 until the preset neural network meets the training end condition, obtaining the trained target model, i.e. the classification model 101 and the text encoder 102.

[0069] Optionally, the training data maintained in the database 120 does not necessarily all come from the data acquisition device 110, but can also be received from other devices, for example, the execution device 140 can also serve as a data acquisition end, directly obtaining the data as new training data and storing it in the database 120. In addition, the training device 130 does not necessarily train the preset neural network based on the training data maintained in the database 120, but can also train the preset neural network based on the training data obtained from the cloud or other devices, for example, the training device 130 can obtain the latest online application software and the text information corresponding to the application software from an application publishing platform, and the above description should not be regarded as a limitation of the embodiments of the present application.

[0070] The training end condition can be that the total loss value of the target loss is less than a preset value, the total loss value of the target loss is within a threshold range, or the number of training reaches a preset number, etc. The target model can be a deep neural network or a network composed of multiple neural networks. Illustratively, the classification model 101 can be a neural network capable of classification such as a support vector machine (SVM), a multilayer perceptron, etc. The text encoder 102 can be a language model: Word2vec, Elmo, BERT, etc. neural network capable of encoding text information to generate vector representation, which is not limited here.

[0071] In the process of executing the related processing of the processing module 141 of the execution device 140, the execution device 140 can call the data, programs, etc. in the data storage system 150 for corresponding calculation processing, and store the processing results, etc. data and instructions obtained by calculation processing in the data storage system 150. For example, the execution device 140 can use the text encoder 102 and the classification model 101 to perform multi-level classification on the target object to be classified, and store the classification results in the data storage system 150.

[0072] The above-mentioned training device 130 and execution device 140 can be a server or a terminal computer device. The server can be a standalone physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content distribution networks (CDN), blockchains, and big data and artificial intelligence platforms, etc. Basic cloud computing services. The terminal can be a notebook computer, a tablet computer, a desktop computer, etc.

[0073] It should be noted that, Figure 1The system architecture provided in the embodiments of the present application is only a schematic diagram, and the system architecture and application scenarios described in the embodiments of the present application are used to more clearly illustrate the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. For example, Figure 1 The data storage system 150 in the above formula is an internal memory relative to the execution device 140, and in other cases, the data storage system 150 can also be placed outside the execution device 140.

[0074] Please refer to Figure 2 , Figure 2 A system deployment schematic diagram is shown. Exemplarily, the training method of the text encoder and the related method provided by the embodiments of the present application can also be deployed in a system as shown in Figure 2 The system includes a terminal 240, an Internet 230, a gateway 220, a server 210, and the like.

[0075] The terminal 240 can include desktop computers, laptop computers, personal digital assistants (PDAs), smart phones, special-purpose terminals, and the like. In addition, it can be a single device or a collection of multiple devices. The terminal 240 can communicate with the Internet 230 in a wired or wireless manner and exchange data. For example, the terminal 240 can communicate with the Internet 230 through a wireless router 250.

[0076] The server 210 refers to a computer system that can provide certain services to the terminal 240. Compared with the ordinary terminal 240, the server 210 has higher requirements in stability, security, performance, and the like. The server 210 can be a high-performance computer in a network platform, a cluster of multiple high-performance computers, a part of a high-performance computer (for example, a virtual machine), a combination of parts of multiple high-performance computers (for example, virtual machines), and the like.

[0077] The gateway 220 is also called an inter-network connector or a protocol converter. The gateway realizes network interconnection at the transport layer and is a computer system or device that acts as a conversion role. In the case of two systems using different communication protocols, data formats or languages, or even completely different architectures, the gateway is a translator. At the same time, the gateway can also provide filtering and security functions. The messages sent by the terminal 240 to the server 210 can be transmitted to the corresponding server 210 through the gateway 220. The messages sent by the server 210 to the terminal 240 can also be transmitted to the corresponding terminal 240 through the gateway 220.

[0078] The training method of the text encoder and the related method can be implemented completely on the terminal 240, completely on the server 210, partially on the terminal 240, and partially on the server 210.

[0079] As an implementation, the training method of the text encoder and the related method can be implemented completely on the terminal 240. For example, the terminal 240 can train the text encoder and the classification model locally as a training device. Then, the terminal 240 can also classify target objects that need to be classified locally as an execution device using the text encoder and the classification model. For example, for a drug classification task, the terminal 240 can perform text encoding on the text in the specification of the drug to be classified by using the text encoder, input the vector representation obtained by encoding into the classification model for drug classification, and obtain the classification result, which can be in the form of a label, such as <prescription drug, dermatology>.

[0080] As another implementation, the training method of the text encoder and the related method can be implemented completely on the server 210. For example, the server 210 can train the text encoder and the classification model locally as a training device. Then, the server 210 can also classify target objects that need to be classified locally as an execution device using the text encoder and the classification model. For example, the terminal 240 stores a plurality of news articles to be classified in a classification list. The terminal 240 can upload the classification list to the server 210, and then the server 210 can perform text encoding on the plurality of news articles in the classification list by using the text encoder to obtain corresponding vector representations, input the vector representations obtained by encoding into the classification model for news classification, and obtain the classification result corresponding to each news article. The server 210 returns the classification result to the terminal 240.

[0081] As another implementation, the training method of the text encoder and the related method can be implemented completely on the server 210. For example, the server 210 can train the text encoder and the classification model locally as a training device. Then, the server 210 can also classify target objects that need to be classified locally as an execution device using the text encoder and the classification model. For example, the terminal 240 stores a plurality of news articles to be classified in a classification list. The terminal 240 can upload the classification list to the server 210, and then the server 210 can perform text encoding on the plurality of news articles in the classification list by using the text encoder to obtain corresponding vector representations, input the vector representations obtained by encoding into the classification model for news classification, and obtain the classification result corresponding to each news article. The server 210 returns the classification result to the terminal 240.

[0082] It should be noted that, Figure 2This is only a system deployment diagram provided by an embodiment of the present application. The system deployment scheme described in the embodiments of the present application is only for more clearly illustrating the technical solutions of the embodiments of the present application, and does not constitute a limitation on the technical solutions provided by the embodiments of the present application. For example, the terminal 240 can generally refer to one of a plurality of terminals, and the embodiments of the present application only take three terminals shown in FIG. 2 as an example. Those skilled in the art can know that, as the system deployment scheme evolves, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems. Figure 2

[0083] The embodiments of the present application can be applied in various scenarios, such as a cloud test service scenario as shown in FIG. 1, Figure 3 a game test service scenario as shown in FIG. 2, and the like. Figure 4

[0084] (1) Product analysis scenario

[0085] Internet product analysis refers to the systematic research and evaluation of Internet products to understand the characteristics, functions, user experience, market performance, and the like of the products. Such analysis usually involves multiple aspects, including technology, business, and users. Internet technology companies usually analyze the Internet products, i.e., application software, developed by them in order to deeply understand the product portfolio and develop effective strategies. In this way, the companies can clarify the positioning, target users, and market demand of each product, which helps to accurately position and allocate resources. In the process of analyzing and counting products, software classification is an important environment, and in order to improve the accuracy of classification, multi-level classification is required. The training method of a text encoder and related methods provided by the present application can be applied in this product analysis scenario.

[0086] Exemplarily, the trained text encoder and classification model in the training method of a text encoder and related methods can develop an application program for multi-level multi-label classification of software, which can be applied in a product analysis system. As shown in FIG. 3, the terminal 240 is installed with a product analysis system 241. The product analysis system 241 can perform product analysis on application software, in which the application software to be classified can be classified. For example, a user can import an application list storing software to be classified into the product analysis system 241, and then the product analysis system 241 can read each software to be classified from the application list, perform multi-level multi-label classification, and statistically analyze the classification results, and finally display the results in the form of visual icons on the interface. Figure 3

[0087] ​​​Specifically, the product analysis system 241 can utilize the text encoder to encode the text information of each software to be classified to obtain a corresponding text representation, and further input the text representation into the classification model for classification processing to obtain a classification result corresponding to each software to be classified. For software classification, the scene can set a multi-level multi-label classification label. The classification result of each software to be classified has at least a first level of classification.

[0088] For example, taking the first level as an example, the first level classification can include finance, entertainment, education and medical treatment. The product analysis system 241 can visualize the classification distribution based on the classification result corresponding to each software to be classified, such as Figure 3 As shown, the pie chart displayed by the product analysis system 241 shows the classification distribution of all application software under the first level classification. Through the pie chart, the user can intuitively see the distribution of each application software, in addition, the product analysis system 241 can also generate a classification tree based on the classification result corresponding to each software to be classified to display the relationship between the classifications. In this way, by classifying the application software, the user can analyze and count the products to determine the product positioning and market demand.

[0089] (2) Auxiliary diagnosis scene

[0090] Intelligent triage is a process that uses artificial intelligence technology and intelligent algorithms to analyze and evaluate information such as patient symptoms and medical history, and provides appropriate medical advice and guidance for patients. Through the intelligent triage system, patients can input symptom information at home or in the hospital outpatient department, and the system will automatically identify possible disease types based on the data provided by the patient and evaluate the severity of the condition. The system will provide medical advice based on the patient's condition, including the urgency of seeking medical treatment, department recommendations, and self-management suggestions. This intelligent system can effectively assist hospitals and doctors in managing the triage process, improve the utilization efficiency of medical resources, and provide more convenient and personalized medical services for patients. The training method of the text encoder and the related method provided in the present application can be applied in this auxiliary diagnosis scene.

[0091] Exemplarily, the trained text encoder and classification model in the training method of the text encoder and the related method based on the text encoder can develop an application program for disease classification, which can be applied in the guidance system. As shown in Figure 4As shown, the terminal 240 is installed with the guidance system 241. The user can input a disease condition and inquire a treatment department at a dialogue interface of the guidance system 241. Further, the guidance system 241 performs content extraction on the input of the user, obtains disease condition content, and performs disease classification on the disease condition content by using a text encoder and a classification model, to obtain a classification result of the disease condition content, for example, "surgery-neurosurgery". Further, the guidance system 241 generates a reply sentence based on the classification result, and inputs the reply sentence to the dialogue interface, so as to complete guidance for the user.

[0092] It should be noted that, Figure 3 and Figure 4 are only two application scenario schematic diagrams provided by the embodiments of the present application. The application scenarios described in the embodiments of the present application are only for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. For example, the training method of the text encoder and the related method proposed in the present application can also be used in a multi-level multi-label classification scene that needs to be based on text information, such as paper classification and title classification. It can be known by those skilled in the art that, with the evolution of application scenarios, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems.

[0093] According to an embodiment of the present application, a training method of a text encoder is provided. The training method of the text encoder can be executed by a computer device (the computer device can be a server, an edge computing device, or other terminal device with certain processing capability), which at least has the functions of storage, calculation and communication. Figure 5 A flowchart of the training method of the text encoder provided by the embodiment is shown. As shown in the figure, Figure 5 The training method of the text encoder can specifically include:

[0094] S110: Obtain a classification training set, and determine a reference sample from a plurality of training samples of the classification training set.

[0095] In the embodiments of the present application, the classification training set refers to a sample set used for training the text encoder or the classification model. The classification training set at least includes a plurality of training samples, each training sample being a target object to be classified. The target object includes, but is not limited to, news content, application software, and drugs, and the like, which are things to be classified through textual description. In different classification tasks, the target object can be various, depending on the actual task and application scenario, which is not limited herein. For example, for the classification task of sentiment polarity in a recommendation system, the target object can be a user's product review, and for the scenario of terminal application security detection, the target object can be an application software installed on a terminal. Each training sample is associated with corresponding text information and a sample hierarchical label. The text information refers to the textual description of the training sample, and the text information can describe the features or attributes of the training sample, which is an expression of semantic understanding of the training sample.

[0096] For example, the training sample can be an application software to be classified, and correspondingly, the text information can include the application name (App Name) of the application software: a name capable of reflecting the function or theme of the application software; application description (App Description): a textual description of the function, characteristics and advantages of the application software; keywords (Keywords): some descriptive words or phrases for searching and discovering the application software; update log (Update Log): update content and improvement description published when the application software is updated, which can include information such as repaired code errors or defects, new functions, performance improvements, etc., to help users understand the latest changes of the application software, and the like. For example, the training sample can also be a drug to be classified, and correspondingly, the text information can be the drug instruction manual of the drug, which provides detailed information about the use of the drug, including the drug name and specification: the generic name and trade name of the drug, and the specification, dosage form, and the like of the drug; composition and action description: the main composition and mechanism of action of the drug, and the influence of the drug on the body; indication: description of the diseases or symptoms that the drug is suitable for treating, and the like.

[0097] In the classification task, the label refers to an identifier used to identify and distinguish different categories. They are usually used to describe the categories of data or the attributes of categories. The sample hierarchical label in the present application refers to an identifier used to divide the training sample into a plurality of levels or categories, and each category is assigned a label. Each label represents a specific category or concept, and the level represents the superior-inferior relationship between the labels. For example, Figure 6As shown in the schematic diagram of the hierarchical classification label, the sample hierarchical label of the training sample includes three levels: a first level, a second level, and a third level, each level corresponding to different categories, and each level of categories is divided into more specific sub-categories by the next level. This hierarchical structure can provide more rich semantic information, making the classification result more interpretable and explainable. At the same time, it can also help to reduce the number of categories, thereby improving the efficiency and accuracy of the classifier.

[0098] As an implementation, when the training method of the text encoder as described above is executed on the server 210 or the terminal 240, the server 210 or the terminal 240 can obtain the classification training set through a data calling interface from a local database or a cloud database storing the classification training set. Alternatively, the original data can also be obtained from the Internet to generate the training sample, for example, the server 210 collects different application software and related text information from the Internet, including targeted collection through specific websites in the field where the application software is published, social media platforms, etc., or using search engines to search for keywords associated with the application software publication to obtain the application software and related text information. Further, the collected text information is preprocessed to meet the standard for training, and the application software is labeled in multiple levels, thereby constructing the classification training set.

[0099] Generally, the classification task uses the text encoder to encode the text information of the object to be classified into a feature vector, and then uses the classification model to classify based on the feature vector, so the quality of the feature vector encoded by the text encoder plays a crucial role in the accuracy of the classification task. Therefore, the present application proposes a training method of a text encoder. This method learns by comparing the similarities and differences between sample data.

[0100] In the present application, the text encoder is required to map similar samples to close positions in the representation space, and to map dissimilar samples to distant positions in the representation space. In this way, the goal of learning is to learn a good feature representation so that similar samples are closer in the feature space, and dissimilar samples are farther apart. By learning the similarities and differences between different samples, the text encoder can generate a feature representation of the text information, which can be used for classification tasks.

[0101] The similarity and difference between different samples are measured by a triple consisting of three different roles of samples. The three roles of samples include anchor, positive and negative samples. The anchor sample is the starting point of sample comparison, also known as anchor point or reference point, and can be regarded as the object to be compared. The positive sample is similar to the anchor sample, and the negative sample is dissimilar to the anchor sample. For example, a piece of financial news is taken as the anchor sample, then a piece of financial planning news can be regarded as the positive sample of the anchor sample, and a piece of sports information can be regarded as the negative sample of the anchor sample. During the training process, the text encoder is required to make the distance between the anchor sample and the positive sample as small as possible, and the distance between the anchor sample and the negative sample as large as possible in the representation space.

[0102] As an implementation, the anchor sample can be determined from the plurality of training samples of the classification training set. Exemplarily, the classification training set can be split into different batches by batch processing operation, and then in the process of training based on each batch, the training samples in each batch are traversed, and each traversed training sample is taken as the anchor sample. In this way, dividing the classification training set into different batches can increase the randomness and diversity of training, and then randomly selecting training samples in each batch can ensure that the text encoder can learn different sample combinations in each iteration, thereby improving the generalization ability of the text encoder and reducing the risk of overfitting.

[0103] For example, the number of samples included in each batch is specified by setting a batch size parameter, and then a data loader is used to iterate through the training dataset and obtain data for each batch. The data loader can be implemented using the corresponding tools and interfaces provided by deep learning frameworks such as PyTorch and TensorFlow. Taking PyTorch as an example, in PyTorch, the torch.utils.data.DataLoader class can be used to create a data loader for batch allocation of the classification training set. Specifically, a custom dataset class CustomDataset is first defined, and the __len__ and __getitem__ methods are implemented in the CustomDataset class to specify the length of the dataset and obtain a single training sample. Further, an instance of the dataset is created. Then, a data loader dataloader is created using the DataLoader class, and the batch size and whether to shuffle the classification training set are set. Finally, by iterating through the dataloader, the training samples in each batch can be obtained, and each batch of training samples is stored in batch_data.

[0104] S120: Comparing the reference level classification label of the reference sample and the sample level classification labels of the plurality of training samples, and determining the positive sample and the negative sample corresponding to the reference sample from the plurality of training samples according to the comparison result.

[0105] Currently, when the related art learns the similarities and differences between different samples, in the same representation space, the vector representations of the positive sample pair are as close as possible, and the vector representations of the negative sample pair are as far apart as possible, wherein the positive sample pair is composed of the reference sample and its data augmented data, and the negative sample pair is composed of the reference sample and the randomly selected sample from the batch except the positive sample. It can be understood that this learning method focuses on learning the features of the sample source, that is, the features of the positive sample and the features of the negative sample are compared with the features of the reference sample, respectively.

[0106] However, this learning method cannot accurately sample positive and negative samples, because the positive and negative sample sampling based on the "sample source" only uses data augmented data as positive samples, does not consider sampling positive samples from the batch, and considers all other training samples in the batch except the reference sample as negative samples.

[0107] For example, taking a hierarchical classification task as an example, assuming that the size of a batch is 2. According to the sampling method of the related art, all data except the positive sample are regarded as negative samples, and after positive and negative sample sampling, two triplets (Anchor1, Positive1, Negative1) and (Anchor2, Positive2, Negative2) can be obtained. Among them, Anchor is the anchor sample, Positive is the positive sample, and Negative is the negative sample.

[0108] Please refer to Figure 7 and Figure 8 , Figure 7 a schematic diagram of a similarity matrix is shown, Figure 8 a schematic diagram of another similarity matrix is shown. Figure 7 and Figure 8 In the similarity matrix of Figure 7 As shown, in the learning process of the text encoder, the goal of learning is to make the similarity of the shaded cells (positive sample pairs) in the similarity matrix tend to 1, and the similarity of other cells (negative sample pairs) tend to 0 (excluding the similarity of the sample and itself on the diagonal position), that is, the feature vectors of the same class should be similar, and the feature vectors of different classes should be dissimilar.

[0109] But the related art regards all data except the positive sample in the batch as negative samples. For the anchor sample, other samples in the batch may contain labels belonging to the same class, which often occurs in a hierarchical label sample pair. As shown in Figure 8 The labels of Anchor1 and Positive1 are finance, so when the labels of Anchor2 and Positive2 are also finance, Figure 7 The other cells in Figure 8 The shaded cells in The anchor2 and positive2 will be regarded as negative samples, which will cause sampling errors of positive and negative samples.

[0110] Therefore, the positive and negative sample sampling based on the “sample origin” of the related art will inevitably cause the error of regarding the same class samples in the batch as positive samples, and regarding the same class samples as negative samples, thereby interfering with the learning of the text encoder. Therefore, the present application proposes a method for determining positive and negative samples based on sample labels, that is, a method for determining positive and negative samples for anchor samples.

[0111] In a multi-level label classification scenario, data is usually organized using hierarchical labels, which show the most specific category information to which each data point belongs. The related art fails to accurately find positive samples and negative samples for a reference sample because it ignores the most direct category information between samples and uses "sample origin" to sample positive and negative samples. Therefore, the present application uses multi-level labels to determine positive samples belonging to the same category as the reference sample and negative samples not belonging to the same category for the reference sample, thereby improving the accuracy of positive and negative sample sampling.

[0112] In some embodiments, by comparing the reference hierarchical classification label of the reference sample with the sample hierarchical classification labels of the plurality of training samples, positive samples whose label categories are the same as each other and negative samples whose label categories are different from each other can be determined for the reference sample from the plurality of training samples according to the comparison result. Alternatively, the comparison can be performed according to the hierarchical classification structure of the hierarchical classification labels of the training samples, or according to the label information of the hierarchical classification labels of the training samples.

[0113] As an implementation, the comparison is performed according to the hierarchical classification structure of the hierarchical classification labels of the training samples, and the positive samples and the negative samples are determined for the reference sample from the plurality of training samples based on the comparison result, which can specifically include the following steps:

[0114] (1) Taking the reference hierarchical classification label of the reference sample as the positive sample classification label corresponding to the reference sample.

[0115] The reference hierarchical classification label refers to the sample hierarchical classification label corresponding to the reference sample, and the positive sample classification label refers to the sample hierarchical classification label corresponding to the positive sample. For example, since the positive sample is a training sample whose sample hierarchical classification label is the same as that of the reference sample, the reference hierarchical classification label of the reference sample can be directly taken as the positive sample classification label of the positive sample corresponding to the reference sample.

[0116] (2) Obtaining a first mapping relationship corresponding to the positive sample classification label, and determining a first sample corresponding to the positive sample classification label based on the first mapping relationship.

[0117] The mapping relationship is used to record the mapping relationship between the sample hierarchical classification label and the training sample, and specifically, a sample hierarchical classification label of a category can have a mapping relationship with at least one training sample. For example, the sample hierarchical classification label <game-online> has a mapping relationship with game software A and game software B, i.e., the sample hierarchical classification labels of game software A and game software B include <game-online>. The mapping relationship corresponding to the positive sample classification label is the first mapping relationship. The first sample refers to a training sample having a mapping relationship with the positive sample classification label.

[0118] For example, the mapping relationship can be stored in a database using a mapping table data format. This mapping table can record the sample hierarchical classification labels and the unique identifiers (IDs) of the training samples. When obtaining the classification label of a positive sample, based on the positive sample classification label in the database, the unique identifier of the first sample corresponding to the positive sample classification label is obtained by querying the mapping table, and then the first sample is obtained based on this unique identifier. Figure 9 The diagram illustrates an example of the mapping relationship. Each sample-level classification label has a corresponding bag-of-words, which records the sample-level classification label and its corresponding training samples. For example, the sample-level classification label includes the bag-of-words corresponding to <game-online>, which records game software A and game software B corresponding to this label. Therefore, game software A and game software B can be obtained based on game software A and game software B.

[0119] (3) Based on the label structure tree corresponding to the classification training set, determine the negative sample classification label corresponding to the benchmark sample.

[0120] In this context, the label structure tree refers to a tree-like structure used to organize and represent the hierarchical relationships between labels. In the multi-class classification task of this application, labels have a hierarchical structure, where some labels are subcategories of other labels, and other labels are higher-level parent categories. The label structure tree is a data structure used to represent this hierarchical relationship.

[0121] A label structure tree typically consists of tree nodes and edges. Nodes represent categories or labels, with each node being the hierarchical classification label for a sample of a specific category in the classification training set. Edges represent the hierarchical relationships between categories. Each node may have one parent node and zero or more child nodes, thus forming a tree structure. The leaf nodes of the tree represent the most specific subcategory. By analyzing the hierarchical relationship of the baseline hierarchical classification labels within the label structure tree, the classification labels of the negative samples corresponding to the baseline samples can be determined. This can involve the following steps:

[0122] (3.1) Find the corresponding lower-level classification label in the label structure tree for the baseline classification label;

[0123] (3.2) Obtain the classification labels of other levels in the label structure tree, excluding the baseline classification label and the lower-level classification labels;

[0124] (3.3) Use the classification labels of other levels as the negative sample classification labels corresponding to the baseline sample.

[0125] In this context, lower-level classification labels refer to classification labels that belong to the same category as the baseline classification label, but whose classification level is lower than that of the baseline classification label. For example... Figure 10A schematic diagram of a label structure tree is shown, which describes the hierarchical relationship between labels related to the classification task of the application software. The label structure tree has three classification levels, and the lower level classification labels of the first level classification label <game> include <game-online>, <game-single> and <game-online-multiplayer>.

[0126] Considering that the training samples corresponding to the sample hierarchical classification labels different from the hierarchical classification labels are directly used as negative samples, this method has the possibility of using training samples of the same category as negative samples, for example, using Figure 10 Taking the label <game-online> as an example, the lower level classification label <game-online-multiplayer> of the label is also regarded as a negative sample, but in fact, the lower level classification label <game-online-multiplayer> and the label <game-online> belong to the same label.

[0127] Therefore, when obtaining the reference hierarchical classification label, the corresponding lower level classification label of the reference hierarchical classification label in the label structure tree can be found first, and other hierarchical classification labels in the label structure tree except the reference hierarchical classification label and the lower level classification label can be obtained, and then the other hierarchical classification labels are classified as the negative sample classification labels corresponding to the reference sample. For example, the lower level classification label <game-online-multiplayer> of the reference hierarchical classification label <game-online> is found, and other hierarchical classification labels <game-single>, <education>, <education-exam-research>, <social-travel> and the like can be obtained in the label structure tree.

[0128] (4) Obtain a second mapping relationship corresponding to the negative sample classification label, and determine a second sample corresponding to the negative sample classification label based on the second mapping relationship.

[0129] The mapping relationship corresponding to the negative sample classification label is the second mapping relationship. The second sample refers to a training sample having a mapping relationship with the negative sample classification label. Exemplarily, the second mapping relationship can be stored in a database through the data format of the mapping table, and when the negative sample classification label is obtained, the second sample corresponding to the negative sample classification label is obtained by querying the mapping table based on the negative sample classification label in the database.

[0130] (5) Compare the plurality of training samples with the first sample and the second sample respectively, and determine the positive sample and the negative sample corresponding to the reference sample from the plurality of training samples according to the comparison results.

[0131] The comparison results include first comparison results and second comparison results. The first comparison results refer to the comparison results obtained by comparing the plurality of training samples with the first sample respectively. The second comparison results refer to the comparison results obtained by comparing the plurality of training samples with the second sample respectively.

[0132] Specifically, the first comparison result is obtained by comparing each of the plurality of training samples with the first sample, and the training sample matching the first sample among the plurality of training samples is determined as the positive sample corresponding to the reference application based on the first comparison result. Further, the second comparison result is obtained by comparing each of the plurality of training samples with the second sample, and the training sample matching the second sample among the plurality of training samples is determined as the negative sample corresponding to the reference application based on the second comparison result. For example, the plurality of training samples {Sam1, Sam2, Sam3, Sam4, Sam5, Sam6, Sam7}, the first sample includes {Sam1, Sam3, Sam4}, and the second sample includes {Sam2, Sam7}. Comparing the plurality of training samples with the first sample, it is found that Sam1, Sam3, and Sam4 in the plurality of training samples are the same as the first sample, and further Sam1, Sam3, and Sam4 in the plurality of training samples are determined as the positive sample corresponding to the reference sample. Comparing the plurality of training samples with the second sample, it is found that Sam2 and Sam7 in the plurality of training samples are the same as the second sample, and further Sam2 and Sam7 in the plurality of training samples are determined as the negative sample corresponding to the reference sample.

[0133] As another implementation, the comparison can be performed according to the hierarchical classification structure of the hierarchical classification label of the training sample, or the comparison can be performed according to the label information of the hierarchical classification label of the training sample, and the positive sample and the negative sample are determined for the reference sample from the plurality of training samples based on the comparison result. Specifically, it can include the following steps:

[0134] (1) obtaining the first label information of the reference hierarchical classification label of the reference sample and the second label information of the sample hierarchical classification label of the plurality of training samples;

[0135] (2) determining the second label information same as the first label information as the positive sample label information, and determining the training sample corresponding to the positive sample label information as the positive sample corresponding to the reference sample;

[0136] (3) determining the second label information different from the first label information as the negative sample label information, and determining the training sample corresponding to the negative sample label information as the negative sample corresponding to the reference sample.

[0137] The label information refers to content information represented by the label. The label can be represented in various ways, depending on the requirements of the application and the characteristics of the data, including but not limited to: textual representation, using text to represent the label, for example, a label can be a word or a phrase, etc., used to describe a certain category or attribute; numerical representation, using numbers to represent the label, which can be continuous or discrete, representing different categories or attributes; One-Hot Encoding: a commonly used representation method in machine learning. For a limited number of labels, each label is encoded as a vector, the length of the vector is equal to the number of labels, and the position corresponding to the label is 1 and the other positions are 0.

[0138] In the embodiments of the present application, the first label information refers to the label information of the reference hierarchical classification label, and the second label information refers to the label information of the sample hierarchical classification label. For example, the first label information can be compared with the second label information corresponding to each training sample, the second label information identical to the first label information is determined as positive sample label information, and the training sample corresponding to the positive sample label information is determined as a positive sample corresponding to the reference sample. Further, the second label information different from the first label information is determined as negative sample label information. Wherein, the same label information refers to the information representing the label content being the same information, for example, taking the label information in the form of One-Hot Encoding as an example, the label information [1, 0, 1, 1] of label A is the same as the label information [1, 0, 1, 1] of label B, and the label information [1, 0, 1, 1] of label A is different from the label information [1, 0, 1, 0] of label D, the label information [1, 0] of label E, and the label information [1, 0, 1, 1, 1] of label E.

[0139] Accordingly, considering that determining the second label information, which is different from the first label information, as the negative sample label information, and determining the training sample corresponding to the negative sample label information as the negative sample corresponding to the benchmark sample, this method may result in the use of training samples of the same category, such as samples at the next level of the same category as the benchmark sample, as negative samples. Therefore, we can first determine the second label information, which is different from the first label information, as the intermediate sample label information. Then, we determine the intermediate sample label information that does not cover the first label information as the negative sample label information, and from multiple training samples, we determine the training sample corresponding to the negative sample label information as the negative sample corresponding to the benchmark sample. Here, "covering" can be understood as the elements at the same element position in the label information being the same. Taking the label information in the form of one-hot encoding as an example, the label information [1,0] of label H and the label information [1,0,1,1] of label G, because the elements "1,0" at the first element position and the second element position are the same, the label information of label G covers the label information of label H. Taking word-based tag information as an example, the tag information of tag T, <game, online>, and the tag information of tag R, <game, online, multiplayer>, are the same because the elements "game, online" in the first and second element positions are identical. Therefore, the tag information of tag R overrides the tag information of tag T.

[0140] like Figure 11 The diagram illustrates the label information for drug classification. The label information for each sample level can be represented by a vector, where the length of the vector equals the number of labels, and each element in the vector corresponds to a label. For example... Figure 11 There are 7 labels: Traditional Chinese Medicine (TCM), Prepared Chinese Medicine, Processed Chinese Herbal Medicine, Western Medicine, Respiratory System, Digestive System, and Nervous System. The corresponding label information vector is [TCM, Prepared Chinese Medicine, Processed Chinese Herbal Medicine, Western Medicine, Respiratory System, Digestive System, Nervous System]. Element values ​​are represented by 0 or 1, where 1 indicates that the training sample belongs to the category of that label, and 0 indicates that the training sample does not belong to the category of that label. For the sample hierarchical classification label: <TCM-Processed Chinese Herbal Medicine>, the corresponding label information is: [1,1,0,0,0,0,0]. For the sample hierarchical classification label: <Western Medicine-Digestive System>, the corresponding label information is: [0,0,0,1,0,1,0].

[0141] Exemplarily, the first label information [1, 0, 0, 0, 0, 0, 0] of the benchmark level category label <Traditional Chinese Medicine> is compared with the second label information corresponding to each training sample, the second label information same as the first label information is determined as positive sample label information, and the training sample corresponding to the positive sample label information is determined as the positive sample corresponding to the benchmark sample. Further, the second label information different from the first label information is determined as intermediate sample label information, for example, the intermediate sample label information [1, 1, 0, 0, 0, 0, 0], [0, 0, 0, 1, 0, 1, 0]. Then, the intermediate sample label information not covering the first label information is determined as negative sample label information, for example, since the first element of the intermediate sample label information [1, 1, 0, 0, 0, 0, 0] overlaps with the first element of the first label information [1, 0, 0, 0, 0, 0, 0], that is, the intermediate sample label information covers the first label information, so the intermediate sample label information [1, 1, 0, 0, 0, 0, 0] cannot be used as negative sample label information. Therefore, the intermediate sample label information [0, 0, 0, 1, 0, 1, 0] can be used as negative sample label information, and then from the plurality of training samples, the training sample corresponding to the negative sample label information is determined as the negative sample corresponding to the benchmark sample.

[0142] The present application determines the positive sample and the negative sample corresponding to the benchmark sample by using the multi-level hierarchical category label, which compensates for the error caused by the "sample source" sampling idea that the same class samples in the training samples are regarded as negative samples instead of positive samples, accurately determines the positive sample and the negative sample corresponding to the benchmark sample, improves the sampling accuracy of the positive sample and the negative sample, and helps the text encoder to more accurately learn the similarities and differences between different samples to generate more accurate feature representation of the text information.

[0143] S130: input the benchmark text information corresponding to the benchmark sample, the positive sample text information corresponding to the positive sample, and the negative sample text information corresponding to the negative sample into the text encoder for text encoding, to obtain the corresponding benchmark sample representation, positive sample representation and negative sample representation.

[0144] The text encoder refers to a neural network that converts text data into a vector representation. This vector representation can be used in text classification tasks. The text encoder can map words, phrases, or entire sentences in the text to points in a representation space, so that computers can better understand and process text data. The text encoder in the embodiments of the present application can include word embedding networks such as Word2Vec, GloVe, and BERT. These encoders can capture text semantic and syntactic information so that the classification model can perform classification tasks based on text semantic and syntactic information.

[0145] As an implementation, the reference text information corresponding to the reference sample, the positive sample text information corresponding to the positive sample, and the negative sample text information corresponding to the negative sample are obtained. Further, the reference text information is input into the text encoder for text encoding to obtain the reference sample representation, the positive sample text information is input into the text encoder for text encoding to obtain the positive sample representation, and the negative sample text information is input into the text encoder for text encoding to obtain the negative sample representation.

[0146] S140: Perform first similarity calculation based on the reference sample representation and the positive sample representation to obtain the first comparison similarity, and perform second similarity calculation based on the reference sample representation and the negative sample representation to obtain the second comparison similarity.

[0147] As an implementation, the similarity calculation refers to calculating the similarity between two sample representations. Specifically, the similarity can at least include cosine similarity (Cosine Similarity): calculating the cosine value of the included angle between two vectors, the larger the value, the more similar the two vectors; Euclidean distance (Euclidean Distance): calculating the distance between two points in Euclidean space, the smaller the distance, the greater the similarity; Pearson correlation coefficient (Pearson Correlation Coefficient): calculating the degree of linear correlation between two vectors, the closer the value to 1, the more relevant the two vectors, and the closer the value to -1, the less relevant the two vectors. The method of similarity calculation is not limited here.

[0148] It should be noted that the first similarity calculation and the second similarity calculation need to use the same type of similarity calculation method. Exemplarily, cosine similarity calculation can be performed on the reference sample representation and the positive sample representation to obtain the cosine similarity between the reference sample representation and the positive sample representation, that is, the first comparison similarity. Cosine similarity calculation is performed based on the reference sample representation and the negative sample representation to obtain the cosine similarity between the reference sample representation and the negative sample representation, that is, the second comparison similarity.

[0149] S150: Construct a target loss function, and train the text encoder through the target loss function to obtain a trained text encoder.

[0150] In deep learning, the role of the loss function is to measure the gap or error between the model's prediction and the actual target, guiding the update of the model parameters to minimize the difference between the predicted value and the true value. Through the optimization of the loss function, the model can continuously learn and improve its performance, making its prediction results more accurate. In the process of training the text encoder, the optimization goal is to minimize the similarity measure of the positive pairs and maximize the similarity measure of the negative pairs, so that the distance between the positive samples is as small as possible, and the distance between the negative samples is as large as possible. To this end, a target loss function can be constructed based on the similarity measure of the positive pairs and the similarity measure of the negative pairs.

[0151] Considering the special nature of the labels of the present application, i.e., multi-level classification labels, the classification labels of each training sample at each level need to be compared to determine whether they have the same label at the corresponding level. To this end, the multi-classification task can be converted into multiple binary classification tasks with the level classification label as the classification standard, so that the target loss function can be constructed using the form of the loss function of the binary classification task: binary cross-entropy loss function. The binary cross-entropy loss is a commonly used loss function for binary classification in the field of deep learning / machine learning, which is used to measure the difference between the model's prediction and the true label. The reason why the present application uses binary cross-entropy loss is that the gradient calculation of the binary cross-entropy loss function is relatively simple, which is convenient for the optimization algorithm of the present application to solve, and the binary cross-entropy loss function is suitable for measuring the difference of the probability distribution in the present application, especially when the probability distribution is close to 0 or 1, which can avoid the problem of gradient vanishing or gradient explosion.

[0152] As an implementation, probability prediction can be performed according to the first comparison similarity to obtain a positive example probability distribution corresponding to the positive sample, and probability prediction can be performed according to the second comparison similarity to obtain a negative example probability distribution corresponding to the negative sample, and then the positive example probability distribution and the negative example probability distribution are used for accumulation calculation to construct the target loss function. The positive example probability distribution represents the probability of the positive sample occurring given the reference sample. The negative example probability distribution represents the probability of the negative sample occurring given the reference sample. The calculation formula of the target loss function is as follows:

[0153]

[0154] where Loss represents the target loss function, ai represents the i-th reference sample, p(a i ,x + ) represents the positive example probability distribution, x + represents the positive sample, X + represents the positive sample set, p(a i ,x - ) represents the negative example probability distribution, x - represents the negative sample, X - represents the negative sample set, p(a i ,x + ) and p(a i ,x - ) ∈ (0, 1).

[0155] Optionally, the positive example probability distribution the negative example probability distribution That is:

[0156]

[0157] wherein sim(a i ,x + ) represents the first contrast similarity, sim(a i ,x - ) represents the second contrast similarity, in the training process, the text encoder is required to make the distance between the reference sample and the positive sample in the representation space as small as possible, and the distance between the reference sample and the negative sample as large as possible, that is, as the first contrast similarity becomes smaller, the second contrast similarity becomes larger, so that the function value of the target loss function becomes smaller, so the target loss function is a decreasing function of the first contrast similarity and an increasing function of the second contrast similarity. In this way, the greater the first contrast similarity, the smaller the loss value of the target loss function, and the smaller the second contrast similarity, the smaller the loss value of the target loss function, by minimizing the target loss function, the distance between the same sample pairs can be reduced or their similarity increased, and the distance between different sample pairs can be increased or their similarity reduced, and further the text encoder can learn a better representation of text information.

[0158] Further, the text encoder is iteratively trained by the target loss function, and whether the target loss function satisfies a training end condition is determined at each iteration. The training end condition can be that the total loss value of the target loss function is less than a preset value, the total loss value of the target loss function no longer changes, or the number of training iterations reaches a preset number, etc. Optionally, an optimizer can be used to optimize the target loss, wherein the optimizer is an algorithm for updating model parameters to minimize the loss function, and is used to find the local minimum or global minimum of the loss function by iteratively adjusting the model parameters. Common optimizers include Stochastic Gradient Descent (SGD), Adam, Adagrad, RMSprop, etc. Each optimizer has its specific update rule and hyperparameters, such as learning rate, etc. Specifically, the learning rate and the training epoch can be set based on experimental experience.

[0159] When the target loss function does not satisfy the training end condition, the parameters of the text encoder are adjusted until the target loss function satisfies the training end condition. When the target loss function satisfies the training end condition, the training of the text encoder is stopped, and the trained text encoder is obtained.

[0160] The embodiment can obtain a classification training set including at least a plurality of training samples. A reference sample is determined from the plurality of training samples of the classification training set. The reference sample is compared with sample hierarchical classification labels of the plurality of training samples according to a reference hierarchical classification label of the reference sample, and positive samples belonging to the same category as the reference sample and negative samples not belonging to the same category as the reference sample are determined for the reference sample. In this way, the positive samples and negative samples are determined for the reference sample based on the more stringent classification standard of sample categories, thereby improving the sampling accuracy of the positive and negative samples. Further, the reference text information corresponding to the reference sample, the positive sample text information corresponding to the positive sample, and the negative sample text information corresponding to the negative sample are respectively input into the text encoder for text encoding, and corresponding reference sample representations, positive sample representations, and negative sample representations are obtained.

[0161] Further, a first similarity calculation is performed based on the reference sample representation and the positive sample representation to obtain a first contrast similarity, and a second similarity calculation is performed based on the reference sample representation and the negative sample representation to obtain a second contrast similarity; a target loss function is constructed, wherein the target loss function is a decreasing function of the first contrast similarity and an increasing function of the second contrast similarity, and the text encoder is trained through the target loss function to obtain a trained text encoder. In this way, the text information of the reference sample and the text information of the positive and negative samples of the reference sample obtained by accurate sampling are respectively encoded by the text encoder, which can more accurately capture the semantic information in the text, improve the accuracy of the text encoder in learning the similarity between the reference sample and its positive sample, and the accuracy of the difference between the reference sample and its negative sample, thereby improving the text encoding performance of the text encoder and generating high-quality feature representations.

[0162] According to an embodiment of the present application, a training method of a classification model is provided. The training of the classification model can be performed after step S150 of the above-mentioned embodiment, and the training method of the classification model can be performed by a computer device (the computer device can be a server, an edge computing device, or other terminal device with certain processing capability), which has at least the functions of storage, calculation, and communication. Figure 12 A flowchart of the training method of the classification model provided by the present embodiment is shown. Figure 13 A flowchart of the training of a classification model is shown. In the following, the training of the classification model will be described in detail with reference to the flowchart. Figure 12 With reference to the flowchart, the training of the classification model includes the following steps: Figure 13 The specific steps of the training method of the classification model will be described as follows:

[0163] It is considered that the quality of the classification data representation required by the multi-level classification task is higher than that of the single-label classification of the prior art, because the multi-level classification involves more classes and sub-classes, and as the classification level increases, the relationship between the labels becomes more complex, so the classification data representation is more complex in multi-level classification. Therefore, the present application proposes to jointly train the classification by using a text encoder trained based on multi-level labels and a classification model. On the one hand, the text encoder trained based on the positive and negative samples determined by the multi-level labels can accurately encode the text information of the object to be classified, obtaining more accurate text representation for classification. On the other hand, the joint training of the two makes the text encoder better adapt to the specific multi-level classification task.

[0164] As Figure 13The training process of the classification model is shown. First, the text encoder is trained based on the classification training set, the text encoder is a pre-trained text encoder, for example, Word2Vec, GloVe and BERT, and then the trained text encoder is obtained. Further, the text encoder and the classification model are trained based on the classification training set, so as to update the parameters of the text encoder and the classification model, and obtain the trained target text encoder and the trained classification model. The specific training steps include:

[0165] S210: The computer device obtains a classification training set.

[0166] As an implementation manner, the computer device can obtain the classification training set through a data calling interface from a local database or a cloud database storing the classification training set. Alternatively, the computer device can also obtain original data from the Internet to generate training samples, for example, the computer device collects different application software and related text information from the Internet, including targeted collection through specific websites in the field where the application software is published, social media platforms and the like, or keyword retrieval associated with the application software publication by using a search engine to obtain the application software and related text information. Further, the computer device preprocesses the collected text information, and the text information can meet the standard for training, and labels the application software in multiple levels, so as to build the classification training set.

[0167] S220: The computer device inputs the text description information of each training sample in the classification training set into the trained text encoder for text encoding, to obtain the text description representation corresponding to each training sample.

[0168] The training process of the trained text encoder can refer to the content of steps 110 to 150 of the above embodiment, which will not be described here. The training sample is associated with corresponding text description information and sample hierarchical classification label. The classification training set S={(x1, y1), (x2, y2), …, (x n , y n ),}. Wherein, n is the number of training samples, n>0&n∈N * . The i-th training sample corresponds to text description information x i and sample hierarchical classification label y i .

[0169] Exemplarily, the computer device can input the text description information of each training sample in the classification training set S into the text encoder for text encoding, to obtain the text description representation corresponding to each training sample, and the calculation formula can be as follows:

[0170] e i =F(x i )

[0171] Where F(·) represents a text encoder, for example, F(x i )=Hx i +b, where H,b are the weight parameters of the text encoder. i Let $e_i$ represent the text description representation corresponding to the $i$-th training sample. Thus, we obtain the text description representation ${e_1, e_2, ..., e_i$ for each training sample. n}

[0172] S230: The computer device classifies each text description representation into a sample using a classification model, and obtains the predicted classification result corresponding to each training sample.

[0173] For example, a computer device can represent the text description of each training sample as {e1, e2, ..., e...} n The input is fed into a classification model, which classifies each text description representation to obtain the predicted classification result for each training sample. The calculation formula can be as follows:

[0174]

[0175] Where C(·) represents the classification model. This represents the predicted classification result for the i-th training sample. Thus, the predicted classification result for each training sample is obtained.

[0176] S240: The computer device determines the classification loss function based on the sample hierarchical classification label and the predicted classification result corresponding to each training sample, and trains the trained text encoder and classification model according to the classification loss function to obtain the trained target text encoder and the trained classification model.

[0177] Specifically, the computer device can obtain the sample-level classification label corresponding to each training sample, and calculate the difference between the predicted classification result and the sample-level classification label corresponding to each training sample to construct a classification loss function. Furthermore, the trained text encoder and classification model are iteratively trained according to the classification loss function until the training termination condition is met, thus obtaining the trained target text encoder and the trained classification model.

[0178] In the embodiments of the application, in order to verify the performance of the text encoder training method and the classification model training method, such as Figure 14The test flowchart shown can be used for performance testing by using the classification test set. The performance testing includes performance testing of the text encoder and testing of the classification task. The test method is a control test, that is, performance testing is performed on the same classification test set corresponding to the prior art and the method of the application respectively:

[0179] (1) Performance testing of the text encoder

[0180] Test metric: cosine similarity

[0181] Test object: open source text encoder; text encoder trained according to the application

[0182] Test implementation: six application software {APP1, APP2, APP3, APP4, APP5, APP6} are selected from the classification training set. APP1 and APP2 are application software of the video playing category. APP3 and APP4 are application software of the game entertainment category, and APP5 and APP6 are application software of the financial application category. APP1 and APP2 have the same secondary label <video playing-online video>, APP5 and APP6 have the same secondary label <financial application-consumption>, and APP3 and APP4 have different secondary labels due to the difference in play and theme.

[0183] The above six application software is encoded by using the open source text encoder and the text encoder trained according to the application respectively, and the cosine similarity is calculated based on the encoded text representation of each application software. The heat map is visualized according to the calculated cosine similarity. In this way, the encoding ability under different text encoding modes for the same input can be intuitively displayed through the heat map, which helps to understand the performance of the text encoder. For example, the cosine similarity calculated according to the encoded text representation can be presented on the heat map in terms of color depth or numerical size, so as to quickly understand the ability of different text encoders to represent text information. For example, Figure 15 is a cosine similarity heat map corresponding to the open source text encoder, Figure 16 is another cosine similarity heat map corresponding to the text encoder trained according to the application.

[0184] Comparative observation Figure 15 and Figure 16 It can be seen that the cosine similarity of software of the same category is obviously higher than that between different categories. This is because the text encoder trained according to the application can generate higher quality text representation than the open source text encoder, and the result of APP classification based on the high-quality text representation is more accurate. Therefore, the cosine similarity of software of the same category is more obvious, and Figure 16 The accuracy of the cosine similarity in Figure 15The accuracy of the cosine similarity in APP3 and APP4, which belong to the game entertainment category, is at a medium level, which is not as high as the cosine similarity between the other two pairs of application software. This is because the application combines the business scenario to perform hierarchical classification. Although APP3 and APP4 belong to the game entertainment category, their secondary labels are different, so the final cosine similarity can still reflect the difference. This also shows that the text encoder and classification model of the application can perform more detailed division and classification when performing a multi-level classification task, so that the final classification category is more specific and more accurate.

[0185] (2) Classification task test

[0186] Metrics: Micro-Precision, Micro-Recall, and Micro-F1 score, where Micro-F1 score = 2 x Micro-Precision x Micro-Recall / (Micro-Precision + Micro-Recall).

[0187] Test object: open source text encoder (prior art) + classification model; text encoder trained by the application + classification model.

[0188] Test implementation: On the classification test set, the open source text encoding + classification model and the text encoder trained by the application + classification model are used to perform the classification task of the test sample, respectively. Based on the classification results, each test index is calculated, and the test results are shown in the following table:

[0189]

[0190] From the above table, it can be observed that the test indicators of the classification task of the application are higher than those of the prior art classification task. In addition, classification tasks can be performed based on different levels of labels to obtain F1 scores. The test results show that the first-level classification F1 is improved by 0.7%, the second-level classification F1 is improved by 7.3%, and the third-level classification F1 is improved by 8.2%.

[0191] The embodiment can obtain a classification training set, and then input text description information of each training sample in the classification training set to a text encoder for text encoding to obtain a text description representation corresponding to each training sample. A classification model is used to classify each text description representation to obtain a predicted classification result corresponding to each training sample. Then, a classification loss function is determined based on the sample hierarchical classification label and the predicted classification result corresponding to each training sample, and the trained text encoder and the classification model are trained according to the classification loss function to obtain a trained target text encoder and a trained classification model. Since the text encoder can generate high-quality text description representations for training samples with multiple hierarchical labels, training the classification model based on the high-quality text description representations for the multi-level classification task can enable the classification model to capture more accurate supervision information, thereby effectively improving the multi-level classification performance of the classification model and obtaining accurate classification results.

[0192] Referring to Figure 17 which shows a structural block diagram of a training device 300 of a text encoder provided by an embodiment of the present application. The device 300 can include a first obtaining module 310, a sample determining module 320, a first encoding module 330, a similarity calculating module 340, and a first training module 350.

[0193] The first obtaining module 310 is configured to obtain a classification training set, wherein the classification training set includes at least a plurality of training samples, and determine a reference sample from the plurality of training samples in the classification training set.

[0194] The sample determining module 320 is configured to compare a reference hierarchical classification label of the reference sample with sample hierarchical classification labels of the plurality of training samples, and determine, according to a comparison result, a positive sample and a negative sample corresponding to the reference sample from the plurality of training samples.

[0195] The first encoding module 330 is configured to input reference text information corresponding to the reference sample, positive sample text information corresponding to the positive sample, and negative sample text information corresponding to the negative sample into a text encoder for text encoding to obtain a reference sample representation, a positive sample representation, and a negative sample representation, respectively.

[0196] The similarity calculating module 340 is configured to perform first similarity calculation based on the reference sample representation and the positive sample representation to obtain a first comparison similarity, and perform second similarity calculation based on the reference sample representation and the negative sample representation to obtain a second comparison similarity.

[0197] The first training module 350 is configured to construct a target loss function which is a decreasing function of the first contrast similarity and an increasing function of the second contrast similarity, and train the text encoder by using the target loss function to obtain a trained text encoder.

[0198] In some embodiments, the sample determination module 320 can include a first label determination unit, a first sample acquisition unit, a second label determination unit, a second sample acquisition unit, and a sample comparison unit.

[0199] The first label determination unit is configured to determine a reference level classification label of the reference sample as a positive sample classification label corresponding to the reference sample.

[0200] The first sample acquisition unit is configured to acquire a first mapping relationship corresponding to the positive sample classification label, and determine a first sample corresponding to the positive sample classification label based on the first mapping relationship.

[0201] The second label determination unit is configured to determine a negative sample classification label corresponding to the reference sample based on a label structure tree corresponding to the classification training set.

[0202] The second sample acquisition unit is configured to acquire a second mapping relationship corresponding to the negative sample classification label, and determine a second sample corresponding to the negative sample classification label based on the second mapping relationship.

[0203] The sample comparison unit is configured to compare the plurality of training samples with the first sample and the second sample respectively, and determine a positive sample and a negative sample corresponding to the reference sample from the plurality of training samples according to comparison results.

[0204] In some embodiments, the second label determination unit can be specifically configured to: find a low-level classification label corresponding to the reference level classification label in the label structure tree; acquire other level classification labels in the label structure tree except the reference level classification label and the low-level classification label; and determine the other level classification labels as the negative sample classification label corresponding to the reference sample.

[0205] In some embodiments, the comparison results include a first comparison result and a second comparison result, and the sample determination unit can be specifically configured to: obtain the first comparison result according to the comparison of the plurality of training samples with the first sample; based on the first comparison result, determine a training sample matching the first sample from the plurality of training samples as the positive sample corresponding to the reference application; obtain the second comparison result according to the comparison of the plurality of training samples with the second sample; and based on the second comparison result, determine a training sample matching the second sample from the plurality of training samples as the negative sample corresponding to the reference application.

[0206] In some embodiments, the sample determination module 320 can further include a label information acquisition unit, a positive sample determination unit, and a negative sample determination unit.

[0207] The label information acquisition unit is configured to acquire first label information of a reference level classification label of the reference sample and second label information of a sample level classification label of the plurality of training samples.

[0208] The positive sample determination unit is configured to determine second label information identical to the first label information as positive sample label information, and determine a training sample corresponding to the positive sample label information as a positive sample corresponding to the reference sample.

[0209] The negative sample determination unit is configured to determine second label information different from the first label information as negative sample label information, and determine a training sample corresponding to the negative sample label information as a negative sample corresponding to the reference sample.

[0210] In some embodiments, the negative sample determination unit can be specifically configured to determine second label information different from the first label information as intermediate sample label information, determine intermediate sample label information that does not cover the first label information as negative sample label information, and determine a training sample corresponding to the negative sample label information as the negative sample corresponding to the reference sample from the plurality of training samples.

[0211] In some embodiments, the first training module 350 can include a first probability prediction unit, a second probability prediction unit, and a loss construction unit.

[0212] The first probability prediction unit is configured to perform probability prediction according to the first contrast similarity to obtain a positive example probability distribution corresponding to the positive sample.

[0213] The second probability prediction unit is configured to perform probability prediction according to the second contrast similarity to obtain a negative example probability distribution corresponding to the negative sample.

[0214] The loss construction unit is configured to perform accumulation calculation based on the positive example probability distribution and the negative example probability distribution to construct a target loss function.

[0215] In some embodiments, the first training module 350 can further include a training judgment unit, a first training unit, and a second training unit.

[0216] The training judgment unit is configured to perform iterative training on the text encoder through the target loss function, and judge whether the target loss function satisfies a training end condition at each iteration.

[0217] The first training unit is configured to adjust the parameters of the text encoder when the target loss function does not satisfy the training end condition, and return to perform iterative training of the text encoder through the target loss function until the target loss function satisfies the training end condition.

[0218] The second training unit is configured to stop training the text encoder when the target loss function satisfies the training end condition, and obtain the trained text encoder.

[0219] In some embodiments, the first encoding module 330 can include a text acquisition unit, a first encoding unit, a second encoding unit, and a third encoding unit.

[0220] The text acquisition unit is configured to acquire reference text information corresponding to the reference sample, positive sample text information corresponding to the positive sample, and negative sample text information corresponding to the negative sample.

[0221] The first encoding unit is configured to input the reference text information into a text encoder for text encoding to obtain a reference sample representation.

[0222] The second encoding unit is configured to input the positive sample text information into the text encoder for text encoding to obtain a positive sample representation.

[0223] The third encoding unit is configured to input the negative sample text information into the text encoder for text encoding to obtain a negative sample representation.

[0224] The embodiment can obtain a classification training set including at least a plurality of training samples, determine a reference sample from the plurality of training samples of the classification training set, and compare a reference level classification label of the reference sample with sample level classification labels of the plurality of training samples, to determine, for the reference sample, positive samples belonging to the same category as the reference sample and negative samples not belonging to the same category as the reference sample. In this way, the positive samples and the negative samples are determined for the reference sample based on the more stringent classification standard of sample categories, thereby improving the sampling accuracy of the positive and negative samples. Further, the reference text information corresponding to the reference sample, the positive sample text information corresponding to the positive sample, and the negative sample text information corresponding to the negative sample are respectively input to a text encoder for text encoding, to obtain corresponding reference sample representations, positive sample representations, and negative sample representations. Further, a first similarity calculation is performed based on the reference sample representation and the positive sample representation to obtain a first comparison similarity, and a second similarity calculation is performed based on the reference sample representation and the negative sample representation to obtain a second comparison similarity; a target loss function is constructed, which is a decreasing function of the first comparison similarity and an increasing function of the second comparison similarity, and the text encoder is trained through the target loss function to obtain a trained text encoder. In this way, the text information of the reference sample and the text information of the positive and negative samples of the reference sample accurately sampled are respectively encoded by the text encoder, which can more accurately capture the semantic information in the text, improve the accuracy of the text encoder in learning the similarity between the reference sample and its positive sample, and the accuracy of the text encoder in learning the difference between the reference sample and its negative sample, thereby improving the text encoding performance of the text encoder and generating high-quality feature representations.

[0225] Referring to Figure 18 which shows a structural block diagram of an application test device 400 provided by an embodiment of the present application. The device 400 can include:

[0226] A second obtaining module 410 is configured to obtain a classification training set.

[0227] A second encoding module 420 is configured to input text description information of each training sample in the classification training set into a trained text encoder for text encoding, to obtain a text description representation corresponding to each training sample; wherein the trained text encoder is obtained according to the training method of the text encoder.

[0228] A sample classification module 430 is configured to perform sample classification on each text description representation through a classification model, to obtain a predicted classification result corresponding to each training sample.

[0229] The second training module 440 is configured to determine a classification loss function based on the sample hierarchical classification label corresponding to each training sample and the predicted classification result, and train the trained text encoder and the classification model according to the classification loss function, to obtain a trained target text encoder and a trained classification model.

[0230] In some embodiments, the second training module 440 can be specifically configured to: obtain the sample hierarchical classification label corresponding to each training sample; perform difference calculation based on the predicted classification result corresponding to each training sample and the sample hierarchical classification label corresponding to each training sample, to construct a classification loss function; and iteratively train the trained text encoder and the classification model according to the classification loss function until a training end condition is met, to obtain a trained target text encoder and a trained classification model.

[0231] The embodiments can obtain a classification training set, and then input the text description information of each training sample in the classification training set into a text encoder for text encoding, to obtain a text description representation corresponding to each training sample. The classification model is used to perform sample classification on each text description representation, to obtain a predicted classification result corresponding to each training sample. Then, a classification loss function is determined based on the sample hierarchical classification label corresponding to each training sample and the predicted classification result, and the trained text encoder and the classification model are trained according to the classification loss function, to obtain a trained target text encoder and a trained classification model. Since the text encoder can generate high-quality text description representations for training samples with multi-level labels, training the classification model based on the high-quality text description representations for the multi-level classification task can enable the classification model to capture more accurate supervision information, thereby effectively improving the multi-level classification performance of the classification model and obtaining accurate classification results.

[0232] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the above-described apparatuses and modules can refer to the corresponding process in the foregoing method embodiments, which will not be described herein.

[0233] In several embodiments provided in the present application, the coupling between the modules can be electrical, mechanical or other forms of coupling.

[0234] In addition, each functional module in each embodiment of the present application can be integrated into one processing module, or each module can exist physically independently, or two or more modules can be integrated into one module. The integrated module can be realized in the form of hardware or in the form of a software functional module.

[0235] As Figure 19As shown, the embodiments of the present application further provide a computer device 500, which comprises a processor 510, a memory 520, a power supply 530 and an input unit 540. The memory 520 stores a computer program, and the computer program can implement various method steps provided by the above embodiments when called by the processor 510. Those skilled in the art can understand that the structure of the computer device shown in the figure does not constitute a limitation on the computer device, and can include more or fewer components than those shown in the figure, or combine certain components, or different component arrangements. Among them:

[0236] The processor 510 can include one or more processing cores. The processor 510 connects various parts within the battery management system through various interfaces and lines, calls data stored in the memory 520 by running or executing instructions, programs, instruction sets or program sets stored in the memory 520, executes various functions of the battery management system and processes data, and executes various functions of the computer device and processes data, thereby overall controlling the computer device. Optionally, the processor 510 can be implemented in at least one of the hardware forms of a digital signal processor (DSP), a field-programmable gate array (FPGA), and a programmable logic array (PLA). The processor 510 can integrate a combination of one or several of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. Among them, the CPU mainly processes operating systems, user interfaces, and application programs; the GPU is responsible for rendering and drawing display content; and the modem is used for processing wireless communication. It can be understood that the above-mentioned modem can also not be integrated into the processor 510, but be implemented by a separate communication chip.

[0237] The memory 520 may include random access memory (RAM) or read-only memory (ROM). The memory 520 can be used to store instructions, programs, instruction sets, or program assemblies. The memory 520 may include a program storage area and a data storage area. The program storage area may store instructions for implementing an operating system, instructions for implementing at least one function (such as touch functionality, sound playback functionality, image playback functionality, etc.), and instructions for implementing the various method embodiments described above. The data storage area may also store data created during the use of the computer device (such as phonebook and audio / video data). Accordingly, the memory 520 may also include a memory controller to provide the processor 510 with access to the memory 520.

[0238] The power supply 530 can be logically connected to the processor 510 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 530 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.

[0239] The input unit 540 can be used to receive input digital or character information, and to generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.

[0240] Although not shown, the computer device 500 may also include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 510 in the computer device loads the executable files corresponding to the processes of one or more computer programs into the memory 520 according to the following instructions, and the processor 510 runs the data stored in the memory 520, such as telephone books and audio and video data, thereby implementing the various method steps provided in the foregoing embodiments.

[0241] like Figure 20 As shown, this application embodiment also provides a computer-readable storage medium 600, which stores a computer program 610. The computer program 610 can be invoked by a processor to execute various method steps provided in this application embodiment.

[0242] The computer readable storage medium can be an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, or any suitable combination of the foregoing. More specific examples of a computer readable storage medium are a portable magnetic disc, a hard disk, a ROM, an EEPROM, a cassette tape or the like. The computer readable storage medium can optionally be non-transitory. The computer readable storage medium 600 has a storage space for computer programs that perform any of the method steps of the above embodiments. These computer programs can be read from or written to one or more computer program products. The computer programs can be compressed in a suitable form.

[0243] According to an aspect of the present application, there is provided a computer program product comprising a computer program stored in a computer readable storage medium. A processor of a computer device reads the computer program from the computer readable storage medium, and the processor executes the computer program to cause the computer device to perform various method steps provided by the above embodiments.

[0244] The terms "comprise", "comprising", "include", "including", "contain", "containing", "have" and "having" and any variations thereof in the specification and in the claims are intended to cover both the express and implicit meaning of the terms, i.e. both the meaning of "consist of" and the meaning of "consist essentially of" and the meaning of "consist of". The terms "module" or "unit" refer to a computer program or a part of a computer program having a predetermined function and working together with other relevant parts to achieve a predetermined object, and can be implemented wholly or partially by using software, hardware (such as a processing circuit or a memory) or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. Furthermore, each module or unit can be a part of an integral module or unit that contains the function of the module or unit.

[0245] It should be understood that, in the present application, "at least one" means one or more, and "multiple" means two or more. "And / or" is used to describe the relationship between associated objects, which means that there can be three relationships, for example, "A and / or B" can mean that there are three cases of A only, B only, and A and B at the same time, where A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects. "At least one of the following" or similar expressions means any combination of these items, including any combination of single or multiple items. For example, at least one of a, b or c can mean a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0246] It should be understood that in the description of the embodiments of the present application, the meaning of multiple (or multiple items) is two or more, greater than, less than, more than, etc. is not included in the number, above, below, etc. is included in the number.

[0247] In several embodiments provided by the present application, it should be understood that the disclosed system, device and method can be implemented in other ways. For example, the device embodiments described above are only schematic, for example, the division of units is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the displayed or discussed mutual ones can be indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0248] The units described as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, that is, they can be located in one place, or they can be distributed on multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the present embodiment. In addition, each functional unit in each embodiment of the present application can be integrated into a processing unit, or each unit can be physically present alone, or two or more units can be integrated into one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.

[0249] The integrated unit, if implemented in the form of a software function unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or say the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the embodiments of the present application. It should also be understood that the various embodiments provided by the present application can be combined in any way to achieve different technical effects.

[0250] The above is only a preferred embodiment of the present application, and does not limit the present application in any form. Although the present application has been disclosed as above with a preferred embodiment, it is not intended to limit the present application. Any person skilled in the art can make slight changes or modifications to the above disclosed technical content without departing from the scope of the technical solutions of the present application, and any equivalent embodiments with equivalent changes are equivalent. Any modification, change, modification, and modification of the above embodiments according to the technical essence of the present application are still within the scope of the technical solutions of the present application.

Claims

1. A method for training a text encoder, characterized in that, The method comprises: obtaining a classification training set, the classification training set comprising at least a plurality of training samples, and determining a reference sample from the plurality of training samples of the classification training set; comparing a reference level classification label of the reference sample with sample level classification labels of the plurality of training samples, and determining, according to a comparison result, a positive sample and a negative sample corresponding to the reference sample from the plurality of training samples; inputting reference text information corresponding to the reference sample, positive sample text information corresponding to the positive sample, and negative sample text information corresponding to the negative sample into a text encoder respectively for text encoding to obtain reference sample representation, positive sample representation, and negative sample representation; performing first similarity calculation based on the reference sample representation and the positive sample representation to obtain a first comparison similarity, and performing second similarity calculation based on the reference sample representation and the negative sample representation to obtain a second comparison similarity; constructing a target loss function, the target loss function being a decreasing function of the first comparison similarity and an increasing function of the second comparison similarity, and training the text encoder through the target loss function to obtain a trained text encoder.

2. The method of claim 1, wherein, The comparison of the reference level classification label of the reference sample with the sample level classification labels of the plurality of training samples, and the determination, according to a comparison result, of a positive sample and a negative sample corresponding to the reference sample from the plurality of training samples, comprise: taking the reference level classification label of the reference sample as a positive sample classification label corresponding to the reference sample; obtaining a first mapping relationship corresponding to the positive sample classification label, and determining a first sample corresponding to the positive sample classification label based on the first mapping relationship; determining a negative sample classification label corresponding to the reference sample based on a label structure tree corresponding to the classification training set; obtaining a second mapping relationship corresponding to the negative sample classification label, and determining a second sample corresponding to the negative sample classification label based on the second mapping relationship; comparing the plurality of training samples with the first sample and the second sample respectively, and determining, according to a comparison result, a positive sample and a negative sample corresponding to the reference sample from the plurality of training samples.

3. The method of claim 2, wherein, The determination of a negative sample classification label corresponding to the reference sample based on a label structure tree corresponding to the classification training set, comprises: finding a low level classification label corresponding to the reference level classification label in the label structure tree; obtaining other level classification labels in the label structure tree except the reference level classification label and the low level classification label; taking the other level classification labels as the negative sample classification label corresponding to the reference sample.

4. The method of claim 3, wherein, The comparison result comprises a first comparison result and a second comparison result, and the comparison of the plurality of training samples with the first sample and the second sample respectively, and the determination, according to a comparison result, of a positive sample and a negative sample corresponding to the reference sample from the plurality of training samples, comprise: comparing the plurality of training samples with the first sample respectively to obtain a first comparison result; based on the first comparison result, the training sample matched with the first sample in the plurality of training samples is determined as the corresponding positive sample of the benchmark application; obtaining a second comparison result by comparing the plurality of training samples with the second sample respectively; based on the second comparison result, the training sample matched with the second sample in the plurality of training samples is determined as the corresponding negative sample of the benchmark application.

5. The method of claim 1, wherein, comparing the benchmark level classification label of the benchmark sample and the sample level classification label of the plurality of training samples, and determining the positive sample and the negative sample corresponding to the benchmark sample from the plurality of training samples according to the comparison result, comprising: obtaining the first label information of the benchmark level classification label of the benchmark sample and the second label information of the sample level classification label of the plurality of training samples; determining the second label information same as the first label information as positive sample label information, and determining the training sample corresponding to the positive sample label information as the positive sample corresponding to the benchmark sample; determining the second label information different from the first label information as negative sample label information, and determining the training sample corresponding to the negative sample label information as the negative sample corresponding to the benchmark sample.

6. The method of claim 5, wherein, determining the second label information different from the first label information as negative sample label information, and determining the training sample corresponding to the negative sample label information as the negative sample corresponding to the benchmark sample, comprising: determining the second label information different from the first label information as intermediate sample label information; determining the intermediate sample label information not covering the first label information as negative sample label information; from the plurality of training samples, the training sample corresponding to the negative sample label information is determined as the negative sample corresponding to the benchmark sample.

7. The method according to any one of claims 1 to 6, characterized in that, the target loss function is constructed, comprising: obtaining the positive example probability distribution corresponding to the positive sample according to the first comparison similarity for probability prediction; obtaining the negative example probability distribution corresponding to the negative sample according to the second comparison similarity for probability prediction; based on the positive example probability distribution and the negative example probability distribution, the target loss function is constructed by accumulation calculation.

8. The method of claim 7, wherein, the text encoder is trained through the target loss function to obtain the trained text encoder, comprising: iteratively training the text encoder through the target loss function, and determining whether the target loss function meets the training end condition at each iteration; when the target loss function does not meet the training end condition, the parameters of the text encoder are adjusted, and the iteration training of the text encoder through the target loss function is returned to be executed until the target loss function meets the training end condition; when the target loss function meets the training end condition, the training of the text encoder is stopped, and the trained text encoder is obtained.

9. The method of claim 8, wherein, The reference sample corresponding reference text information, the positive sample corresponding positive sample text information and the negative sample corresponding negative sample text information are respectively input into a text encoder for text encoding to obtain corresponding reference sample representation, positive sample representation and negative sample representation, comprising: Obtaining the reference sample corresponding reference text information, the positive sample corresponding positive sample text information and the negative sample corresponding negative sample text information; The reference text information is input into a text encoder for text encoding to obtain reference sample representation; The positive sample text information is input into a text encoder for text encoding to obtain positive sample representation; The negative sample text information is input into a text encoder for text encoding to obtain negative sample representation. 10.A method for training a classification model, the method comprising: The method comprises: Obtaining a classification training set; Inputting the text description information of each training sample in the classification training set into the trained text encoder for text encoding to obtain the text description representation corresponding to each training sample; Wherein, the trained text encoder is obtained according to the training method of the text encoder of any one of claims 1 to 9; Classifying each text description representation through a classification model to obtain the predicted classification result corresponding to each training sample; Determine the classification loss function based on the sample level classification label and the predicted classification result corresponding to each training sample, and train the trained text encoder and the classification model according to the classification loss function to obtain the trained target text encoder and the trained classification model.

11. The method of claim 10, wherein, The classification loss function is determined based on the sample level classification label and the predicted classification result corresponding to each training sample, and the trained text encoder and the classification model are trained according to the classification loss function to obtain the trained target text encoder and the trained classification model, comprising: Obtaining the sample level classification label corresponding to each training sample; Based on the difference calculation of the predicted classification result corresponding to each training sample and the sample level classification label corresponding to each training sample, the classification loss function is constructed; According to the classification loss function, the trained text encoder and the classification model are iteratively trained until the training end condition is met, and the trained target text encoder and the trained classification model are obtained.

12. A training device for a text encoder, characterized in that, The device comprises: A first obtaining module for obtaining a classification training set, the classification training set comprising at least a plurality of training samples, and determining a reference sample from the plurality of training samples in the classification training set; A sample determining module for comparing the reference level classification label of the reference sample with the sample level classification label of the plurality of training samples, and determining the positive sample and the negative sample corresponding to the reference sample from the plurality of training samples according to the comparison result; A first encoding module for inputting the reference text information corresponding to the reference sample, the positive sample text information corresponding to the positive sample and the negative sample text information corresponding to the negative sample into a text encoder for text encoding to obtain the corresponding reference sample representation, positive sample representation and negative sample representation; The similarity calculation module is configured to perform a first similarity calculation based on the reference sample representation and the positive sample representation to obtain a first contrast similarity, and perform a second similarity calculation based on the reference sample representation and the negative sample representation to obtain a second contrast similarity. The first training module is configured to construct a target loss function, the target loss function being a decreasing function of the first contrast similarity and an increasing function of the second contrast similarity, and train the text encoder based on the target loss function to obtain a trained text encoder.

13. An apparatus for training a classification model, the apparatus comprising: The apparatus comprises: The second obtaining module is configured to obtain a classification training set. The second encoding module is configured to input text description information of each training sample in the classification training set into the trained text encoder for text encoding to obtain a text description representation corresponding to each training sample. The trained text encoder is obtained according to the training method of the text encoder in any one of claims 1 to 9. The sample classification module is configured to perform sample classification on each text description representation based on the classification model to obtain a predicted classification result corresponding to each training sample. The second training module is configured to determine a classification loss function based on the sample-level classification label and the predicted classification result corresponding to each training sample, and train the trained text encoder and the classification model based on the classification loss function to obtain a trained target text encoder and a trained classification model.

14. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer readable instructions, and when the computer readable instructions are executed by a processor, the training method of the text encoder in any one of claims 1 to 9 or the training method of the classification model in any one of claims 10 to 11 is implemented.

15. A computer device, comprising: comprises: a memory; a processor, and the memory stores a computer program, and the computer program is executed by the processor to implement the training method of the text encoder in any one of claims 1 to 9 or the training method of the classification model in any one of claims 10 to 11.

16. A computer program product comprising computer programs or instructions, characterized in that, The computer program or instructions are executed by the processor to implement the training method of the text encoder in any one of claims 1 to 9 or the training method of the classification model in any one of claims 10 to 11.