Text-based Model Training Method, Related Devices and Storage Media

Through multiple feature perturbation coding and comparative learning strategies, the feature coding model is optimized, and the problem of insufficient text classification accuracy in the neural network model in object search scenarios is solved, and more accurate text classification is achieved.

CN115130668BActive Publication Date: 2025-07-25TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210356995.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-06
Publication Date
2025-07-25
Estimated Expiration
2042-04-06

AI Technical Summary

Technical Problem

In the prior art, it is difficult for neural network models to accurately predict the correlation between object description text and search text in object search scenarios, resulting in insufficient text classification accuracy.

Method used

By calling the feature encoding model multiple times, the text in the target sample is characterized perturbated, and positive samples with semantic similarity are constructed, and the model parameters of the feature encoding model are optimized using comparative learning strategies to reduce the gap in classification results and improve the accuracy of text classification.

Benefits of technology

It improves the text classification accuracy of neural network models in object search scenarios, ensures that the classification results are more consistent, and reduces the probability gap in the object description text belonging to multiple text categories.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115130668B_ABST
    Figure CN115130668B_ABST
Patent Text Reader

Abstract

An embodiment of the present application discloses a text-based model training method, related devices, and a storage medium. The method includes: obtaining a target sample, where the target sample includes an object search text and an object description text; repeatedly calling a feature encoding model to perform feature perturbation encoding on each text in the target sample to obtain a first feature vector and a second feature vector; then respectively classifying the object description text in the target sample according to the first feature vector and the second feature vector to obtain two classification results; and finally optimizing the model parameters of the feature encoding model according to the two classification results according to a contrastive learning strategy. By adopting the embodiment of the present application, the model performance of a neural network model can be optimized to improve the accuracy of text classification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technologies, and in particular, to a text-based model training method, related devices, and storage media. Background Art

[0002] In an object search scenario, after a user inputs a search text (query), a computer device can determine the relevance between the object description text (doc) of each object in a database and the search text, and thus output and display each object based on the relevance to provide feedback on the search text. Currently, the relevance between the object description text and the search text is usually obtained by a neural network model classifying and predicting the corresponding object description text. It can be seen that the model performance of the neural network model is closely related to the prediction result of the relevance. Based on this, how to obtain a neural network model with better model performance through model training to improve the accuracy of text classification is an urgent problem to be solved currently. Summary of the Invention

[0003] Embodiments of this application provide a text-based model training method, related devices, and storage media, which can optimize the model performance of a neural network model to improve the accuracy of text classification.

[0004] On the one hand, embodiments of this application provide a text-based model training method, and the method includes:

[0005] Obtain a target sample, where the target sample includes an object search text and an object description text; the object described by the object description text is the object output for providing feedback on the object search text;

[0006] Call a feature encoding model to perform first feature perturbation encoding on each text in the target sample to obtain a first feature vector; and call the feature encoding model to perform second feature perturbation encoding on each text in the target sample to obtain a second feature vector;

[0007] Classify the object description text in the target sample respectively according to the first feature vector and the second feature vector to obtain two classification results; any classification result includes: the probability that the object description text in the target sample belongs to at least one text category, and the text category is used to indicate the relevance between the object description text and the object search text;

[0008] Optimize the model parameters of the feature encoding model according to the two classification results according to a contrastive learning strategy.

[0009] On the other hand, an embodiment of the present application provides a text-based model training device, where the text-based model training device includes an acquisition unit and a processing unit, and specifically:

[0010] The acquisition unit is configured to acquire a target sample, where the target sample includes an object search text and an object description text; the object described by the object description text is the object output in response to the object search text;

[0011] The processing unit is configured to call a feature encoding model to perform first feature perturbation encoding on each text in the target sample to obtain a first feature vector; and call the feature encoding model to perform second feature perturbation encoding on each text in the target sample to obtain a second feature vector;

[0012] The processing unit is further configured to perform classification processing on the object description text in the target sample according to the first feature vector and the second feature vector respectively to obtain two classification results; any classification result includes: the probability that the object description text in the target sample belongs to at least one text category, and the text category is used to indicate the correlation between the object description text and the object search text;

[0013] The processing unit is further configured to optimize the model parameters of the feature encoding model according to the two classification results according to a contrast learning strategy.

[0014] On the other hand, an embodiment of the present application provides a computer device, where the computer device includes an input interface and an output interface, and the computer device further includes:

[0015] A processor, adapted to implement one or more computer programs; and,

[0016] A computer storage medium storing one or more computer programs, where the one or more computer programs are adapted to be loaded and executed by the processor to perform the following steps:

[0017] Acquire a target sample, where the target sample includes an object search text and an object description text; the object described by the object description text is the object output in response to the object search text;

[0018] Call a feature encoding model to perform first feature perturbation encoding on each text in the target sample to obtain a first feature vector; and call the feature encoding model to perform second feature perturbation encoding on each text in the target sample to obtain a second feature vector;

[0019] Classify the object description text in the target sample according to the first feature vector and the second feature vector respectively to obtain two classification results; any classification result includes: the probability that the object description text in the target sample belongs to at least one text category, and the text category is used to indicate the correlation between the object description text and the object search text;

[0020] Optimize the model parameters of the feature encoding model according to the two classification results according to the contrast learning strategy.

[0021] On the other hand, an embodiment of the present application provides a computer storage medium, which stores one or more computer programs, and the one or more computer programs are adapted to be loaded and executed by a processor to perform the following steps:

[0022] Obtain a target sample, where the target sample includes an object search text and an object description text; the object described by the object description text is the object output by giving feedback on the object search text;

[0023] Call the feature encoding model to perform first feature perturbation encoding on each text in the target sample to obtain a first feature vector; and call the feature encoding model to perform second feature perturbation encoding on each text in the target sample to obtain a second feature vector;

[0024] Classify the object description text in the target sample according to the first feature vector and the second feature vector respectively to obtain two classification results; any classification result includes: the probability that the object description text in the target sample belongs to at least one text category, and the text category is used to indicate the correlation between the object description text and the object search text;

[0025] Optimize the model parameters of the feature encoding model according to the two classification results according to the contrast learning strategy.

[0026] On the other hand, an embodiment of the present application provides a computer program product or a computer program, which includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the above-mentioned text-based model training method.

[0027] In the embodiments of the present application, by calling the feature encoding model multiple times and respectively performing first feature perturbation encoding and second feature perturbation encoding on each text in the target sample, two positive samples semantically similar to the object search text and the object description text in the target sample can be constructed, and the first feature vector and the second feature vector of the two positive samples can be obtained; and by respectively classifying the object description text in the target sample according to the first feature vector and the second feature vector, two classification results can be obtained, which is equivalent to obtaining the positions of the two semantically similar positive samples in the representation space. Then, through the contrastive learning strategy, in the direction of reducing the distance between the two classification results, the model parameters of the feature encoding model are optimized, so that after the feature encoding model encodes two semantically similar samples, the two feature vectors obtained will be more consistent when respectively used for classification processing, thereby reducing the probability that the classification results simultaneously point to multiple text categories or belong to multiple text categories by about the same, so as to improve the accuracy of text classification. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0029] Figure 1a is a schematic diagram of a contrastive learning idea provided by an embodiment of the present application;

[0030] Figure 1b is a schematic diagram of another contrastive learning idea provided by an embodiment of the present application;

[0031] Figure 2 is a schematic diagram of the architecture of a model training system based on text provided by an embodiment of the present application;

[0032] Figure 3 is a schematic diagram of the flow of a model training method based on text provided by an embodiment of the present application;

[0033] Figure 4 is a schematic diagram of the structure of a bert two-tower model provided by an embodiment of the present application;

[0034] Figure 5 is a schematic diagram of the architecture of a contrastive learning strategy provided by an embodiment of the present application;

[0035] Figure 6 is a schematic diagram of the flow of another model training method based on text provided by an embodiment of the present application;

[0036] Figure 7 It is a schematic diagram of outputting a description text of a target object provided by an embodiment of the present application;

[0037] Figure 8 It is a schematic structural diagram of a text-based model training device provided by an embodiment of the present application;

[0038] Figure 9 It is a schematic structural diagram of a computer device provided by an embodiment of the present application. Specific Embodiments

[0039] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention.

[0040] With the continuous development of Internet technology, artificial intelligence (AI) technology has also been better developed. The so-called artificial intelligence technology refers to the theory, method, technology and application system that uses a digital computer or a machine controlled by a digital computer to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science; it mainly produces a new intelligent machine that can react in a way similar to human intelligence by understanding the essence of intelligence, so that the intelligent machine has multiple functions such as perception, reasoning and decision-making. Correspondingly, AI technology is an interdisciplinary subject, which mainly includes several major directions such as computer vision technology (CV), speech processing technology, natural language processing technology, and machine learning (ML) / deep learning.

[0041] Among them, machine learning is an interdisciplinary subject involving multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of AI and the fundamental way to endow computer devices with intelligence. The so-called machine learning is an interdisciplinary subject involving multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computer devices simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Deep learning, on the other hand, is a technology for machine learning that uses deep neural network systems. Machine learning / deep learning usually includes various technologies such as artificial neural networks, reinforcement learning (RL), supervised learning, unsupervised learning, and contrastive learning. The so-called supervised learning refers to the processing method of training a model using training samples with known categories (with labeled categories), and unsupervised learning refers to the processing method of training a model using training samples with unknown categories (without being labeled).

[0042] Among them, the so-called contrastive learning is a discriminative representation learning method based on the contrastive idea. It mainly compares a sample with samples that are semantically similar to it (i.e., the positive samples of the sample) and samples that are semantically dissimilar to it (i.e., the negative samples of the sample), so that the sample representation corresponding to the sample is closer to the sample representation corresponding to the positive sample in the representation space and farther away from the sample representation corresponding to the negative sample in the representation space. Specifically, please refer to the appendix Figure 1a , which shows a schematic diagram of the contrastive learning idea. The sample representation corresponding to sample A is at the center position in the representation space. The black circles in the representation space refer to the sample representations of samples that are semantically similar to sample A in the representation space, and the white circles in the representation space refer to the sample representations of samples that are semantically dissimilar to sample A in the representation space. After contrastive learning, it can be seen that the sample representations corresponding to samples that are semantically similar to sample A (i.e., the black circles) will cluster in the representation space, and the sample representations corresponding to samples that are semantically dissimilar to sample A (i.e., the white circles) will be far away from the sample representation corresponding to sample A in the representation space. At this time, by simply judging the distances of the sample representations corresponding to each sample in the sample space, it is possible to accurately and clearly distinguish which samples are semantically similar and which samples are semantically dissimilar. In addition, please refer to the appendix Figure 1b, showing a schematic diagram of another contrastive learning idea. The black circles represent samples semantically similar to sample A, the white circles represent samples semantically similar to sample B, and the gray circles represent samples that are not semantically similar to either sample A or sample B. Then, through contrastive learning, only the sample representations corresponding to semantically similar samples are made closer in the representation space. It will be found that the sample representations corresponding to samples semantically similar to sample A cluster together, and the sample representations corresponding to samples semantically similar to sample B cluster together. At the same time, relatively speaking, samples semantically similar to sample B are also samples that are not semantically similar to sample A. After the above contrastive learning, the sample representations corresponding to samples semantically similar to sample B are also relatively far from the sample representation corresponding to sample A in the representation space, thus achieving the effect that the sample representation corresponding to sample A is farther from the sample representations corresponding to samples that are not semantically similar to sample A in the representation space. At this time, it is also possible to accurately and clearly distinguish which are semantically similar samples and which are semantically dissimilar samples by judging the distances of the sample representations corresponding to each sample in the sample space.

[0043] In addition, in text classification, text is generally divided into at least one text category. The text category is used to indicate the correlation between the object description text and the object search text. Exemplarily, five text categories can be determined, and the correlations between the object description texts and the object search texts indicated by different text categories are different. For example, the first text category is used to indicate that the object description text and the object search text are not relevant both in text and semantics; the second text category is used to indicate that the object description text and the object search text are relevant or partially relevant in text but not relevant in semantics; the third text category is used to indicate that the object description text and the object search text are relevant or partially relevant in text and partially relevant in semantics; the fourth text category is used to indicate that the object description text and the object search text are highly relevant both in text and semantics; the fifth text category is used to indicate that the object description text and the object search text are completely relevant both in text and semantics. The text relevance in this embodiment may be that the characters in the object description text are similar or identical to the characters in the object search text. For example, if the object search text is "food, video", the object description text relevant to this object search text may be "food exploration video"; semantic relevance means that the semantics of the object description text are similar or identical to the semantics of the object search text. For example, if the object search text is "solving method of binary equation", the object description texts semantically relevant to this object search semantics may be "solving ideas of high school equations", "how to solve binary equations", etc. Generally speaking, the irrelevance, partial relevance, high relevance, and complete relevance of the object description text and the object search text in text can be accurately determined through character comparison. At the same time, the irrelevance and complete relevance of the object description text and the object search text in semantics can also be classified relatively accurately by comparing semantics. However, the semantic gap indicated between the partial relevance and high relevance of the object description text and the object search text in semantics is relatively vague. Therefore, it is difficult to accurately classify the object description texts belonging to the third text category and the fourth text category whether by manually annotating the text category or by determining the text category through a neural network model. In addition, since classification can often be referred to as grading, the text category can also be called the relevance grade or grade. For example, the first text category to the fifth text category in the above example can be called grade 1, grade 2, grade 3, grade 4, and grade 5 respectively.

[0044] Based on the above-mentioned contrastive learning and text classification, the present application proposes a text-based model training method. During the process of optimizing the feature encoding model, this method will, by repeatedly invoking the feature encoding model, perform multiple feature perturbation encodings on the object search text and object description text in the same target sample, so as to obtain multiple different feature vectors. That is to say, it is equivalent to constructing multiple positive samples semantically similar to the target sample and inputting them into the feature encoding model, thereby obtaining multiple different feature vectors. Then, according to each feature vector, classification processing is performed on the object description text in the target sample, such that one target sample corresponds to multiple different classification results (at this time, the classification results are equivalent to the sample representations of the samples mentioned in the above-mentioned contrastive learning in the representation space), where the classification results include the probabilities that the object description text in the target sample belongs to at least one text category. Finally, according to the contrastive learning strategy, in the direction of reducing the classification result gap between multiple classification results, the model parameters of the feature encoding model are optimized. That is to say, this method first constructs multiple semantically similar positive samples, and then, through the contrastive learning strategy, optimizes the model parameters of the feature encoding model, such that the classification results of the object description text in the multiple positive samples obtained after classification processing according to the feature vectors obtained from the optimized feature encoding model will tend to be consistent. Therefore, if the probabilities that the object description text of the target sample belongs to multiple text categories are the same in the original classification results and it is difficult to determine a unique text category through the classification results, but since the classification results will approach the classification results of the samples semantically similar to it, and the classification results of the samples semantically similar to it can clearly indicate that the object description text belongs to a specific text category, then the classification results do not need to waver among multiple text categories, and the text category to which the object description text belongs can be determined according to the classification results of the samples semantically similar to it, thereby achieving the purpose of accurately classifying the text. It is not difficult to see that by adopting the above-mentioned text-based model training method, the feature vectors generated by the feature encoding model optimized through the contrastive learning strategy for classification processing will cluster the classification results of the object description text in the semantically similar samples, thereby achieving the purpose of accurately classifying the text.

[0045] Among them, the feature encoding model refers to a model that can encode natural language, which can be a BERT model (Bidirectional Encoder Representations from Transformer, a bidirectional encoder representation based on Transformer), a BERT twin tower model (a twin network model composed of two branches of BERT models with exactly the same parameters), XLNet (an autoregressive and autoencoding model), LayoutLM (Pre-training of Text and Layout for Document Image Understanding), etc. The process of the model encoding the text is a commonly used technical means by those skilled in the art and will not be elaborated here. Feature perturbation encoding refers to randomly discarding some features during the encoding process to generate the final feature vector through the undiscarded features. In practical applications, feature perturbation encoding can be dropout (a method that randomly makes some nodes in the network not work (i.e., sets the output to zero) during model training). Since features are used to represent the semantics of text, performing feature perturbation encoding on the same sample multiple times is equivalent to encoding multiple samples with little semantic difference (i.e., semantically similar). In addition, since the feature vector can represent the semantics of text, the feature vector obtained by calling the feature encoding model can represent the semantics of the object search text and the object description text.

[0046] In addition, a target sample includes an object search text and an object description text. The object search text refers to the text used as a search condition, and the object description text refers to the text used as a search result. The object described by the object description text is the object output by the feedback on the object search text. Here, the object can be anything, such as a specific commodity, a person, a restaurant, a tourist attraction, etc., and is not limited here. Specifically, the object search text can be the text received in response to the input operation during the user's search, or the text received in response to the selection operation during the user's search, etc.; while the object description text can be the text that has been stored in a module with a storage function and is used to describe the object. The module with a storage function can be a local or remote database, a local or remote memory, etc., and is not limited here. For example, when the user wants to get a haircut, they will enter "haircut" in the search bar of the group-buying software to search for barbershops. At this time, "haircut" is an object search text, and all the texts in the database of the group-buying software that are used to describe the merchants related to "haircut", such as "AA Hair Salon", "BB Haircut & Blow Dry", "CC Styling", "DD Beauty & Hair", etc., are object description texts that can provide feedback on the object search text of "haircut".

[0047] Based on the above text-based model training method, an embodiment of the present application provides a text-based model training system, as shown in Figure 2 , Figure 2 The text-based model training system shown can include multiple terminal devices 201 and multiple servers 202, and a communication connection is established between any terminal device and any server. The terminal device 201 can include any one or more of a smart phone, a tablet computer, a laptop computer, a desktop computer, an intelligent vehicle, and a smart wearable device. A variety of clients (applications, APPs) can run in the terminal device, such as a multimedia playback client, a social client, a browser client, an information flow client, an education client, and so on. The server 202 can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. The terminal device 201 and the server 202 can be directly or indirectly communicatively connected by wired or wireless communication methods, and the present application does not limit this here.

[0048] In one embodiment, the above text-based model training method can be completed only by Figure 2 the terminal device 201 in the text-based model training system shown. The specific execution process is as follows: The terminal device first obtains a target sample, and then calls the feature encoding model to perform first feature perturbation encoding on each text in the target sample to obtain a first feature vector, and calls the feature encoding model to perform second feature perturbation encoding on each text in the target sample to obtain a second feature vector; then, the terminal device respectively classifies the object description text in the target sample according to the first feature vector and the second feature vector to obtain two classification results; finally, the terminal device optimizes the model parameters of the feature encoding model according to the two classification results according to the contrast learning strategy. Optionally, the above text-based model training method can also be executed only by Figure 2 the server 202 in the text-based model training system shown, and its specific execution process can refer to the specific execution process of the terminal device and will not be elaborated here.

[0049] In another embodiment, the above text-based model training method can run in a text-based model training system. The text-based model training system can include a terminal device and a server. Among them, the text-based model training method can be performed by Figure 2It is completed jointly by the terminal device 201 and the server 202 included in the text-based model training system shown. The specific execution process is as follows: The server transmits the target sample to the terminal device; the terminal device calls the feature encoding model to perform first feature perturbation encoding on each text in the target sample to obtain a first feature vector, and calls the feature encoding model to perform second feature perturbation encoding on each text in the target sample to obtain a second feature vector; then, the terminal device classifies the object description text in the target sample according to the first feature vector and the second feature vector respectively to obtain two classification results; finally, the terminal device optimizes the model parameters of the feature encoding model according to the two classification results according to the contrast learning strategy.

[0050] Please refer to Figure 3 , Figure 3 is a schematic flowchart of a text-based model training method provided by an embodiment of the present application. This text-based model training method can be executed by the above-mentioned terminal device or server. For example, Figure 3 as shown, this text-based model training method includes steps S301 - S304:

[0051] S301, obtain a target sample.

[0052] In an embodiment of the present application, the target sample includes an object search text and an object description text. The object described by the object description text is the object output by the feedback on the object search text. Specifically, the object search text is generally also a text describing an object. It's just that the object described by the object search text is usually not as accurate and specific as the object described by the object description text. Generally speaking, the object described by the object description text and the object described by the object search text are the same or similar. For example, if the object search text is "male star", the object description text may be "post-90s popular drama male star".

[0053] In addition, there are mainly two ways to determine whether the object described in the object description text is the object output in response to the object search text. One way is to pre-annotate that the object described in a certain object description text is related to the object described in a certain object search text. The other way is to analyze historical data to determine relevant object search texts and object description texts. In practical applications, since a large number of samples usually need to be obtained, the latter way is usually adopted. Specifically, search click logs can be obtained from the historical database of the terminal device or the server. The search click logs are generated based on historical object search behaviors and object click behaviors. Each historical object search behavior corresponds to an object search text, and each object click behavior indicates an object description text that is clicked. Therefore, relevant object search texts and object description texts can be determined through the object click behaviors corresponding to each historical object search behavior recorded in the search click logs.

[0054] S302, call the feature encoding model to perform first feature perturbation encoding on each text in the target sample to obtain a first feature vector; and, call the feature encoding model to perform second feature perturbation encoding on each text in the target sample to obtain a second feature vector.

[0055] In the embodiments of the present application, both the first feature perturbation encoding and the second feature perturbation encoding belong to feature perturbation encoding. It's just that each time the feature encoding model is called to perform feature perturbation encoding on each text in the target sample, different features are randomly discarded. Therefore, each feature perturbation encoding needs to be distinguished. Since the target sample includes an object search text and an object description text, the feature encoding model can be called to perform feature extraction on the object search text and the object description text respectively. After obtaining the vector representations corresponding to the two texts (i.e., the feature vectors corresponding to the two texts), the vector representations of the two texts can be fused or concatenated in terms of features to obtain the first feature vector or the second feature vector corresponding to the target sample. Feature fusion and feature concatenation are common technical means for those skilled in the art and will not be elaborated here.

[0056] Specifically, the process of calling the feature encoding model to perform the first feature perturbation encoding on each text in the target sample to obtain the first feature vector can be as follows: First, call the feature encoding model to perform feature representation on each text in the target sample according to the first feature perturbation strategy to obtain the first vector representation of each text (i.e., the feature vector of the object search text and the feature vector of the object description text); then perform interaction processing on the first vector representations of each text to obtain the cross features between the texts in the target sample; finally, perform integration processing on the first vector representations of each text and the cross features between the texts to obtain the first feature vector. Similarly, the process of calling the feature encoding model to perform the second feature perturbation encoding on each text in the target sample to obtain the second feature vector can be as follows: First, call the feature encoding model to perform feature representation on each text in the target sample according to the second feature perturbation strategy to obtain the second vector representation of each text; then perform interaction processing on the second vector representations of each text to obtain the cross features between the texts in the target sample; finally, perform integration processing on the second vector representations of each text and the cross features between the texts to obtain the second feature vector.

[0057] Among them, since the feature perturbation encoding is achieved by randomly setting some network nodes in the network layer (such as the convolutional layer, etc.) in the feature encoding model not to work (i.e., setting the output to zero), the first feature perturbation strategy corresponding to the first feature perturbation encoding includes randomly generated parameters for determining which network nodes' outputs are set to zero. In addition, the first vector representation of the object search text is used to characterize the semantics of the object search text, and the first vector representation of the object description text is used to characterize the semantics of the object description text. Therefore, by performing interaction processing on the first vector representation of the object search text and the first vector representation of the object description text, semantic interaction between the two texts can be achieved, which will generate a cross-semantics. Since the semantics of two sentences that are more similar will not change much when combined, by constructing cross features that represent the cross-semantics between the object search text and the object description text, it is convenient to confirm the correlation between the object search text and the object description text, which is beneficial to determining the text category of the object description text in the subsequent process. In addition, the integration processing method can be feature fusion, feature splicing, etc. of the first vector representations of each text and the cross features between the texts, which will not be elaborated here.

[0058] Optionally, the feature encoding model may be composed of a first feature representation network and a second feature representation network; the first vector representation of the object search text in the target sample is obtained through the first feature representation network, and the first vector representation of the object description text in the target sample is obtained through the second feature representation network. Then, the way to call the feature encoding model to perform feature representation on each text in the target sample according to the first feature perturbation strategy to obtain the first vector representation of each text may be as follows: call the first feature representation network to extract features from the object search text to obtain the text features of the object search text; according to the first random parameter in the first feature perturbation strategy, discard some features in the text features of the object search text, and then perform pooling on the remaining features that are not discarded to obtain the first vector representation of the object search text. At the same time, call the second feature representation network to extract features from the object description text to obtain the text features of the object description text; according to the second random parameter in the first feature perturbation strategy, discard some features in the text features of the object description text, and then perform pooling on the remaining features that are not discarded to obtain the first vector representation of the object description text. Among them, the pooling process may be to perform normalization fusion on the remaining features that are not discarded, which is a commonly used technical means for those skilled in the art and will not be elaborated here.

[0059] In a possible implementation, the feature encoding model may also contain only one third feature representation network. Then, the method of calling the feature encoding model to perform feature representation on each text in the target sample according to the first feature perturbation strategy to obtain the first vector representation of each text can be as follows: First, input the object search text into the third feature representation network of the feature encoding model, and the third feature representation network extracts features from the object search text to obtain the text features of the object search text; then, according to the first random parameter in the first feature perturbation strategy, discard some features in the text features of the object search text, and then perform pooling processing on the remaining features that are not discarded to obtain the first vector representation of the object search text. After obtaining the first vector representation of the object search text, input the object description text into the third feature representation network of the feature encoding model, and the third feature representation network extracts features from the object description text to obtain the text features of the object description text; then, according to the second random parameter in the first feature perturbation strategy, discard some features in the text features of the object description text, and then perform pooling processing on the remaining features that are not discarded to obtain the first vector representation of the object description text. In a specific implementation, the feature encoding model can be a bert model with dropout. It should be noted that the bert model or the bert model branch in the bert two-tower model mentioned in this embodiment can be a publicly used bert model such as googlebert, or a bert model obtained by fine-tuning the publicly used bert model in advance by annotating the relevance between the object search text and the object description text in a specific application scenario such as the e-commerce field, and then using the object search text, object description text, and relevance in this specific application scenario.

[0060] Optionally, since the first vector representation is a feature vector and is also used to represent the text semantics, the semantics of the object search text and the object description text can be crossed by performing difference and sum operations on the first vector representations of each text. Then, the method of performing interaction processing on the first vector representations of each text to obtain the cross features between each text in the target sample can be specifically as follows: Perform a difference operation on the first vector representations of each text to obtain the difference operation result; at the same time, perform a sum operation on the first vector representations of each text to obtain the sum operation result; finally, use the difference operation result and the sum operation result to construct the cross features between each text in the target sample.

[0061] For example, the feature encoding model can be a bert two-tower model with dropout. Please refer to the appendix Figure 4, which shows a schematic structural diagram of a BERT two-tower model, including a first feature representation network 401 (composed of a BERT model branch and a pooling layer), a second feature representation network 402 (composed of a BERT model branch and a pooling layer), and an interaction layer 403. In addition, the target sample includes an object search text Query and an object description text Title. First, the Query is input into the first feature representation network 401. The BERT model branch in the first feature representation network 401 extracts features from the Query to obtain four text features a, b, c, d of the Query. Then, the first random parameter is determined to be 3 through the first feature perturbation strategy in this BERT model branch, that is, the output of the third text feature is set to zero. Therefore, only three text features a, b, d of the Query need to be input into the pooling layer in the first feature representation network 401 for pooling processing to obtain the first vector representation q of the Query.

[0062] At the same time, the Title can be input into the second feature representation network 402. The BERT model branch in the second feature representation network 402 extracts features from the Title to obtain five text features x, y, z, m, n of the Query. Then, the second random parameters are determined to be 1 and 4 through the second feature perturbation strategy in this BERT model branch, that is, the outputs of the first and fourth text features are set to zero. Therefore, only three text features y, z, n of the Title need to be input into the pooling layer in the second feature representation network 402 for pooling processing to obtain the first vector representation t of the Title.

[0063] Finally, q and t are input into the interaction layer 403. The interaction layer 403 performs a difference operation on q and t to obtain a difference operation result q - t, and performs a sum operation on q and t to obtain a sum operation result q + t; then, it is determined that the cross features between the Query and the Title include q - t and q + t. Finally, the interaction layer 403 concatenates q - t, q + t, q, and t to obtain a first feature vector. Similarly, the BERT two-tower model can be called again to perform second feature perturbation encoding on the object search text Query and the object description text Title in the target sample to obtain a second feature vector, which will not be elaborated here.

[0064] S303: Classify the object description text in the target sample according to the first feature vector and the second feature vector respectively to obtain two classification results.

[0065] In an embodiment of the present application, any classification result includes the probabilities that the object description text in the target sample belongs to at least one text category. That is to say, the classification result can be a probability vector. At the same time, if the probability that the object description text belongs to a certain text category in the classification result is the largest, then it can be said that the text category indicated by the classification result is the certain text category. For example, it can be set that there are three text categories A, B, and C in total. The probability that the object description text in the target sample belongs to A is 0.6, the probability that it belongs to B is 0.3, and the probability that it belongs to C is 0.1. Then, the classification result can be determined as [0.6, 0.3, 0.1], and the text category indicated by this classification result is text category A. It should be noted that for each probability in the subsequent classification results corresponding to the text categories, it is the same as in this example. If the order of description of the text categories is A, B, and C, and the classification result is [0.6, 0.3, 0.1], then 0.6 corresponds to text category A, 0.3 corresponds to text category B, and 0.1 corresponds to text category C, which will not be elaborated hereinafter. It should be noted that since the first feature vector and the second feature vector obtained through the feature encoding model are used to classify the object description text in the target sample to obtain the classification result of the object description text, and thus determine the text category of the object description text, the feature vectors such as the first feature vector and the second feature vector obtained through the feature encoding model can also be called relevance binning vectors.

[0066] In addition, the specific manner of the classification process may specifically be to call a classification loss function (such as softmax, sigmoid, etc.) to perform classification processing according to the first feature vector and the second feature vector to obtain two classification results, or to call a machine learning model for classification such as a gradient boosting decision tree (i.e., GBDT model) to perform classification processing according to the first feature vector and the second feature vector to obtain two classification results, which is not limited herein. Since the process of performing classification processing through the classification function and the GBDT model is a common technical means for those skilled in the art, it will not be elaborated herein.

[0067] S304, optimize the model parameters of the feature encoding model according to the two classification results according to the contrast learning strategy.

[0068] In the embodiments of the present application, the optimization of the model parameters of the feature encoding model according to the two classification results based on the contrast learning strategy specifically refers to optimizing the model parameters of the feature encoding model in the direction of reducing the gap between the two classification results. Specifically, the method of optimizing the model parameters of the feature encoding model according to the two classification results based on the contrast learning strategy can be as follows: obtain the target loss function constructed based on the contrast learning strategy; then call the target loss function to perform a loss value operation according to the two classification results to obtain the target loss value; finally, optimize the model parameters of the feature encoding model in the direction of reducing the target loss value. Among them, the classification results mentioned in step S303 can be represented by a probability vector, and contrast learning is to optimize the model parameters of the feature encoding model in the direction of reducing the gap between the two classification results. Therefore, the target loss function in step S304 can be a loss function that can calculate the distance between two probability vectors. Exemplarily, the target loss function can be the Kullback-Leibler divergence, and the specific formula is as follows:

[0069]

[0070] Among them, since multiple samples are often used for training in the model optimization process, i in the formula refers to the i-th sample, and KL refers to that this loss function is the Kullback-Leibler divergence, refers to one of the two classification results (specifically a probability vector), refers to the other classification result of the two classification results (specifically a probability vector). The distance between the two classification results can be calculated through the Kullback-Leibler divergence. Therefore, optimizing the model parameters of the feature encoding model in the direction of reducing the target loss value means that during the model optimization process, the distance between the two classification results obtained after the first feature vector and the second feature vector obtained by processing the target sample by the updated and optimized feature encoding model are used for classification processing will become smaller and smaller.

[0071] Further, the method for optimizing the model parameters of the feature encoding model in the direction of reducing the target loss value can specifically be: obtaining the class label of the object description text in the target sample, and calling the cross-entropy loss function to calculate the cross-entropy loss value according to the class label and the two classification results; then, integrating the cross-entropy loss value and the target loss value to obtain the integrated loss value; and optimizing the model parameters of the feature encoding model in the direction of reducing the integrated loss value. Among them, the class label of the object description text is used to indicate the text class to which the object description text belongs, which can be manually labeled or calculated through relevant data of the object search text and the object description text, and is not limited here. In addition, since the aforementioned text class can be referred to as the relevance gear, the class label can also be the gear label. Exemplarily, the specific formula of the cross-entropy loss function is as follows:

[0072]

[0073] Among them, since multiple samples are often used for training in the model optimization process, i in the formula refers to the i-th sample, ce refers to that this loss function is the cross-entropy loss function, y i refers to the class label of the object description text. In the cross-entropy loss function, the type label can be expressed in the form of a probability vector. For example, if there are a total of 3 text classes and the class label is determined to be the second text class, then at this time y i can be [0, 1, 0], that is, the probability that the object description text belongs to the second text class is determined to be 1. And The same as above, which will not be elaborated here.

[0074] In specific implementation, the method for integrating the cross-entropy loss value and the target loss value to obtain the integrated loss value can be to directly add the cross-entropy loss value and the target loss value, or to add the cross-entropy loss value and the target loss value with weights, etc., which is not limited here. For example, the cross-entropy loss value and the target loss value can be integrated through the formula where refers to the cross-entropy loss value, refers to the target loss value, and α is a hyperparameter used to adjust the weights of the cross-entropy loss value and the target loss value, which can be flexibly set according to the model training situation.

[0075] For example, please refer to the appendix Figure 5, which shows a schematic diagram of the architecture of a contrastive learning strategy. The target sample i contains an object search text Query and an object description text Doc. The bert two-tower model with dropout is called to perform the first feature perturbation encoding on the target sample i to obtain the first feature vector, and then the classification function softmax is called to determine the first classification result of Doc according to the first feature vector The bert two-tower model with dropout is called again to perform the second feature perturbation encoding on the target sample i to obtain the second feature vector, and the classification function softmax is called again to determine the second classification result of Doc according to the second feature vector After obtaining the classification result and the classification result After that, the classification result and the classification result can be substituted into the KL divergence function to obtain the target loss value, and the classification result classification result and the class label y of the object description text i are substituted into the cross-entropy loss function to obtain the cross-entropy loss value; furthermore, the cross-entropy loss value and the target loss value are integrated, and the model parameters of the bert two-tower model with dropout are optimized in the direction of reducing the integrated loss value.

[0076] As can be seen from the above description, the feature encoding model optimized by the above text-based model training method can make the classification results of samples with similar semantics tend to be consistent, thereby achieving the purpose of improving the accuracy of text classification. For example, the object search text Q1 is "5-point short sleeve", and the object description text T1 is "AAA spring and autumn knitted women's 5-point sleeve V-neck half sleeve ice silk top bottoming shirt mid sleeve sweater for outer wear blue XL". Although the semantics of T1 include "5-point short sleeve", due to text stacking in T1, there are also more semantics unrelated to "5-point short sleeve", making it difficult to determine the relevance between T1 and Q1. However, the classification result predicted by the optimized feature encoding model will approach the classification results of other samples with semantics similar to T1 and Q1, so that the relevance between Q2 and T2 can be accurately determined to be relatively high. Therefore, it can be determined that T1 is in grade 1. Another example is that the object search text Q2 is "iPhone 12 Pro", and the object description text T2 is "In stock and shipped quickly iPhone Fruit 11 Pro midnight green 256GB". Since the semantics of T2 contain "11" which is not in the semantics of Q2, the classification result predicted by the optimized encoding model will approach the classification results of other samples with "11" in their semantics, so that the relevance between Q2 and T2 can be determined to be relatively low. Therefore, it can be accurately determined that T2 is in grade 3. It should be noted that grade 3 and grade 4 in this embodiment correspond to the relevance grades in the foregoing example of text categories, which will not be elaborated here.

[0077] In the embodiment of the present application, by calling the feature encoding model multiple times to perform first feature perturbation encoding and second feature perturbation encoding on each text in the target sample respectively, two positive samples semantically similar to the object search text and the object description text in the target sample can be constructed, and the first feature vector and the second feature vector of the two positive samples can be obtained; and by classifying the object description text in the target sample according to the first feature vector and the second feature vector respectively, two classification results can be obtained, which is equivalent to obtaining the positions of two semantically similar positive samples in the representation space. Then, through the contrastive learning strategy, in the direction of reducing the distance between the two classification results, the model parameters of the feature encoding model are optimized, so that after the feature encoding model encodes two semantically similar samples, the two feature vectors obtained will be more consistent when they are respectively used for classification processing, thus reducing the probability that the classification results point to multiple text categories at the same time or belong to multiple text categories with similar probabilities. In addition, in the contrastive learning strategy, the cross-entropy loss value can also be calculated according to the category label of the object description text in the target sample and the two classification results, and the model parameters of the feature encoding model are optimized in the direction of reducing the integrated loss value obtained by integrating the cross-entropy loss value and the target loss value, so that when the feature encoding model encodes two semantically similar samples, the two feature vectors obtained will not only be more consistent when they are respectively used for classification processing, but also approach the category label indicating the correct text category, thus realizing that the text category indicated by the classification result is more accurate while reducing the fuzziness of the text category indicated by the classification result.

[0078] Please refer to Figure 6 , Figure 6 is a schematic flowchart of another text-based model training method provided by the embodiment of the present application. This text-based model training method can also be executed by the above-mentioned terminal device or server. As Figure 6 shown, this text-based model training method includes steps S601 - S609:

[0079] S601, Obtain a target sample.

[0080] S602, Call the feature encoding model to perform first feature perturbation encoding on each text in the target sample to obtain a first feature vector; and call the feature encoding model to perform second feature perturbation encoding on each text in the target sample to obtain a second feature vector.

[0081] S603, Classify the object description text in the target sample according to the first feature vector and the second feature vector respectively to obtain two classification results.

[0082] S604. Optimize the model parameters of the feature encoding model according to the two classification results based on the contrastive learning strategy.

[0083] During the model training process, if samples with a high correlation between the object search text and the object description text (i.e., samples with text categories that are not easily blurred) are first input into the feature encoding model for training, it can enable the feature encoding model to better learn the ability to perform feature encoding on text during the early stage of model training. This is beneficial for the subsequent feature encoding model to better learn the ability to cluster semantically similar samples through the contrastive learning strategy when facing samples with a relatively high or general correlation between the object search text and the object description text (i.e., samples with text categories that are easily blurred).

[0084] Therefore, in steps S601 - S604, the method for obtaining the target samples can be specifically as follows: First, obtain a sample sequence for model optimization, where the sample sequence includes multiple samples arranged in sequence; then, traverse the sample sequence in turn and determine the currently traversed sample as the target sample. That is, according to the arrangement order of each sample in the sample sequence, each sample is input into the feature encoding model.

[0085] Specifically, since the correlation between the object search text and the object description text can be measured by the probability that the object description text is clicked when the object search text is searched, the sample sequence can be constructed as follows: First, obtain the search click log, which is generated based on historical object search behaviors and historical object click behaviors; then parse the search click log to obtain the parsing result, where the parsing result includes at least one object search text input historically, one or more object description texts corresponding to each object search text, and the click-through rate of the object described by each object description text; then, based on each object search text and each object description text in the parsing result, construct multiple samples, where one sample includes one object search text and one object description text corresponding to the object search text; finally, sort the multiple samples according to the click-through rate of the object described by the object description text in each sample from high to low to obtain the sample sequence. Specifically, it is to construct a partial order relationship or a pre-order relationship of one or more object description texts corresponding to the object search text through the click-through rate.

[0086] Among them, the way to parse the search click log can be to find out the object search text corresponding to each historical object search behavior to obtain at least one object search text of the historical input, and to find out the object description text indicated by one or more historical object click behaviors corresponding to at least one object search text, so as to determine one or more object description texts corresponding to each object search text. Finally, count the search times of each object search text, and during the search process of each object search text, count the number of times each object description text is clicked, so as to obtain the click-through rate of the object described by each object description text. Then, after parsing the search click log, according to each object search text and each object description text in the parsing result, the format of each sample can be <object search text, object description text>: search times, click times, click-through rate.

[0087] For example, from the historical search click log in the search engine background of an application with a product search function, filter out the log containing the product brands existing in the product database of the product search application and the product dictionary data (such as the publicly available product database dictionary across the network), and use it as the target search click log. Then parse the target search click log to obtain the parsing result. Since in the parsing result, the object search text Query 1 corresponds to 3 object description texts, namely Doc1, Doc2, and Doc3, and the click-through rate of Doc1 is 0.35, the click-through rate of Doc2 is 0.59, and the click-through rate of Doc3 is 0.06, 3 samples of Query 1 can be constructed as follows:

[0088] Sample 1: <Query 1, Doc1> 0.35;

[0089] Sample 2: <Query 1, Doc2> 0.59;

[0090] Sample 3: <Query 1, Doc3> 0.06;

[0091] Then, sort these three samples in descending order according to the click-through rate, and the obtained sample sequence is Sample 2, Sample 1, Sample 3. Then, the model will be optimized in the order of first determining Sample 2 as the target sample, then determining Sample 1 as the target sample, and finally determining Sample 3 as the target sample.

[0092] Further, since in steps S601 - S604, the correlation between the object search text and the object description text is measured by the click - through rate, and the correlations between the object search text and the object description text indicated by different text categories are different, the text category to which the object description text belongs can be determined by setting click - through rate intervals. In addition, according to the description of the contrastive learning strategy in step S304, it can be determined that the category label of the object description text in the target sample will be used in the contrastive learning strategy, where the category label refers to the text category to which the object description text belongs. Then, from the above description, the method for determining the category label of the object description text can be specifically as follows: Determine multiple click - through rate intervals, where one click - through rate interval corresponds to one text category; then for any sample, use the click - through rate of the object described by the object description text in the sample to perform a hit match on the multiple click - through rate intervals; finally, construct the category label of the object description text in any sample according to the text category corresponding to the hit click - through rate interval.

[0093] For example, as shown in Table 1, the following five click - through rate intervals can be set, and each click - through rate interval corresponds to a click - through rate range.

[0094] Table 1

[0095] Click-through rate range 0 ≤ click-through rate < 0.1 0.1 ≤ Click-Through Rate < 0.4 0.4 ≤ Click-through rate < 0.6 0.6 ≤ click-through rate < 0.8 0.8 ≤ Click-through rate Click-through rate level 1 2 3 4 5

[0096] Since one click - through rate interval corresponds to one text category, one click - through rate range is also one text category. At the same time, because the category label refers to the text category to which the object description text belongs, if the click - through rate range of the object description text is 1, then it can be determined that the category label of the object description text is 1. Therefore, if the click - through rate of the object described by the object description text in sample A is 0.75, then it can be determined that the click - through rate is between 0.6 and 0.8, and thus it can be determined that the category label of the object description text in sample A is 4. Similarly, if the click - through rate of the object described by the object description text in sample B is 0.59, then it can be determined that the click - through rate is between 0.4 and 0.6, and thus it can be determined that the category label of the object description text in sample B is 3.

[0097] It should be noted that for other specific implementation manners of steps S601 - S604, reference can be made to the specific implementation manners in steps S301 - S304, which will not be elaborated herein in this application.

[0098] S605, if a target object search text is received, obtain the target object description texts of each object used to provide feedback for the target object search text.

[0099] In the embodiment of the present application, the target object search text may be the text received by a terminal device or a server in response to a user's search editing operation. The search editing operation may be a search text input operation, a search text selection operation, a search text modification operation, etc., which is not limited herein. In addition, the manner of obtaining the target object description text of each object for providing feedback on the target object search text may be to obtain all object description texts stored in the local storage module of the terminal device or the server for storing description objects, or to obtain the object description texts corresponding to other historical object search texts that are relatively similar to the received target object search text stored in the local storage module of the terminal device or the server, which is not limited herein.

[0100] S606. Use the target object search text and the obtained target object description texts of each object to construct multiple text pairs.

[0101] In the embodiment of the present application, each sample pair includes a target object search text and a target object description text. The specific implementation manner of constructing the text pair may refer to the specific implementation manner of constructing the sample in steps S601 to S604, which will not be elaborated herein.

[0102] S607. Call the optimized feature encoding model to perform feature encoding on each text in each text pair to obtain the target feature vector of each text pair.

[0103] Among them, the specific implementation manner of calling the feature encoding model to perform feature encoding on the text is a commonly used technical means for those skilled in the art, which will not be elaborated herein. It should be emphasized that the optimized feature encoding model does not need to perform feature perturbation encoding on each text in each text pair as in the model training process. Feature perturbation encoding is performed in the model training process to construct positive samples with similar semantics and to facilitate subsequent use of the contrast learning strategy to make the classification results corresponding to the two positive samples more consistent, so as to improve the stability of the feature encoding model for minor perturbations while also making the feature encoding model more suitable for supervised tasks such as text classification.

[0104] S608. Classify the target object description text in each text pair according to the target feature vector of each text pair to obtain the classification result of each target object description text.

[0105] It should be noted that other implementation manners of step S608 may refer to the specific implementation manner in step S303, which will not be elaborated in the present application.

[0106] S609. Determine the output order of the objects described by each target object description text according to the classification result of each target object description text, and output each object in the corresponding output order.

[0107] In the embodiment of the present application, the determination of the output order of the objects described by each target object description text according to the classification result of each target object description text means: output the objects described by the target object description text in the order from high to low of the relevance indicated by the text category according to the text category of the target object description text indicated by the classification result of each target object description text.

[0108] For example, please refer to the appendix Figure 7 , which shows a schematic diagram of outputting a target object description text. Five text categories 1, 2, 3, 4, and 5 can be preset in advance. Among them, the object search text indicated by text category 5 has the highest relevance to the object description text, and the relevance of the object search text indicated by text category 4, text category 3, text category 3, and text category 1 to the object description text decreases in turn. As shown in the product search interface 701, it can be determined that the received target object search text is "small h bottle". Then, the target object description texts of each object for feedback on "small h bottle" can be obtained from the product database corresponding to the product search, which are target object description text 702, target object description text 703, target object description text 704, and target object description text 705 respectively. At this time, "small h bottle" can be constructed into a text pair with target object description text 702, "small h bottle" can be constructed into a text pair with target object description text 703, "small h bottle" can be constructed into a text pair with target object description text 704, and "small h bottle" can be constructed into a text pair with target object description text 705 to obtain 4 text pairs.

[0109] After obtaining 4 text pairs, call the optimized feature encoding model to perform feature encoding on each text in each text pair to obtain the target feature vector of each text pair; and classify the target object description text in each text pair according to the target feature vector of each text pair respectively, so that the classification result of the target object description text 702 can be obtained as [0, 0, 0.03, 0.17, 0.8]. Through the classification result of the target object description text 702, it can be determined that the probability that the text category of the target object description text 702 is text category 5 is the highest. Therefore, the text category indicated by the classification result of the target object description text 702 is text category 5. Similarly, it can be determined that the text categories of the target object description text 703, the target object description text 704, and the target object description text 705 are text category 4, text category 3, and text category 3 respectively. Finally, as shown in the product search interface 706, output each target object description text and the object described by each target object description text (that is, the picture next to each target object description text) in descending order according to the text category corresponding to each target object description text.

[0110] In another embodiment, steps S601 to S604 may be executed by Figure 3 the server in the text-based model training system shown in. After the server obtains the optimized feature encoding model, it can receive the target object search text sent from Figure 3 the terminal device in the text-based model training system shown in. Then the server obtains the target object description texts of each object that gives feedback on the target object search text, and executes steps S605 to S608. After the server obtains the classification results of each target object description text, the server will send the classification results of each target object description text to the terminal device, and finally the terminal device executes step S609.

[0111] In the embodiments of the present application, in order to enable the feature vectors obtained by the feature encoding model to be better used to determine the classification result of the object description text, during the training process of the feature encoding model, a sample sequence can be obtained first, where each sample in the sample sequence is arranged in descending order of the correlation between the object search text and the object description text in each sample. Then, since the correlation between the object search text and the object description text can be measured by the click-through rate of the object described by the object description text, the category label of the object description text can be quickly determined by setting a click-through rate interval corresponding to a text category, avoiding the tediousness and low efficiency of manual annotation. Finally, by calling the target feature vectors of each text pair obtained by the optimized feature encoding model, it is convenient to more accurately predict the classification result of the target object description text in each text pair in the future, which is conducive to preferentially outputting the target object text more relevant to the target object search text, thereby improving the user experience.

[0112] Based on the description of the related embodiments of the above model training method, the embodiments of the present application also propose a text-based model training device, which can be a computer program (including program code) running in a computer device. The text-based model training device can execute Figure 3 or Figure 6 the model training method shown; please refer to Figure 8 , and the text-based model training device can run the following units:

[0113] An acquisition unit 801, configured to acquire a target sample, where the target sample includes an object search text and an object description text; the object described by the object description text is the object output by giving feedback on the object search text;

[0114] The processing unit 802 is further configured to call the feature encoding model to perform first feature perturbation encoding on each text in the target sample to obtain a first feature vector; and call the feature encoding model to perform second feature perturbation encoding on each text in the target sample to obtain a second feature vector;

[0115] The processing unit 802 is further configured to perform classification processing on the object description text in the target sample respectively according to the first feature vector and the second feature vector to obtain two classification results; any classification result includes: the probability that the object description text in the target sample belongs to at least one text category, and the text category is used to indicate the correlation between the object description text and the object search text;

[0116] The processing unit 802 is further configured to optimize the model parameters of the feature encoding model according to the two classification results according to the contrast learning strategy.

[0117] In one embodiment, the processing unit 802 may further be configured to:

[0118] Call the feature encoding model to perform feature representation on each text in the target sample according to the first feature perturbation strategy, and obtain the first vector representation of each text;

[0119] Perform interaction processing on the first vector representations of each text to obtain the cross features between each text in the target sample;

[0120] Integrate and process the first vector representations of each text and the cross features between each text to obtain the first feature vector.

[0121] In another embodiment, when the processing unit 802 is used to perform interaction processing on the first vector representations of each text to obtain the cross features between each text in the target sample, it may specifically be configured to:

[0122] Perform a difference operation on the first vector representations of each text to obtain a difference operation result; and perform a summation operation on the first vector representations of each text to obtain a summation operation result;

[0123] Use the difference operation result and the summation operation result to construct the cross features between each text in the target sample.

[0124] In another embodiment, the feature encoding model includes a first feature representation network and a second feature representation network; the first vector representation of the object search text in the target sample is obtained through the first feature representation network, and the first vector representation of the object description text in the target sample is obtained through the second feature representation network; when the processing unit 802 is used to call any one of the feature representation networks to perform feature representation on the corresponding text according to the first feature perturbation strategy to obtain the corresponding first vector representation, it may specifically be configured to:

[0125] Call any one of the feature representation networks to perform feature extraction on the corresponding text to obtain the text features of the corresponding text;

[0126] Discard some features in the text features according to the first random parameter in the first feature perturbation strategy;

[0127] Perform pooling processing on the remaining features in the text features that are not discarded to obtain the first vector representation of the corresponding text.

[0128] In another embodiment, when the processing unit 802 is used to optimize the model parameters of the feature encoding model according to the contrast learning strategy based on two classification results, it may specifically be configured to:

[0129] Obtain the target loss function constructed based on the contrast learning strategy;

[0130] Call the target loss function to calculate the loss value based on the two classification results, and obtain the target loss value;

[0131] Optimize the model parameters of the feature encoding model in the direction of reducing the target loss value.

[0132] In another implementation manner, when the processing unit 802 is used to optimize the model parameters of the feature encoding model in the direction of reducing the target loss value, it can specifically be used for:

[0133] Obtain the category label of the object description text in the target sample, and call the cross-entropy loss function to calculate the cross-entropy loss value according to the category label and the two classification results;

[0134] Integrate the cross-entropy loss value and the target loss value to obtain the integrated loss value; and optimize the model parameters of the feature encoding model in the direction of reducing the integrated loss value.

[0135] In another implementation manner, when the obtaining unit 801 obtains the target sample, it can specifically be used for:

[0136] Obtain a sample sequence for model optimization, where the sample sequence includes a plurality of samples arranged in sequence; traverse the sample sequence in turn, and determine the currently traversed sample as the target sample.

[0137] Among them, the sample sequence is constructed as follows:

[0138] Obtain the search click log, which is generated according to the historical object search behavior and the historical object click behavior;

[0139] Parse the search click log to obtain the parsing result; the parsing result includes: at least one object search text input historically, one or more object description texts corresponding to each object search text, and the click-through rate of the object described by each object description text;

[0140] Construct a plurality of samples according to each object search text and each object description text in the parsing result; one sample includes: one object search text and the corresponding one object description text;

[0141] Sort the plurality of samples according to the click-through rate of the object described by the object description text in each sample in descending order of click-through rate to obtain the sample sequence.

[0142] In another implementation manner, the obtaining unit 801 is further used for:

[0143] Determine a plurality of click-through rate intervals, and one click-through rate interval corresponds to one text category;

[0144] For any sample, use the click-through rate of the object described by the object description text in the sample to perform hit matching for multiple click-through rate intervals;

[0145] According to the text category corresponding to the hit click-through rate interval, construct a category label for the object description text in any sample.

[0146] In another embodiment, after obtaining the optimized feature encoding model, the processing unit 802 is further configured to:

[0147] If a target object search text is received, obtain the target object description texts of each object used to feedback on the target object search text;

[0148] Use the target object search text and the obtained target object description texts to construct multiple text pairs; a text pair includes the target object search text and a target object description text;

[0149] Call the optimized feature encoding model to perform feature encoding on each text in each text pair to obtain the target feature vector of each text pair;

[0150] Respectively classify the target object description texts in each text pair according to the target feature vector of each text pair to obtain the classification results of each target object description text;

[0151] According to the classification results of each target object description text, determine the output order of the objects described by each target object description text, and output each object in the corresponding output order.

[0152] According to an embodiment of the present invention, Figure 3 or Figure 6 Each step involved in the method shown can be executed by each unit in the text-based model training device shown. For example, Figure 8 The step S301 shown in can be executed by the obtaining unit 801 shown in, and the steps S302 to S304 can all be executed by Figure 3 The processing unit 802 shown in. Another example, Figure 8 The step S601 shown in can be executed by the obtaining unit 801 shown in, and the steps S602 to S609 can all be executed by Figure 8 The processing unit 802 shown in. And so on. Figure 6 The step S601 shown in can be executed by the obtaining unit 801 shown in, and the steps S602 to S609 can all be executed by Figure 8 The obtaining unit 801 shown in, and the steps S602 to S609 can all be executed by Figure 5 The updating unit 502 shown in, and so on.

[0153] According to another embodiment of the present invention, Figure 8Each unit in the text-based model training device shown can be separately or all combined into one or several other units to form, or a certain one (or some) of the units can be further split into multiple smaller units with more specific functions to form. This can achieve the same operations without affecting the realization of the technical effects of the embodiments of the present invention. The above units are divided based on logical functions. In practical applications, the function of one unit can also be realized by multiple units, or the functions of multiple units can be realized by one unit. In other embodiments of the present invention, the text-based model training device can also include other units. In practical applications, these functions can also be assisted by other units and can be realized through the cooperation of multiple units.

[0154] According to another embodiment of the present invention, it can be achieved by running a computer program (including program code) that can execute the respective steps involved in the corresponding methods shown in Figure 3 or Figure 6 on a general computing device such as a computer including processing elements and storage elements such as a central processing unit (CPU), a random access storage medium (RAM), and a read-only storage medium (ROM), to construct a text-based device as shown in Figure 8 and to implement the text-based model training method of the embodiments of the present invention. The computer program can be recorded on, for example, a computer-readable storage medium, loaded into the above computing device through the computer-readable storage medium, and run therein.

[0155] In the embodiments of the present application, by repeatedly calling the feature encoding model and respectively performing first feature perturbation encoding and second feature perturbation encoding on each text in the target sample, it is possible to construct two positive samples that are semantically similar to the object search text and the object description text in the target sample, and obtain the first feature vector and the second feature vector of the two positive samples; and by respectively classifying the object description text in the target sample according to the first feature vector and the second feature vector, two classification results can be obtained, which is equivalent to obtaining the positions of the two semantically similar positive samples in the representation space. Then, through a contrastive learning strategy, in the direction of reducing the distance between the two classification results, the model parameters of the feature encoding model are optimized, so that after the feature encoding model encodes two semantically similar samples, the two feature vectors obtained will be more consistent when used for classification processing respectively, thereby reducing the probability that the classification results simultaneously point to multiple text categories or belong to multiple text categories with similar probabilities, so as to improve the accuracy of text classification.

[0156] Based on the descriptions of the above method embodiments and device embodiments, the embodiments of the present invention also provide a computer device. Please refer to Figure 9, the computer device at least includes a processor 901, an input interface 902, an output interface 903, and a computer storage medium 904. Among them, the processor 901, the input interface 902, the output interface 903, and the computer storage medium 904 in the computer device can be connected through a bus or other means.

[0157] The computer storage medium 904 can be stored in the memory of the computer device. The computer storage medium 904 is used to store a computer program, and the computer program includes program instructions. The processor 901 is used to execute the program instructions stored in the computer storage medium 904. The processor 901 (or CPU (Central Processing Unit, central processor)) is the computing core and control core of the computer device, and is suitable for implementing one or more instructions. Specifically, it is suitable for loading and executing one or more instructions to implement the corresponding method flow or corresponding function.

[0158] In one embodiment, one or more computer programs stored in the computer storage medium 904 can be loaded and executed by the processor 901 to implement the corresponding method steps in the above-mentioned Figure 3 and Figure 6 shown method embodiments; in specific implementation, one or more computer programs in the computer storage medium 904 can be loaded and executed by the processor 901 as follows: obtaining a target sample, where the target sample includes an object search text and an object description text; the object described by the object description text is the object output by giving feedback on the object search text; calling a feature encoding model to perform first feature perturbation encoding on each text in the target sample to obtain a first feature vector; and calling the feature encoding model to perform second feature perturbation encoding on each text in the target sample to obtain a second feature vector; respectively classifying the object description text in the target sample according to the first feature vector and the second feature vector to obtain two classification results; any classification result includes: the probability that the object description text in the target sample belongs to at least one text category, and the text category is used to indicate the relevance between the object description text and the object search text; optimizing the model parameters of the feature encoding model according to the two classification results according to the contrast learning strategy.

[0159] In one embodiment, when the processor 901 invokes the feature encoding model to perform first feature perturbation encoding on each text in the target sample to obtain a first feature vector, it is specifically configured to load and execute: invoking the feature encoding model to perform feature representation on each text in the target sample according to a first feature perturbation strategy to obtain a first vector representation of each text; performing interaction processing on the first vector representations of each text to obtain cross features between each text in the target sample; integrating the first vector representations of each text and the cross features between each text to obtain a first feature vector.

[0160] In another embodiment, when the processor 901 performs interaction processing on the first vector representations of each text to obtain cross features between each text in the target sample, it is specifically configured to load and execute: performing a difference operation on the first vector representations of each text to obtain a difference operation result; and performing a summation operation on the first vector representations of each text to obtain a summation operation result; constructing cross features between each text in the target sample by using the difference operation result and the summation operation result.

[0161] In another embodiment, the feature encoding model includes a first feature representation network and a second feature representation network; the first vector representation of the object search text in the target sample is obtained through the first feature representation network, and the first vector representation of the object description text in the target sample is obtained through the second feature representation network; wherein, when the processor 901 invokes any one of the feature representation networks to perform feature representation on the corresponding text according to the first feature perturbation strategy to obtain the corresponding first vector representation, it is specifically configured to load and execute: invoking any one of the feature representation networks to perform feature extraction on the corresponding text to obtain a text feature of the corresponding text; discarding some features in the text feature according to the first random parameter in the first feature perturbation strategy; performing pooling processing on the remaining features in the text feature that are not discarded to obtain a first vector representation of the corresponding text.

[0162] In another embodiment, when the processor 901 optimizes the model parameters of the feature encoding model according to the two classification results according to the contrast learning strategy, it is specifically configured to load and execute: obtaining a target loss function constructed according to the contrast learning strategy; invoking the target loss function to perform a loss value operation according to the two classification results to obtain a target loss value; optimizing the model parameters of the feature encoding model in the direction of reducing the target loss value.

[0163] In yet another embodiment, when optimizing the model parameters of the feature encoding model in the direction of reducing the target loss value, the processor 901 may specifically be configured to load and execute: obtain the category label of the object description text in the target sample, and call the cross-entropy loss function to calculate the cross-entropy loss value according to the category label and the two classification results; integrate the cross-entropy loss value and the target loss value to obtain the integrated loss value; and optimize the model parameters of the feature encoding model in the direction of reducing the integrated loss value.

[0164] In yet another embodiment, when obtaining the target sample, the processor 901 may specifically be configured to load and execute: obtain a sample sequence for model optimization, where the sample sequence includes a plurality of samples arranged in sequence; traverse the sample sequence in sequence, and determine the currently traversed sample as the target sample. The sample sequence is constructed as follows: obtain the search click log, which is generated according to the historical object search behavior and the historical object click behavior; parse the search click log to obtain the parsing result; the parsing result includes: at least one object search text input historically, one or more object description texts corresponding to each object search text, and the click-through rate of the object described by each object description text; construct a plurality of samples according to each object search text and each object description text in the parsing result; one sample includes: one object search text and the corresponding one object description text; sort the plurality of samples according to the click-through rate of the object described by the object description text in each sample in descending order of the click-through rate to obtain the sample sequence.

[0165] In yet another embodiment, the processor 901 may further be configured to load and execute: determine a plurality of click-through rate intervals, where one click-through rate interval corresponds to one text category; for any sample, perform a hit match on the plurality of click-through rate intervals by using the click-through rate of the object described by the object description text in the sample; construct the category label of the object description text in any sample according to the text category corresponding to the hit click-through rate interval.

[0166] In yet another embodiment, after obtaining the optimized feature encoding model, the processor 901 can also be used to load and execute: if a target object search text is received, obtain the target object description texts of each object for providing feedback on the target object search text; use the target object search text and the obtained target object description texts to construct multiple text pairs; one text pair includes the target object search text and a target object description text; call the optimized feature encoding model to perform feature encoding on each text in each text pair respectively to obtain the target feature vectors of each text pair; classify the target object description texts in each text pair respectively according to the target feature vectors of each text pair to obtain the classification results of each target object description text; determine the output order of the objects described by each target object description text according to the classification results of each target object description text, and output each object in the corresponding output order.

[0167] In the embodiments of the present application, by calling the feature encoding model multiple times to perform first feature perturbation encoding and second feature perturbation encoding on each text in the target sample respectively, two positive samples semantically similar to the object search text and the object description text in the target sample can be constructed, and the first feature vectors and the second feature vectors of the two positive samples can be obtained; and by classifying the object description text in the target sample according to the first feature vector and the second feature vector respectively, two classification results can be obtained, which is equivalent to obtaining the positions of two semantically similar positive samples in the representation space. Then, through the contrastive learning strategy, in the direction of reducing the distance between the two classification results, the model parameters of the feature encoding model are optimized, so that after the feature encoding model performs feature encoding on two semantically similar samples, the two feature vectors obtained are used for classification processing respectively, and the two classification results obtained will be more consistent, thereby reducing the probability that the classification results point to multiple text categories at the same time or belong to multiple text categories with similar probabilities, so as to improve the accuracy of text classification.

[0168] The present application also provides a computer storage medium, in which one or more computer programs corresponding to the above multimedia processing method are stored. When one or more processors load and execute the one or more computer programs, the description of the method for training a text-based model in the embodiments can be implemented, which will not be elaborated here. The description of the beneficial effects of using the same method will not be elaborated here. It can be understood that the computer program can be deployed on one or more devices capable of communicating with each other for execution.

[0169] It should be noted that according to one aspect of the present application, a computer program product or a computer program is further provided. The computer program product includes a computer program, and the computer program is stored in a computer storage medium. The processor in the computer device reads the computer program from the computer storage medium and then executes the computer program, so that the computer device can execute the above-mentioned Figure 3 and Figure 6 methods provided in various alternative ways in the embodiments of the text-based model training method shown.

[0170] It can be understood that in the specific implementation manner of the present application, some embodiments involve data related to user information such as search click logs, historical object click behaviors, and historical object search behaviors; therefore, when the above embodiments of the present application are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.

[0171] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned data processing methods. Among them, the computer storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.

[0172] The above-disclosed are only partial embodiments of the present application. Of course, the scope of rights of the present application cannot be limited thereby. Those of ordinary skill in the art can understand all or part of the processes of the above embodiments and the equivalent changes made according to the claims of the present application still fall within the scope covered by the invention.

Claims

1. A text-based model training method, characterized in that, Including: Obtain a target sample, where the target sample includes an object search text and an object description text; the object description text is a pre-stored text for describing an object, and the object described by the object description text is the object output by providing feedback on the object search text; Call a feature encoding model to perform first feature perturbation encoding on each text in the target sample to obtain a first feature vector; and call the feature encoding model to perform second feature perturbation encoding on each text in the target sample to obtain a second feature vector; the first feature vector represents the cross-semantics between the texts after the first feature perturbation encoding in the target sample, and the second feature vector represents the cross-semantics between the texts after the second feature perturbation encoding in the target sample; Classify the object description text in the target sample respectively according to the first feature vector and the second feature vector to obtain two classification results; any classification result includes: the probability that the object description text in the target sample belongs to at least one text category, and the text category is used to indicate the correlation level corresponding to the correlation between the object description text and the object search text; Obtain a target loss function constructed based on a contrast learning strategy; Call the target loss function to perform a loss value operation according to the two classification results to obtain a target loss value; Obtain the category label of the object description text in the target sample, and the category label is used to indicate the correlation level between the object description text and the object search text in the target sample; Call a cross-entropy loss function to calculate a cross-entropy loss value according to the category label and the two classification results; Integrate the cross-entropy loss value and the target loss value to obtain an integrated loss value; and optimize the model parameters of the feature encoding model in the direction of reducing the integrated loss value.

2. The method according to claim 1, wherein The step of calling the feature encoding model to perform first feature perturbation encoding on each text in the target sample to obtain a first feature vector includes: Call the feature encoding model to perform feature representation on each text in the target sample respectively according to a first feature perturbation strategy to obtain a first vector representation of each text; Perform an interaction process on the first vector representations of each text to obtain cross-features between the texts in the target sample; Integrate the first vector representations of each text and the cross-features between the texts to obtain a first feature vector.

3. The method according to claim 2, characterized in that, The step of performing an interaction process on the first vector representations of each text to obtain cross-features between the texts in the target sample includes: Perform a difference operation on the first vector representations of each text to obtain a difference operation result; and perform a sum operation on the first vector representations of each text to obtain a sum operation result; Construct cross-features between the texts in the target sample by using the difference operation result and the sum operation result.

4. The method according to claim 2, wherein The feature encoding model includes a first feature representation network and a second feature representation network; the first vector representation of the object search text in the target sample is obtained through the first feature representation network, and the first vector representation of the object description text in the target sample is obtained through the second feature representation network; Among them, the method of calling any one of the feature representation networks to perform feature representation on the corresponding text according to the first feature perturbation strategy to obtain the corresponding first vector representation includes: Calling the any one of the feature representation networks to extract features from the corresponding text to obtain the text features of the corresponding text; Discarding some features in the text features according to the first random parameter in the first feature perturbation strategy; Performing pooling processing on the remaining features in the text features that are not discarded to obtain the first vector representation of the corresponding text.

5. The method according to claim 1, wherein The obtaining of the target sample includes: Obtaining a sample sequence for model optimization, where the sample sequence includes a plurality of samples arranged in sequence; traversing the sample sequence in turn and determining the currently traversed sample as the target sample; Among them, the sample sequence is constructed as follows: Obtaining a search click log, which is generated according to historical object search behaviors and historical object click behaviors; Parsing the search click log to obtain a parsing result; the parsing result includes: at least one object search text input historically, one or more object description texts corresponding to each object search text, and the click-through rate of the object described by each object description text; Constructing a plurality of samples according to each object search text and each object description text in the parsing result; one sample includes: one object search text and the corresponding one object description text; Sorting the plurality of samples according to the click-through rate of the object described by the object description text in each sample in descending order of the click-through rate to obtain a sample sequence.

6. The method according to claim 5, wherein The method further includes: Determining a plurality of click-through rate intervals, and one click-through rate interval corresponds to one text category; For any sample, using the click-through rate of the object described by the object description text in the any sample to perform hit matching on the plurality of click-through rate intervals; Constructing the category label of the object description text in the any sample according to the text category corresponding to the hit click-through rate interval.

7. The method according to claim 1, characterized in that After obtaining the optimized feature encoding model, the method further includes: If a target object search text is received, obtaining the target object description texts of each object for providing feedback on the target object search text; Constructing a plurality of text pairs using the target object search text and the obtained target object description texts; one text pair includes the target object search text and a target object description text; Calling the optimized feature encoding model to perform feature encoding on each text in each text pair to obtain the target feature vectors of each text pair; Classifying the target object description texts in each text pair respectively according to the target feature vectors of each text pair to obtain the classification results of each target object description text; Based on the classification results of the description texts of the respective target objects, determine the output order of the objects described by the description texts of the respective target objects, and output the respective objects in the corresponding output order.

8. A text-based model training device, characterized in that, The text-based model training apparatus includes an acquisition unit and a processing unit, where: The acquisition unit is configured to acquire a target sample, where the target sample includes an object search text and an object description text; the object description text is a pre-stored text for describing an object, and the object described by the object description text is the object output by giving feedback on the object search text; The processing unit is configured to call a feature encoding model to perform first feature perturbation encoding on each text in the target sample to obtain a first feature vector; and call the feature encoding model to perform second feature perturbation encoding on each text in the target sample to obtain a second feature vector; the first feature vector represents the cross-semantics between the texts in the target sample after the first feature perturbation encoding, and the second feature vector represents the cross-semantics between the texts in the target sample after the second feature perturbation encoding; The processing unit is further configured to perform classification processing on the object description text in the target sample respectively according to the first feature vector and the second feature vector to obtain two classification results; any one of the classification results includes: the probability that the object description text in the target sample belongs to at least one text category, and the text category is used to indicate the correlation gear corresponding to the correlation between the object description text and the object search text; The processing unit is further configured to optimize the model parameters of the feature encoding model according to the two classification results according to a contrast learning strategy; When the processing unit optimizes the model parameters of the feature encoding model according to the two classification results according to a contrast learning strategy, it is specifically configured to: Obtain a target loss function constructed according to a contrast learning strategy; Call the target loss function to perform a loss value operation according to the two classification results to obtain a target loss value; Obtain the category label of the object description text in the target sample, where the category label is used to indicate the correlation gear between the object description text and the object search text in the target sample; Call a cross-entropy loss function to calculate a cross-entropy loss value according to the category label and the two classification results; Integrate the cross-entropy loss value and the target loss value to obtain an integrated loss value; and optimize the model parameters of the feature encoding model in the direction of reducing the integrated loss value.

9. A computer device, characterized in that, Including: A processor, where the processor is adapted to implement one or more computer programs; A computer storage medium storing one or more computer programs, where the one or more computer programs are adapted to be loaded and executed by the processor to perform the text-based model training method according to any one of claims 1-7.

10. A computer storage medium, characterized in that, The computer storage medium stores one or more computer programs, where the one or more computer programs are adapted to be loaded and executed by the processor to perform the text-based model training method according to any one of claims 1-7.

11. A computer program product, characterized in that, The computer program product includes a computer program, and the computer program is adapted to be loaded and executed by a processor to perform the text-based model training method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Search result acquiring method and device and storage medium

    CN112115347A

  • Intention classification method and device, electronic equipment and computer readable storage medium

    CN113792818A

  • Semantic data processing method and device and semantic data searching method and device

    CN113869060A