Text processing method and device, equipment, storage medium and program product
By obtaining auxiliary text similar to the text to be classified from the preset text library and classifying texts with auxiliary text features, the problem of degradation of classification accuracy caused by the single feature of text to be classified is solved, and a higher text classification accuracy is achieved.
Patent Information
- Application Number
- CN202311548087.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-17
- Publication Date
- 2025-05-27
AI Technical Summary
During the text classification process, the classification accuracy decreases due to the single characteristics of the text to be classified.
Get M auxiliary texts that are most similar to the text to be classified from the preset text library, and combine them with the text to be classified, perform feature extraction and text classification, and combine auxiliary text features and category features to improve classification accuracy.
By utilizing the characteristics of auxiliary text, the accuracy of text classification is effectively improved and the text to be classified can be described more accurately.
Smart Images

Figure CN120045647A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to text processing technology in the field of computer applications, and in particular to a text processing method, device, equipment, storage medium and program product. Background Art
[0002] In the process of text processing, text classification is often involved; generally speaking, in order to achieve text classification, text classification is usually performed directly based on the features of the text to be classified; since the features of the text to be classified are single, when text classification is performed based on the features of the text to be classified, the accuracy of text classification is affected. Summary of the invention
[0003] The embodiments of the present application provide a text processing method, apparatus, device, storage medium and program product, which can improve the accuracy of text classification.
[0004] The technical solution of the embodiment of the present application is implemented as follows:
[0005] The present application provides a text processing method, the method comprising:
[0006] From the preset text library, obtain M auxiliary texts that are most similar to the text to be classified, where M is a positive integer;
[0007] Combining the text to be classified with the M auxiliary texts respectively to obtain M texts to be processed;
[0008] Perform feature extraction on the M texts to be processed respectively to obtain M features to be processed;
[0009] Combining the M features to be processed and the C category features to perform text classification, obtaining a text classification score, wherein the category feature is determined by a sub-support set corresponding to a text category, and the text classification score represents the C category scores of the text to be classified belonging to the C text categories, where C is a positive integer;
[0010] Based on the text classification score, the target text category to which the to-be-classified text belongs is determined from the C text categories.
[0011] The present application provides a text processing device, the text processing device comprising:
[0012] The text acquisition module is used to acquire M auxiliary texts that are most similar to the text to be classified from a preset text library, where M is a positive integer;
[0013] A text combination module, used for combining the text to be classified with the M auxiliary texts respectively to obtain M texts to be processed;
[0014] A feature extraction module is used to extract features from the M texts to be processed respectively to obtain M features to be processed;
[0015] A text classification module, used for combining the M features to be processed and C category features to perform text classification to obtain a text classification score, wherein the category feature is determined by a sub-support set corresponding to a text category, and the text classification score represents the C category scores of the text to be classified belonging to the C text categories, where C is a positive integer;
[0016] A category determination module is used to determine the target text category to which the to-be-classified text belongs from the C text categories based on the text classification score.
[0017] In an embodiment of the present application, the text acquisition module is also used to perform feature extraction on the text to be classified to obtain features of the text to be classified; for the preset text features of each preset text in the preset text library, obtain the feature correlation between the preset text features and the features of the text to be classified; determine the feature correlation as the text similarity between the text to be classified and the preset text; and determine the M preset texts corresponding to the M largest text similarities in the preset text library as the M auxiliary texts most similar to the text to be classified.
[0018] In an embodiment of the present application, the feature extraction module is also used to perform the following processing on the sub-support set corresponding to each of the C text categories: extract at least one supporting text feature corresponding to at least one supporting text in the sub-support set; integrate at least one supporting text feature to obtain the category feature; and obtain C category features corresponding to the C text categories from the category features corresponding to each of the text categories.
[0019] In an embodiment of the present application, the text classification module is also used to traverse the M features to be processed and perform the following processing on each of the traversed features to be processed: combine the features to be processed and the C category features to perform text classification to obtain a sub-text classification score, wherein the sub-text classification score represents C auxiliary scores that the text to be processed belongs to the C text categories; obtain M sub-text classification scores corresponding to the M features to be processed from the sub-text classification scores corresponding to each of the features to be processed; and integrate the M sub-text classification scores into the text classification score.
[0020] In an embodiment of the present application, the text classification module is also used to obtain C feature similarities corresponding to the feature to be processed and the C category features; fuse the C feature similarities with the feature to be processed to obtain a target classification feature; and perform text classification based on the target classification feature to obtain the sub-text classification score.
[0021] In an embodiment of the present application, the text classification module is also used to obtain M auxiliary scores corresponding to each text category from the M sub-text classification scores; select the maximum auxiliary score from the M auxiliary scores; and determine the C maximum auxiliary scores corresponding to the C text categories as the text classification scores.
[0022] In an embodiment of the present application, the text classification module is also used to obtain M auxiliary weights corresponding to the M features to be processed, wherein the auxiliary weights are C classification weights corresponding to the features to be processed on the C text categories; from the M auxiliary weights, M classification weights corresponding to each text category are obtained; the M classification weights are fused with the M auxiliary scores in a one-to-one correspondence to obtain the category score; and the C category scores corresponding to the C text categories are determined as the text classification score.
[0023] In an embodiment of the present application, the text classification module is further used to obtain the average score of the M auxiliary scores to obtain C average scores corresponding to the C text categories; and determine the C average scores as the text classification score.
[0024] In an embodiment of the present application, the text classification is implemented through a text classification model, and the text processing device also includes a model training module, which is used to obtain M auxiliary samples that are most similar to the text sample from the preset text library; the text sample is combined with the M auxiliary samples respectively to obtain M samples to be processed; feature extraction is performed on the M samples to be processed respectively to obtain M features of the samples to be processed; a first model to be trained is used to perform text classification on the M features of the samples to be processed and C category features to obtain an estimated classification score, wherein the first model to be trained is a neural network model to be trained for classifying text; the first model to be trained is trained in combination with the estimated classification score and the classification label of the text sample to obtain the text classification model.
[0025] In an embodiment of the present application, the model training module is also used to select a category estimated score corresponding to the classification label of the text sample from the estimated classification scores; normalize the category estimated score to obtain a classification probability; obtain a classification probability set corresponding to a text sample set from the classification probability of the text sample, wherein the text sample is any text in the text sample set; integrate the classification probability set to obtain a classification loss function value; and train the first model to be trained based on the classification loss function value to obtain the text classification model.
[0026] In an embodiment of the present application, the M auxiliary texts are obtained through a retrieval model, and the model training module is also used to determine the retrieval loss function value based on the difference between the M sample similarities and the M auxiliary training weights, wherein the sample similarity represents the similarity between the auxiliary sample and the text sample, and the auxiliary training weight represents the C weights corresponding to the C text categories determined by the features of the sample to be processed; the second model to be trained is trained based on the retrieval loss function value to obtain the retrieval model, wherein the M auxiliary samples are obtained through the second model to be trained, and the second model to be trained is a neural network model to be trained for obtaining the text most similar to the target text.
[0027] An embodiment of the present application provides an electronic device for text processing, the electronic device comprising:
[0028] A memory for storing computer executable instructions or computer programs;
[0029] The processor is used to implement the text processing method provided in the embodiment of the present application when executing the computer executable instructions or computer programs stored in the memory.
[0030] An embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions or a computer program. When the computer-executable instructions or the computer program are executed by a processor, the text processing method provided in the embodiment of the present application is implemented.
[0031] An embodiment of the present application provides a computer program product, including computer executable instructions or a computer program. When the computer executable instructions or the computer program are executed by a processor, the text processing method provided by the embodiment of the present application is implemented.
[0032] The embodiments of the present application have at least the following beneficial effects: when performing text classification on a text to be classified, first retrieve M auxiliary texts that are most similar to the text to be classified from a preset text library, and then use the M auxiliary texts as non-parametric knowledge to assist in determining C classification scores of the text to be classified corresponding to C text categories; this enables effective use of non-parametric knowledge in the text classification process, thereby achieving accurate description of the text to be classified, thereby improving the accuracy of text classification. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 is an exemplary schematic diagram of model training;
[0034] Figure 2 is another exemplary model training schematic diagram;
[0035] Figure 3 It is a schematic diagram of the architecture of the text processing system provided in the embodiment of the present application;
[0036] Figure 4 This is a method provided by the embodiment of the present application. Figure 3 A schematic diagram of the structure of the terminal in FIG.
[0037] Figure 5 This is a flowchart of the text processing method provided in the embodiment of the present application. Figure 1 ;
[0038] Figure 6 This is a flowchart of the text processing method provided in the embodiment of the present application. Figure 2 ;
[0039] Figure 7 This is a flowchart of the text processing method provided in the embodiment of the present application. Figure 3 ;
[0040] Figure 8 It is a flowchart of the model training provided in the embodiment of the present application;
[0041] Fig. 9 is an exemplary model training schematic diagram provided in an embodiment of the present application;
[0042] Fig.10 is an exemplary attention processing schematic diagram provided in an embodiment of the present application;
[0043] Fig.11 It is an exemplary performance change schematic diagram provided in an embodiment of the present application. DETAILED DESCRIPTION
[0044] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the accompanying drawings. The described embodiments should not be regarded as limiting the present application. All other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of this application.
[0045] In the following description, reference is made to “some embodiments”, which describe a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0046] In the following description, the terms "first\second" are used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0047] Unless otherwise defined, all technical and scientific terms used in the embodiments of the present application have the same meanings as those commonly understood by those skilled in the art. The terms used in the embodiments of the present application are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.
[0048] Before further describing the embodiments of the present application in detail, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations.
[0049] 1) Artificial Intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. This application embodiment describes the application of AI in the field of text classification.
[0050] 2) Machine Learning (ML) is a multi-disciplinary subject involving probability theory, statistics, approximation theory, convex analysis, and algorithmic complexity theory. It is used to study how computers simulate or implement human learning behaviors to acquire new knowledge or skills; and to reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Machine learning applications are spread across all areas of artificial intelligence. Machine learning usually includes technologies such as artificial neural networks, meta-learning, belief networks, reinforcement learning, transfer learning, and inductive learning.
[0051] 3) Artificial neural network is a mathematical model that imitates the structure and function of biological neural networks. The exemplary structures of artificial neural networks in the embodiments of the present application include graph convolutional networks (Graph Convolutional Network, GCN, a neural network for processing graph-structured data), deep neural networks (Deep Neural Networks, DNN), convolutional neural networks (Convolutional Neural Network, CNN) and recurrent neural networks (Recurrent Neural Network, RNN), neural state machines (Neural State Machine, NSM) and phase function neural networks (Phase-Functioned Neural Network, PFNN), etc. The first model to be trained, the second model to be trained, the text classification model and the retrieval model involved in the embodiments of the present application are all models corresponding to artificial neural networks (referred to as neural network models).
[0052] 4) N-sample (Shot), which means that in the model training process, there are N corresponding text samples for a text category, and N category labels corresponding to the N text samples, that is, the text category recognition task is trained by N text samples; N is, for example, 0, 1, 5, etc. The support set of the embodiment of the present application is a data set of N-sample mode.
[0053] 5) Meta-learning refers to first training a hyperparameter through a learning task, and then using the hyperparameter to adjust the parameters for another learning task. In addition, the training unit of meta-learning is a learning task, which usually includes a training task and a test task, wherein the training task includes multiple subtasks for learning hyperparameters, and the test task is a specific task for training based on the learned hyperparameters to determine the model parameters; and the data of each subtask in the training task is divided into a support set and a query set, and the data of the test task is a training set and a test set. The text samples in the embodiments of the present application can be texts in the query set or texts in the training set; and the text to be classified in the embodiments of the present application can be texts in the test set or texts in the application process.
[0054] Generally speaking, in order to achieve text classification, text classification is usually performed directly based on the features of the text to be classified; since the features of the text to be classified are single, the accuracy of text classification is affected when the text is classified based on the features of the text to be classified. Here, the features of the text to be classified can be extracted through a neural network model, and the extracted features can be used for text classification; wherein the neural network model is pre-trained based on the support set and the knowledge base (KB), however, the amount of data of the text samples used to train the neural network model is limited, which affects the generalization ability of the neural network model, and further affects the accuracy of text classification.
[0055] It should be noted that in order to train the neural network model, the support set and the knowledge base are usually used to generate task-related network parameters, and the task-related network performs text classification on the text samples in the query set based on the generated network parameters.
[0056] For example, see Figure 1 , Figure 1 is an exemplary model training diagram; Figure 1 As shown, the text samples 1-31 in the query set 1-21 and the text samples in the support set 1-22 (at least one text sample 1-321, at least one text sample 1-322 and at least one text sample 1-323 corresponding to three text categories are exemplarily shown) are processed by the embedding model 1-1 to obtain the corresponding embedding representations (the embedding representation 1-41 corresponding to the text sample 1-31, the embedding representation 1-421 corresponding to the at least one text sample 1-321, the embedding representation 1-422 corresponding to the at least one text sample 1-322, and the embedding representation 1-423 corresponding to the at least one text sample 1-323); the obtained embedding representations are fused to obtain the features 1-51 corresponding to the embedding representations 1-41 and 1-421, and the features 1-52 corresponding to the embedding representations 1-41 and 1-421. The feature 1-52 corresponding to the embedded representation 1-422 and the feature 1-53 corresponding to the embedded representation 1-41 and the embedded representation 1-423 are obtained; in addition, knowledge retrieval is performed in the knowledge graph 1-61 (referred to as the knowledge base) based on the support set 1-22 to obtain matching information 1-62, and then the relationship network parameters of the task-related network 1-71 are determined based on the parameterized network 1-63 and the matching information 1-62; finally, the task-independent network 1-72 and the task-related network 1-71 are used to calculate the category scores of the features 1-51 to 1-53 respectively to obtain scores 1-81 and 1-82; the scores 1-81 and 1-82 are combined to obtain the final score 1-83; and the loss function value for training the task-independent network 1-72 and the task-related network 1-71 is determined based on the final score 1-83 and the label vector 1-9.
[0057] It should be noted that Figure 1 The model training process shown uses the feature representation of the knowledge base to generate task-related network parameters; then, the task-related network uses the generated network parameters to obtain the category score of the text. However, due to the limitations of the features extracted by the parameterized neural network, using only the parameterized network for text classification affects the accuracy of text classification.
[0058] In addition, in order to use the neural network model for text classification, the characteristics of the string of the query text can also be enhanced. For example, see Figure 2 , Figure 2 is another exemplary model training diagram; Figure 2 As shown, the feature 2-21 of the query text 2-1 is first obtained through the embedding model 2-3, the knowledge graph 2-5 related to each character string in the query text 2-1 is extracted from the knowledge base 2-4, and the graph embedding representation of the knowledge graph 2-5 is learned through the graph neural network (GNN) 2-6 to obtain the additional feature representation 2-22 of each character string; and the feature 2-21 and the additional feature representation 2-22 are aggregated through the aggregation model 2-81, and then the classification result 2-9 corresponding to the aggregation result is obtained through the adaptation layer 2-82.
[0059] It should be noted that Figure 2 Although the model training process shown enhances the features of the query text 2-1, the implicit use of additional information for text classification still cannot accurately represent the features of the query text, affecting the accuracy of text classification.
[0060] Based on this, the embodiments of the present application provide a text classification method, apparatus, device, computer-readable storage medium and computer program product, which can improve the accuracy of text classification. The following describes an exemplary application of the text classification device provided in the embodiments of the present application. The text classification device provided in the embodiments of the present application can be implemented as various types of terminals such as smart phones, smart watches, laptops, tablet computers, desktop computers, smart home appliances, set-top boxes, smart car devices, portable music players, personal digital assistants, dedicated messaging devices, intelligent voice interaction devices, portable game devices and smart speakers. It can also be implemented as a server, and can also be implemented as a terminal and a server. Below, an exemplary application of the text processing device when it is implemented as a terminal will be described.
[0061] See also Figure 3 , Figure 3 Schematic diagram of the text processing system provided in the embodiment of the present application; Figure 3As shown, to support a text processing application, in the text processing system 100, the terminal 400 (terminal 400-1 and terminal 400-2 are shown as examples) is connected to the server 200 via the network 300; the network 300 can be a wide area network or a local area network, or a combination of the two; the server 200 is used to provide computing services to the terminal 400. For example, when the terminal 400 uses a retrieval model to obtain M auxiliary texts and uses a text classification model to classify text, the server 200 is used to train the retrieval model and the text classification model, and provide the functional services included in the retrieval model and the text classification model to the terminal 400 via the network 300. In addition, the text processing system 100 also includes a database 500 for providing data support to the server 200; and, Figure 3 What is shown in FIG. 5 is a case where the database 500 is independent of the server 200. In addition, the database 500 may also be integrated in the server 200, which is not limited in the embodiment of the present application.
[0062] Terminal 400 is used to respond to a text classification request for a text to be classified, obtain M auxiliary texts that are most similar to the text to be classified from a preset text library; combine the text to be classified with the M auxiliary texts respectively to obtain M texts to be processed; perform feature extraction on the M texts to be processed respectively to obtain M features to be processed; perform text classification based on the M features to be processed and C category features to obtain a text classification score; based on the text classification score, determine a target text category to which the text to be classified belongs from the C text categories, and display the text to be classified based on the target text category (graphic interface 410-1 and graphical interface 410-2 are shown as examples).
[0063] In some embodiments, the server 200 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms. The terminal and the server may be directly or indirectly connected via wired or wireless communication, which is not limited in the embodiments of the present application.
[0064] See also Figure 4 , Figure 4 This is a method provided by the embodiment of the present application. Figure 3 The schematic diagram of the terminal structure in FIG. Figure 4As shown, the terminal 400 includes: at least one processor 410, a memory 450, at least one network interface 420 and a user interface 430. The various components in the terminal 400 are coupled together through a bus system 440. It can be understood that the bus system 440 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 440 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, Figure 4 Various buses are labeled as bus system 440 .
[0065] The processor 410 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor, etc.
[0066] The user interface 430 includes one or more output devices 431 that enable presentation of media content, including one or more speakers and / or one or more visual display screens. The user interface 430 also includes one or more input devices 432, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.
[0067] The memory 450 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard disk drives, optical disk drives, etc. The memory 450 may optionally include one or more storage devices that are physically remote from the processor 410.
[0068] The memory 450 includes a volatile memory or a non-volatile memory, and may also include both volatile and non-volatile memories. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 450 described in the embodiments of the present application is intended to include any suitable type of memory.
[0069] In some embodiments, memory 450 can store data to support various operations, examples of which include programs, modules, and data structures, or a subset or superset thereof, as exemplarily described below.
[0070] Operating system 451, including system programs for processing various basic system services and performing hardware-related tasks, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks;
[0071] A network communication module 452, for reaching other electronic devices via one or more (wired or wireless) network interfaces 420, exemplary network interfaces 420 include: Bluetooth, Wireless Fidelity (Wi-Fi), and Universal Serial Bus (USB), etc.;
[0072] a presentation module 453 for enabling presentation of information via one or more output devices 431 (e.g., display screen, speaker, etc.) associated with the user interface 430 (e.g., a user interface for operating peripherals and displaying content and information);
[0073] The input processing module 454 is used to detect one or more user inputs or interactions from one of the one or more input devices 432 and translate the detected inputs or interactions.
[0074] In some embodiments, the text processing device provided in the embodiments of the present application can be implemented in software. Figure 4 The text processing device 455 stored in the memory 450 is shown, which can be software in the form of a program and a plug-in, etc., including the following software modules: a text acquisition module 4551, a text combination module 4552, a feature extraction module 4553, a text classification module 4554, a category determination module 4555 and a model training module 4556. These modules are logical, so they can be arbitrarily combined or further split according to the functions implemented. The functions of each module will be explained below.
[0075] In some embodiments, the text processing device provided in the embodiments of the present application can be implemented in hardware. As an example, the text processing device provided in the embodiments of the present application can be a processor in the form of a hardware decoding processor, which is programmed to execute the text processing method provided in the embodiments of the present application. For example, the processor in the form of a hardware decoding processor can adopt one or more application specific integrated circuits (Application Specific Integrated Circuit, ASIC), DSP, programmable logic device (Programmable Logic Device, PLD), complex programmable logic device (Complex Programmable Logic Device, CPLD), field programmable gate array (Field-Programmable Gate Array, FPGA) or other electronic components.
[0076] In some embodiments, the terminal or server can implement the text processing method provided by the embodiment of the present application by running various computer executable instructions or computer programs. For example, computer executable instructions can be commands, machine instructions or software instructions at the microprogram level. The computer program can be a native program or software module in the operating system; it can be a local (Native) application (APPlication, APP), that is, a program that needs to be installed in the operating system to run, such as a text processing APP; it can also be a small program that can be embedded in any APP, that is, a program that can be run only by downloading it to a browser environment. In short, the above-mentioned computer executable instructions can be instructions in any form, and the above-mentioned computer program can be an application, module or plug-in in any form.
[0077] Below, the text processing method provided by the embodiment of the present application will be described in combination with the exemplary application and implementation of the text processing device provided by the embodiment of the present application. In addition, the text processing method provided by the embodiment of the present application is applied to various text classification scenarios such as cloud technology, artificial intelligence, smart transportation, resource transfer, maps and vehicle-mounted.
[0078] See also Figure 5 , Figure 5 This is a flowchart of the text processing method provided in the embodiment of the present application. Figure 1 ,in, Figure 5 The execution subject of each step in the text processing device is the text processing device; Figure 5 The steps shown are explained.
[0079] Step 101: Obtain M auxiliary texts that are most similar to the text to be classified from a preset text library.
[0080] In an embodiment of the present application, the text processing device can obtain a preset text library from its own storage module or other devices (for example, storage devices such as databases), and the preset text library includes various preset texts, each of which is a corpus text. When the text processing device responds to a text classification request and starts text classification processing, the text to be classified is the text to be classified, such as the name of the asset transfer account, etc. The text processing device first retrieves at least one preset text that is most similar to the text to be classified from the preset text library, and calls the retrieved at least one preset text that is most similar to the text to be classified M auxiliary texts; thus, the auxiliary text is one of the M preset texts in the preset text library that are most similar to the text to be classified. Wherein, M is a positive integer, such as 5, 8, etc.
[0081] It should be noted that the preset text library is a text library composed of various corpus texts; for example, a news title library, a search database, an organization name library, a combination of the above, etc.; the business type corresponding to the text to be classified may belong to the business type included in the preset text library, or it may be different from the business type included in the preset text library, and the embodiment of the present application does not limit this. In addition, the auxiliary text and the text to be classified can be of the same text granularity, for example, they are both strings, sentences, paragraphs, or articles, etc.; the auxiliary text and the text to be classified can also be of different text granularities. At this time, in order to enhance the auxiliary role of the auxiliary text, the text granularity level of the auxiliary text can be equal to or greater than the text granularity level of the text to be classified; for example, when the text granularity level of the text to be classified is a sentence, the text granularity level of the text to be classified can be a sentence or a paragraph or an article; of course, the text granularity level of the auxiliary text can also be a fixed text granularity, for example, fixed to a paragraph.
[0082] It should also be noted that the text processing device can obtain M auxiliary texts based on the similarity of feature dimensions; thus, in the preset text library, the features of the M auxiliary texts are most similar to the features of the text to be classified.
[0083] See also Figure 6 , Figure 6 This is a flowchart of the text processing method provided in the embodiment of the present application. Figure 2 ,in, Figure 6 The execution subject of each step in is the text processing device; Figure 6 As shown, in the embodiment of the present application, step 101 can be implemented through steps 1011 to 1014; that is, the text processing device obtains M auxiliary texts that are most similar to the text to be classified from the preset text library, including steps 1011 to 1014, and each step is explained below.
[0084] Step 1011: extract features of the text to be classified to obtain features of the text to be classified.
[0085] In the embodiment of the present application, the text processing device extracts features from the text to be classified, and the extracted features are the features of the text to be classified; thus, the features of the text to be classified represent the features of the text to be classified.
[0086] Step 1012: for each preset text feature in the preset text library, obtain the feature correlation between the preset text feature and the feature of the text to be classified.
[0087] It should be noted that the preset text library can store preset text features corresponding to each preset text, and the text processing device can also extract features of each preset text in real time. The extracted features are preset text features, and the embodiments of the present application do not limit this; thus, the text processing device can obtain the preset text features of each preset text in the preset text library, and then can compare the correlation between the preset text features and the features of the text to be classified, and the correlation between the preset text features and the features of the text to be classified is called feature correlation. Here, the correlation between the preset text features and the features of the text to be classified is positively correlated with the feature correlation degree; the preset text features are the features of the preset text.
[0088] It can be understood that when the preset text library stores preset text features corresponding to each preset text, the text processing device can directly acquire feature correlations based on the stored preset text features, thereby simplifying the extraction process of the preset text features, thereby improving the acquisition efficiency of M auxiliary texts, and further improving the text classification efficiency.
[0089] Step 1013: determine the feature correlation as the text similarity between the text to be classified and the preset text.
[0090] It should be noted that, since the feature correlation represents the correlation between the preset text feature and the text feature to be classified, and the preset text feature is the feature of the preset text, and the text feature to be classified is the feature of the text to be classified, the text processing device can determine the feature correlation as the similarity between the text to be classified and the preset text; here, the similarity between the text to be classified and the preset text is called text similarity; it is easy to know that the degree of similarity between the text to be classified and the preset text is positively correlated with the text similarity.
[0091] Step 1014: determine the M preset texts corresponding to the M maximum text similarities in the preset text library as the M auxiliary texts most similar to the text to be classified.
[0092] In an embodiment of the present application, a text processing device selects preset texts from a preset text library based on text similarity; when selecting, the text processing device first determines M maximum text similarities, then selects M preset texts corresponding to the M maximum text similarities from the preset text library, and finally uses the selected M preset texts as M auxiliary texts.
[0093] Step 102: Combine the text to be classified with the M auxiliary texts respectively to obtain M texts to be processed.
[0094] In the embodiment of the present application, the text processing device combines the text to be classified with each of the M auxiliary texts into one text to be processed, so that M texts to be processed can be combined for the M auxiliary texts.
[0095] It should be noted that the text to be processed is the result of combining the text to be classified and the auxiliary text. The combination method can be to combine the text to be classified and the auxiliary text into one text, or to connect the text to be classified and the auxiliary text based on a specified mark (such as a start mark and a spacing mark, etc.), etc., which is not limited in the embodiment of the present application. Among them, the M texts to be processed correspond to the M auxiliary texts one by one.
[0096] Step 103: extract features from the M texts to be processed respectively to obtain M features to be processed.
[0097] In an embodiment of the present application, the text processing device performs feature extraction on each of the M texts to be processed, and the features extracted for each text to be processed are the features to be processed; thus, for the M texts to be processed, the text processing device can obtain M features to be processed.
[0098] It should be noted that the M texts to be processed correspond to the M features to be processed one by one. In addition, feature extraction can be implemented based on a specified embedding representation method, or can be implemented based on a neural network model for embedding representation, etc., which is not limited in the embodiments of the present application.
[0099] Step 104: Combine the M features to be processed and the C category features to perform text classification and obtain a text classification score.
[0100] In an embodiment of the present application, C text categories are preset for the classification result of the text; the text processing device can obtain the corresponding category feature for each text category, thereby, for the C text categories, C category features can be obtained; wherein the category feature is used to represent the feature corresponding to the text belonging to the text category. Here, the text processing device obtains the association relationship between each of the M features to be processed and the C category features to determine the C auxiliary scores of the text to be processed corresponding to each feature to be processed in the C text categories; thereby, M auxiliary score sets can be obtained for the M features to be processed, and the auxiliary score sets include C auxiliary scores; the M auxiliary score sets are integrated to obtain the category score for each text category, thereby obtaining C category scores corresponding to the C text categories, and the C category scores corresponding to the C text categories are the text classification scores.
[0101] It should be noted that the category feature is determined by the sub-support set corresponding to the text category, and the sub-support set refers to at least one text belonging to the text category (also called at least one supporting text); it is easy to know that C text categories correspond to C sub-support sets, and the C sub-support sets are the support sets of C text categories. The text classification score represents the C category scores of the text to be classified belonging to C text categories, that is, the text classification score is the C category scores of the text to be classified corresponding to C text categories, and each category score represents the possibility that the text to be classified belongs to the corresponding text category; wherein, the C text categories correspond to the C category scores one by one, and C is a positive integer (for example, 1, 3, etc.). In addition, the auxiliary score refers to the possibility that the text to be processed belongs to the corresponding text category.
[0102] It should also be noted that the category features can be obtained before the text to be classified is processed, or can be obtained in real time, or can be a combination of the above two, etc., which is not limited in the embodiments of the present application. In addition, the extraction of category features can be implemented based on a specified embedding representation method, or can be implemented based on a neural network model for embedding representation, etc., which is not limited in the embodiments of the present application.
[0103] It is understandable that when the category features corresponding to each text category are obtained in advance, the text processing device can directly classify the text based on the category features, which simplifies the step of extracting the category features, thereby improving the efficiency of text classification.
[0104] See also Figure 7 , Figure 7 This is a flowchart of the text processing method provided in the embodiment of the present application. Figure 3 ,in, Figure 7 The execution subject of each step in is the text processing device; Figure 7 As shown, in the embodiment of the present application, step 104 can be implemented through steps 1041 to 1043; that is, the text processing device combines M features to be processed and C category features to perform text classification to obtain a text classification score, including steps 1041 to 1043, and each step is explained below.
[0105] In the embodiment of the present application, the text processing device traverses M features to be processed and performs the following processing on each traversed feature to be processed (step 1041).
[0106] It should be noted that step 1041 describes obtaining, based on each feature to be processed, C auxiliary scores of the corresponding text to be processed belonging to C text categories.
[0107] Step 1041: Combine the feature to be processed and the C category features to perform text classification to obtain a sub-text classification score.
[0108] It should be noted that, since the feature to be processed is the feature of the text to be processed, and the C category features are the features corresponding to the C text categories respectively, the text processing device will compare the feature to be processed with the C category features respectively to determine the C auxiliary scores corresponding to the C text categories for the feature to be processed; and determine the C auxiliary scores corresponding to the C text categories as the sub-text classification scores of the traversed feature to be processed; thus, the sub-text classification scores represent the C auxiliary scores of the text to be processed belonging to the C text categories.
[0109] In an embodiment of the present application, the text processing device can directly obtain a sub-text classification score by comparing the feature to be processed with the C category features, and can also perform attention processing on the feature to be processed and the C category features, and then compare the feature to be processed after the attention processing with the C category features to obtain the sub-text classification score, which is not limited in the embodiment of the present application. When the text processing device combines attention processing to obtain a sub-text classification score, the text processing device combines the feature to be processed and the C category features to perform text classification to obtain a sub-text classification score, including: obtaining C feature similarities corresponding to the feature to be processed and the C category features; fusing the C feature similarities with the feature to be processed to obtain a target classification feature; and performing text classification based on the target classification feature to obtain a sub-text classification score.
[0110] It should be noted that the text processing device compares the feature to be processed with each of the C category features to obtain the similarity between the feature to be processed and each category feature, and the similarity between the feature to be processed and each category feature is called feature similarity; thus, the text processing device can obtain C feature similarities corresponding to the C category features, and the C category features correspond to the C feature similarities one by one. Here, the text processing device can perform a linear transformation on the feature to be processed and the C category features through the preset linear transformation parameters in the attention mechanism, and then determine the dot product of the transformation results of the two as the feature similarity.
[0111] It should also be noted that the more similar the feature to be processed is to the category feature, the higher the contribution of the feature to be processed to the text classification. Therefore, the feature similarity represents the weight of the feature to be processed in the text classification. Here, the text processing device can fuse each feature similarity with the feature to be processed (for example, dot product, etc.) to obtain C fusion results; then fuse the C fusion results (for example, summation, etc.) to obtain the target classification feature. Among them, the target classification feature refers to the feature obtained after the feature to be processed is subjected to attention processing in combination with the category feature.
[0112] In the embodiment of the present application, when the attention processing is single-head attention processing, the text processing device obtains the similarities of the feature to be processed and the C feature corresponding to the C category features based on the single-head attention parameters; the target classification feature is the processing result of the single-head attention processing on the feature to be processed. The single-head attention parameters include preset linear transformation parameters.
[0113] In the embodiment of the present application, when the attention processing is multi-head attention processing, the multi-head attention processing corresponds to H attention parameters; at this time, the text processing device obtains C feature similarities corresponding to the feature to be processed and the C category features respectively, including: the text processing device obtains the C feature similarities corresponding to the feature to be processed and the C category features respectively based on each attention parameter in the H attention parameters; thus, H feature similarity sets corresponding to the H attention parameters can be obtained, and the feature similarity set includes C feature similarities. Wherein, H is a positive integer greater than 1, and each attention parameter includes a preset linear transformation parameter.
[0114] Correspondingly, the text processing device fuses the C feature similarities with the features to be processed respectively to obtain the target classification feature, including: the text processing device fuses the C feature similarities with the features to be processed respectively to obtain the attention feature; obtains H attention features corresponding to the H attention parameters from the attention features corresponding to each attention parameter; and fuses the H attention features and the category features into the target classification feature.
[0115] It should be noted that the text processing device fuses each feature similarity with the feature to be processed to obtain C fusion results, and then fuses the C fusion results to obtain the attention feature. The attention feature refers to the feature obtained after performing an attention process on the feature to be processed in combination with the category feature.
[0116] It can be understood that by paying attention to the features to be processed and the category features, the features to be processed that have a greater auxiliary effect on text classification among the M texts to be processed can achieve semantic retention, thereby improving the accuracy of text classification.
[0117] Step 1042: Obtain M sub-text classification scores corresponding to the M features to be processed from the sub-text classification scores corresponding to each feature to be processed.
[0118] It should be noted that when the text processing device obtains the corresponding sub-text classification score for each traversed feature to be processed, M sub-text classification scores can be obtained for M features to be processed; it is easy to know that the M features to be processed correspond one-to-one to the M sub-text classification scores.
[0119] Step 1043: Integrate the M sub-text classification scores into a text classification score.
[0120] In the embodiment of the present application, the text processing device can integrate the M sub-text classification scores into the text classification score by at least one of the following methods: averaging, maximizing, and weighted summing. In addition, when integrating the M sub-text classification scores, the text processing device integrates them at the granularity of each text category to obtain the category score corresponding to each text category; thus, the text classification score includes C category scores corresponding to C text categories.
[0121] In an embodiment of the present application, the text processing device integrates M sub-text classification scores into a text classification score, including: the text processing device obtains M auxiliary scores corresponding to each text category from the M sub-text classification scores; selects the maximum auxiliary score from the M auxiliary scores; and determines the C maximum auxiliary scores corresponding to the C text categories as the text classification score.
[0122] It should be noted that the M auxiliary scores correspond to the M texts to be processed one by one; after the text processing device selects a maximum auxiliary score corresponding to the text category from the M auxiliary scores, the maximum auxiliary score is determined as the possibility that the text to be processed belongs to the text category; here, the text processing device can obtain C maximum auxiliary scores corresponding to C text categories from the maximum auxiliary score selected for each text category. In other words, when the text processing device integrates the M sub-text classification scores by finding the maximum value, the text classification score is the C maximum auxiliary scores corresponding to the C text categories.
[0123] In an embodiment of the present application, after the text processing device obtains M auxiliary scores corresponding to each text category from the M sub-text classification scores, the text processing method also includes: the text processing device obtains M auxiliary weights corresponding to the M features to be processed, wherein the auxiliary weights are C classification weights corresponding to the features to be processed on C text categories; from the M auxiliary weights, obtains M classification weights corresponding to each text category; performs one-to-one correspondence fusion of the M classification weights and the M auxiliary scores to obtain a category score; and determines the C category scores corresponding to the C text categories as the text classification score.
[0124] It should be noted that the classification weight may be a fusion result of the feature to be processed and a specified parameter. The specified parameter may be a parameter obtained through model training or a fixed parameter, which is not limited in the embodiments of the present application.
[0125] In the embodiment of the present application, after the text processing device obtains M auxiliary scores corresponding to each text category from the M sub-text classification scores, the text processing method further includes: the text processing device obtains the average score of the M auxiliary scores to obtain C average scores corresponding to C text categories; and determines the C average scores as the text classification score. At this time, the text classification score is the C average scores corresponding to the C text categories.
[0126] Step 105: Based on the text classification score, determine the target text category to which the to-be-classified text belongs from the C text categories.
[0127] In an embodiment of the present application, the text processing device can select at least one highest category score from the text classification scores, and select at least one text category corresponding to the at least one highest category score from C text categories, and finally use the selected at least one text category as the target text category.
[0128] It should be noted that after obtaining the target text category, the text processing device can also display the target text category of the text to be classified. For example, the text processing device stores the text to be classified and the target text category in correspondence, and in response to a text display request for the target text category, displays at least one text to be displayed including the text to be classified.
[0129] It can be understood that by retrieving M auxiliary texts that are most similar to the text to be classified from a preset text library, and using the M auxiliary texts as non-parametric knowledge, C classification scores of the text to be classified corresponding to C text categories are assisted in determining; in the text classification process, effective use of non-parametric knowledge is achieved, which can accurately describe the text to be classified, and therefore, the accuracy of text classification can be improved.
[0130] Before step 104 of the embodiment of the present application, the text processing device also includes a process of obtaining category features; that is, the text processing device combines M features to be processed and C category features to perform text classification, and before obtaining the text classification score, the text processing method also includes: the text processing device performs the following processing on the sub-support set corresponding to each text category in the C text categories: extracting at least one supporting text feature corresponding to at least one supporting text in the sub-support set; integrating at least one supporting text feature to obtain category features; and obtaining C category features corresponding to the C text categories from the category features corresponding to each text category.
[0131] It should be noted that the sub-support set corresponding to each text category includes at least one supporting text, and the text processing device extracts features from each supporting text, and the extracted results are supporting text features. Here, the text processing device can integrate at least one supporting text feature by using at least one of pooling, concatenation, and weighted summation.
[0132] In the embodiment of the present application, text classification can be implemented by a neural network model. The neural network model used for text classification is a text classification model; see Figure 8 , Figure 8 is a flow chart of the model training provided in the embodiment of the present application, wherein: Figure 8 The execution subject of each step in is the text processing device; Figure 8 As shown, the text classification model can be trained through steps 106 to 110, and each step is described below.
[0133] Step 106: Obtain M auxiliary samples that are most similar to the text sample from a preset text library.
[0134] It should be noted that the acquisition of M auxiliary samples is similar to the acquisition of M auxiliary texts, and the embodiments of the present application will not be described again here. Here, the number of auxiliary samples and the number of auxiliary texts obtained by the text processing device can be both M or different, and the embodiments of the present application do not limit this. Among them, the text sample is a sample to be classified and is text.
[0135] Step 107: Combine the text sample with the M auxiliary samples respectively to obtain M samples to be processed.
[0136] It should be noted that the process of combining the text sample and the M auxiliary samples by the text processing device is similar to the process of combining the text to be classified and the M auxiliary texts, and the embodiment of the present application will not be repeated here.
[0137] Step 108: extract features of the M samples to be processed respectively to obtain features of the M samples to be processed.
[0138] It should be noted that the process of the text processing device obtaining M features of samples to be processed is similar to the process of obtaining M features to be processed, and the embodiment of the present application will not be repeated here.
[0139] Step 109: Use the first model to be trained to perform text classification on the M sample features to be processed and the C category features to obtain an estimated classification score.
[0140] It should be noted that the first model to be trained is a neural network model to be trained for classifying text; it can be an original neural network model constructed, or it can be a pre-trained neural network model, etc., and the embodiments of the present application are not limited to this. In addition, the estimated classification score represents the estimated scores of C categories of the text sample belonging to C text categories. Since the process of text classification of M sample features to be processed and C category features is similar to the process of text classification of M features to be processed and C category features, the embodiments of the present application will not be repeated here.
[0141] Step 110: Train the first model to be trained by combining the estimated classification scores and the classification labels of the text samples to obtain a text classification model.
[0142] It should be noted that the classification label represents the text category to which the text sample actually belongs; the text processing device determines the loss function value based on the estimated category score corresponding to the classification label in the estimated classification score, and then, based on the determined loss function value, backpropagates in the first model to be trained to adjust the model parameters in the first model to be trained; in addition, the training of the first model to be trained can be performed iteratively, and when the iterative training is completed, the first model to be trained trained by the current iteration is the text classification model.
[0143] It should be noted that when the text processing device determines that the iterative training meets the first training end condition, the iterative training is determined to be ended; otherwise, the iterative training continues. The first training end condition may be reaching the first accuracy index threshold, or reaching the first iteration number threshold, or reaching the first iteration duration threshold, or a combination of the above, etc., which is not limited in the embodiments of the present application.
[0144] In an embodiment of the present application, a text processing device combines the estimated classification score and the classification label of the text sample to train a first model to be trained to obtain a text classification model, including: the text processing device selects a category estimated score corresponding to the classification label of the text sample from the estimated classification score; normalizes the category estimated score to obtain a classification probability; obtains a classification probability set corresponding to the text sample set from the classification probability of the text sample; integrates the classification probability set to obtain a classification loss function value; and trains the first model to be trained based on the classification loss function value to obtain a text classification model.
[0145] It should be noted that the text sample is any text in a text sample set. The text sample set can be a query set, a training set, or a combination of the two, etc. This embodiment of the present application does not limit this.
[0146] In an embodiment of the present application, M auxiliary texts can be obtained through a retrieval model, and the retrieval model is trained through the following steps: the text processing device determines the retrieval loss function value based on the difference between the M sample similarities and the M auxiliary training weights; and trains the second model to be trained based on the retrieval loss function value to obtain the retrieval model.
[0147] It should be noted that the sample similarity represents the similarity between the auxiliary sample and the text sample, and the auxiliary training weight represents the C weights corresponding to the C text categories determined by the features of the sample to be processed, wherein the M auxiliary samples are obtained through the second model to be trained. The second model to be trained is a neural network model to be trained for obtaining the text most similar to the target text, which can be the constructed original neural network model, or a pre-trained neural network model, etc., which is not limited in the embodiments of the present application; wherein the target text is any specified text, such as a text to be classified, a text sample, etc.
[0148] It should also be noted that the text processing device performs back propagation in the second model to be trained based on the retrieval loss function value to adjust the model parameters in the second model to be trained; in addition, the training of the second model to be trained can be iterative, and when the iterative training ends, the second model to be trained by the current iterative training is the retrieval model. When the text processing device determines that the iterative training meets the second training end condition, it determines that the iterative training ends; otherwise, iterative training continues. Here, the second training end condition can be reaching the second accuracy index threshold, or reaching the second iteration number threshold, or reaching the second iteration duration threshold, or a combination of the above, etc., which is not limited in the embodiments of the present application.
[0149] In an embodiment of the present application, when the text processing device is used to train a neural network model, the text processing device can be various servers; when the text processing device is used to perform text classification processing of text to be classified, the text processing device can be various servers or various terminals, and the embodiment of the present application does not limit this.
[0150] Below, an exemplary application of the embodiment of the present application in an actual application scenario will be described. The exemplary application describes the text classification process corresponding to the resource transfer object in the resource transfer scenario. It is easy to know that the text classification method provided by the embodiment of the present application is applicable to any text classification scenario. Here, the text classification corresponding to the resource transfer object in the resource transfer scenario is used as an example for description.
[0151] See also Fig. 9 , Fig. 9 is an exemplary model training schematic diagram provided in the embodiment of the present application; Fig. 9As shown, on the one hand, M auxiliary paragraphs (including auxiliary paragraphs 9-31 to auxiliary paragraphs 9-3M, called M auxiliary samples) corresponding to the query text 9-2 (called text samples) are retrieved from the knowledge base 9-1 (called the preset text base), and through the encoder 9-71, the query text 9-2 and the M auxiliary paragraphs are combined to obtain M query features (including query features 9-41 to query features 9-4M, called M sample features to be processed). On the other hand, for the sampled training tasks, the support sets (supporting texts 9-521 and 9-522 for category 9-511, supporting texts 9-523 and 9-524 for category 9-512, supporting texts 9-525 and 9-526 for category 9-513, where C is 3) corresponding to the C prototypes 9-52 (called C category features) are obtained in sequence through the encoder 9-72, the average pooling layer 9-73 and the fusion layer 9-74. Finally, the scores of the M query features and C prototypes 9-52 are predicted through the information fusion network 9-6 (called the first model to be trained), and M classification scores 9-7 corresponding to the M query features are obtained. The M classification scores 9-7 are pooled to obtain the final classification score 9-8 (called the estimated classification score); and, based on the final classification score 9-8 and the label category, the loss function value 9-9 (called the classification loss function value) is determined, and then, the information fusion network 9-6 is trained based on the loss function value 9-9.
[0152] Next, the process of retrieving M auxiliary paragraphs is explained.
[0153] When retrieving auxiliary paragraphs, the maximum inner product search (MIPS) method can be used to obtain the first M paragraphs with the greatest relevance to the query text from the knowledge base, thus obtaining M auxiliary paragraphs.
[0154] It should be noted that each paragraph text z in the knowledge base i The correlation f(x, z) with the query text x i ) can be determined by calculating the corresponding inner product, as shown in formula (1).
[0155] f(x,z i )=Embed input (x) T Embed aoc (z i ) (1);
[0156] Among them, Embed input() is used to obtain the vector embedding of the query text, T represents the transposition processing; Embed doc () belongs to the retriever (called the retrieval model), which is used to obtain the vector embedding of each paragraph text in the knowledge base; “·” represents the inner product processing.
[0157] The retriever is also used to determine the text z of each paragraph based on relevance i The retrieval probability p(z i |x), as shown in formula (2).
[0158]
[0159] Among them, z j Represents the knowledge base except paragraph text z i Any paragraph text other than j ) represents text text z j The correlation between the query text x and exp represents the exponential function.
[0160] It should be noted that the M auxiliary paragraphs are M f(x, z) selected from the knowledge base. i ) The largest paragraph text. Based on auxiliary paragraphs z i and query text x to determine the probability p(y|z i , x), and the retrieval probability p(z i |x), we can obtain the probability p(y|x) that the query text x belongs to category y, as shown in formula (3).
[0161]
[0162] Here, m represents a variable from 1 to M.
[0163] The following describes the process of obtaining the final classification score through the information fusion network.
[0164] It should be noted that, in order to achieve semantic retention, the embodiment of the present application uses an attention mechanism to process query features and C prototypes. First, the query Q, key K and value V in the cross-attention mechanism are calculated based on linear transformation, as shown in formula (4).
[0165] Q=PW Q , K = XW K , V = XW V (4);
[0166] Among them, W Q , W K and W V is the linear transformation parameter in the cross attention mechanism; P is the C prototype, that is, d represents the dimension; X is the query feature, (l represents the length of the query text and the auxiliary paragraph), each query feature is the embedding vector corresponding to the concatenation result of the query text and an auxiliary paragraph.
[0167] Then, calculate the cth Q in Q c , c∈[1,C] and K m , m∈[1,M] (when X is the query feature corresponding to the mth auxiliary paragraph and the query text, the corresponding key K is K m ) Among them, Q c is the query corresponding to the cth prototype in Q, K m is the key corresponding to the mth query feature; here, the similarity score is calculated by Q c and K m The dot product of is obtained by normalizing the dot product result, as shown in formula (5).
[0168]
[0169] Among them, m1∈[1,M].
[0170] It should be noted that if K m The more contributions to the classification results of the query text, and the classification label is Q c The corresponding prototype category, in the iterative training process, the similarity score The higher the , the more semantic preservation can be achieved.
[0171] Finally, based on the above semantic preservation processing, the classification scores S of the query text and the mth auxiliary paragraph relative to C categories are calculated. m , as shown in formula (6).
[0172]
[0173] Among them, r is a learnable parameter, It is obtained by formula (7), which is as follows.
[0174]
[0175] Among them, Z (called the target classification feature) is obtained by formula (8), which is as follows.
[0176] Z=Cross attn (P,X,X)=Concat(head 1 ,head 1 ,……,head h )W 0 (8);
[0177] Among them, Cross attn () indicates cross attention processing; Concat() indicates concatenation processing; is a learnable parameter; head h (called attention feature) is obtained by formula (9), which is as follows.
[0178]
[0179] Among them, h represents the h-th attention processing module, with a total of H; Softmax() is the normalization function.
[0180] For example, see Fig.10 , Fig.10 is an exemplary attention processing diagram provided in an embodiment of the present application; Fig.10 As shown, first, query feature 10-1 is transformed by linear transformation parameter 10-21 to obtain value 10-31; query feature 10-1 is transformed by linear transformation parameter 10-22 to obtain key 10-32; C prototypes 10-3 are transformed by linear transformation parameter 10-23 to obtain query 10-33. Next, the similarity score 10-4 between query 10-33 and key 10-32 is calculated. Then, the similarity score 10-4 is fused with value 10-31, and the residual is added to the fusion result (that is, it is spliced with C prototypes 10-3) to obtain the final feature 10-5. Finally, the final feature 10-5 obtains the classification score 10-7 through the learnable parameter 10-6.
[0181] It should be noted that when there are multiple attention processing modules, multiple final features 10-5 will be obtained; at this time, the multiple final features 10-5 are connected, and the connection results are used to obtain classification scores through learnable parameters.
[0182] It should be noted that S = [S 1 , S 2 , ..., S M ]∈(M*C) (called the classification scores of M subtexts), thus, S is pooled to obtain the final classification score; wherein, the pooling process can be achieved by aggregating S through three strategies to obtain the final classification score S pool (called text classification score); among them, the three strategies include averaging, maximum value and weighted summation.
[0183] When the final classification score is obtained by averaging S, it is shown in formula (10).
[0184]
[0185] When the final classification score is obtained by weighted summing of S, it is as shown in formula (11).
[0186]
[0187] in, is the score weight (called auxiliary weight) corresponding to the mth query feature, which is obtained by formula (12). Formula (12) is as follows.
[0188]
[0189] Among them, the initial weight β m As shown in formula (13).
[0190] β m =W proportion BERT CLS (join BERT (x,z m )) (13);
[0191] Among them, W proportion is a learnable parameter; BERT CLS Indicates the process of obtaining the embedded vector; join BERT (x,z m ) is obtained by formula (14), which is as follows.
[0192] join BERT (x,z m )=[CLS]x[SEP]z m [SEP] (14);
[0193] Among them, CLS represents the start character, and SEP represents the interval character.
[0194] It should be noted that for the query set Q T Each query text and the corresponding category label in (x, y) can obtain the final classification score S corresponding to C categories. pool ; At this time, the e-th query set Q can be obtained T The corresponding loss L e , as shown in formula (15).
[0195]
[0196] Where q represents the query set Q T The qth query text in p(y=y q |x q ) represents the query text x q The predicted category y is category y q The probability of , as shown in formula (16).
[0197]
[0198] Where m2 represents S pool In category y q The corresponding sequence value.
[0199] Therefore, the final loss L is shown in formula (17).
[0200]
[0201] Where E represents the total number of query sets.
[0202] It should be noted that for the query text x, if the auxiliary paragraph z m The contribution to the current classification task is greater than that of other auxiliary paragraphs, so the corresponding will be greater; thus, As the supervision information for training the retriever, the trained retriever can retrieve the auxiliary paragraphs that contribute most to the current classification task. m |x)align to To obtain the loss L for training the retriever θ (called the retrieval loss function value), as shown in formula (18).
[0203]
[0204] Where Z represents the M auxiliary paragraphs retrieved by the retriever. θ is the total number of parameters of the retriever, including the parameters used to obtain the embedding vector of the auxiliary paragraph.
[0205] The text classification effect provided by the embodiment of the present application is described below.
[0206] See Table 1, which shows the accuracy of seven baseline models (baseline model 1 to baseline model 7) and the Retrieval-Augmented Meta Learning for Low-Resource Text Classification (RAML) model provided by an embodiment of the present application (hereinafter referred to as the model of the present application) on dataset 1 (resource classification dataset) and dataset 2 (news headline dataset), respectively.
[0207] Table 1
[0208]
[0209]
[0210] As can be seen from Table 1, the model of this application outperforms the baseline model in various data sets and various sample sizes by utilizing the neural network model and the non-parametric knowledge obtained from the knowledge base (i.e., M auxiliary paragraphs). In addition, in the 5-sample scenario, the accuracy of data set 1 is significantly improved compared to data set 2, indicating that the more relevant the knowledge base is to the business field, the higher the accuracy. Among them, although baseline model 6 and baseline model 7 both introduce external knowledge, the accuracy of the model of this application is higher in various data sets and various sample sizes, indicating that the model of this application can more effectively utilize external knowledge.
[0211] In addition, the model of the present application can improve the generalization in few-sample (the number of samples is less than the specified number) learning. Continuing to refer to Table 1, compared with the strongest baseline model, in the 1-sample and 5-sample scenarios of Dataset 1, the accuracy of the model of the present application is improved by 4.46% and 0.97% respectively; in the 1-sample and 5-sample scenarios of Dataset 2, the accuracy of the model of the present application is improved by 2.33% and 2.78% respectively; thus, it is shown that the classification accuracy in few-sample learning can be improved through neural network models and non-parametric knowledge.
[0212] Here, the effectiveness of the combination of the neural network model and non-parametric knowledge is also determined through ablation experiments; that is, for the query text, the corresponding accuracy is determined by changing the number of auxiliary paragraphs retrieved.
[0213] See also Fig.11 , Fig.11 is an exemplary performance change schematic diagram provided in the embodiment of the present application; Fig.11 As shown, as the number of retrieved auxiliary paragraphs 11-1 increases, the accuracy 11-2 of text classification becomes higher; and as the number of retrieved auxiliary paragraphs 11-1 decreases, when the number of retrieved auxiliary paragraphs 11-1 decreases to 0, the accuracy 11-2 is the lowest, and the accuracy at this time is still better than the accuracy of the baseline model. This shows that when text classification is performed on query features and prototypes, the cross-attention mechanism is used to determine different weights for different query features, so that the semantic features are retained, thereby improving the accuracy of text classification.
[0214] It can be understood that by explicitly using auxiliary paragraphs for text classification, the semantic enhancement of the query text is achieved, thereby improving the accuracy of text classification in the resource transfer scenario.
[0215] The following is a description of an exemplary structure of a text processing device 455 provided in an embodiment of the present application implemented as a software module. In some embodiments, Figure 4 As shown, the software modules stored in the text processing device 455 of the memory 450 may include:
[0216] The text acquisition module 4551 is used to acquire M auxiliary texts that are most similar to the text to be classified from a preset text library, where M is a positive integer;
[0217] A text combination module 4552, used for combining the text to be classified with the M auxiliary texts respectively to obtain M texts to be processed;
[0218] A feature extraction module 4553 is used to extract features from the M texts to be processed respectively to obtain M features to be processed;
[0219] A text classification module 4554 is used to perform text classification by combining the M features to be processed and C category features to obtain a text classification score, wherein the category feature is determined by a sub-support set corresponding to a text category, and the text classification score represents the C category scores of the text to be classified belonging to the C text categories, where C is a positive integer;
[0220] The category determination module 4555 is used to determine the target text category to which the to-be-classified text belongs from the C text categories based on the text classification score.
[0221] In an embodiment of the present application, the text acquisition module 4551 is also used to perform feature extraction on the text to be classified to obtain features of the text to be classified; for the preset text features of each preset text in the preset text library, obtain the feature correlation between the preset text features and the features of the text to be classified; determine the feature correlation as the text similarity between the text to be classified and the preset text; and determine the M preset texts corresponding to the M largest text similarities in the preset text library as the M auxiliary texts most similar to the text to be classified.
[0222] In an embodiment of the present application, the feature extraction module 4553 is also used to perform the following processing on the sub-support set corresponding to each of the C text categories: extract at least one supporting text feature corresponding to at least one supporting text in the sub-support set; integrate at least one supporting text feature to obtain the category feature; and obtain C category features corresponding to the C text categories from the category features corresponding to each of the text categories.
[0223] In an embodiment of the present application, the text classification module 4554 is also used to traverse the M features to be processed and perform the following processing on each of the traversed features to be processed: combine the features to be processed and the C category features to perform text classification to obtain a sub-text classification score, wherein the sub-text classification score represents the C auxiliary scores of the text to be processed belonging to the C text categories; obtain M sub-text classification scores corresponding to the M features to be processed from the sub-text classification scores corresponding to each of the features to be processed; and integrate the M sub-text classification scores into the text classification score.
[0224] In an embodiment of the present application, the text classification module 4554 is also used to obtain C feature similarities corresponding to the feature to be processed and the C category features; fuse the C feature similarities with the feature to be processed to obtain a target classification feature; perform text classification based on the target classification feature to obtain the sub-text classification score.
[0225] In an embodiment of the present application, the text classification module 4554 is also used to obtain M auxiliary scores corresponding to each text category from the M sub-text classification scores; select the maximum auxiliary score from the M auxiliary scores; and determine the C maximum auxiliary scores corresponding to the C text categories as the text classification scores.
[0226] In an embodiment of the present application, the text classification module 4554 is also used to obtain M auxiliary weights corresponding to the M features to be processed, wherein the auxiliary weights are C classification weights corresponding to the features to be processed on the C text categories; from the M auxiliary weights, M classification weights corresponding to each text category are obtained; the M classification weights are fused with the M auxiliary scores in a one-to-one correspondence to obtain the category score; and the C category scores corresponding to the C text categories are determined as the text classification score.
[0227] In the embodiment of the present application, the text classification module 4554 is also used to obtain the average score of the M auxiliary scores, obtain C average scores corresponding to the C text categories; and determine the C average scores as the text classification score.
[0228] In an embodiment of the present application, the text classification is implemented through a text classification model, and the text processing device 455 also includes a model training module 4556, which is used to obtain M auxiliary samples that are most similar to the text sample from the preset text library; combine the text sample with the M auxiliary samples respectively to obtain M samples to be processed; perform feature extraction on the M samples to be processed respectively to obtain M features of the samples to be processed; use a first model to be trained to perform text classification on the M features of the samples to be processed and C category features to obtain an estimated classification score, wherein the first model to be trained is a neural network model to be trained for classifying text; combine the estimated classification score and the classification label of the text sample to train the first model to be trained to obtain the text classification model.
[0229] In an embodiment of the present application, the model training module 4556 is also used to select a category estimated score corresponding to the classification label of the text sample from the estimated classification scores; normalize the category estimated score to obtain a classification probability; obtain a classification probability set corresponding to a text sample set from the classification probability of the text sample, wherein the text sample is any text in the text sample set; integrate the classification probability set to obtain a classification loss function value; and train the first model to be trained based on the classification loss function value to obtain the text classification model.
[0230] In an embodiment of the present application, the M auxiliary texts are obtained through a retrieval model, and the model training module 4556 is also used to determine the retrieval loss function value based on the difference between the M sample similarities and the M auxiliary training weights, wherein the sample similarity represents the similarity between the auxiliary sample and the text sample, and the auxiliary training weight represents the C weights corresponding to the C text categories determined by the features of the sample to be processed; based on the retrieval loss function value, a second model to be trained is trained to obtain the retrieval model, wherein the M auxiliary samples are obtained through the second model to be trained, and the second model to be trained is a neural network model to be trained for obtaining the text most similar to the target text.
[0231] The embodiment of the present application provides a computer program product, which includes computer executable instructions or computer programs, and the computer executable instructions or computer programs are stored in a computer readable storage medium. The processor of the text processing device reads the computer executable instructions or computer programs from the computer readable storage medium, and the processor executes the computer executable instructions or computer programs, so that the text processing device performs the text processing method described above in the embodiment of the present application.
[0232] The present application embodiment provides a computer-readable storage medium in which computer executable instructions or computer programs are stored. When the computer executable instructions or computer programs are executed by a processor, the processor will be caused to execute the text processing method provided by the present application embodiment, for example, Figure 5 The text processing method shown.
[0233] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or may be various devices including one or any combination of the above memories.
[0234] In some embodiments, computer executable instructions may be in the form of a program, software, software module, script or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine or other unit suitable for use in a computing environment.
[0235] As an example, computer-executable instructions may, but need not, correspond to a file in a file system, may be stored as part of a file that stores other programs or data, such as in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files storing one or more modules, subroutines, or code portions).
[0236] As an example, computer-executable instructions may be deployed to be executed on one electronic device (in which case, the one electronic device is a text processing device), or on multiple electronic devices located at one location (in which case, the multiple electronic devices located at one location are text processing devices), or on multiple electronic devices distributed at multiple locations and interconnected by a communication network (in which case, the multiple electronic devices distributed at multiple locations and interconnected by a communication network are text processing devices).
[0237] It is understandable that in the embodiments of the present application, when the embodiments of the present application are applied to specific products or technologies, it is necessary to obtain user permission or consent, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions. In addition, in the present application, for the implementation of the data capture technical solution involved in the sample text, when the above embodiments of the present application are applied to specific products or technologies, the relevant data collection, use and processing process should comply with the requirements of national laws and regulations, comply with the principles of legality, legitimacy and necessity, do not involve the acquisition of data types prohibited or restricted by laws and regulations, and will not hinder the normal operation of the target website.
[0238] In summary, the embodiment of the present application retrieves M auxiliary texts that are most similar to the text to be classified from a preset text library, and uses the M auxiliary texts as non-parametric knowledge to assist in determining C classification scores of the text to be classified corresponding to C text categories; in the text classification process, effective use of non-parametric knowledge is achieved, and the text to be classified can be accurately described, thereby improving the accuracy of text classification.
[0239] The above is only an embodiment of the present application and is not intended to limit the protection scope of the present application. Any modifications, equivalent substitutions and improvements made within the spirit and scope of the present application are included in the protection scope of the present application.
Claims
1. A text processing method, It is characterized in that The method comprises: From the preset text library, obtain M auxiliary texts that are most similar to the text to be classified, where M is a positive integer; Combining the text to be classified with the M auxiliary texts respectively to obtain M texts to be processed; Perform feature extraction on the M texts to be processed respectively to obtain M features to be processed; Combining the M features to be processed and the C category features to perform text classification, obtaining a text classification score, wherein the category feature is determined by a sub-support set corresponding to a text category, and the text classification score represents the C category scores of the text to be classified belonging to the C text categories, where C is a positive integer; Based on the text classification score, the target text category to which the to-be-classified text belongs is determined from the C text categories.
2. The method according to claim 1, It is characterized in that The step of obtaining M auxiliary texts that are most similar to the text to be classified from the preset text library includes: Extracting features of the text to be classified to obtain features of the text to be classified; For each preset text feature in the preset text library, obtaining a feature correlation between the preset text feature and the feature of the text to be classified; Determine the feature correlation as the text similarity between the text to be classified and the preset text; The M preset texts corresponding to the M largest text similarities in the preset text library are determined as the M auxiliary texts that are most similar to the text to be classified.
3. The method according to claim 1, It is characterized in that Before combining the M features to be processed and the C category features to perform text classification and obtain a text classification score, the method further includes: For the sub-support set corresponding to each of the C text categories, the following processing is performed: extracting at least one supporting text feature corresponding to at least one supporting text in the sub-support set; Integrating at least one of the supporting text features to obtain the category feature; C category features corresponding to the C text categories are obtained from the category features corresponding to each of the text categories.
4. The method according to claim 1, It is characterized in that Combining the M features to be processed and the C category features to perform text classification to obtain a text classification score includes: Traverse the M features to be processed, and perform the following processing on each of the traversed features to be processed: Combining the feature to be processed and the C category features to perform text classification, and obtaining a sub-text classification score, wherein the sub-text classification score represents C auxiliary scores that the text to be processed belongs to the C text categories; Obtaining M sub-text classification scores corresponding to the M features to be processed from the sub-text classification scores corresponding to each of the features to be processed; The M sub-text classification scores are integrated into the text classification score.
5. The method according to claim 4, It is characterized in that The combining the feature to be processed and the C category features to perform text classification to obtain a sub-text classification score includes: Obtaining C feature similarities corresponding to the feature to be processed and the C category features respectively; The C feature similarities are respectively fused with the feature to be processed to obtain the target classification feature; Text classification is performed based on the target classification feature to obtain the sub-text classification score.
6. The method according to claim 4, It is characterized in that The step of integrating the M sub-text classification scores into the text classification score comprises: From the M sub-text classification scores, obtain M auxiliary scores corresponding to each of the text categories; Selecting the maximum auxiliary score from the M auxiliary scores; The C maximum auxiliary scores corresponding to the C text categories are determined as the text classification scores.
7. The method according to claim 6, It is characterized in that After obtaining the M auxiliary scores corresponding to each of the text categories from the M sub-text classification scores, the method further includes: Obtaining M auxiliary weights corresponding to the M features to be processed, wherein the auxiliary weights are C classification weights corresponding to the features to be processed on the C text categories; From the M auxiliary weights, obtain the M classification weights corresponding to each of the text categories; The M classification weights are fused with the M auxiliary scores in a one-to-one correspondence to obtain the category score; The C category scores corresponding to the C text categories are determined as the text classification scores.
8. The method according to claim 6, It is characterized in that After obtaining the M auxiliary scores corresponding to each of the text categories from the M sub-text classification scores, the method further includes: Obtaining an average score of the M auxiliary scores to obtain C average scores corresponding to the C text categories; The C average scores are determined as the text classification score.
9. The method according to any one of claims 1 to 8, It is characterized in that The text classification is achieved by a text classification model, and the text classification model is trained by the following steps: Obtaining M auxiliary samples that are most similar to the text sample from the preset text library; Combining the text sample with the M auxiliary samples respectively to obtain M samples to be processed; Extracting features of the M samples to be processed respectively to obtain features of the M samples to be processed; Using a first model to be trained to perform text classification on the M sample features to be processed and the C category features to obtain an estimated classification score, wherein the first model to be trained is a neural network model to be trained for classifying text; The first model to be trained is trained in combination with the estimated classification score and the classification label of the text sample to obtain the text classification model.
10. The method according to claim 9, It is characterized in that The step of training the first model to be trained by combining the estimated classification score and the classification label of the text sample to obtain the text classification model includes: From the estimated classification scores, selecting a category estimated score corresponding to the classification label of the text sample; Normalizing the estimated scores of the categories to obtain classification probabilities; Obtaining a classification probability set corresponding to a text sample set from the classification probability of the text sample, wherein the text sample is any text in the text sample set; Integrating the classification probability set to obtain a classification loss function value; The first model to be trained is trained based on the classification loss function value to obtain the text classification model.
11. The method according to claim 9, It is characterized in that The M auxiliary texts are obtained through a retrieval model, and the retrieval model is trained through the following steps: Determine a retrieval loss function value based on the difference between M sample similarities and M auxiliary training weights, wherein the sample similarity represents the similarity between the auxiliary sample and the text sample, and the auxiliary training weight represents C weights corresponding to the C text categories determined by the features of the sample to be processed; The second model to be trained is trained based on the retrieval loss function value to obtain the retrieval model, wherein the M auxiliary samples are obtained through the second model to be trained, and the second model to be trained is a neural network model to be trained for obtaining the text most similar to the target text.
12. A text processing device, It is characterized in that The text processing device comprises: The text acquisition module is used to acquire M auxiliary texts that are most similar to the text to be classified from a preset text library, where M is a positive integer; A text combination module, used for combining the text to be classified with the M auxiliary texts respectively to obtain M texts to be processed; A feature extraction module is used to extract features from the M texts to be processed respectively to obtain M features to be processed; A text classification module, used for combining the M features to be processed and C category features to perform text classification to obtain a text classification score, wherein the category feature is determined by a sub-support set corresponding to a text category, and the text classification score represents the C category scores of the text to be classified belonging to the C text categories, where C is a positive integer; A category determination module is used to determine the target text category to which the to-be-classified text belongs from the C text categories based on the text classification score.
13. An electronic device for text processing, It is characterized in that The electronic device comprises: A memory for storing computer executable instructions or computer programs; A processor, configured to implement the text processing method according to any one of claims 1 to 11 when executing the computer executable instructions or computer programs stored in the memory.
14. A computer-readable storage medium storing computer-executable instructions or a computer program, It is characterized in that When the computer executable instructions or computer programs are executed by a processor, the text processing method according to any one of claims 1 to 11 is implemented.
15. A computer program product comprising computer executable instructions or a computer program, It is characterized in that When the computer executable instructions or computer programs are executed by a processor, the text processing method according to any one of claims 1 to 11 is implemented.