Knn retrieval enhanced text classification method and device, equipment and medium

CN118820927BActive Publication Date: 2026-09-25BEIJING UNIV OF POSTS & TELECOMM +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410816275.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-24
Publication Date
2026-09-25
Estimated Expiration
2044-06-24

AI Technical Summary

Technical Problem

[0005]1.训练数据集在推理阶段的利用不足:在传统的文本分类方法中,训练数据集仅在训练阶段被用来训练模型,而在推理阶段很少或根本没有被利用

Benefits of technology

[0055]1.解决训练数据集在推理阶段利用不足的问题:在训练阶段构建了一个表示存储,用于保存训练样本的嵌入表示。在推理阶段,模型通过检索测试文本的最近邻居,并结合文本增强技术,使得训练数据集的语义信息得以在分类决策中发挥作用。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118820927B_ABST
    Figure CN118820927B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a text classification method and device based on KNN retrieval enhanced, equipment and medium. The method comprises the following steps: constructing a function f(·), mapping text sequences of a training data set to a fixed length vector representation form by using the function f(·) in response to an input training set, and storing vector representations of all text sequences and corresponding labels in the training data set; constructing a text enhancement module, enhancing the training data set by using the text enhancement module to obtain an enhanced training data set; constructing a K-nearest neighbor classifier, training the K-nearest neighbor classifier by using the enhanced training data set, and realizing text classification by using the trained K-nearest neighbor classifier. The present application significantly improves the performance of various deep learning models (such as CNN, LSTM, BERT and RoBERTa) in the text classification task, and simultaneously enhances the generalization ability and classification accuracy of the model by using the training data set information without additional training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to a text classification method, apparatus, device, and medium based on KNN retrieval enhancement, belonging to the field of natural language processing technology. Background Technology

[0002] Text classification is one of the fundamental tasks in Natural Language Processing (NLP), involving the automatic categorization of given text into one or more predefined categories. This technology is of great significance in various application scenarios such as spam detection, sentiment analysis, and relation extraction. Traditional text classification methods rely on machine learning models, requiring tedious feature engineering, while deep learning methods, especially models based on recurrent neural networks (RNNs), convolutional neural networks (CNNs), and Transformer architectures (such as BERT and RoBERTa), are favored because they can automatically learn high-level features from raw data.

[0003] Deep learning-based text classification methods typically involve two phases: training and inference. During training, the model optimizes its parameters by learning features from the training dataset. However, during inference, the training dataset is usually only used to generate the model's initial parameters and is no longer involved in the classification decision process. This means that the semantic information of the training dataset is not fully utilized during inference. Deep learning models, such as CNNs, LSTMs, BERTs, or RoBERTa, are used to extract feature representations from text data. Structurally, these models typically consist of multiple layers, including input layers, hidden layers, and output layers. The input layer receives text data, the hidden layers process the data through various network structures (such as convolutional layers, recurrent layers, or self-attention layers), and the output layer classifies the text based on the learned features. In practice, the model first learns the feature representations of the dataset through forward propagation during training and then adjusts the model parameters using backpropagation to minimize the loss function. During inference, the model uses the trained parameters to classify new text data, typically predicting the text's category based on the learned feature representations and a softmax function. However, these existing techniques fail to fully utilize the semantic information in the training dataset during the inference phase, potentially leading to poor performance when processing new texts with distributions similar to the training data. To address this issue, this invention proposes a novel text classification method, the K-Nearest Neighbor Retrieval Augmentation Model (KRA), which significantly improves the model's classification performance and generalization ability by introducing retrieval and text augmentation techniques based on the training dataset during the inference phase.

[0004] The main shortcomings and problems of existing deep learning-based text classification methods include the following:

[0005] 1. Insufficient utilization of training datasets during inference: In traditional text classification methods, training datasets are only used to train the model during the training phase, and are rarely or never used during the inference phase. This results in the rich semantic information contained in the training dataset not being fully utilized, limiting the model's classification performance on new texts.

[0006] 2. Limited generalization ability: Since the model relies solely on feature representations learned during training, it may not generalize well to new text outside the training set. This is especially true when the dataset is small or the classes are imbalanced, further limiting the model's generalization ability.

[0007] 3. Dependence on large datasets: Some advanced text classification methods, such as Transformer-based models, have achieved significant performance improvements, but they usually require a large amount of labeled data for training, which is not only costly but may also be difficult to obtain.

[0008] The reasons for these shortcomings and problems include:

[0009] 1. Static use of datasets: In traditional methods, the role of training datasets is static, used only for the initial training of the model, rather than dynamically participating in the inference process, which limits the model's ability to utilize the data.

[0010] 2. Limitations of feature representation: Although deep learning models can automatically extract features, these feature representations may not be sufficient to cover all possible text variations and categories, especially when faced with new texts that differ from the distribution of the training data.

[0011] 3. Fixed model structure: Existing deep learning model structures are usually fixed, and they are designed to handle and represent text data, which may not be suitable for all types of text classification tasks. Summary of the Invention

[0012] To address the aforementioned technical problems, embodiments of this application provide a text classification method, apparatus, device, and medium based on KNN retrieval enhancement. By dynamically utilizing the training dataset during the inference phase and combining it with text enhancement techniques, the generalization ability and classification accuracy of the model are improved. This method not only better utilizes the semantic information in the training dataset but also enhances the model's adaptability to new texts by expanding the retrieval set.

[0013] Other features and advantages of this application will become apparent from the following detailed description, or may be learned in part from practice of this application.

[0014] According to one aspect of the embodiments of this application, a text classification method based on KNN retrieval enhancement is provided, the method comprising:

[0015] Construct a function f(·) in response to the input training set, and use the function f(·) to map the text sequences of the training dataset to a fixed-length vector representation, and store the vector representations of all text sequences and their corresponding labels in the training dataset;

[0016] A text enhancement module is constructed, and the training dataset is enhanced using the text enhancement module to obtain the enhanced training dataset;

[0017] Construct a K-nearest neighbor classifier, and train the K-nearest neighbor classifier using the enhanced training dataset to achieve text classification.

[0018] Furthermore, in response to the input training set, the text sequences of the training dataset are mapped to fixed-length vector representations using the function f(·), and the vector representations of all text sequences and their corresponding labels are stored in the training dataset, specifically including:

[0019] Get the i-th example (c) from the training set D i , l i )∈D, where c i Let l represent the i-th text sequence. i This represents the i-th label;

[0020] Determine a key-value pair (k i v i ), key k i A vector representation of a text sequence, with value v i Labels representing text sequences;

[0021] Data storage (K, V) is performed in the training dataset using the following formula: (K, V represents all text sequences as vector representations and their labels.)

[0022] (K, V) = {(f(c)} i ), l i )|(c i , l i )∈D}

[0023] Where f(c) i ) is a vector representation of a text sequence.

[0024] Furthermore, the text enhancement module enhances the training dataset by synonym replacement, random insertion, random swapping, random deletion, and reverse translation to obtain an enhanced training dataset;

[0025] The synonym replacement involves replacing some words in the original text with synonyms.

[0026] The random insertion refers to randomly inserting one or more additional words or phrases into the original text to simulate noise or interference in a real-world scenario;

[0027] The random swapping refers to randomly swapping words or phrases in the original text;

[0028] The random deletion refers to deleting words or phrases from the original text with a set probability to simulate the situation of missing information.

[0029] The reverse translation involves translating the original text into another language and then translating the translated text back into the original language.

[0030] Furthermore, based on the trained K-nearest neighbor classifier, text classification is achieved through the following steps:

[0031] In response to the input text x, the model generates a vector f(x);

[0032] The model is used to generate a vector f(x) for searching data storage, and the nearest neighbor K is obtained based on the Euclidean distance;

[0033] Input the nearest neighbor K corresponding to the original text into the text enhancement module to obtain the enhanced neighbor vector K′;

[0034] The weights of the enhanced neighbor vector K′ are calculated based on the Euclidean distance between each enhanced neighbor vector f(C′) and f(x);

[0035] The probability distribution of neighboring labels is calculated based on the weights of the enhanced neighbor vector K′ to obtain the KNN prediction results;

[0036] The final prediction result is determined based on the KNN prediction results.

[0037] Furthermore, based on the Euclidean distance between each enhanced neighbor vector f(C′) and f(x), the weight of the enhanced neighbor vector K′ is calculated using the following formula:

[0038] W′=Softmax(-|D′|)

[0039] Where W′ is the weight of the enhanced neighbor vector K′, D′ is the Euclidean distance between each enhanced neighbor vector f(C′) and f(x), and Softmax is the normalized exponential function.

[0040] Furthermore, based on the weights of the enhanced neighbor vectors K′, the probability distribution of neighboring labels is calculated using the following formula to obtain the KNN prediction results:

[0041] KNN(x) = W'L'

[0042] Where KNN(x) is the KNN prediction result, W′ is the weight of the enhanced neighbor vector K′, and L′ is the label set of the enhanced neighbors.

[0043] Furthermore, based on the KNN prediction results, the final prediction result is determined using the following formula:

[0044] s=λKNN(x)+(1-λ)Softmax(f(x))

[0045] Where s is the final prediction result, the ratio λ is a hyperparameter used to adjust the probability distribution of the model and the probability distribution obtained from neighboring vectors, Softmax(f(x)) is the model prediction result, and the model is a deep learning model, including one of CNN, LSTM, BERT and RoBERTa.

[0046] According to one aspect of the embodiments of this application, a text classification apparatus based on KNN retrieval enhancement is provided, comprising:

[0047] The storage unit is configured to construct a function f(·) in response to the input training set, using the function f(·) to map the text sequences of the training dataset to a fixed-length vector representation, and storing the vector representations of all text sequences and their corresponding labels in the training dataset;

[0048] The text enhancement unit is configured to construct a text enhancement module and use the text enhancement module to enhance the training dataset to obtain an enhanced training dataset;

[0049] The inference and prediction unit is configured to construct a K-nearest neighbor classifier, train the K-nearest neighbor classifier using the enhanced training dataset, and use the trained K-nearest neighbor classifier to achieve text classification.

[0050] Upon receiving the corresponding data, the data is encrypted and then sent to the data requesting terminal.

[0051] According to one aspect of the embodiments of this application, an electronic device is provided, including: a controller; and a memory for storing one or more programs, which, when executed by the controller, cause the controller to implement the KNN-based retrieval-enhanced text classification method described above.

[0052] According to one aspect of the embodiments of this application, a computer-readable storage medium is also provided, on which computer-readable instructions are stored, which, when executed by a computer's processor, cause the computer to perform the above-described text classification method based on KNN retrieval enhancement.

[0053] According to one aspect of the embodiments of this application, a computer program product or computer program is also provided, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the above-described KNN-based retrieval-enhanced text classification method.

[0054] The technical solutions provided in the embodiments of this application have at least the following advantages:

[0055] 1. Addressing the issue of insufficient utilization of the training dataset during the inference phase: A representation store is constructed during the training phase to save the embedding representations of training samples. During the inference phase, the model retrieves the nearest neighbors of the test text and incorporates text augmentation techniques, enabling the semantic information of the training dataset to play a role in classification decisions.

[0056] 2. Improve the model's generalization ability: During the inference phase, the model not only relies on the feature representations learned in the training phase, but also makes classification decisions based on the retrieved nearest neighbors and their enhanced versions. This approach enables the model to handle more diverse texts, improving its adaptability and generalization ability to new texts.

[0057] 3. Reduced reliance on large datasets: The size of the retrieval set is expanded through text augmentation techniques, which means that even with a small training dataset, the model can generate more training samples through augmentation techniques, thereby reducing the need for large-scale labeled datasets.

[0058] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this application. Attached Figure Description

[0059] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application. It is obvious that the drawings described below are merely some embodiments of this application, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort. In the drawings:

[0060] Figure 1 This is a flowchart illustrating a text classification method based on KNN retrieval enhancement, as shown in an exemplary embodiment of this application;

[0061] Figure 2 This is a structural diagram of a text classification model (KRA) based on KNN retrieval enhancement, as illustrated in an exemplary embodiment of this application.

[0062] Figure 3 This is a schematic diagram illustrating the structure of a text classification device based on KNN retrieval enhancement, as shown in an exemplary embodiment of this application. Detailed Implementation

[0063] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.

[0064] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.

[0065] The flowcharts shown in the accompanying drawings are merely illustrative and do not necessarily include all content and operations / steps, nor do they necessarily have to be performed in the described order. For example, some operations / steps can be broken down, while others can be combined or partially combined; therefore, the actual execution order may change depending on the specific circumstances.

[0066] In this application, "multiple" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects have an "or" relationship.

[0067] Please see Figure 1 , Figure 1 This is a flowchart illustrating an exemplary embodiment of a text classification method based on KNN retrieval enhancement. Figure 1 As shown, the method includes at least steps S100 to S300, which are described in detail below:

[0068] Step S100: Construct a function f(·) in response to the input training set. Use the function f(·) to map the text sequences of the training dataset to a fixed-length vector representation. Store the vector representations of all text sequences and their corresponding labels in the training dataset.

[0069] For example, suppose f(·) is a function that maps the target text to a fixed-length vector representation computed by a pre-trained language model. For instance, let c be the current text sequence, which in BERT represents mapping the text sequence c to a vector representation of length 768. Then, for the i-th example (c) in the training set D... i , l i For each ∈ D, we define a key-value pair (k) i v i Key k i The vector representation c of a text sequence i That is, f(c) i ); value v i Label c representing a text sequence i Therefore, the data storage (K, V) stores the vector representations and labels of all text sequences in the training dataset, as shown in the following equation:

[0070] (K, V)={(f(c i ), l i )|(c i , l i )∈D}

[0071] It should be noted that representation storage is a crucial component in this application. Through the construction of representation storage, the embedding representations of training samples and their corresponding labels are preserved during the training phase. This storage mechanism enables the information from the training dataset to be effectively utilized during the inference phase, which is one of the keys to improving model performance.

[0072] Step S200: Construct a text enhancement module and use the text enhancement module to enhance the training dataset to obtain an enhanced training dataset.

[0073] In one exemplary embodiment, to enhance the diversity of the retrieved neighbors, traditional text enhancement methods and LLM-based text enhancement methods are employed to expand the retrieved neighbors. Table 1 shows examples of different enhancement methods.

[0074] Table 1. Text Enhancement Effects Demonstration

[0075]

[0076]

[0077] This embodiment of the application utilizes GPT-3.5 for text rewriting, and Table 2 summarizes the prompts for this embodiment.

[0078] Table 2 shows the enhanced text prompts generated using GPT-3.5.

[0079]

[0080] In this application embodiment, the following five methods are used to implement text enhancement: synonym replacement (SR), random insertion (RI), random swapping (RS), random deletion (RD), and reverse translation (BT).

[0081] SR (Semantic Representation) technology expands the lexical expressiveness of the dataset by replacing some words in the original text with synonyms, introducing semantic variations. RI (Reference Injection) technology randomly inserts additional words or phrases into the original text, simulating noise or interference in real-world scenarios. RS (Reference Synthesis) technology randomly swaps words or phrases in the original text, disrupting grammar and semantics to some extent. RD (Reference Deletion) technology deletes words or phrases from the original text with a certain probability, simulating situations where some information is missing. BT (Browser Translation) technology is a machine translation-based method that generates new samples with semantic perturbations by translating the original text into another language and then translating the translation back into the original language, thus introducing greater semantic differences and enriching the diversity of the dataset.

[0082] By combining these methods, we can effectively enhance the training dataset and improve the richness and diversity of adjacent data.

[0083] It should be noted that the text enhancement module in this application expands the retrieved neighbors, increasing the diversity and richness of the data. This step improves the model's adaptability to different types of text by generating new text samples, and is a crucial step in enhancing model performance.

[0084] Step S300: Construct a K-nearest neighbor classifier and train the K-nearest neighbor classifier using the enhanced training dataset to achieve text classification.

[0085] In an exemplary embodiment, the process of text classification using a trained K-nearest neighbor classifier is shown. It is understood that the principle of text classification using a trained K-nearest neighbor classifier is the same as the principle of training a K-nearest neighbor classifier using a training dataset; therefore, the process of step S300 is described in detail here using the process of text classification using a trained K-nearest neighbor classifier.

[0086] Specifically, in response to the input text x, the model generates a vector f(x); this vector is then used to search the data store to find the nearest neighbors K based on Euclidean distance. Euclidean distance is used here instead of cosine similarity to obtain the K nearest neighbors. These K nearest neighbors corresponding to the original text C are then input into the text enhancement module to obtain the enhanced neighbor vector K′. The weight of each vector is calculated based on distance as follows:

[0087] W′=Softmax(-|D′|)

[0088] Here, D′ is the Euclidean distance between each enhanced neighbor vector f(C′) and f(x). Then, the probability distribution of neighboring labels is calculated to obtain the KNN prediction result KNN(x):

[0089] KNN(x)=W′L′

[0090] Here, L′ is the label set of the enhanced neighbors. Finally, combining the model prediction result Softmax(f(x)), we obtain the final prediction result s:

[0091] s=λKNN(x)+(1-λ)Softmax(f(x))

[0092] The ratio λ is a hyperparameter used to adjust the probability distribution of the model itself and the probability distribution obtained from neighboring vectors.

[0093] It should be noted that the method of obtaining the model prediction results is based on the known deep learning models applied to text classification. It is the existing knowledge in this field that they can achieve text classification prediction and obtain model prediction results. Therefore, the specific process of obtaining model prediction results will not be described in detail here.

[0094] The embodiments of this application used four commonly used deep learning models to verify the invention on a text classification dataset, and the results are shown in Table 3.

[0095] Table 3 shows the enhanced text prompts generated using GPT-3.5.

[0096]

[0097] In another exemplary embodiment, a text classification method based on K-Nearest Neighbor Retrieval Enhancement (KRA) is provided, aiming to address the problems of insufficient utilization of training datasets during the inference stage, limited generalization ability, and dependence on large datasets in existing technologies. To better utilize the semantic information of the training dataset and improve model performance, a KRA model is proposed. The KRA model consists of a classic K-Nearest Neighbor classifier and a text enhancement module. The purpose of the K-Nearest Neighbor classifier is to utilize the semantic information of the training dataset to assist the model's classification decisions. The text enhancement module aims to improve the quality of the training dataset, enhance the generalization ability of the K-Nearest Neighbor classifier, and ensure the effectiveness of KRA even when the training dataset is not large enough. Figure 2 The overall structure of the model KRA is shown.

[0098] like Figure 2As shown, the text is first represented by a language model, which includes one or a combination of LSTM, CNN, and BERT; this embodiment is not specifically limited to these models. The obtained text representation is then sequentially fed into two modules / models for training. One module / model is the KRA module / model proposed in this application, and the other is an existing module / model, such as the text representation being processed by a linear classifier and then by Softmax to obtain a model prediction result. The process of processing the text representation through the KRA module / model to obtain the KNN prediction result has been detailed in the above embodiment and will not be repeated here. Finally, the KNN prediction result and the model prediction result are combined to obtain the final prediction result.

[0099] Therefore, the core advancements of this application are reflected in:

[0100] 1) K-Nearest Neighbor Retrieval Mechanism: The K-Nearest Neighbor retrieval mechanism adopted in this application retrieves the k training samples most similar to the test text during the inference stage. The accuracy of this retrieval method directly affects the quality of the final classification decision and is a key technology for improving the model's generalization ability.

[0101] 2) Weight Calculation and Integration Mechanism: This application assigns weights by calculating the distance between the augmented samples and the target text, and combines this with the prediction results of the original classifier for the final decision. This weight calculation and integration method ensures that the model considers the relevance and importance of samples during classification, which is a key technical point for improving classification accuracy.

[0102] 3) Applicability of deep learning models: The method proposed in this application can be used in combination with various deep learning models, such as CNN, LSTM, BERT and RoBERTa. This compatibility and flexibility is an improvement of this invention, ensuring that this invention can be widely applied to different text classification tasks.

[0103] 4) Enhancement capabilities without additional training: A significant advantage of this application is that it can improve performance during the inference stage without requiring additional training on the pre-trained deep learning model. This plug-and-play feature lowers the barrier to entry and cost of implementing this application.

[0104] Another aspect of this application provides a text classification device based on KNN retrieval enhancement, such as... Figure 3 As shown, the device 300 includes:

[0105] Storage unit 301 is configured to construct a function f(·) in response to the input training set, using the function f(·) to map the text sequences of the training dataset to a fixed-length vector representation, and storing the vector representations of all text sequences and their corresponding labels in the training dataset;

[0106] The text enhancement unit 302 is configured to construct a text enhancement module and use the text enhancement module to enhance the training dataset to obtain an enhanced training dataset.

[0107] The inference prediction unit 303 is configured to construct a K-nearest neighbor classifier, train the K-nearest neighbor classifier using the enhanced training dataset, and use the trained K-nearest neighbor classifier to achieve text classification.

[0108] In another embodiment, the representation storage unit is further configured to retrieve the i-th example (c) from the training set D. i , l i )∈D, where c i Let l represent the i-th text sequence. i This represents the i-th label;

[0109] Determine a key-value pair (k i v i ), key k i A vector representation of a text sequence, with value v i Labels representing text sequences;

[0110] Data storage (K, V) is performed in the training dataset using the following formula: (K, V represents all text sequences as vector representations and their labels.)

[0111] (K, V) = {(f(c)} i ), l i )|(c i , l i )∈D}

[0112] Where f(c) i ) is a vector representation of a text sequence.

[0113] In another embodiment, the text enhancement unit is further configured to use the text enhancement module to enhance the training dataset through synonym replacement, random insertion, random swapping, random deletion, and reverse translation to obtain an enhanced training dataset;

[0114] The synonym replacement involves replacing some words in the original text with synonyms.

[0115] The random insertion refers to randomly inserting one or more additional words or phrases into the original text to simulate noise or interference in a real-world scenario;

[0116] The random swapping refers to randomly swapping words or phrases in the original text;

[0117] The random deletion refers to deleting words or phrases from the original text with a set probability to simulate the situation of missing information.

[0118] The reverse translation involves translating the original text into another language and then translating the translated text back into the original language.

[0119] In another embodiment, the text enhancement unit is further configured to perform text classification based on a trained K-nearest neighbor classifier through the following steps:

[0120] In response to the input text x, the model generates a vector f(x);

[0121] The model is used to generate a vector f(x) for searching data storage, and the nearest neighbor K is obtained based on the Euclidean distance;

[0122] Input the nearest neighbor K corresponding to the original text into the text enhancement module to obtain the enhanced neighbor vector K′;

[0123] The weights of the enhanced neighbor vector K′ are calculated based on the Euclidean distance between each enhanced neighbor vector f(C′) and f(x);

[0124] The probability distribution of neighboring labels is calculated based on the weights of the enhanced neighbor vector K′ to obtain the KNN prediction results;

[0125] The final prediction result is determined based on the KNN prediction results.

[0126] In another embodiment, the text enhancement unit is further configured to calculate the weight of the enhanced neighbor vector K′ based on the Euclidean distance between each enhanced neighbor vector f(C′) and f(x) using the following formula:

[0127] W′=Softmax(-|D′|)

[0128] Where W′ is the weight of the enhanced neighbor vector K′, D′ is the Euclidean distance between each enhanced neighbor vector f(C′) and f(x), and Softmax is the normalized exponential function.

[0129] In another embodiment, the text enhancement unit is further configured to calculate the probability distribution of neighboring labels based on the weights of the enhanced neighbor vectors K to obtain the KNN prediction result using the following formula:

[0130] KNN(x)=W′L′

[0131] Where KNN(x) is the KNN prediction result, W′ is the weight of the enhanced neighbor vector K′, and L′ is the label set of the enhanced neighbors.

[0132] In another embodiment, the text enhancement unit is further configured to determine the final prediction result based on the KNN prediction result using the following formula:

[0133] s=λKNN(x)+(1-λ)Softmax(f(x))

[0134] Where s is the final prediction result, the ratio λ is a hyperparameter used to adjust the probability distribution of the model and the probability distribution obtained from neighboring vectors, Softmax(f(x)) is the model prediction result, and the model is a deep learning model, including one of CNN, LSTM, BERT and RoBERTa.

[0135] It should be noted that the text classification device based on KNN retrieval enhancement provided in the above embodiments and the text classification method based on KNN retrieval enhancement provided in the aforementioned embodiments belong to the same concept. The specific way in which each module and unit performs operations has been described in detail in the method embodiments, and will not be repeated here.

[0136] Another aspect of this application provides an electronic device, including: a controller; and a memory for storing one or more programs, which, when executed by the controller, perform the KNN-based retrieval-enhanced text classification described in the various embodiments above.

[0137] Specifically, according to embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program including a computer program for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication component, and / or installed from a removable medium. When the computer program is executed by the central processing unit (CPU) 701, it performs various functions defined in the system of this application.

[0138] It should be noted that the computer-readable medium shown in the embodiments of this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying a computer-readable computer program. The transmitted data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The computer program contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to wireless, wired, etc., or any suitable combination thereof.

[0139] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. Each block in a flowchart or block diagram may represent a module, segment, or portion of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0140] The units described in the embodiments of this application can be implemented in software or hardware, and the described units can also be located in a processor. The names of these units do not necessarily limit the specific unit itself.

[0141] Another aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the aforementioned KNN-based retrieval-enhanced text classification method. This computer-readable storage medium may be included in the electronic device described in the above embodiments, or it may exist independently and not incorporated into the electronic device.

[0142] Another aspect of this application provides a computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the KNN-based retrieval-enhanced text classification method provided in the various embodiments described above.

[0143] According to one aspect of the embodiments of this application, a computer system is also provided, including a Central Processing Unit (CPU), which can perform various appropriate actions and processes based on a program stored in read-only memory (ROM) or a program loaded from storage into random access memory (RAM), such as performing the methods described above. Various programs and data required for system operation are also stored in the RAM. The CPU, ROM, and RAM are interconnected via a bus. Input / output (I / O) interfaces are also connected to the bus.

[0144] The following components are connected to the I / O interface: input components including keyboards, mice, etc.; output components including cathode ray tubes (CRTs), liquid crystal displays (LCDs), and speakers; storage components including hard drives; and communication components including network interface cards such as LAN (Local Area Network) cards and modems. The communication components perform communication processing via networks such as the Internet. Drives are also connected to the I / O interface as needed. Removable media, such as disks, optical discs, magneto-optical discs, semiconductor memories, etc., are installed on the drive as needed so that computer programs read from them can be installed into the storage components as required.

[0145] The above description is merely a preferred exemplary embodiment of this application and is not intended to limit the implementation of this application. Those skilled in the art can easily make corresponding modifications or alterations based on the main concept and spirit of this application. Therefore, the scope of protection of this application should be determined by the scope of protection claimed in the claims.

Claims

1. A text classification method based on KNN retrieval enhancement, characterized in that, The method includes: Construct a function f(·) in response to the input training dataset, and use the function f(·) to map the text sequences of the training dataset to a fixed-length vector representation, and store the vector representations of all text sequences and their corresponding labels in the training dataset; A text enhancement module is constructed, and the training dataset is enhanced using the text enhancement module to obtain the enhanced training dataset; Construct a K-nearest neighbor classifier, and train the K-nearest neighbor classifier using the enhanced training dataset to achieve text classification; In response to the input training dataset, the function is used Mapping the text sequences in the training dataset to fixed-length vector representations, and storing the vector representations of all text sequences and their corresponding labels in the training dataset, specifically includes: Obtain the training dataset The first in An example ,in Indicates the first A text sequence, Indicates the first One tag; Determine a key-value pair ,key A vector representation of a text sequence, with values... Labels representing text sequences; Data storage The vector representations and labels of all text sequences are stored in the training dataset using the following formula: in It is a vector representation of a text sequence.

2. The text classification method based on KNN retrieval enhancement according to claim 1, characterized in that, The text enhancement module enhances the training dataset by synonym replacement, random insertion, random swapping, random deletion, and reverse translation to obtain the enhanced training dataset; The synonym replacement involves replacing some words in the original text with synonyms. The random insertion refers to randomly inserting one or more additional words or phrases into the original text to simulate noise or interference in a real-world scenario; The random swapping refers to randomly swapping words or phrases in the original text; The random deletion refers to deleting words or phrases from the original text with a set probability to simulate the situation of missing information. The reverse translation involves translating the original text into another language and then translating the translated text back into the original language.

3. The text classification method based on KNN retrieval enhancement according to claim 1, characterized in that, Based on a trained K-nearest neighbor classifier, text classification is achieved through the following steps: Responding to the input text , to obtain the model generated vector ; Generate vectors using the model Search data storage and find the nearest neighbor based on Euclidean distance. ; The nearest neighbor corresponding to the original text The input is fed into the text enhancement module to obtain the enhanced neighbor vectors. ; Based on each enhanced neighbor vector and Calculate the enhanced neighbor vectors by Euclidean distance between them. The weights; Based on the enhanced neighbor vector The weights are used to calculate the probability distribution of adjacent labels to obtain... KNN Prediction results; according to KNN The prediction results determine the final prediction result.

4. The text classification method based on KNN retrieval enhancement according to claim 3, characterized in that, Based on each enhanced neighbor vector and The Euclidean distance between them is calculated using the following formula to determine the enhanced neighbor vectors. Weights: in, It is an enhanced neighbor vector The weight, It is the vector of each enhanced neighbor. and The Euclidean distance between them It is a normalized exponential function.

5. The text classification method based on KNN retrieval enhancement according to claim 3, characterized in that, Based on the enhanced neighbor vector The weights are calculated using the following formula to determine the probability distribution of adjacent labels. KNN Prediction results: in, yes KNN Prediction results It is an enhanced neighbor vector The weight, It is a set of tags that enhance neighbors.

6. The text classification method based on KNN retrieval enhancement according to claim 3, characterized in that, according to KNN The final prediction result is determined using the following formula: in, This is the final prediction result, the ratio. It is a hyperparameter used to adjust the probability distribution of the model and the probability distribution obtained from neighboring vectors. It is the model prediction result, and the model is a deep learning model, including one of CNN, LSTM, BERT and RoBERTa.

7. A text classification device based on KNN retrieval enhancement, used to implement the method according to any one of claims 1 to 6, characterized in that, include: Represents a storage unit, configured as a construction function. In response to the input training dataset, the function is used. The text sequences in the training dataset are mapped to fixed-length vector representations, and the vector representations of all text sequences and their corresponding labels are stored in the training dataset. The text enhancement unit is configured to construct a text enhancement module and use the text enhancement module to enhance the training dataset to obtain an enhanced training dataset; The inference and prediction unit is configured to construct a K-nearest neighbor classifier, train the K-nearest neighbor classifier using the enhanced training dataset, and use the trained K-nearest neighbor classifier to achieve text classification.

8. An electronic device, characterized in that, include: Controller; A memory for storing one or more programs that, when executed by the controller, cause the controller to implement the text classification method based on KNN retrieval enhancement as described in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, It stores computer-readable instructions that, when executed by the computer's processor, cause the computer to perform the text classification method based on KNN retrieval enhancement as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Text classification model training method and device, equipment and readable medium

    CN112883193A

  • Continuous learning method based on nearest neighbor retrieval enhancement

    CN117669682A