Lightweight sample data-based text classification method and apparatus
By semantic segmentation and feature extraction of classified text, local and global features are generated, and the problem of high training cost of traditional text classification models is solved, and efficient text classification and accurate prediction are achieved.
Patent Information
- Application Number
- CN202510205243.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-05-13
AI Technical Summary
Traditional text classification models need to collect a large amount of training data and manually label the tags during training, resulting in higher training costs.
The text classification method based on lightweight sample data is adopted. By semantic segmentation of the classified text, word vectors are extracted, and local features and global features are analyzed, multiple local features and global features are generated, and input them into the text classification model for processing to realize text type recognition.
This reduces the training cost of training text classification models, improves the accuracy of model prediction, makes up for the defect of small amount of training data, and improves the accuracy of model prediction.
Smart Images

Figure CN119988613A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing, and in particular to a text classification method and device based on lightweight sample data. Background Art
[0002] With the rise and popularity of small sample learning in the field of natural language processing, small sample learning models are roughly divided into three types: metric-based, model-based, and optimization-based. Small sample learning methods do not rely on large-scale training samples, thus avoiding the high cost of data preparation in certain specific applications and achieving low-cost and fast model deployment.
[0003] However, the traditional training method and text classification method of text classification models input the text samples under each task in multiple tasks into the corresponding private feature extractor and public feature extractor, and train the private feature extractors and classifiers under multiple different tasks at the same time to obtain the trained text classification model; for example, the common Bert model has a fixed internal feature dimension of 768. However, this method cannot be trained under the conditions of small data volume and incomplete data labels. It requires the collection of a large amount of training data and manual labeling, which has high training costs.
[0004] In the above-mentioned related technologies, traditional text classification models still need to collect a large amount of training data and manually mark labels during training, resulting in high training costs. No effective solution has been proposed yet. Summary of the invention
[0005] The embodiments of the present invention provide a text classification method and device based on lightweight sample data, so as to at least solve the technical problem that the traditional text classification model in the related art still needs to collect a large amount of training data and manually mark labels during training, resulting in high training costs.
[0006] According to one aspect of an embodiment of the present invention, a text classification method based on lightweight sample data is provided, comprising: when receiving a text to be classified, segmenting the text to be classified according to semantics to obtain a plurality of word vectors, wherein the word vector is a vector representation of each word segmentation fragment in the text to be classified, and each word segmentation fragment includes at least one character in the text to be classified; performing local feature association analysis on at least two adjacent word vectors in turn to obtain a plurality of local features of the text to be classified; performing global feature association analysis on each word vector according to the semantic association between all the word vectors to obtain a global feature corresponding to each word vector; inputting the plurality of local features and the plurality of global features into a text classification model for processing to obtain a text type of the text to be classified, wherein the text classification model is trained by deep learning using a plurality of sets of training data, and each of the plurality of sets of training data includes: a sample input feature and a sample text type corresponding to the sample input feature, and the sample input feature includes: a plurality of sample local features and a plurality of sample global features corresponding to the sample input text, and the number of the sample input texts is not higher than a predetermined value.
[0007] Optionally, when receiving a text to be classified, the text to be classified is segmented according to semantics to obtain multiple word vectors, including: when receiving the text to be classified, the text to be classified is input into a BERT model for processing, so as to use the BERT model to segment the text to be classified using a word segmentation strategy to obtain multiple word segmentation fragments, wherein the word segmentation strategy is a strategy embedded in the BERT model and based on character fragments for word segmentation; position codes are added to each of the word segmentation fragments according to the word order of the text to be classified; and the word vector corresponding to each of the word segmentation fragments is generated in sequence by using the BERT model according to the semantics and in the order of the position codes to obtain multiple word vectors.
[0008] Optionally, after adding position codes to each of the word segmentation fragments according to the word order of the text to be classified, the text classification method based on lightweight sample data also includes: determining the word segmentation fragment at the starting position as the starting word segmentation fragment according to the word order of the text to be classified, and determining the word segmentation fragment at the ending position as the ending word segmentation fragment; adding a start mark to the position code of the starting word segmentation fragment, and adding an end mark to the position code of the ending word segmentation fragment.
[0009] Optionally, local feature association analysis is performed on at least two adjacent word vectors in sequence to obtain multiple local features of the text to be classified, including: determining that the number of text spaces in the convolution kernel is a predetermined number, wherein the convolution kernel is a feature extractor for local feature extraction in a Text-CNN model, and the Text-CNN model is a model for text classification, and each of the text spaces in the convolution kernel is used to place a word vector; selecting multiple adjacent word vectors according to the predetermined number and placing them in the convolution kernel for local feature association analysis until all the word vectors are placed in the convolution kernel at least once, thereby obtaining multiple initial local features; and performing a pooling operation on each of the initial local features using the Text-CNN model to obtain multiple local features, wherein the pooling operation is used to make the feature dimension of the local feature lower than a dimensionality threshold.
[0010] Optionally, a global feature association analysis is performed on each of the word vectors according to the semantic associations between all the word vectors to obtain a global feature corresponding to each of the word vectors, including: determining any of the word vectors as a target word vector; calculating the degree of influence of each candidate word vector on the target word vector based on the semantics of the text to be classified to obtain an attention weight for each of the candidate word vectors, wherein the candidate word vectors are the word vectors other than the target word vector, the degree of influence refers to the similarity between the meaning expressed after the candidate word vector is combined with the target word vector and the semantics, and the attention weight is used to quantify the degree of influence; a global feature association analysis is performed according to the attention weights of each of the candidate word vectors using a multi-head attention mechanism to obtain the global feature corresponding to the target word vector.
[0011] Optionally, before inputting the multiple local features and the multiple global features into the text classification model for processing to obtain the text type of the text to be classified, the text classification method based on lightweight sample data also includes: obtaining multiple groups of sample text data, wherein each of the multiple groups of sample text data includes: the sample input text, the sample text type corresponding to the sample input text, and the number of the sample text data is not higher than the predetermined value; performing word segmentation processing and feature analysis on the sample input texts in the multiple groups of sample text data respectively to obtain sample input features corresponding to each sample input text, and the sample input features include: multiple sample local features and multiple sample global features corresponding to the sample input text; using a bidirectional long short-term memory network to perform sequence modeling on the multiple groups of sample input features and the sample text types to obtain the text classification model.
[0012] Optionally, the text classification method based on lightweight sample data also includes: in the process of using the text classification model to process multiple local features and multiple global features to identify the type of the text to be classified, optimizing the text classification model using the multiple local features and multiple global features of the text to be classified and the text type of the text to be classified, and updating the text classification model to the optimized text classification model.
[0013] According to another aspect of an embodiment of the present invention, a text classification device based on lightweight sample data is also provided, comprising: a first acquisition unit, for, upon receiving a text to be classified, segmenting the text to be classified according to semantics to obtain a plurality of word vectors, wherein the word vector is a vector representation of each word segmentation fragment in the text to be classified, and each of the word segmentation fragments includes at least one character in the text to be classified; a second acquisition unit, for performing local feature association analysis on at least two adjacent word vectors in turn to obtain a plurality of local features of the text to be classified; and a third acquisition unit, for performing local feature association analysis on each word vector according to the semantic associations between all the word vectors. A global feature association analysis is performed on the word vector to obtain a global feature corresponding to each of the word vectors; a fourth acquisition unit is used to input the multiple local features and the multiple global features into a text classification model for processing to obtain the text type of the text to be classified, wherein the text classification model is trained by deep learning using multiple sets of training data, each of the multiple sets of training data includes: sample input features, and sample text types corresponding to the sample input features, the sample input features include: multiple sample local features and multiple sample global features corresponding to sample input texts, and the number of sample input texts is not higher than a predetermined value.
[0014] Optionally, the first acquisition unit includes: a first acquisition module, which is used to input the text to be classified into a BERT model for processing when the text to be classified is received, so as to use the BERT model to segment the text to be classified using a word segmentation strategy to obtain a plurality of word segmentation fragments, wherein the word segmentation strategy is a strategy embedded in the BERT model and based on character fragments for word segmentation; a first adding module, which is used to add position codes to each of the word segmentation fragments according to the word order of the text to be classified; and a second acquisition module, which is used to use the BERT model to generate the word vector corresponding to each of the word segmentation fragments in sequence according to the semantics and in the order of the position codes, so as to obtain a plurality of word vectors.
[0015] Optionally, the text classification device based on lightweight sample data also includes: a first determination module, which is used to determine the word segmentation segment at the starting position as the starting word segmentation segment, and determine the word segment at the ending position as the ending word segmentation segment according to the word order of the text to be classified after adding position codes to each of the word segmentation segments according to the word order of the text to be classified; and a second adding module, which is used to add a start mark to the position code of the starting word segmentation segment, and add an end mark to the position code of the ending word segment segment.
[0016] Optionally, the second acquisition unit includes: a second determination module, used to determine that the number of text spaces in the convolution kernel is a predetermined number, wherein the convolution kernel is a feature extractor for local feature extraction in the Text-CNN model, and the Text-CNN model is a model for text classification, and each of the text spaces in the convolution kernel is used to place a word vector; a third acquisition module, used to select multiple adjacent word vectors according to the predetermined number and put them into the convolution kernel for local feature association analysis, until all the word vectors are placed in the convolution kernel at least once, so as to obtain multiple initial local features; a fourth acquisition module, used to use the Text-CNN model to perform a pooling operation on each of the initial local features to obtain multiple local features, wherein the pooling operation is used to make the feature dimension of the local feature lower than the dimension threshold.
[0017] Optionally, the third acquisition unit includes: a third determination module, used to determine that any one of the word vectors is a target word vector; a fifth acquisition module, used to calculate the degree of influence of each candidate word vector on the target word vector based on the semantics of the text to be classified, and obtain the attention weight of each candidate word vector, wherein the candidate word vector is the word vector other than the target word vector, and the degree of influence refers to the similarity between the meaning expressed after the candidate word vector is combined with the target word vector and the semantics, and the attention weight is used to quantify the degree of influence; a sixth acquisition module, used to use a multi-head attention mechanism to perform global feature association analysis according to the attention weights of each candidate word vector to obtain the global feature corresponding to the target word vector.
[0018] Optionally, the text classification device based on lightweight sample data also includes: a fifth acquisition unit, which is used to acquire multiple groups of sample text data before inputting multiple local features and multiple global features into the text classification model for processing to obtain the text type of the text to be classified, wherein each of the multiple groups of sample text data includes: the sample input text, the sample text type corresponding to the sample input text, and the number of the sample text data is not higher than the predetermined value; a sixth acquisition unit, which is used to perform word segmentation processing and feature analysis on the sample input texts in the multiple groups of sample text data respectively to obtain sample input features corresponding to each sample input text, and the sample input features include: multiple sample local features and multiple sample global features corresponding to the sample input text; a seventh acquisition unit, which is used to perform sequence modeling on multiple groups of sample input features and the sample text types using a bidirectional long short-term memory network to obtain the text classification model.
[0019] Optionally, the text classification device based on lightweight sample data also includes: an optimization unit, which is used to optimize the text classification model by using the multiple local features and the multiple global features of the text to be classified and the text type of the text to be classified, and update the text classification model to the optimized text classification model.
[0020] According to another aspect of an embodiment of the present invention, a text classification system based on lightweight sample data is further provided. The text classification system based on lightweight sample data uses any of the above-mentioned text classification methods based on lightweight sample data.
[0021] According to another aspect of an embodiment of the present invention, a computer-readable storage medium is further provided, wherein the computer-readable storage medium includes a stored program, wherein the program executes any one of the above-mentioned text classification methods based on lightweight sample data.
[0022] According to another aspect of an embodiment of the present invention, a processor is further provided, the processor being used to run a program, wherein the program executes any one of the above-mentioned text classification methods based on lightweight sample data when running.
[0023] According to another aspect of an embodiment of the present invention, a computer program product is provided, including computer instructions, which, when executed by a processor, perform any of the above-mentioned text classification methods based on lightweight sample data.
[0024] In an embodiment of the present invention, when a text to be classified is received, the text to be classified is segmented according to semantics to obtain multiple word vectors, wherein the word vector is a vector representation of each word segmentation fragment in the text to be classified, and each word segmentation fragment includes at least one character in the text to be classified; local feature association analysis is performed on at least two adjacent word vectors in turn to obtain multiple local features of the text to be classified; global feature association analysis is performed on each word vector according to the semantic association between all word vectors to obtain a global feature corresponding to each word vector; the multiple local features and the multiple global features are input into a text classification model for processing to obtain a text type of the text to be classified, wherein the text classification model is trained by deep learning using multiple sets of training data, and each of the multiple sets of training data includes: sample input features, sample text types corresponding to the sample input features, and the sample input features include: multiple sample local features and multiple sample global features corresponding to the sample input text, and the number of sample input texts is not higher than a predetermined value. Through the above technical scheme, the purpose of semantically parsing the text input by the user, extracting features locally and globally respectively, to extract multiple local features and global features corresponding to the text, and using the text classification model to process the extracted features, and finally outputting the text type is achieved. The technical effect of enriching the data volume by performing feature analysis and extraction on the text to make up for the defect of small sample data volume when training the text classification model, thereby improving the model prediction accuracy, reducing the training cost of training the text classification model, and improving the model prediction accuracy, thereby solving the technical problem in the related technology that the traditional text classification model still needs to collect a large amount of training data and manually mark labels during training, resulting in high training costs. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0026] Figure 1 It is a hardware structure block diagram of a mobile terminal of a text classification method based on lightweight sample data according to an embodiment of the present invention;
[0027] Figure 2 is a flowchart of a text classification method based on lightweight sample data according to an embodiment of the present invention;
[0028] Figure 3 is a flowchart of an optional text classification method based on lightweight sample data according to an embodiment of the present invention;
[0029] Figure 4is a schematic diagram of a text classification device based on lightweight sample data according to an embodiment of the present invention.
[0030] The above drawings include the following reference numerals:
[0031] 102, processor; 104, memory; 106, transmission device; 108, input and output devices. DETAILED DESCRIPTION
[0032] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.
[0033] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0034] As introduced in the background technology, the traditional text classification model in the related art still needs to collect a large amount of training data and manually mark labels during training, resulting in high training costs. In view of the above defects, a text classification method and device based on lightweight sample data are provided in an embodiment of the present invention.
[0035] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the accompanying drawings in the embodiments of the present invention.
[0036] The method embodiments provided in the embodiments of the present invention can be executed in a mobile terminal, a computer terminal or a similar computing device. Taking running on a mobile terminal as an example, Figure 1 1 is a hardware structure block diagram of a mobile terminal of a text classification method based on lightweight sample data according to an embodiment of the present invention. Figure 1 As shown, the mobile terminal may include one or more ( Figure 1Only one is shown in the figure) a processor 102 (the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data, wherein the mobile terminal may also include a transmission device 106 and an input / output device 108 for communication functions. It can be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the mobile terminal. Figure 1 More or fewer components as shown, or with Figure 1 Different configurations are shown.
[0037] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the text classification method based on lightweight sample data in the embodiment of the present invention. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, the above method is implemented. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 104 may further include a memory remotely arranged relative to the processor 102, and these remote memories can be connected to the mobile terminal via a network. Examples of the above network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof. The transmission device 106 is used to receive or send data via a network. The specific example of the above network may include a wireless network provided by a communication provider of the mobile terminal. In one example, the transmission device 106 includes a network adapter (Network Interface Controller, referred to as NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 106 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0038] According to an embodiment of the present invention, a method embodiment of a text classification method based on lightweight sample data is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0039] Figure 2 is a flowchart of a text classification method based on lightweight sample data according to an embodiment of the present invention. Figure 2 As shown, the method comprises the following steps:
[0040] Step S202, when receiving the text to be classified, segment the text to be classified according to semantics to obtain multiple word vectors, wherein the word vector is a vector representation of each word segmentation fragment in the text to be classified, and each word segmentation fragment includes at least one character in the text to be classified.
[0041] Optionally, the text to be classified refers to text to be classified.
[0042] Specifically, in NLP methods based on deep neural networks, characters / words in a text are usually represented by one-dimensional vectors (generally referred to as "word vectors"); therefore, when receiving a text to be classified input by a user, it can first be segmented to be divided into multiple word vectors; in particular, characters / words with similar semantics are usually expected to be close in distance in the feature vector space, so that the text vector converted from the character / word vector can also contain more accurate semantic information.
[0043] According to the above embodiment of the present invention, in the above step S202, when the text to be classified is received, the text to be classified is segmented according to semantics to obtain multiple word vectors, including: when the text to be classified is received, the text to be classified is input into the BERT model for processing, so as to use the BERT model to adopt a word segmentation strategy to segment the text to be classified, and obtain multiple word segmentation fragments, wherein the word segmentation strategy is a strategy embedded in the BERT model and based on character fragments for word segmentation; position codes are added to each word segmentation fragment according to the word order of the text to be classified; and the word vectors corresponding to each word segmentation fragment are generated in sequence according to the semantics and in the order of the position codes using the BERT model to obtain multiple word vectors.
[0044] Combine the following Figure 3 The above embodiments of the present invention are described in detail. Figure 3 is a flowchart of an optional text classification method based on lightweight sample data according to an embodiment of the present invention; Figure 3 As shown in the figure, after receiving the text input by the user, it can be input into the Bert model (i.e., Bert model) for processing, so as to use the word segmentation strategy in the Bert model to segment the input text and encode it to obtain the vector representation of each word, i.e., the word vector; the Bert model can generate a vector containing contextual information for each word according to the context. For example, for the word "good" in "very delicious", Bert will generate a vector that not only contains the inherent semantics of the word "good", but also combines the influence of "very" and "delicious" on "good".
[0045] In a specific embodiment of the present invention, after adding position codes to each word segmentation fragment according to the word order of the text to be classified, the text classification method based on lightweight sample data also includes: determining the word segmentation fragment at the starting position as the starting word segmentation fragment according to the word order of the text to be classified, and determining the word segmentation fragment at the ending position as the ending word segmentation fragment; adding a start mark to the position code of the starting word segmentation fragment, and adding an end mark to the position code of the ending word segmentation fragment.
[0046] Specifically, when encoding each word segment, it is also necessary to add start and end tags to the words at the start and end positions respectively according to the word order of the original input text, so that the start and end positions can be clearly identified when the text is subsequently processed using the relevant feature analysis model; if the text input by the user consists of multiple paragraphs, other tags can be added to the beginning and end of each paragraph to mark the beginning and end of a paragraph, thereby improving efficiency.
[0047] In particular, since the word vectors obtained by processing the classified text using the Bert model do not incorporate the information of other words in the sentence, in order to improve the accuracy of the analysis of the classified text, the understanding of the classified text can be more comprehensively improved by extracting local and global features to ensure the accuracy of the final classification of the classified text; in addition, such feature extraction method enriches the sample features to a certain extent and avoids overfitting, which makes up for the small amount of training data in the training stage of the text classification model. For some application scenarios where it is difficult to obtain a large amount of labeled data, it effectively reduces the cost of data preparation and model training, and improves the accuracy and efficiency of text classification.
[0048] Step S204, performing local feature association analysis on at least two adjacent word vectors in sequence to obtain multiple local features of the text to be classified.
[0049] like Figure 3 As shown, the Text-CNN model can be used to extract local features from the extracted word vectors to obtain multiple local features of the text to be classified.
[0050] According to the above embodiment of the present invention, in the above step S204, local feature association analysis is performed on at least two adjacent word vectors in turn to obtain multiple local features of the text to be classified, including: determining that the number of text spaces in the convolution kernel is a predetermined number, wherein the convolution kernel is a feature extractor for local feature extraction in the Text-CNN model, and the Text-CNN model is a model for text classification, and each text space in the convolution kernel is used to place a word vector; selecting multiple adjacent word vectors according to a predetermined number and putting them into the convolution kernel for local feature association analysis until all word vectors are placed in the convolution kernel at least once, and multiple initial local features are obtained; using the Text-CNN model to perform a pooling operation on each initial local feature to obtain multiple local features, wherein the pooling operation is used to make the feature dimension of the local feature lower than the dimension threshold.
[0051] Specifically, the convolution kernel in the Text-CNN model can be used to perform a sliding window operation on the word vector. The convolution kernel slides on the sequence of word vectors and performs convolution operations on the word vectors in the window to capture the combined features of the word vectors in the window, reflecting the local relationship between words, thereby obtaining multiple local features; convolution kernels of different sizes can capture n-grams of different lengths. For example, a convolution kernel of size 3 can capture the local features of three-word groups (3-grams) between words; then, the local features initially obtained (i.e., initial local features) can be pooled (such as maximum pooling or average pooling) to further reduce the feature dimension while retaining important local information, i.e., key local features in the text, while ignoring unimportant details; finally, the local features extracted from convolution kernels of different sizes can be fused together to form a comprehensive feature vector, which contains local pattern information of multiple different lengths in the text, providing rich detailed descriptions for subsequent text classification.
[0052] It should be noted that here, the multiple local features obtained can be first subjected to simple feature fusion (such as splicing), or can be subjected to feature fusion together with the global features when they are to be input into the text classification model, without specific limitation.
[0053] Step S206, performing global feature association analysis on each word vector according to the semantic associations between all word vectors, and obtaining global features corresponding to each word vector.
[0054] like Figure 3 As shown, a multi-head attention mechanism can be used to perform global feature analysis on each word vector based on the semantic associations between all word vectors to obtain multiple global features of the text to be classified.
[0055] According to the above embodiment of the present invention, in the above step S206, a global feature association analysis is performed on each word vector according to the semantic association between all word vectors to obtain a global feature corresponding to each word vector, including: determining any word vector as a target word vector; calculating the degree of influence of each candidate word vector on the target word vector based on the semantics of the text to be classified, and obtaining the attention weight of each candidate word vector, wherein the candidate word vector is a word vector other than the target word vector, the degree of influence refers to the similarity between the meaning and semantics expressed after the candidate word vector is combined with the target word vector, and the attention weight is used to quantify the degree of influence; using a multi-head attention mechanism to perform a global feature association analysis according to the attention weights of each candidate word vector to obtain a global feature corresponding to the target word vector.
[0056] Specifically, compared to the use of the Text-CNN model for local features, which can only perform feature association analysis on word vectors in adjacent positions, the multi-head attention mechanism can perform feature association analysis on word vectors that are far away, so as to perform semantic and type analysis on the classified text more comprehensively; the multi-head attention mechanism performs global feature association analysis on each word vector based on the semantic association between all word vectors, and obtains the global features corresponding to each word vector. The specific process can be as follows:
[0057] First, the word vector at each position in the input sequence is linearly transformed (projected) to generate three vectors: query vector (Query, Q), key vector (Key, K) and value vector (V), and the Q, K, and V vectors are split into H heads respectively. Each head has an independent weight matrix, that is, each head performs a different linear transformation on the input vector to obtain H groups of query vectors (Q1, Q2, ..., Q H ), H groups of key vectors (K1, K2, …, K H ) and H sets of value vectors (V1, V2, …, V H ); then, in each head, the dot product (or scaled dot product) is used to calculate the similarity score between the query vector and the key vector, which usually involves a normalization step (such as using the softmax function) to ensure that the sum of the scores is 1, forming an attention weight matrix; the calculated attention weight matrix can then be used to weighted sum the corresponding value vectors to generate the context vector of each head (that is, the global features corresponding to each word vector), and the context vectors of the H heads are spliced together to form a larger vector, which contains information from different attention heads; finally, the spliced or merged vector will pass through a linear transformation (fully connected layer) again to adjust it back to the feature dimension required by the model. This transformation helps to fuse the outputs of different heads into a unified representation for subsequent processing.
[0058] It should be noted that the multiple global features obtained here can be first fused, or they can be fused together with local features when they are input into the text classification model, without specific limitation.
[0059] Step S208: Input multiple local features and multiple global features into a text classification model for processing to obtain the text type of the text to be classified, wherein the text classification model is trained by deep learning using multiple sets of training data, and each of the multiple sets of training data includes: sample input features, and sample text types corresponding to the sample input features. The sample input features include: multiple sample local features and multiple sample global features corresponding to the sample input texts, and the number of sample input texts is not higher than a predetermined value.
[0060] like Figure 3 As shown, the local features and global features obtained in the above steps S204 and S206 can be spliced and input into the BiLSTM model (i.e., the text classification model in the embodiment of the present invention) for type recognition to obtain the text type of the text to be classified.
[0061] For example, the original sentence is input into the Bert model to obtain text vector 1, which contains word vectors, but the vector representation of each word does not integrate the information of other words in the sentence. The global vector of text vector 1 is extracted through the implementation of the multi-head attention mechanism to obtain text vector 2 (equivalent to the global feature in the embodiment of the present invention). At the same time, text vector 1 is input into the text-CNN model to extract the local text vector to obtain text vector 3 (equivalent to the local feature in the embodiment of the present invention). Finally, text vector 2 and text vector 3 are concatenated and sent to the BiLSTM model for text classification to obtain the text type of the original sentence.
[0062] In a specific embodiment of the present invention, before inputting multiple local features and multiple global features into a text classification model for processing to obtain the text type of the text to be classified, the text classification method based on lightweight sample data also includes: obtaining multiple groups of sample text data, wherein each of the multiple groups of sample text data includes: sample input text, sample text type corresponding to the sample input text, and the number of sample text data is not higher than a predetermined value; performing word segmentation processing and feature analysis on the sample input texts in the multiple groups of sample text data respectively to obtain sample input features corresponding to each sample input text, and the sample input features all include: multiple sample local features and multiple sample global features corresponding to the sample input text; using a bidirectional long short-term memory network to perform sequence modeling on the multiple groups of sample input features and sample text types to obtain a text classification model.
[0063] Specifically, feature analysis can be performed on a smaller amount of sample text data to obtain multiple sample local features and multiple sample global features corresponding to each sample input text, so as to enrich the sample features of the training model. Then, a bidirectional long short-term memory network is used to perform sequence modeling and model training on the multiple groups of sample input features obtained and their true types (sample text types) to obtain the corresponding text classification model.
[0064] It should be noted that the advantage of the text classification model proposed in the embodiment of the present invention is that it can be trained by a small amount of data samples and tested using its own data set. The number of test set samples is 2K, and the categories are 5. Finally, after data set evaluation, the data set F1 value (the harmonic mean of precision and recall) reaches 85%. The experimental results show that the model has a good classification effect.
[0065] In a preferred embodiment of the present invention, the text classification method based on lightweight sample data also includes: in the process of using the text classification model to process multiple local features and multiple global features to identify the type of the text to be classified, the text classification model is optimized using the multiple local features and multiple global features of the text to be classified and the text type of the text to be classified, and the text classification model is updated to the optimized text classification model.
[0066] Specifically, in the process of using the trained text classification model to identify the type of text input by the user, these data can be continuously used to adaptively optimize the model to improve the performance and accuracy of the text classification model.
[0067] Through the technical solution provided by the above-mentioned embodiment of the present invention, when a text to be classified is received, the text to be classified can be segmented according to semantics to obtain multiple word vectors, wherein the word vector is a vector representation of each word segmentation fragment in the text to be classified, and each word segmentation fragment includes at least one character in the text to be classified; local feature association analysis is performed on at least two adjacent word vectors in turn to obtain multiple local features of the text to be classified; global feature association analysis is performed on each word vector according to the semantic association between all word vectors to obtain global features corresponding to each word vector; multiple local features and multiple global features are input into a text classification model for processing to obtain the text type of the text to be classified, wherein the text classification model is trained by deep learning using multiple sets of training data, and the multiple sets of training data are trained by deep learning. Each group in the data includes: sample input features, sample text types corresponding to the sample input features, the sample input features include: multiple sample local features and multiple sample global features corresponding to the sample input text, the number of sample input texts is not higher than a predetermined value, and the purpose of semantically parsing the text input by the user, extracting features from local and global aspects respectively, extracting multiple local features and global features corresponding to the text, and using the text classification model to process the extracted features, and finally outputting the text type is achieved, and the technical effect of enriching the data volume by feature analysis and extraction of the text is achieved to make up for the defect of small sample data volume when training the text classification model, thereby improving the model prediction accuracy, reducing the training cost of training the text classification model, and improving the model prediction accuracy.
[0068] Therefore, the technical solution provided by the above-mentioned embodiment of the present invention solves the technical problem in the related art that the traditional text classification model still needs to collect a large amount of training data and manually mark labels during training, resulting in high training costs.
[0069] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the present application is not limited by the described order of actions, because according to the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present application.
[0070] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in each embodiment of the present application.
[0071] According to an embodiment of the present invention, there is also provided a text classification device based on lightweight sample data for implementing the above-mentioned text classification method based on lightweight sample data. Figure 4 is a schematic diagram of a text classification device based on lightweight sample data according to an embodiment of the present invention. Figure 4 As shown, the device comprises: a first acquisition unit 41, a second acquisition unit 43, a third acquisition unit 45 and a fourth acquisition unit 47. The text classification device based on lightweight sample data is described in detail below.
[0072] The first acquisition unit 41 is used to segment the text to be classified according to semantics when receiving the text to be classified, so as to obtain multiple word vectors, wherein the word vector is a vector representation of each word segmentation fragment in the text to be classified, and each word segmentation fragment includes at least one character in the text to be classified.
[0073] The second acquisition unit 43 is used to perform local feature association analysis on at least two adjacent word vectors in sequence to obtain multiple local features of the text to be classified.
[0074] The third acquisition unit 45 is used to perform global feature association analysis on each word vector according to the semantic association between all word vectors to obtain the global feature corresponding to each word vector.
[0075] The fourth acquisition unit 47 is used to input multiple local features and multiple global features into the text classification model for processing to obtain the text type of the text to be classified, wherein the text classification model is trained by deep learning using multiple groups of training data, and each of the multiple groups of training data includes: sample input features, and sample text types corresponding to the sample input features. The sample input features include: multiple sample local features and multiple sample global features corresponding to the sample input text, and the number of sample input texts is not higher than a predetermined value.
[0076] It should be noted here that the above-mentioned first acquisition unit 41, second acquisition unit 43, third acquisition unit 45 and fourth acquisition unit 47 correspond to steps S202 to S208 in the above-mentioned embodiments, and the four units are the same as the instances and application scenarios implemented by the corresponding steps, but are not limited to the contents disclosed in the above-mentioned embodiments.
[0077] As can be seen from the above, in the scheme recorded in the above-mentioned embodiment of the present invention, the first acquisition unit can be used to segment the text to be classified according to semantics when receiving the text to be classified, so as to obtain multiple word vectors, wherein the word vector is a vector representation of each word segmentation fragment in the text to be classified, and each word segmentation fragment includes at least one character in the text to be classified; then the second acquisition unit is used to perform local feature association analysis on at least two adjacent word vectors in turn to obtain multiple local features of the text to be classified; then the third acquisition unit is used to perform global feature association analysis on each word vector according to the semantic association between all word vectors to obtain the global feature corresponding to each word vector; finally, the fourth acquisition unit is used to input the multiple local features and the multiple global features into the text classification model for processing to obtain the text type of the text to be classified, wherein the text classification model is a text classification model that uses multiple sets of training data to obtain the text type of the text to be classified. The data is obtained by training in a deep learning manner, and each of the multiple groups of training data includes: sample input features, and sample text types corresponding to the sample input features. The sample input features include: multiple sample local features and multiple sample global features corresponding to the sample input texts. The number of sample input texts is not higher than a predetermined value, and the purpose of semantically parsing the text input by the user, extracting features from local and global perspectives respectively, extracting multiple local features and global features corresponding to the text, and processing the extracted features with a text classification model, and finally outputting the text type is achieved. The technical effect of enriching the data volume by performing feature analysis and extraction on the text is achieved to make up for the defect of a small amount of sample data when training the text classification model, thereby improving the model prediction accuracy, reducing the training cost of training the text classification model, and improving the model prediction accuracy.
[0078] Therefore, the technical solution provided by the above-mentioned embodiment of the present invention solves the technical problem in the related art that the traditional text classification model still needs to collect a large amount of training data and manually mark labels during training, resulting in high training costs.
[0079] In an optional embodiment of the present invention, the first acquisition unit includes: a first acquisition module, which is used to input the text to be classified into the Burt model for processing when receiving the text to be classified, so as to use the Burt model to adopt a word segmentation strategy to segment the text to be classified to obtain multiple word segmentation fragments, wherein the word segmentation strategy is a strategy embedded in the Burt model and based on character fragments for word segmentation; a first adding module, which is used to add position codes to each word segmentation fragment according to the word order of the text to be classified; and a second acquisition module, which is used to use the Burt model to generate word vectors corresponding to each word segmentation fragment in sequence according to the semantics and the order of the position codes to obtain multiple word vectors.
[0080] In an optional embodiment of the present invention, the text classification device based on lightweight sample data also includes: a first determination module, which is used to determine the word segmentation segment at the starting position as the starting word segmentation segment, and determine the word segment at the ending position as the ending word segmentation segment according to the word order of the text to be classified after adding position codes to each word segmentation segment according to the word order of the text to be classified; and a second adding module, which is used to add a start mark to the position code of the starting word segmentation segment, and to add an end mark to the position code of the ending word segment segment.
[0081] In an optional embodiment of the present invention, the second acquisition unit includes: a second determination module, used to determine that the number of text spaces in the convolution kernel is a predetermined number, wherein the convolution kernel is a feature extractor used for local feature extraction in the Text-CNN model, and the Text-CNN model is a model for text classification, and each text space in the convolution kernel is used to place a word vector; a third acquisition module, used to select multiple adjacent word vectors according to a predetermined number and put them into the convolution kernel for local feature association analysis, until all word vectors are placed in the convolution kernel at least once, and multiple initial local features are obtained; a fourth acquisition module, used to use the Text-CNN model to perform a pooling operation on each initial local feature to obtain multiple local features, wherein the pooling operation is used to make the feature dimension of the local feature lower than the dimension threshold.
[0082] In an optional embodiment of the present invention, the third acquisition unit includes: a third determination module, which is used to determine that any word vector is a target word vector; a fifth acquisition module, which is used to calculate the degree of influence of each candidate word vector on the target word vector based on the semantics of the text to be classified, and obtain the attention weight of each candidate word vector, wherein the candidate word vector is a word vector other than the target word vector, and the degree of influence refers to the similarity between the meaning and semantics expressed after the candidate word vector is combined with the target word vector, and the attention weight is used to quantify the degree of influence; a sixth acquisition module, which is used to use a multi-head attention mechanism to perform global feature association analysis according to the attention weights of each candidate word vector to obtain a global feature corresponding to the target word vector.
[0083] In an optional embodiment of the present invention, the text classification device based on lightweight sample data also includes: a fifth acquisition unit, which is used to acquire multiple groups of sample text data before inputting multiple local features and multiple global features into the text classification model for processing to obtain the text type of the text to be classified, wherein each of the multiple groups of sample text data includes: sample input text, sample text type corresponding to the sample input text, and the number of sample text data is not higher than a predetermined value; a sixth acquisition unit, which is used to perform word segmentation processing and feature analysis on the sample input texts in the multiple groups of sample text data respectively, to obtain sample input features corresponding to each sample input text, and the sample input features include: multiple sample local features and multiple sample global features corresponding to the sample input text; a seventh acquisition unit, which is used to use a bidirectional long short-term memory network to perform sequence modeling on multiple groups of sample input features and sample text types to obtain a text classification model.
[0084] In an optional embodiment of the present invention, the text classification device based on lightweight sample data also includes: an optimization unit, which is used to optimize the text classification model by using multiple local features and multiple global features of the text to be classified and the text type of the text to be classified in the process of using the text classification model to process multiple local features and multiple global features to identify the type of the text to be classified, and update the text classification model to the optimized text classification model.
[0085] According to another aspect of an embodiment of the present invention, a text classification system based on lightweight sample data is further provided. The text classification system based on lightweight sample data uses any of the above-mentioned text classification methods based on lightweight sample data.
[0086] According to another aspect of an embodiment of the present invention, a computer-readable storage medium is further provided. The computer-readable storage medium includes a stored program, wherein the program executes any one of the above-mentioned text classification methods based on lightweight sample data.
[0087] Optionally, in this embodiment, the computer-readable storage medium may be located in any one of the computer terminals in a computer terminal group in a computer network, or in any one of the communication devices in a communication device group.
[0088] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for executing the following steps: when receiving a text to be classified, segmenting the text to be classified according to semantics to obtain multiple word vectors, wherein the word vector is a vector representation of each word segmentation fragment in the text to be classified, and each word segmentation fragment includes at least one character in the text to be classified; performing local feature association analysis on at least two adjacent word vectors in turn to obtain multiple local features of the text to be classified; performing global feature association analysis on each word vector according to the semantic association between all word vectors to obtain a global feature corresponding to each word vector; inputting the multiple local features and the multiple global features into a text classification model for processing to obtain a text type of the text to be classified, wherein the text classification model is trained by deep learning using multiple sets of training data, and each of the multiple sets of training data includes: sample input features, and sample text types corresponding to the sample input features, and the sample input features include: multiple sample local features and multiple sample global features corresponding to the sample input text, and the number of sample input texts is not higher than a predetermined value.
[0089] Optionally, in this embodiment, the computer-readable storage medium is configured to store program codes for executing the following steps: when receiving a text to be classified, inputting the text to be classified into a BERT model for processing, so as to segment the text to be classified using a word segmentation strategy using the BERT model to obtain a plurality of word segmentation fragments, wherein the word segmentation strategy is a strategy embedded in the BERT model and based on character fragments for word segmentation; adding position codes to each word segmentation fragment according to the word order of the text to be classified; and using the BERT model to sequentially generate word vectors corresponding to each word segmentation fragment according to semantics and in the order of the position codes to obtain a plurality of word vectors.
[0090] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for executing the following steps: determining the word segmentation fragment at the starting position as the starting word segmentation fragment according to the word order of the text to be classified, and determining the word segmentation fragment at the ending position as the ending word segmentation fragment; adding a start mark to the position code of the starting word segmentation fragment, and adding an end mark to the position code of the ending word segmentation fragment.
[0091] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: determining that the number of text spaces in the convolution kernel is a predetermined number, wherein the convolution kernel is a feature extractor for local feature extraction in the Text-CNN model, the Text-CNN model is a model for text classification, and each text space in the convolution kernel is used to place a word vector; selecting multiple adjacent word vectors according to a predetermined number and putting them into the convolution kernel for local feature association analysis until all word vectors are placed in the convolution kernel at least once, thereby obtaining multiple initial local features; using the Text-CNN model to perform a pooling operation on each initial local feature to obtain multiple local features, wherein the pooling operation is used to make the feature dimension of the local feature lower than the dimension threshold.
[0092] Optionally, in this embodiment, a computer-readable storage medium is configured to store program code for performing the following steps: determining any word vector as a target word vector; calculating the degree of influence of each candidate word vector on the target word vector based on the semantics of the text to be classified, and obtaining the attention weight of each candidate word vector, wherein the candidate word vector is a word vector other than the target word vector, and the degree of influence refers to the similarity between the meaning and semantics expressed after the candidate word vector is combined with the target word vector, and the attention weight is used to quantify the degree of influence; using a multi-head attention mechanism to perform global feature association analysis based on the attention weights of each candidate word vector to obtain a global feature corresponding to the target word vector.
[0093] Optionally, in this embodiment, a computer-readable storage medium is configured to store a program code for executing the following steps: obtaining multiple groups of sample text data, wherein each of the multiple groups of sample text data includes: sample input text, a sample text type corresponding to the sample input text, and the amount of sample text data is not higher than a predetermined value; performing word segmentation processing and feature analysis on the sample input texts in the multiple groups of sample text data respectively to obtain sample input features corresponding to each sample input text, wherein the sample input features include: multiple sample local features and multiple sample global features corresponding to the sample input text; using a bidirectional long short-term memory network to perform sequence modeling on the multiple groups of sample input features and sample text types to obtain a text classification model.
[0094] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: in the process of using a text classification model to process multiple local features and multiple global features to perform type identification on the text to be classified, the text classification model is optimized using multiple local features and multiple global features of the text to be classified and the text type of the text to be classified, and the text classification model is updated to the optimized text classification model.
[0095] According to another aspect of an embodiment of the present invention, a processor is further provided, and the processor is used to run a program, wherein the program executes any one of the above-mentioned text classification methods based on lightweight sample data when running.
[0096] According to another aspect of an embodiment of the present invention, a computer program product is provided, including computer instructions, which, when executed by a processor, execute any one of the above-mentioned text classification methods based on lightweight sample data.
[0097] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.
[0098] In the above embodiments of the present invention, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0099] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the units can be a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
[0100] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
[0101] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0102] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, magnetic disk or optical disk and other media that can store program codes.
[0103] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A text classification method based on lightweight sample data, characterized in that: include: When receiving a text to be classified, segment the text to be classified according to semantics to obtain multiple word vectors, wherein the word vector is a vector representation of each word segment in the text to be classified, and each word segment includes at least one character in the text to be classified; Performing local feature association analysis on at least two adjacent word vectors in sequence to obtain multiple local features of the text to be classified; Performing global feature association analysis on each of the word vectors according to the semantic associations between all the word vectors to obtain a global feature corresponding to each of the word vectors; Input the multiple local features and the multiple global features into a text classification model for processing to obtain the text type of the text to be classified, wherein the text classification model is trained by deep learning using multiple groups of training data, and each of the multiple groups of training data includes: sample input features, and sample text types corresponding to the sample input features, and the sample input features include: multiple sample local features and multiple sample global features corresponding to sample input texts, and the number of the sample input texts is not higher than a predetermined value.
2. The text classification method based on lightweight sample data according to claim 1, characterized in that: When receiving the text to be classified, the text to be classified is segmented according to semantics to obtain multiple word vectors, including: When the text to be classified is received, the text to be classified is input into the BERT model for processing, so as to segment the text to be classified using the BERT model using a word segmentation strategy to obtain a plurality of word segmentation fragments, wherein the word segmentation strategy is a strategy embedded in the BERT model and based on character fragments for word segmentation; Adding position codes to each of the word segments according to the word order of the text to be classified; The BERT model is used to generate the word vector corresponding to each word segment in sequence according to the semantics and in the order of the position encoding to obtain a plurality of word vectors.
3. The text classification method based on lightweight sample data according to claim 2 is characterized in that: After adding position codes to each of the word segments according to the word order of the text to be classified, the method further includes: Determine the word segment at the starting position as the starting word segment according to the word order of the text to be classified, and determine the word segment at the ending position as the ending word segment; A start tag is added to the position code of the start word segment segment, and an end tag is added to the position code of the end word segment segment.
4. The text classification method based on lightweight sample data according to claim 1, characterized in that: Performing local feature association analysis on at least two adjacent word vectors in sequence to obtain multiple local features of the text to be classified, including: Determining that the number of text spaces in a convolution kernel is a predetermined number, wherein the convolution kernel is a feature extractor for performing local feature extraction in a Text-CNN model, the Text-CNN model is a model for performing text classification, and each of the text spaces in the convolution kernel is used to place one of the word vectors; Selecting a plurality of adjacent word vectors according to the predetermined number and placing them into the convolution kernel for local feature association analysis, until all the word vectors are placed into the convolution kernel at least once, thereby obtaining a plurality of initial local features; The Text-CNN model is used to perform a pooling operation on each of the initial local features to obtain a plurality of the local features, wherein the pooling operation is used to make the feature dimension of the local feature lower than a dimension threshold.
5. The text classification method based on lightweight sample data according to claim 1, characterized in that: Performing global feature association analysis on each of the word vectors according to the semantic associations between all the word vectors to obtain global features corresponding to each of the word vectors, including: Determine any of the word vectors as a target word vector; Based on the semantics of the text to be classified, the influence degree of each candidate word vector on the target word vector is calculated to obtain the attention weight of each candidate word vector, wherein the candidate word vector is the word vector other than the target word vector, the influence degree refers to the similarity between the meaning expressed by the candidate word vector after combining with the target word vector and the semantics, and the attention weight is used to quantify the influence degree; A multi-head attention mechanism is used to perform global feature association analysis based on the attention weights of each candidate word vector to obtain the global feature corresponding to the target word vector.
6. The text classification method based on lightweight sample data according to claim 1, characterized in that: Before inputting the plurality of local features and the plurality of global features into a text classification model for processing to obtain the text type of the text to be classified, the method further includes: Acquire multiple groups of sample text data, wherein each of the multiple groups of sample text data includes: the sample input text, the sample text type corresponding to the sample input text, and the amount of the sample text data is not greater than the predetermined value; Performing word segmentation processing and feature analysis on the sample input texts in the multiple groups of sample text data respectively, to obtain sample input features corresponding to each sample input text, wherein the sample input features include: multiple sample local features and multiple sample global features corresponding to the sample input text; A bidirectional long short-term memory network is used to perform sequence modeling on multiple groups of sample input features and the sample text types to obtain the text classification model.
7. The text classification method based on lightweight sample data according to any one of claims 1 to 6, characterized in that: The method further comprises: In the process of using the text classification model to process the multiple local features and the multiple global features to identify the type of the text to be classified, the text classification model is optimized using the multiple local features and the multiple global features of the text to be classified and the text type of the text to be classified, and the text classification model is updated to the optimized text classification model.
8. A text classification device based on lightweight sample data, characterized in that: include: A first acquisition unit is used for, when receiving a text to be classified, segmenting the text to be classified according to semantics to obtain a plurality of word vectors, wherein the word vector is a vector representation of each word segmentation fragment in the text to be classified, and each word segmentation fragment includes at least one character in the text to be classified; A second acquisition unit is used to perform local feature association analysis on at least two adjacent word vectors in sequence to obtain multiple local features of the text to be classified; A third acquisition unit is used to perform global feature association analysis on each of the word vectors according to the semantic associations between all the word vectors to obtain a global feature corresponding to each of the word vectors; A fourth acquisition unit is used to input the multiple local features and the multiple global features into a text classification model for processing to obtain the text type of the text to be classified, wherein the text classification model is trained by deep learning using multiple groups of training data, and each of the multiple groups of training data includes: sample input features, and sample text types corresponding to the sample input features, and the sample input features include: multiple sample local features and multiple sample global features corresponding to sample input texts, and the number of sample input texts is not higher than a predetermined value.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored program, wherein the program executes the text classification method based on lightweight sample data according to any one of claims 1 to 7.
10. A computer program product comprising computer instructions, characterized in that: When the computer instructions are executed by a processor, the text classification method based on lightweight sample data described in any one of claims 1 to 7 is performed.