Text processing method, device, electronic device and storage medium
Through the proposed text processing methods, including feature extraction, encoding and decoding and text classification/clustering, the problem of poor text processing accuracy caused by ignoring the category label relationship in the prior art is solved, and higher text classification and clustering accuracy are achieved.
Patent Information
- Application Number
- CN202111350160.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-15
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2041-11-15
AI Technical Summary
The prior art ignores the relationship between category labels when processing text, resulting in poor accuracy of text classification and clustering.
A text processing method is proposed, by obtaining the original text, performing feature extraction, encoding and decoding processing, obtaining text implicit feature vectors, and label classification or clustering is performed through preset text classification or clustering models.
Improve the accuracy of text classification and clustering, and enhance the accuracy of text processing by capturing the relationship between category labels.
Smart Images

Figure CN114064894B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and particularly to a text processing method, apparatus, electronic device, and storage medium. Background Art
[0002] Currently, when processing text, the multi-label text classification / clustering task is often divided into multiple single-label binary classification / clustering tasks, and the relationship between the text to be processed and the category labels is used for classification / clustering. Although this method can capture the relationship between the text to be processed and the category labels, it ignores the relationship between the category labels, resulting in poor accuracy of text processing. Therefore, how to provide a text processing method that can improve the accuracy of text classification or text clustering has become an urgent technical problem to be solved. Summary of the Invention
[0003] The main objective of the embodiments of this application is to propose a text processing method, apparatus, electronic device, and storage medium, aiming to improve the accuracy of text classification or text clustering.
[0004] To achieve the above objective, a first aspect of the embodiments of this application proposes a text processing method, and the method includes:
[0005] Obtain the original text to be processed;
[0006] Extract features from the original text to obtain target text data;
[0007] Perform encoding processing on the target text data to obtain a text implicit feature vector;
[0008] Perform decoding processing on the text implicit feature vector to obtain a target text vector;
[0009] The method further includes:
[0010] Perform label classification processing on the target text vector through a preset text classification model and text category labels to obtain a target classification text including the text category labels;
[0011] Or,
[0012] Perform clustering processing on the target text vector through a preset text clustering model and text clustering labels to obtain a target clustering text set.
[0013] In some embodiments, the step of extracting features from the original text to obtain target text data includes:
[0014] Identify the text entity features in the original text;
[0015] Use a pre-trained sequence classifier to perform feature classification processing on the text entity features to obtain first text features;
[0016] Perform feature extraction on the first text features to obtain target text data.
[0017] In some embodiments, the step of decoding the text implicit feature vector to obtain a target text vector includes:
[0018] Perform data resampling processing on the text implicit feature vector to obtain an intermediate text vector;
[0019] Perform decoding processing on the intermediate text vector to obtain a target text vector.
[0020] In some embodiments, the step of encoding the target text data to obtain a text implicit feature vector includes:
[0021] Map the target text data to a preset vector space to obtain target text features;
[0022] According to a preset encoding order and encoding dimension, perform encoding processing on the target text features to obtain a text implicit feature vector.
[0023] In some embodiments, the step of performing label classification processing on the target text vector through a preset text classification model and text category labels to obtain a target classification text including the text category labels includes:
[0024] Perform label classification processing on the target text vector according to a preset classification function and text category labels to obtain a label text vector;
[0025] Perform semantic analysis processing on the label text vector to obtain a target classification text.
[0026] In some embodiments, the step of performing semantic analysis processing on the label text vector to obtain a target classification text includes:
[0027] Calculate the similarity between the label text vector and a reference text vector;
[0028] According to the similarity, perform screening processing on text segments in a preset text word library to obtain standard text segments;
[0029] Perform splicing processing on the standard text segments to obtain a target classification text.
[0030] In some embodiments, the step of encoding the target text data to obtain a text implicit feature vector includes:
[0031] Map the target text data to a preset vector space to obtain target text features;
[0032] Encode the target text features according to a preset encoding order and encoding dimension to obtain a text hidden vector with a preset feature dimension;
[0033] Perform weighted processing on the text hidden vector according to a preset weight ratio to obtain a text implicit feature vector.
[0034] In some embodiments, the step of clustering the target text vector by a preset text clustering model and text clustering label to obtain a target clustering text set includes:
[0035] Cluster the target text vector according to a preset clustering algorithm and text clustering label to obtain target clustering texts including text clustering labels;
[0036] Incorporate the target clustering texts with the same text clustering label into the same set to obtain a target clustering text set.
[0037] To achieve the above object, a second aspect of the embodiments of the present application proposes a text processing device, the device includes:
[0038] An original text acquisition module, configured to acquire an original text to be processed;
[0039] A feature extraction module, configured to extract features from the original text to obtain target text data;
[0040] An encoding processing module, configured to perform encoding processing on the target text data to obtain a text implicit feature vector;
[0041] A decoding processing module, configured to perform decoding processing on the text implicit feature vector to obtain a target text vector;
[0042] A text processing module, configured to perform label classification processing on the target text vector by a preset text classification model and text category label to obtain a target classification text including the text category label; or configured to cluster the target text vector by a preset text clustering model and text clustering label to obtain a target clustering text set.
[0043] To achieve the above object, a third aspect of the embodiments of the present application proposes an electronic device, the electronic device includes a memory, a processor, a computer program stored on the memory and executable on the processor, and a data bus for realizing connection communication between the processor and the memory, and when the computer program is executed by the processor, it realizes the method described in the first aspect above.
[0044] To achieve the above object, a fourth aspect of the embodiments of the present application provides a computer-readable storage medium for computer-readable storage. The computer-readable storage medium stores one or more computer programs, and the one or more computer programs can be executed by one or more processors to implement the method described in the first aspect above.
[0045] The text processing method, device, electronic device, and storage medium provided by the present application obtain the original text to be processed; extract features from the original text to obtain target text data, which can effectively eliminate the data with low relevance in the original text and reduce the total amount of data. Furthermore, encode the target text data to obtain a text latent feature vector; decode the text latent feature vector to obtain a target text vector; finally, either perform label classification processing on the target text vector through a preset text classification model and text category labels to obtain a target classification text containing text category labels, which can classify the target text according to the text category and improve the relevance of the target classification text in each text category, thereby improving the accuracy of text classification; or perform clustering processing on the target text vector through a preset text clustering model and text clustering labels to obtain a target clustering text set. By clustering the target text, the target texts with high relevance can be grouped into one category according to the preset text clustering labels to obtain the target clustering text set, thereby improving the accuracy of text clustering. Description of the Drawings
[0046] Figure 1 is a flowchart of the text processing method provided by the embodiments of the present application;
[0047] Figure 2 is Figure 1 a flowchart of step S102 in
[0048] Figure 3 is Figure 1 a flowchart of step S103 in
[0049] Figure 4 is Figure 1 a flowchart of step S104 in
[0050] Figure 5 is Figure 1 a flowchart of step S105 in
[0051] Figure 6 is Figure 5 a flowchart of step S502 in
[0052] Figure 7 is Figure 1 another flowchart of step S103 in
[0053] Figure 8 is Figure 1 Another flowchart of step S105 in
[0054] Figure 9 is a schematic structural diagram of a text processing device provided by an embodiment of the present application;
[0055] Figure 10 is a schematic hardware structure diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0056] In order to make the objectives, technical solutions and advantages of the present application more clear and understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0057] It should be noted that although the functional modules are divided in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from the module division in the device or the order in the flowchart. Terms such as "first" and "second" in the specification, claims and the above-mentioned drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence.
[0058] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.
[0059] First, several nouns involved in the present application are analyzed:
[0060] Artificial intelligence (AI): It is a new technical science that studies, develops theories, methods, technologies and application systems for simulating, extending and expanding human intelligence; artificial intelligence is a branch of computer science. Artificial intelligence attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. The research in this field includes robots, speech recognition, image recognition, natural language processing and expert systems, etc. Artificial intelligence can simulate the information process of human consciousness and thinking. Artificial intelligence also uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results of theories, methods, technologies and application systems.
[0061] Natural Language Processing (NLP): NLP uses computers to process, understand, and apply human languages (such as Chinese, English, etc.). NLP is a branch of artificial intelligence and an interdisciplinary field of computer science and linguistics, often referred to as computational linguistics. Natural language processing includes syntactic analysis, semantic analysis, discourse understanding, etc. Natural language processing is commonly used in technical fields such as machine translation, handwritten and printed character recognition, speech recognition and text-to-speech conversion, information retrieval, information extraction and filtering, text classification and clustering, public opinion analysis, and opinion mining. It involves data mining, machine learning, knowledge acquisition, knowledge engineering, artificial intelligence research related to language processing, and linguistic research related to language computing, etc.
[0062] Information Extraction (NER): A text processing technology that extracts factual information such as entities, relationships, events of a specified type from natural language texts and forms structured data output. Information extraction is a technology for extracting specific information from text data. Text data is composed of some specific units, such as sentences, paragraphs, and passages. Text information is precisely composed of some small specific units, such as characters, words, phrases, sentences, paragraphs, or combinations of these specific units. Extracting noun phrases, personal names, place names, etc. from text data are all text information extraction. Of course, the information extracted by text information extraction technology can be various types of information.
[0063] Variational Auto-Encoder (VAE): The variational auto-encoder is an important type of generative model. During training, the variational auto-encoder adds regularization to prevent overfitting, ensuring that the latent space has sufficient capacity for the generation process in the auto-encoder. The distribution generated by the encoder is chosen as the normal distribution, and the encoder can be trained to return statistics such as the mean and covariance matrix that describe these normal distributions. An input is encoded into a distribution because it can naturally express the global and local regularization of the latent space. Locally, it is due to the control of variance, and globally, it is due to the control of the mean. The loss function of the variational auto-encoder consists of a reconstruction term (last layer) and a regularization term (latent layer). The regularization term is represented by the KL divergence between the generated distribution and the normal distribution. Among them, the role of regularization is to enable the latent space to perform the generation process, so it needs to meet the following two characteristics: continuity and integrity. Continuity can be understood as that two adjacent points in the latent layer should be approximately the same after decoding; integrity can be understood as that the content decoded from the points sampled from the distribution should be meaningful. Simply turning the points in the latent layer into a distribution is not sufficient to meet the above two characteristics. Therefore, a good regularization term needs to be defined, that is, the distribution generated by the encoder is close to the standard normal distribution, the covariance matrix is close to the identity matrix, and the mean is 0. This regularization term can prevent the model from encoding data remotely in the latent space and encourage as much "overlap" of the returned distributions as possible, thus meeting the expected continuity and integrity conditions. The regularization term will increase the reconstruction loss, so the two losses need to be balanced during training.
[0064] Batch: The batch size is a hyperparameter used to define the number of samples to be processed before updating the internal model parameters, that is, to control the number of training samples before updating the internal parameters of the model. The training dataset can be divided into one or more batches. Among them, when all training samples are used to create one batch, the learning algorithm is called batch gradient descent; when the batch size is the size of one sample, the learning algorithm is called stochastic gradient descent; when the batch size is more than one sample and less than the size of the training dataset, the learning algorithm is called mini-batch gradient descent. The batch size is multiple samples processed before updating the model.
[0065] Encoding: It is to transform the input sequence into a vector of a fixed length;
[0066] Decoding: It is to transform the previously generated fixed vector back into an output sequence; among them, the input sequence can be text, speech, image, video; the output sequence can be text, image.
[0067] Latent variable: A latent variable is an unobservable random variable. Usually, inferences about latent variables are made based on samples of observable variables. Take the Gaussian mixture model as an example. In a GMM, the latent variable refers to the Gaussian component corresponding to each observation. Since the generation process is unobservable (or hidden), it is named the latent variable. We can make inferences about the latent variable by collecting samples.
[0068] Upsampling: Upsampling refers to enlarging an image, also known as image interpolation. Its main purpose is to enlarge the original image so that it can be displayed on a display device with a higher resolution. Principle of upsampling: Almost all image enlargements use interpolation methods, that is, on the basis of the original image pixels, appropriate interpolation algorithms are used to insert new elements between pixel points. Interpolation algorithms mainly include edge-based image interpolation algorithms and region-based image interpolation algorithms.
[0069] Subsampling: Subsampling refers to reducing the size of an image, also known as downsampling. Its main purpose is to make the image fit the size of the display area and to generate a thumbnail of the corresponding image. Principle of subsampling: For an image I with size M*N, if it is subsampled by a factor of s, a lower-resolution image of size (M / s)*(N / s) is obtained. Of course, s should be a common divisor of M and N. If considering the image in matrix form, it means turning the image within an s*s window of the original image into a single pixel, and the value of this pixel is the average of all pixels within the window.
[0070] Batch: The batch size is a hyperparameter used to define the number of samples to be processed before updating the internal model parameters, that is, to control the number of training samples before updating the internal parameters of the model. The training dataset can be divided into one or more batches. Among them, when all training samples are used to create one batch, the learning algorithm is called batch gradient descent; when the batch size is the size of one sample, the learning algorithm is called stochastic gradient descent; when the batch size is more than one sample and less than the size of the training dataset, the learning algorithm is called mini-batch gradient descent. The batch size is multiple samples processed before updating the model.
[0071] Backpropagation: The general principle of backpropagation is as follows: Input the training set data into the input layer of the neural network, pass through the hidden layer of the neural network, and finally reach the output layer of the neural network and output the result. Since there is an error between the output result of the neural network and the actual result, calculate the error between the estimated value and the actual value, and backpropagate this error from the output layer to the hidden layer until it reaches the input layer. During the backpropagation process, adjust the values of various parameters according to the error. Continuously iterate the above process until convergence.
[0072] Collaborative filtering algorithm: A relatively well-known and commonly used recommendation algorithm. It discovers users' preference biases based on the mining of users' historical behavior data, predicts the products that users may like for recommendation, or finds similar users (user-based) or items (item-based). The implementation of the user-based collaborative filtering algorithm mainly needs to solve two problems. One is how to find people with similar hobbies to you, that is, to calculate the similarity of data.
[0073] Text categorization: Given a classification system, each text in the text collection is assigned to one or several categories, and this process is called text categorization. Text categorization is a supervised learning process. The text categorization process can be divided into manual categorization and automatic categorization. The most famous example of the former is the web page classification system of Yahoo. The classification system was defined by experts, and then the web pages were classified manually. This method requires a large amount of manpower and is rarely used in reality. Automatic text categorization algorithms can be roughly divided into two categories: knowledge engineering methods and machine learning methods. The knowledge engineering method means that experts define some rules for each category, and these rules represent the characteristics of this category, and automatically classify the documents that meet the rules into the corresponding categories. The most famous system in this regard is CONSTRUE. After the 1990s, machine learning methods became dominant. Compared with the knowledge engineering method, machine learning methods can achieve similar accuracy, but reduce a large amount of manual participation.
[0074] Text clustering: Grouping a text collection into multiple classes or clusters so that the text content in the same cluster has a high degree of similarity, while the text content in different clusters is quite different. This process is called text clustering. Text clustering is an unsupervised learning process. Text clustering has many applications, such as improving the recall rate of IR systems, navigating / organizing electronic resources, etc. According to the characteristics of the clusters formed, clustering techniques are usually divided into hierarchical clustering and partitional clustering. A typical example of the former is the agglomerative hierarchical clustering algorithm, and a typical example of the latter is the k-means algorithm. In recent years, some new clustering algorithms have emerged, which are based on different theories or technologies, such as graph theory, fuzzy set theory, neural networks, and kernel techniques, etc.
[0075] In the related art, text that may belong to multiple categories simultaneously is called multi-label text. With the development of artificial intelligence technology, multi-label text classification and text clustering methods based on machine learning have been widely used. Currently, when processing text, the multi-label text classification / clustering task is often divided into multiple single-label binary classification / clustering tasks, and the relationship between the text to be processed and the category labels is used for classification / clustering. Although this method can capture the relationship between the text to be processed and the category labels, it ignores the relationship between category labels, resulting in poor accuracy of text processing. Therefore, how to provide a text processing method that can improve the accuracy of text classification and text clustering has become a technical problem to be solved urgently.
[0076] Based on this, the embodiments of the present application provide a text processing method, apparatus, electronic device, and storage medium, which can improve the accuracy of text classification and text clustering.
[0077] The text processing method, apparatus, electronic device, and storage medium provided by the embodiments of the present application are specifically described through the following embodiments. First, the text processing method in the embodiments of the present application is described.
[0078] The embodiments of the present application can obtain and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is a theory, method, technology, and application system that uses a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.
[0079] Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0080] The text processing method provided by the embodiments of this application relates to the field of artificial intelligence technology. The text processing method provided by the embodiments of this application can be applied to a terminal, or to a server side, or can also be software running on a terminal or a server side. In some embodiments, the terminal can be a smart phone, a tablet computer, a notebook computer, a desktop computer, etc.; the server side can be configured as an independent physical server, or can be configured as a server cluster or a distributed system composed of multiple physical servers, or can also be configured as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application that implements the text processing method, etc., but is not limited to the above forms.
[0081] This application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet-type devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This application can also be practiced in a distributed computing environment, where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0082] Figure 1 is an optional flowchart of the text processing method provided by the embodiments of this application, Figure 1 The method in may include but is not limited to steps S101 to S105:
[0083] Step S101, obtain the original text to be processed;
[0084] Step S102, perform feature extraction on the original text to obtain target text data;
[0085] Step S103, perform encoding processing on the target text data to obtain a text implicit feature vector;
[0086] Step S104, perform decoding processing on the text implicit feature vector to obtain a target text vector;
[0087] Step S105: Perform label classification processing on the target text vector through a preset text classification model and text category labels to obtain a target classification text containing text category labels; or perform clustering processing on the target text vector through a preset text clustering model and text clustering labels to obtain a target clustering text set.
[0088] In step S101 of some embodiments, data can be crawled purposefully after setting up a data source by writing a web crawler to obtain the original text to be processed. It should be noted that the original text is a natural language text.
[0089] Please refer to Figure 2 , in some embodiments, step S102 may include but is not limited to steps S201 to S203:
[0090] Step S201: Identify the text entity features in the original text;
[0091] Step S202: Perform feature classification processing on the text entity features using a pre-trained sequence classifier to obtain the first text feature;
[0092] Step S203: Extract features from the first text feature to obtain the target text data.
[0093] Specifically, in step S201, a preset lexical analysis model is used to identify the text entity features in the original text. For example, a text data lexicon is pre-constructed, and the text data lexicon can include proper nouns, terms, non-proper names, etc. related to various text types. Through this text data lexicon, the preset lexical analysis model can identify the text entity features in the original text according to the specific text corpus contained in the text data lexicon and the preset part-of-speech categories. The text entity features can include entity words in multiple dimensions such as text proper nouns, terms, non-proper names, modifiers, and time information related to the current requirements.
[0094] To more accurately extract text entity features, in step S202, a pre-trained sequence classifier can also be used to label the text entity features, so that these text entity features can all carry preset labels to improve the classification efficiency. Specifically, the pre-trained sequence classifier can be a maximum entropy Markov model (MEMM model), or a model based on the conditional random field algorithm (CRF), or a model based on the bidirectional long short-term memory algorithm (bi-LSTM). For example, a sequence classifier can be constructed based on the bi-LSTM algorithm. In the model based on the bi-LSTM algorithm, the input word wi and character embeddings are passed through the long short-term memory from left to right and the long short-term memory from right to left, so that a single output layer is generated at the position where the outputs are connected. Through this output layer, the sequence classifier can directly transfer the input text entity features to the softmax classifier, and the softmax classifier creates a probability distribution over the preset part-of-speech category labels, thereby classifying and labeling the text entity features according to the probability distribution to obtain the first text features, and these first text features are text entity features including the target text parameters.
[0095] Finally, step S203 is executed to perform convolution processing on the first text features to extract the first text features and obtain the required target text data.
[0096] In some embodiments, before step S103, the method further includes pre-constructing and training a text processing model, where the text processing model is a variational auto-decoding model. The text processing model includes a plurality of dense layers and a plurality of convolutional layers, and there are skip connections between the dense layers and the convolutional layers. Through the skip connections between the dense layers and the convolutional layers, the gradient loss can be reduced and the fitting performance of the text processing model can be improved.
[0097] To implement the encoding and decoding processing of the target text data, the above text processing model includes an encoding module and a decoding module, where the encoding module includes at least one downsampling network, and the decoding module includes at least one upsampling network. The training process of the text processing model includes, but is not limited to, the following steps a to f:
[0098] Step a, obtaining sample texts;
[0099] Step b, performing multiple mapping processes on the sample texts to obtain sample text features;
[0100] Step c, inputting the sample text features into the initial model;
[0101] Step d, performing normalization processing on the sample text features through the initial model to obtain a batch normalization matrix and a channel normalization matrix;
[0102] Step e, perform matrix multiplication on the batch normalization matrix and the channel normalization matrix according to the preset weight vector to obtain a normalization value;
[0103] Step f, optimize the loss function of the initial model according to the normalization value to update the initial model and obtain a text processing model.
[0104] Specifically, the initial model is a variational auto-decoding model. First, perform step a to obtain a sample text containing text category labels. Furthermore, perform step b, and use the MLP network to perform multiple mapping processes on the sample text to obtain sample text features. Among them, the size of the sample text features is [N, C, H, W], where N represents the quantity, C represents the number of channels, H represents the height, and W represents the width.
[0105] Furthermore, perform step c, and input the sample text features into the initial model.
[0106] When performing step d, normalize the quantity, height, and width of the sample text features in the batch hierarchical dimension (batch). The algorithm process is as follows: calculate the mean u of each batch along the channel direction; calculate the variance σ of each batch along the channel direction 2 ; perform normalization processing on the input sample text features x. The specific calculation formula is shown in formula (1); introduce a scaling variable γ and a translation variable β to obtain the batch normalization matrix as y = γ * x + β. In addition, it is also necessary to perform normalization processing on the number of channels, height, and width of the sample text features in the channel dimension to obtain a channel normalization matrix; among them, the mean of each channel and variance are shown in formulas (2) and (3):
[0107]
[0108]
[0109]
[0110] When performing step e, perform matrix multiplication on the batch normalization matrix and the channel normalization matrix according to the preset weight vector to obtain a normalization value.
[0111] Finally, perform step f. Calculate the model loss of the initial model, i.e., the loss value, according to the normalized value and the preset loss function. Then, use the gradient descent method to perform backpropagation on the loss value, feedback the loss value back to the initial model, and modify the model parameters of the initial model. Repeat the above process until the loss value meets the preset iteration condition. The preset iteration condition can be that the number of iterations reaches a preset value, or the variance of the change in the loss function is less than a preset threshold. When the loss value meets the preset iteration condition, backpropagation can be stopped, and the final model parameters are used as the final model parameters to complete the update of the initial model and obtain the text processing model.
[0112] It should be noted that in the embodiments of the present application, the above model loss may include reconstruction loss, KL divergence loss, and regularization loss. That is, the absolute difference between the original text and the reconstructed text is defined by the reconstruction loss; the difference between the prior distribution and the posterior distribution in the latent variable dimension is defined by the KL divergence loss; the regularization loss can better control the problem of KL divergence, making the entire model smoother. The calculation and optimization of the above model loss are beneficial to the stability of model training.
[0113] Please refer to Figure 3 , in some embodiments, step S103 may include but is not limited to steps S301 to S302:
[0114] Step S301, map the target text data to a preset vector space to obtain target text features;
[0115] Step S302, perform encoding processing on the target text features according to the preset encoding order and encoding dimension to obtain a text latent feature vector.
[0116] Specifically, in step S301, an MLP network can be used to perform multiple mapping processes on the target text data from the semantic space to the vector space, map the target text data into a preset vector space, and obtain target text features. The target text features can be text features or image features.
[0117] Furthermore, step S302 can be executed. Through the encoding module of the above text processing model, the target text features can be encoded according to the bottom-up encoding order and encoding dimension. For example, the target text features are initially encoded to obtain the bottom-layer text latent feature vector z1, and then downsampling processing is performed layer by layer upward to obtain the corresponding text latent feature vectors [z2, z3..., zk] for each layer.
[0118] Further, to improve the encoding quality, the encoding module includes an encoder and a downsampling unit. The stride of the convolutional layer of the encoder is 1, and the input feature and output feature sizes of the convolutional layer of the encoder are the same. The stride of the convolutional layer of the downsampling unit is 2, and the output feature size of the downsampling unit is half of the input feature size. Herein, the input feature can be an image feature or a text feature.
[0119] Through the encoding process of the target text data by the encoding module of the above-mentioned pre-trained text processing model, the obtained text latent feature vector z is no longer a distribution, but multiple distributions [z1, z2, …, zk] in different dimensions. Compared with the method of mapping high-dimensional text information to the low-dimensional latent variable layer z in the traditional technology, this method can effectively avoid the loss of target text data and effectively improve the text quality of the reconstructed text.
[0120] Please refer to Figure 4 , in some embodiments, step S104 may include but is not limited to steps S401 to S402:
[0121] Step S401, perform data resampling on the text latent feature vector to obtain an intermediate text vector;
[0122] Step S402, perform decoding on the intermediate text vector to obtain a target text vector.
[0123] Specifically, in step S401, at least one of the nearest neighbor interpolation method, bilinear interpolation method, and cubic convolution interpolation method can be used to perform data resampling on each text latent feature vector, that is, the gray values of the text latent feature vector are collected at a certain interval, and the collected gray values are analyzed. When the collected gray value is not within the numerical set of the original function at the sampling point, the nearest neighbor interpolation method, bilinear interpolation method, or cubic convolution interpolation method is used to perform interpolation on the sampled points to obtain multiple distributions [Y1, Y2, …, Yk] of the target text data in different dimensions, that is, multiple intermediate text vectors.
[0124] Furthermore, step S402 can be executed to perform decoding and upsampling on the intermediate text vector through the decoding module of the above-mentioned text processing model. This decoding process is symmetric to the aforementioned encoding process. For example, decoding is performed on the intermediate text vectors in different dimensions, and then upsampling is performed layer by layer upward to implement the decoding and upsampling of all-dimensional intermediate text vectors, thereby obtaining the target text vector.
[0125] Further, to improve the decoding quality, the decoding module includes a decoder and an upsampling unit. The stride of the convolutional layer of the decoder is 1, and the input feature and output feature sizes of the decoder are the same. The stride of the convolutional layer of the upsampling unit is 2, and the output feature size of the upsampling unit is twice the input feature size. Herein, the input feature may be an image feature or a text feature.
[0126] Through the decoding process of the target text data by the decoding module of the above-mentioned pre-trained text processing model, the obtained target text vector is also multiple distributions in different dimensions. This method can effectively avoid the loss of target text data compared with the text processing methods in the traditional technology, thereby improving the text quality of the reconstructed text.
[0127] Please refer to Figure 5 , in some embodiments, to implement text classification, step S105 may further include but is not limited to steps S501 to S502:
[0128] Step S501: Perform label classification processing on the target text vector according to a preset classification function and text category label to obtain a labeled text vector;
[0129] Step S502: Perform semantic analysis processing on the labeled text vector to obtain the target classification text.
[0130] It should be noted that the text classification model may be a textCNN model. The text classification model includes an Embedding layer, a convolutional layer, a pooling layer, and an output layer. Generally, through the Embedding layer of the text classification model, algorithms such as ELMO, GLOVE, Word2Vector, and Bert can be used to generate a dense vector from the input text. Then, through the convolutional layer and pooling layer of the text classification model, convolutional processing and pooling processing are performed on the dense vector to obtain a target feature vector. Then, the feature vector is input into the output layer, and through the preset function in the output layer, classification operations can be performed on the target feature vector to achieve text classification.
[0131] Specifically, in step S501, the preset classification function may be a softmax function. Through the softmax function, label classification processing is performed on the target text vector according to the preset text category label, creating a probability distribution for each text category, and marking the target text vector according to the probability distribution of each text category, so that each target text vector carries the corresponding text category label with entropy, thereby obtaining the labeled text vector.
[0132] Furthermore, step S502 is executed. By comparing the label text vector with the reference text vector, a comparison result is obtained. According to the comparison result, as well as the number of characters and part-of-speech categories of the text segments, etc., the text segments in the preset text library are filtered to obtain standard text segments. Finally, the standard text segments are concatenated to obtain the target classification text. This method can filter the label text vector, removing text segments with low relevance or inappropriate part-of-speech, improving the rationality of the target classification text.
[0133] Please refer to Figure 6 , in some embodiments, step S502 may further include but is not limited to steps S601 to S603:
[0134] Step S601, calculate the similarity between the label text vector and the reference text vector;
[0135] Step S602, according to the similarity, filter the text segments in the preset text library to obtain standard text segments;
[0136] Step S603, concatenate the standard text segments to obtain the target classification text.
[0137] Specifically, in step S601, a collaborative filtering algorithm such as the cosine similarity algorithm can be used to calculate the similarity between each label text vector and the reference text vector. For example, assuming the label text vector is u and the reference text vector is v, then according to the cosine similarity algorithm (as shown in formula (4)), calculate the similarity between the label text vector and the reference text vector, where, is the transpose of u.
[0138]
[0139] Furthermore, step S602 can be executed. Then, according to the relationship between the similarity and the preset similarity threshold, the required text fields are filtered from the preset text library. For example, the text segments with similarity greater than or equal to the similarity threshold are filtered from the preset text library, and these text segments are used as the standard text segments.
[0140] Finally, step S603 is executed. These standard text segments are converted into SQL statements, and the SQL statements are concatenated and fused through the database platform to obtain the target classification text that meets the requirements.
[0141] Through the above steps S101 to S105, the target text can be classified according to the text category, improving the relevance of the target classification text in each text category, thereby improving the accuracy of text classification.
[0142] Please refer to Figure 7, in some other embodiments, step S103 may include but is not limited to steps S701 to S703:
[0143] Step S701, map the target text data to a preset vector space to obtain target text features;
[0144] Step S702, perform encoding processing on the target text features according to a preset encoding order and encoding dimension to obtain a text hidden vector with a preset feature dimension;
[0145] Step S703, perform weighted processing on the text hidden vector according to a preset weight ratio to obtain a text implicit feature vector.
[0146] Specifically, in step S701, an MLP network can be used to perform multiple mapping processes on the target text data from the semantic space to the vector space, map the target text data into a preset vector space, and obtain target text features. The target text features can be text features or image features.
[0147] Furthermore, execute step S702. Through the encoding module of the above text processing model, the target text features can be encoded and downsampled according to the bottom-up encoding order and encoding dimension. For example, perform an initial encoding on the target text features to obtain the bottommost text hidden vector, and then perform downsampling processing layer by layer upward to obtain the text hidden vector corresponding to each layer. By identifying the text hidden vector corresponding to each layer according to the preset feature dimension, it is relatively convenient to obtain the text hidden vector of each feature dimension. It should be noted that the preset feature dimensions can include emotional dimension, text semantic dimension, text theme dimension, etc. Identify the text hidden vector corresponding to each preset feature dimension according to the keywords corresponding to each preset feature dimension.
[0148] Finally, execute step S703. According to different clustering requirements, set different weight ratios for each feature dimension, and perform weighted processing, masking processing, etc. on the text hidden vector of each feature dimension through this weight ratio to change the proportion of each feature dimension on each layer, thereby changing the angle of text clustering. For example, assume that the feature dimension is 3, that is, the dimension of the hidden variable layer is 3. The meaning represented by each feature dimension can be obtained by uniformly sampling each hidden variable layer. If the first feature dimension is the emotional dimension, the second feature dimension is the text semantic dimension, and the third feature dimension is the text theme dimension. If the current clustering task is for the emotional aspect, the weight ratio of the first feature dimension can be increased, and the weight ratio can be set to 8:1:1. So that the obtained text implicit feature vector includes more text features in the emotional dimension.
[0149] Further, to improve the encoding quality, the encoding module includes an encoder and a downsampling unit. The stride of the convolutional layer of the encoder is 1, and the input feature and output feature sizes of the convolutional layer of the encoder are the same; the stride of the convolutional layer of the downsampling unit is 2, and the output feature size of the downsampling unit is half of the input feature size, where the input feature can be an image feature or a text feature.
[0150] Through the encoding process of the target text data by the encoding module of the above-mentioned pre-trained text processing model, the obtained text latent feature vector z is no longer a distribution, but multiple distributions [z1, z2, …, zk] in different dimensions. Through this text processing model, different meanings represented by each feature dimension of different latent variable layers can be observed according to the distribution of latent variables, so that different weight ratios can be set according to the actual clustering task to change the clustering angle and improve the accuracy of text clustering.
[0151] Further, step S104 is executed, where step S104 may include but is not limited to the above steps S401 to S402, which will not be elaborated here.
[0152] Finally, step S105 is executed to cluster the target text, group the target texts with high relevance into one category, and obtain the target text set.
[0153] It should be noted that the text clustering model of the present application may include a partitioning-based clustering algorithm. By giving a data set with N tuples or records, K groups are constructed, and each group represents a cluster, where K < N. And these K groups satisfy the following conditions: (1) Each group contains at least one data record; (2) Each data record belongs to and only belongs to one group; for a given K, the partitioning-based clustering algorithm first gives an initial grouping method, and then changes the grouping through an iterative method, so that the grouping scheme after each improvement is better than the previous one, that is, the records in the same group are closer, and the records in different groups are farther apart. Please refer to Figure 8 , in some other embodiments, to implement text clustering, step S105 may further include but is not limited to steps S801 to S802:
[0154] Step S801, cluster the target text vector according to a preset clustering algorithm and text clustering labels to obtain target clustered texts containing text clustering labels;
[0155] Step S802, incorporate the target clustered texts with the same text clustering labels into the same set to obtain the target clustered text set.
[0156] Specifically, step S801 is executed. The preset clustering algorithm may include the kmeans algorithm, the TF-IDF weighting algorithm, and so on. For example, the difference between each target text vector and the reference vector corresponding to each text clustering label is calculated by the TF-IDF weighting algorithm. This difference can be characterized by similarity or other. The TF-IDF weighting algorithm is used to evaluate the importance of each text to be processed for the preset text set. According to the difference between each target text vector and the reference vector, the text clustering set to which the target text vector belongs is determined, and thus, according to the text clustering set to which each target text vector belongs, the text to be processed is labeled to obtain the target clustered text containing the text clustering label.
[0157] Furthermore, step S802 is executed. The text clustering labels of the target text are identified, and the target clustered texts containing the same text clustering label are incorporated into the same set. According to different text clustering labels, multiple different text clustering sets can be obtained, thus achieving the purpose of text clustering.
[0158] Through the above steps S101 to S105, the target text can be clustered, and the target texts with high relevance can be grouped into one category according to the preset text clustering label, thereby improving the accuracy of text clustering.
[0159] In the embodiment of the present application, the original text to be processed is obtained; feature extraction is performed on the original text to obtain target text data, which can effectively eliminate the data with low relevance in the original text and reduce the total amount of data. Furthermore, the target text data is encoded to obtain a text implicit feature vector; the text implicit feature vector is decoded to obtain a target text vector; finally, either label classification processing can be performed on the target text vector according to the preset text category label to obtain the target classification text containing the text category label, which can classify the target text according to the text category and improve the relevance of the target classification text in each text category, thereby improving the accuracy of text classification; or clustering processing can be performed on the target text vector according to the preset text clustering label to obtain a target clustered text set. By clustering the target text, the target texts with high relevance can be grouped into one category according to the preset text clustering label to obtain the target clustered text set, thereby improving the accuracy of text clustering.
[0160] Please refer to Figure 9 , the embodiment of the present application also provides a text processing device that can implement the above text processing method. The device includes:
[0161] An original text acquisition module 901, configured to acquire the original text to be processed;
[0162] A feature extraction module 902 is configured to extract features from the original text to obtain target text data;
[0163] An encoding processing module 903 is configured to perform encoding processing on the target text data to obtain a text latent feature vector;
[0164] A decoding processing module 904 is configured to perform decoding processing on the text latent feature vector to obtain a target text vector;
[0165] A text processing module 905 is configured to perform label classification processing on the target text vector through a preset text classification model and text category labels to obtain a target classification text including text category labels; or to perform clustering processing on the target text vector through a preset text clustering model and text clustering labels to obtain a target clustering text set.
[0166] The specific implementation manner of this text processing device is basically the same as the specific embodiments of the above text processing method, and will not be elaborated here.
[0167] An embodiment of this application further provides an electronic device, which includes: a memory, a processor, a computer program stored on the memory and executable on the processor, and a data bus for implementing connection communication between the processor and the memory. When the computer program is executed by the processor, the above text processing method is implemented. This electronic device can be any intelligent terminal including a tablet computer, an in-vehicle computer, etc.
[0168] Please refer to Figure 10 , Figure 10 which shows the hardware structure of an electronic device in another embodiment. The electronic device includes:
[0169] A processor 1001 can be implemented in a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is configured to execute relevant programs to implement the technical solutions provided by the embodiments of this application;
[0170] A memory 1002 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 1002 can store an operating system and other application programs. When implementing the technical solutions provided by the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1002 and are called by the processor 1001 to execute the text processing method of the embodiments of this application;
[0171] An input / output interface 1003 for implementing information input and output;
[0172] A communication interface 1004 for implementing communication interaction between this device and other devices, which can achieve communication through wired means (such as USB, network cable, etc.) or through wireless means (such as mobile network, WIFI, Bluetooth, etc.);
[0173] A bus 1005 for transmitting information between various components of the device (such as a processor 1001, a memory 1002, an input / output interface 1003, and a communication interface 1004);
[0174] Among them, the processor 1001, the memory 1002, the input / output interface 1003, and the communication interface 1004 achieve communication connections with each other inside the device through the bus 1005.
[0175] The embodiment of the present application also provides a computer-readable storage medium for computer-readable storage. The computer-readable storage medium stores one or more computer programs, and the one or more computer programs can be executed by one or more processors to implement the above text processing method.
[0176] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory can include high-speed random access memory, and can also include non-transitory memory, such as at least one disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory optionally includes a memory remotely set relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above networks include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0177] The embodiments described in the embodiments of the present application are for more clearly explaining the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.
[0178] Those skilled in the art can understand that Figure 1-8 the technical solutions shown do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown, or combine certain steps, or different steps.
[0179] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0180] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and appropriate combinations thereof.
[0181] As used in the specification of this application and the above-mentioned drawings, the terms "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily limit to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0182] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that three relationships can exist. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or a similar expression means any combination of these items, including any combination of single item (one) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0183] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the above units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections between each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical or other forms.
[0184] The units described above as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0185] In addition, each functional unit in various embodiments of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0186] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store programs.
[0187] The preferred embodiments of the embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the rights of the embodiments of the present application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall be within the scope of the rights of the embodiments of the present application.
Claims
1. A text processing method, characterized in that, The method includes: Obtain the original text to be processed; Extract features from the original text to obtain target text data; Perform encoding processing on the target text data to obtain a text implicit feature vector; Perform decoding processing on the text implicit feature vector to obtain a target text vector; The method further includes: Perform label classification processing on the target text vector through a preset text classification model and text category labels to obtain a target classification text including the text category labels; Or, Perform clustering processing on the target text vector through a preset text clustering model and text clustering labels to obtain a target clustering text set; The step of performing encoding processing on the target text data to obtain a text implicit feature vector includes: Map the target text data to a preset vector space based on the encoding module of a pre-trained text processing model to obtain target text features; According to the bottom-up encoding order and encoding dimension, perform primary encoding on the target text features to obtain the bottommost text hidden vector, and perform downsampling processing layer by layer upward. According to the keywords corresponding to each preset feature dimension, identify the text hidden vector corresponding to each feature dimension; Perform weighted processing on the text hidden vectors of multiple feature dimensions according to a preset weight ratio to obtain a text implicit feature vector, where the weight ratio is determined according to clustering requirements; The step of performing decoding processing on the text implicit feature vector to obtain a target text vector includes: Collect the grayscale values of the text implicit feature vector at fixed intervals, and analyze the collected grayscale values; when the collected grayscale value is not within the numerical set of the original function at the sampling points, then use the nearest neighbor interpolation method, bilinear interpolation method or cubic convolution interpolation method to perform interpolation processing on the sampled points to obtain an intermediate text vector; Perform decoding processing on the intermediate text vector based on the decoding module of the text processing model to obtain a target text vector.
2. The text processing method according to claim 1, wherein The step of extracting features from the original text to obtain target text data includes: Identify the text entity features in the original text; Perform feature classification processing on the text entity features using a pre-trained sequence classifier to obtain first text features; Extract features from the first text features to obtain target text data.
3. The text processing method according to claim 1, characterized in that The step of performing encoding processing on the target text data to obtain a text implicit feature vector includes: Map the target text data to a preset vector space to obtain target text features; Perform encoding processing on the target text features according to a preset encoding order and encoding dimension to obtain a text implicit feature vector.
4. The text processing method according to claim 1, characterized in that The step of performing label classification processing on the target text vector through a preset text classification model and text category labels to obtain a target classification text including the text category labels includes: Perform label classification processing on the target text vector according to a preset classification function and text category labels to obtain a label text vector; Perform semantic analysis processing on the label text vector to obtain a target classification text.
5. The text processing method according to claim 4, characterized in that, The step of performing semantic analysis processing on the label text vector to obtain a target classification text includes: Calculate the similarity between the label text vector and the reference text vector; According to the similarity, screen the text segments in the preset text word library to obtain standard text segments; Perform splicing processing on the standard text segments to obtain the target classification text.
6. The text processing method according to claim 1, wherein The step of clustering the target text vector through a preset text clustering model and text clustering labels to obtain a target clustering text set includes: Cluster the target text vector according to a preset clustering algorithm and text clustering labels to obtain a target clustering text including text clustering labels; Incorporate the target clustering texts with the same text clustering labels into the same set to obtain a target clustering text set.
7. A text processing device, characterized in that, The device includes: An original text acquisition module for acquiring the original text to be processed; A feature extraction module for extracting features from the original text to obtain target text data; An encoding processing module for encoding the target text data to obtain a text implicit feature vector; A decoding processing module for decoding the text implicit feature vector to obtain a target text vector; A text processing module for performing label classification processing on the target text vector through a preset text classification model and text category labels to obtain a target classification text including the text category labels; or for clustering the target text vector through a preset text clustering model and text clustering labels to obtain a target clustering text set; The encoding the target text data to obtain a text implicit feature vector includes: The encoding module of the pre-trained text processing model maps the target text data to a preset vector space to obtain target text features; According to the bottom-up encoding order and encoding dimension, perform primary encoding on the target text features to obtain the bottommost text hidden vector, and perform downsampling processing layer by layer upward. According to the keywords corresponding to each preset feature dimension, identify the text hidden vector corresponding to each feature dimension; Perform weighted processing on the text hidden vectors of multiple feature dimensions according to a preset weight ratio to obtain a text implicit feature vector, and the weight ratio is determined according to clustering requirements; The decoding the text implicit feature vector to obtain a target text vector includes: Collect the grayscale values of the text implicit feature vector at fixed intervals and analyze the collected grayscale values; when the collected grayscale value is not within the numerical set of the original function at the sampling point, then use the nearest neighbor interpolation method, bilinear interpolation method or cubic convolution interpolation method to interpolate the sampled points to obtain an intermediate text vector; The decoding module of the text processing model decodes the intermediate text vector to obtain a target text vector.
8. An electronic device, characterized in that, The electronic device includes a memory, a processor, a computer program stored on the memory and executable on the processor, and a data bus for realizing the connection and communication between the processor and the memory. When the computer program is executed by the processor, it realizes the steps of the text processing method according to any one of claims 1 to 6.
9. A computer-readable storage medium for computer-readable storage, characterized in that, The computer-readable storage medium stores one or more computer programs, and the one or more computer programs can be executed by one or more processors to implement the steps of the text processing method described in any one of claims 1 to 6.
Citation Information
Patent Citations
Customer service robot session text classification method, device and apparatus, and storage medium
CN109543030A
Text classification method and device and computer equipment
CN110362684A