A text classification method, a review sentiment analysis method and a device
The text classification method uses pre-trained models with labeled source data and unlabeled target data to adapt to new languages and domains, reducing data collection efforts and enhancing classification accuracy.
Patent Information
- Application Number
- CN202010561456.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-06-18
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2040-06-18
AI Technical Summary
In the cross-language and cross-field text classification, the prior art requires obtaining a large amount of labeled text data in different languages and fields, which makes it expensive and difficult to effectively classify text.
The text classification model is trained using the source language of the marked category, the text data of the source field, the text data of the unlabeled category, and the text data of the target field. The multi-language embedding module is used to generate cross-language representation vectors, and the unsupervised feature decomposition module is used to extract the domain invariant features and domain-specific features to realize cross-language and cross-domain text classification.
It reduces text data acquisition and labeling work, improves the model's adaptability in the target language and target fields, realizes cross-language and cross-domain text classification, and reduces costs.
Smart Images

Figure CN113821629B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing, and in particular, to a text classification method, a review sentiment analysis method and an apparatus. Background Art
[0002] The text classification technology is an important branch in the natural language processing technology, and can be applied to multiple scenarios such as sentiment analysis, spam filtering, news classification, named entity recognition, etc.
[0003] With the development of economic globalization, the languages and fields involved in the text to be classified are increasing. For example, in the application scenario of sentiment analysis of user reviews, users in different countries and regions may use different languages to evaluate products or services in different fields. Correspondingly, the text to be classified is the user review data in these different languages and different fields, such as Chinese evaluations of electronic products, English evaluations of movies, French evaluations of books, Japanese evaluations of music, etc.
[0004] Currently, for the text classification problem of different languages / domains, the labeled text data in that language / domain is usually used as training samples to train a machine learning model, and then the trained model is used to classify the text to be classified in the same language / domain to determine its category. For different languages / domains, different machine learning models need to be trained. However, it is difficult to obtain a large amount of text data from different languages / domains and label enough training samples, which requires a high cost. Therefore, a cross - language and cross - domain text classification method is needed to reduce the acquisition and annotation work of text data. Summary of the Invention
[0005] For this reason, the present invention provides a text classification method and an apparatus to try to solve or at least alleviate the problems existing above.
[0006] According to a first aspect of the present invention, a text classification method is provided, including: obtaining a text to be classified, where the language of the text to be classified is any one of a source language and a set of target languages, and the field to which the text to be classified belongs is any one of a source field and a set of target fields; inputting the text to be classified into a preset text classification model, so that the text classification model outputs the category to which the text to be classified belongs, where the text classification model is trained with the text data in the source language and source field with labeled categories and the text data in the source language and target field without labeled categories as training samples.
[0007] According to a second aspect of the present invention, there is provided a method for comment sentiment analysis, including: obtaining a comment text to be analyzed, where the language of the comment text is any one of the source language and the target language set, and the field to which the comment text belongs is any one of the source field and the target field set; inputting the comment text into a preset sentiment analysis model so that the sentiment analysis model outputs the sentiment polarity of the comment text, where the sentiment analysis model is trained with comment data in the source language and source field with marked sentiment polarity and comment data in the source language and target field without marked sentiment polarity as training samples.
[0008] According to a third aspect of the present invention, there is provided a text classification device, including: a text acquisition module adapted to obtain a text to be classified, where the language of the text to be classified is any one of the source language and the target language set, and the field to which the text to be classified belongs is any one of the source field and the target field set; a category determination module adapted to input the text to be classified into a preset text classification model so that the text classification model outputs the category to which the text to be classified belongs, where the text classification model is trained with text data in the source language and source field with marked categories and text data in the source language and target field without marked categories as training samples.
[0009] According to a fourth aspect of the present invention, there is provided a comment sentiment analysis device, including: a comment acquisition module adapted to obtain a comment text to be analyzed, where the language of the comment text is any one of the source language and the target language set, and the field to which the comment text belongs is any one of the source field and the target field set; a sentiment polarity determination module adapted to input the comment text into a preset sentiment analysis model so that the sentiment analysis model outputs the sentiment polarity of the comment text, where the sentiment analysis model is trained with comment data in the source language and source field with marked sentiment polarity and comment data in the source language and target field without marked sentiment polarity as training samples.
[0010] According to a fifth aspect of the present invention, there is provided a computing device, including: at least one processor and a memory storing program instructions; when the program instructions are read and executed by the processor, the computing device is caused to execute the above text classification method and / or the above comment sentiment analysis method.
[0011] According to a sixth aspect of the present invention, there is provided a readable storage medium storing program instructions, when the program instructions are read and executed by a computing device, the computing device is caused to execute the above text classification method and / or the above comment sentiment analysis method.
[0012] According to the text classification method of the present invention, a text classification model is pre-trained, and then the trained text classification model is applied to classify the text to be classified to determine the category to which the text to be classified belongs.
[0013] The text classification model of the present invention is trained using text data of source language and source domain with labeled categories and text data of source language and target domain with unlabeled categories as training samples. The trained text classification model can be migrated and applied to target language and target domain, which greatly reduces the work of acquiring and labeling text data.
[0014] The text classification model of the present invention includes a multilingual embedding module, an unsupervised feature decomposition module and a classification module. The multilingual embedding module is suitable for generating cross-language representation vectors of text data, and the cross-language representation vectors are features shared across languages. Therefore, the text classification model trained with text data of the source language can also be applied to the target language.
[0015] The unsupervised feature decomposition module can complete the training using only a small amount of unlabeled source language and target domain text data, and decompose the cross-language representation vector output by the multilingual embedding module into domain-invariant features and domain-specific features. The classification module uses labeled source language and source domain text data for training, and is suitable for determining the category to which the text to be classified belongs based on the domain-invariant features and domain-specific features output by the unsupervised feature decomposition module. Since domain-invariant features are cross-domain invariant features, the classification module trained using text data from the source domain can also be applied to the target domain.
[0016] The text classification method of the present invention can be applied to the comment sentiment analysis scenario, that is, the present invention also provides a comment sentiment analysis method. The method pre-trains a sentiment analysis model, and the sentiment analysis model has the same structure as the aforementioned classification model, except that the data in the comment sentiment analysis scenario (i.e., comment data) is used for training. Specifically, the source language and source domain comment data with marked sentiment polarity and the source language and target domain comment data with unmarked sentiment polarity are used for training. The trained sentiment analysis model can also classify the comment text in the target language and target domain, and determine the sentiment polarity (positive or negative) of the comment text.
[0017] The above description is only an overview of the technical solution of the present invention. In order to more clearly understand the technical means of the present invention, it can be implemented according to the contents of the specification. In order to make the above and other purposes, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are listed below. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] To achieve the above and related purposes, certain illustrative aspects are described herein in connection with the following description and drawings, which indicate various ways in which the principles disclosed herein can be practiced, and all aspects and their equivalent aspects are intended to fall within the scope of the claimed subject matter. The above and other purposes, features, and advantages of the present disclosure will become more apparent by reading the following detailed description in conjunction with the drawings. Throughout the present disclosure, the same reference numerals generally refer to the same components or elements.
[0019] Figure 1 A schematic diagram of a computer system 9100 according to an embodiment of the present invention is shown;
[0020] Figure 2 A schematic diagram of a deep neural network of a machine learning model 9120 according to an embodiment of the present invention is shown;
[0021] Figure 3 A structural diagram of a text classification model 300 according to an embodiment of the present invention is shown;
[0022] Figure 4 A flowchart of a text classification method 400 according to an embodiment of the present invention is shown;
[0023] Figure 5 A flowchart of a comment sentiment analysis method 500 according to an embodiment of the present invention is shown;
[0024] Figure 6 A comparison graph of the classification effects between the text classification model (sentiment analysis model) of the present invention and other text classification models is shown;
[0025] Figure 7 A result graph of an ablation study of the text classification model according to the present invention is shown;
[0026] Figure 8 A schematic diagram of a computing device 600 according to an embodiment of the present invention is shown;
[0027] Figure 9 A schematic diagram of a text classification device 700 according to an embodiment of the present invention is shown;
[0028] Figure 10 A schematic diagram of a comment sentiment analysis device 800 according to an embodiment of the present invention is shown. Detailed Description of the Invention
[0029] Exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully conveyed to those skilled in the art.
[0030] In view of the problems existing in the prior art, the present invention provides a cross - language and cross - domain text classification method. The method uses the text data of the source language and source domain that has been labeled and a small amount of unlabeled text data of the source language and target domain to train a text classification model. The trained model has strong adaptability and can also be used to classify the text data of the target language and target domain, thus avoiding a large amount of data acquisition and annotation work in application scenarios of different languages and different domains.
[0031] The text classification method of the present invention is implemented by training and applying a text classification model. Figure 1 The schematic diagram of a computer system 9100 for training and applying a text classification model is shown. As Figure 1 shown, the system 9100 includes a user computing device 9110, a server computing system 9130, and a training computing system 9150 that are communicatively connected via a network 9180.
[0032] The user computing device 9110 can be any type of computing device, including but not limited to, for example, a personal computing device (e.g., a laptop or desktop computer), a mobile computing device (a smart phone or a tablet), a game console or controller, a wearable computing device, an embedded computing device, an edge computing device, or any other type of computing device. The user computing device 9110 can be deployed as an edge intelligent device at the user's site and interact with the user to process user input.
[0033] The user computing device 9110 may store or include one or more machine learning models 9120. The machine learning models 9120 may be designed to perform various tasks, such as image classification, object detection, speech recognition, machine translation, content filtering, and so on. The machine learning models 9120 may be neural networks (e.g., deep neural networks) or other types of machine learning models including non-linear models and / or linear models. Examples of the machine learning models 9120 include but are not limited to various types of deep neural networks (DNNs), such as feedforward neural networks, recurrent neural networks (RNNs, e.g., long short-term memory recurrent neural networks (LSTMs), Transformer neural networks (with or without attention mechanisms)), convolutional neural networks (CNNs), or other forms of neural networks. The machine learning models 9120 may include one machine learning model or may be a combination of multiple machine learning models.
[0034] Figure 2 A neural network that is a machine learning model 9120 according to some embodiments is shown. The neural network has a hierarchical architecture, and each network layer has one or more processing nodes (referred to as neurons or filters) for processing. In a deep neural network, the output of the previous layer after processing is the input of the next layer, where the first layer in the architecture receives the network input for processing, and the output of the last layer is provided as the network output. As Figure 2 shown, the machine learning model 9120 includes network layers 9122, 9124, 9126, etc., where the network layer 9122 receives the network input and the network layer 9126 provides the network output.
[0035] In a deep neural network, the main processing operations within the network are interleaved linear and non-linear transformations. These processes are distributed among the respective processing nodes. Figure 2 An enlarged view of a node 9121 in the model 9120 is also shown. The node 9121 receives a plurality of input values a1, a2, a3, etc., and processes the input values based on corresponding processing parameters (such as weights w1, w2, w3, etc.) to generate an output z. The node 9121 may be designed to process the input using an activation function, which can be expressed as:
[0036] z = σ(w T α + b) (1)
[0037] Where α represents the input vector of node 9121 (including elements a1, a2, a3, etc.); w represents the weight vector in the processing parameters used by node 9121 (including elements w1, w2, w3, etc.), and each weight is used to weight the corresponding input; b represents the bias vector in the processing parameters used by node 9121 (including elements b1, b2, b3, etc.), and each bias is used to bias the corresponding input and the weighted result; σ() represents the activation function used by node 9121, and the activation function can be a linear function or a non-linear function. Commonly used activation functions in neural networks include the sigmoid function, ReLu function, tanh function, maxout function, etc. The output of node 9121 can also be referred to as the activation value. Depending on the network design, the output (i.e., the activation value) of each network layer can be provided as input to one, multiple, or all nodes of the next layer.
[0038] Each network layer in the machine learning model 9120 can include one or more nodes 9121. When viewing the processing in the machine learning model 9120 in terms of network layers, the processing of each network layer can also be similarly represented in the form of formula (1). At this time, α represents the input vector of the network layer, and w represents the weight of the network layer.
[0039] It should be understood that Figure 2 The architecture of the illustrated machine learning model and the number of network layers and processing nodes therein are all illustrative. In different applications, according to needs, the machine learning model can be designed to have other architectures.
[0040] Continuing to refer to Figure 1 , in some implementation manners, the user computing device 9110 can receive the machine learning model 9120 from the server computing system 130 through the network 9180, store it in the memory of the user computing device, and use or implement it by an application in the user computing device.
[0041] In other implementation manners, the user computing device 9110 can call the machine learning module 9140 stored and implemented in the server computing system 9130. For example, the machine learning model 9140 can be implemented by the server computing system 9130 as part of a web service, so that the user computing device 9110 can call the machine learning model 9140 implemented as a web service, for example, through the network 9180 and according to the client-server relationship. Therefore, the machine learning models that can be used at the user computing device 9110 include the machine learning model 9120 stored and implemented at the user computing device 9110 and / or the machine learning model 9140 stored and implemented at the server computing system 9130.
[0042] The user computing device 9110 may also include one or more user input components 9122 that receive user input. For example, the user input component 9122 may be a touch-sensitive component (such as a touch-sensitive display screen or a touchpad) that is sensitive to a user input object (such as a finger or a stylus). The touch-sensitive component can be used to implement a virtual keyboard. Other example user input components include microphones, traditional keyboards, cameras, or other devices through which a user can provide user input.
[0043] The server computing system 9130 may include one or more server computing devices. In cases where the server computing system 9130 includes multiple server computing devices, these server computing devices may operate according to a sequential computing architecture, a parallel computing architecture, or some combination thereof.
[0044] As described above, the server computing system 9130 may store or include one or more machine learning models 9140. Similar to the machine learning model 9120, the machine learning model 9140 may be designed to perform various tasks, such as image classification, object detection, speech recognition, machine translation, content filtering, and so on. The model 9140 may include various machine learning models. Example machine learning models include neural networks or other multi-layer non-linear models. Example neural networks include feedforward neural networks, deep neural networks, recurrent neural networks, and convolutional neural networks.
[0045] The user computing device 9110 and / or the server computing system 9130 may train the model 9120 and / or 9140 via interaction with a training computing system 9150 communicatively coupled through a network 9180. The training computing system 9150 may be separate from the server computing system 9130 or may be a part of the server computing system 9130.
[0046] Similar to the server computing system 9130, the training computing system 9150 may include one or more server computing devices or may otherwise be implemented by one or more server computing devices.
[0047] The training computing system 9150 may include a model trainer 9160 that trains the machine learning models 9120 and / or 9140 stored at the user computing device 9110 and / or the server computing system 9130 using various training or learning techniques such as, for example, backpropagation of error. In some implementations, performing backpropagation of error may include performing truncated backpropagation through time. The model trainer 9160 may perform various generalization techniques (such as weight decay, dropout, etc.) to improve the generalization ability of the model being trained.
[0048] Specifically, the model trainer 9160 can train the machine learning models 9120 and / or 9140 based on a collection of training data 9162. The training data 9162 can include multiple different sets of training data, and each set of training data, for example, can respectively contribute to training the machine learning models 9120 and / or 9140 to perform multiple different tasks. For example, the sets of training data include data sets that contribute to the machine learning models 9120 and / or 9140 performing object detection, object recognition, object segmentation, image classification, and / or other tasks.
[0049] In some implementations, if the user has explicitly consented, the training examples can be provided by the user computing device 9110. Thus, in such implementations, the model 9120 provided to the user computing device 9110 can be trained by the training computing system 9150 on user-specific data received from the user computing device 9110. In some cases, this process can be referred to as a personalized model.
[0050] Additionally, in some implementations, the model trainer 9160 can modify the machine learning model 9140 in the server computing system 9130 to obtain a machine learning model 9120 suitable for use in the user computing device 9110. These modifications include, for example, reducing the number of various parameters in the model, storing parameter values with lower precision, etc., so that the trained machine learning models 9120 and / or 9140 are suitable for running considering the different processing capabilities of the server computing system 9130 and the user computing device 9110.
[0051] The model trainer 9160 includes computer logic for providing the desired functionality. The model trainer 9160 can be implemented using hardware, firmware, and / or software that controls a general-purpose processor. For example, in some implementations, the model trainer 9160 includes program files stored on a storage device, loaded into memory, and executed by one or more processors. In other implementations, the model trainer 9160 includes a set of one or more computer-executable instructions stored in a tangible computer-readable storage medium such as RAM, a hard disk, or an optical or magnetic medium. In some implementations, the model trainer 9160 can be replicated and / or distributed across multiple different devices.
[0052] The network 9180 can be any type of communication network, such as a local area network (e.g., an intranet), a wide area network (e.g., the Internet), or some combination thereof, and can include any number of wired or wireless links. Generally, communication over the network 9180 can be carried via any type of wired and / or wireless connection, using various communication protocols (e.g., TCP / IP, HTTP, SMTP, FTP), encoding or formats (e.g., HTML, XML, and JSON), and / or security schemes (e.g., VPN, HTTPS, SSL).
[0053] In an embodiment of the present invention, the training computing system 9150 can train to generate the text classification model 9144 of the present invention. The model 9144 takes text data as input and outputs the category to which the text data belongs. The trained text classification model 9144 can be deployed in the server computing system 9130.
[0054] The model 9144 can be implemented as part of a web service by the server computing system 9130, so that the user computing device 9110 can call the text classification model 9144 implemented as a web service, for example, via the network 9180 and according to the client-server relationship. Specifically, the user can initiate a text classification request to the server computing system 9130 through the user computing device 9110 and specify the text to be classified. In response to the request, the server computing system 9130 calls the text classification model 9144, inputs the text to be classified into the model 9144, and the model 9144 outputs the category label of the text to be classified. Subsequently, the server computing system 9130 can return the category label of the text to be classified to the user computing device 9110.
[0055] In some other embodiments of the present invention, the text classification model 9144 of the present invention trained and generated by the training computing system 9150 can be modified, for example, quantizing the model parameter values (i.e., storing the parameter values with a smaller precision, such as 8-bit integer), and deploying the modified model in the user computing device 9110. As Figure 1 shown, the text classification model deployed in the user computing device 9110 is denoted as the text classification model 9124. Those skilled in the art can understand that the text classification model 9144 deployed in the server computing system 9130 and the text classification model 9124 deployed in the user computing device 9110 have the same function. The user can call the text classification model 9124 in the user computing device 9110, input the text to be classified into the model, and the model will output the category to which the text to be classified belongs.
[0056] The text classification model of the present invention can be applied to various scenarios, such as comment sentiment analysis, spam filtering, fraud SMS identification, news classification, intelligent customer service, etc. Those skilled in the art can understand that the text classification models adopted in different application scenarios have the same structure, but only the data used for training the models are different. For example, when applying the text classification model to the comment sentiment analysis scenario, the text classification model is trained using user comment data with labeled sentiment polarities (positive or negative); when applying the text classification model to the spam filtering scenario, the text classification model is trained using email data with labeled discriminant labels (whether it is spam); when applying the text classification model to the fraud SMS identification scenario, the text classification model is trained using SMS data with labeled discriminant labels (whether it is a fraud SMS); when applying the text classification model to the news classification scenario, the text classification model is trained using news data with labeled news categories (society, finance, entertainment, etc.); when applying the text classification model to the intelligent customer service scenario, the text classification model is trained using user question data with labeled question types (shipping issues, invoice issues, after-sales issues, etc.); and so on.
[0057] The text classification model of the present invention is trained using text data in the source language and source domain with labeled categories and text data in the source language and target domain without labeled categories. The trained text classification model has good transferability and can be used to classify text data in different target languages and target domains. For example, we can use an English-language electronic product review dataset with labeled sentiment polarities (positive or negative) and an English-language book and music review dataset without labeled sentiment polarities to train the text classification model. The trained model can be used to classify the sentiment polarities of review data such as German book reviews, French book reviews, and Japanese music reviews.
[0058] Figure 3 FIG. shows the structural diagram of a text classification model 300 according to an embodiment of the present invention. As Figure 3 shown, the text classification model 300 includes a multilingual embedding module 310, an unsupervised feature decomposition module 320, and a classification module 330.
[0059] The multilingual embedding module 310 is adapted to process the text to be classified to generate a cross-language representation vector of the text to be classified.
[0060] The multilingual embedding module 310 is trained on an open-domain dataset from multiple languages (e.g., more than 100 languages). The cross-language representation vector extracted by it is a feature that is independent of the language type and shared across languages in the text data to be classified.
[0061] It should be noted that the present invention does not limit the specific structure of the multilingual embedding module 310, which can be implemented, for example, as a neural network including multiple hidden layers. According to one embodiment, in order to improve the training efficiency of the text classification model 300, during the training process of the text classification model 300, the pre-trained multilingual embedding module 310 is directly used without re-training the multilingual embedding module 310. The pre-trained multilingual embedding module 310 can be, for example, the XLM model, but is not limited thereto (for the model structure, see the paper Guillaume Lample and Alexis Conneau. Cross-lingual language model pretraining. NeurIPS, 2019).
[0062] The unsupervised feature decomposition module 320 (Unsupervised Feature Decomposition, UFD) is adapted to extract domain-invariant features and domain-specific features from the cross-lingual representation vectors output by the multilingual embedding module 310.
[0063] According to one embodiment, the unsupervised feature decomposition module 320 further includes a domain-invariant feature extractor 322 and a domain-specific feature extractor 324, wherein the domain-invariant feature extractor 322 is adapted to extract domain-invariant features from the cross-lingual representation vectors, and the domain-specific feature extractor 324 is adapted to extract domain-specific features from the cross-lingual representation vectors.
[0064] Domain-invariant features are the feature parts shared by text data in different domains. The domain-invariant feature extractor 322 extracts domain-invariant features from the cross-lingual representation vectors in an unsupervised manner.
[0065] As Figure 3 shown, the domain-invariant feature extractor 322 includes at least two first processing units 326 ( Figure 3 two first processing units 326 are exemplarily shown in), and each first processing unit 326 includes at least one feed-forward processing layer ( Figure 3 one feed-forward processing layer is exemplarily shown in each first processing unit 326) and a residual connection layer, wherein the residual connection layer is adapted to add the input of the first feed-forward processing layer in the corresponding first processing unit to the output of the last feed-forward processing layer, so as to better retain the domain-invariant features in the cross-lingual representation vectors.
[0066] According to one embodiment, each feed-forward processing layer in the domain-invariant feature extractor 322 is activated by using the ReLU activation function.
[0067] Since the multilingual embedding module 310 is pre-trained on open-domain datasets from over 100 languages, we believe that the cross-lingual representation vectors generated by this multilingual embedding module 310 should contain certain cross-domain feature information, and among the domain-invariant features extracted by the domain-invariant feature extractor 322, this part of the feature information should be retained to the greatest extent. Therefore, according to one embodiment, the mutual information between the input and output of the domain-invariant feature extractor 322 can be maximized as the training objective of the extractor 322, so that the useful information in the cross-lingual representation vectors can be transmitted to the domain-invariant features.
[0068] Mutual Information (MI) refers to the amount of information of one random variable about another random variable, and is used to measure the degree of mutual dependence between random variables. The greater the mutual information, the stronger the dependence between the two variables.
[0069] It should be noted that in practice, the calculation of mutual information is usually very difficult, especially in the case of continuous and high-dimensional variables. Therefore, in the embodiments of the present invention, instead of directly calculating the mutual information according to Equation (3), an estimation algorithm is used to estimate the value of the mutual information. For example, the neural network gradient descent algorithm proposed by Belghazi et al. in 2018 (see Mohamed Ishmael Belghazi, Aristide Baratin, Sai Rajeshwar, Sherjil Ozair, Yoshua Bengio, Devon Hjelm, and Aaron Courville. Mutual information neural estimation. In ICML, pages 530–539, 2018) can be used to estimate the mutual information, and the estimation formula is as follows:
[0070]
[0071] where I(X; Y) represents the mutual information between variable X and variable Y, and the symbol := means defined as, is the joint probability distribution function of X and Y, is the product of the marginal probability distribution function of X and the marginal probability distribution function of Y, represents and the KL divergence of, is the DV lower bound of the mutual information, T ω is a discriminator, which is implemented as a neural network with parameters ω to be trained, and it can map the sample (x, y) in the X*Y space into a numerical value T ω (x, y). Denote \(T\) ω (x, y) under the expectation of Denote \(T\) ω (x, y) under the expectation of, where \(e\) is the natural constant. By maximizing the discriminator \(T\) ω will be trained to map samples from to a larger value and map samples from to a smaller value, so as to distinguish samples from and samples.
[0072] In an embodiment of the present invention, the training objective of the domain-invariant feature extractor 322 is to maximize the mutual information between its input and output. Since the output \(F\) s (X) of the domain-invariant feature extractor 322 depends on its input X (\(F\) s is the domain-invariant feature extractor 322), therefore, when calculating the mutual information between the input and output of the domain-invariant feature extractor 322 in the present invention, the above neural estimation method can be further simplified to a Jensen-Shannon mutual information estimator (see Guillaume Lample and Alexis Conneau. Cross-lingual language model pretraining. NeurIPS, 2019), so as to accelerate the calculation speed of the mutual information. The calculation formula is as follows:
[0073]
[0074] where represents the mutual information between X and \(F\) s (X), \(x\) is the input sample with the empirical probability distribution , since \(F\) s (x) is calculated from \(x\), so \((x, F\) s (x)) can be regarded as a sample in the joint distribution of X and \(F\) s (X). \(x'\) is a random sample with the empirical probability distribution , where Therefore \((x', F\) s (x)) is a sample in the product of the marginal probability distributions. \(sp(z)=\log(1 + e\) z ) is the softplus activation function.
[0075] Based on the above training objective of maximizing the mutual information between the input and output, the loss function \(L\) of the domain-invariant feature extractor 322 sIt can be expressed in the following form:
[0076]
[0077] where ψ s is the set of parameters to be trained in the domain-invariant feature extractor 322 (F s ), ω s is the set of parameters to be trained in the discriminator T ω in the mutual information estimator of the domain-invariant feature extractor 322, and the training objective is to minimize the value of the above loss function L s .
[0078] Domain-specific features are the characteristic parts unique to the text data in different domains. The domain-specific feature extractor 324 extracts domain-specific features from the cross-lingual representation vectors in an unsupervised manner.
[0079] As Figure 3 shown, the domain-specific feature extractor 324 includes at least two second processing units 328 ( Figure 3 two second processing units 328 are exemplarily shown in Figure 3 ), and each second processing unit 328 includes at least one feed-forward processing layer ( Figure 3 one feed-forward processing layer is exemplarily shown in each second processing unit 328). According to one embodiment, each feed-forward processing layer in the domain-specific feature extractor 324 is activated using the ReLU activation function.
[0080] According to one embodiment, the number of first processing units 326 included in the domain-invariant feature extractor 322 is the same as the number of second processing units 328 included in the domain-specific feature extractor 324, and the number of feed-forward processing layers included in the first processing unit 326 is also the same as the number of feed-forward processing layers included in the second processing unit 328. That is to say, the structural difference between the domain-invariant feature extractor 322 and the domain-specific feature extractor 324 is only that the domain-invariant feature extractor 322 is provided with a residual connection layer in each first processing unit 326.
[0081] Considering that domain-specific features are the features unique to the text data in different domains, they should be mutually independent and exclusive with domain-invariant features. Therefore, according to one embodiment, the training objective of the domain-specific feature extractor 324 can be set to minimize the mutual information between the output of each second processing unit 328 and the output of the corresponding first processing unit 326. Based on this training objective, taking the structure where the extractors 322 and 324 shown in Figure 3 each include two processing units as an example, the loss function of the domain-specific feature extractor 324 includes the following two parts:
[0082]
[0083]
[0084] Among them, L m is the loss function corresponding to the first pair of processing units, and F s ’(X), F p ’(X) respectively represent the outputs of the first processing unit of the domain-invariant feature extractor 322 (i.e., F s ), and the first processing unit of the domain-specific feature extractor 324 (i.e., F p ). represents the mutual information between F s ’(X) and F p ’(X), and ω m is the set of parameters to be trained in the discriminator T ω in the mutual information estimator of the first pair of processing units. ψ s , ψ p are respectively the sets of parameters to be trained in the domain-invariant feature extractor 322 (F s ), and the domain-specific feature extractor 324 (F p ).
[0085] L p is the loss function corresponding to the second pair of processing units, and F s (X), F p (X) respectively represent the outputs of the second processing unit of the domain-invariant feature extractor 322 (F s ), and the second processing unit of the domain-specific feature extractor 324 (F p ), that is, the output of F s and the output of F p . represents the mutual information between F s (X) and F p (X), and ω p is the set of parameters to be trained in the discriminator T ω in the mutual information estimator of the second pair of processing units. The training objective is to minimize the values of the above loss functions L m , L p .
[0086] Combining the loss function L s of the domain-invariant feature extractor 322 shown in the above formula (4), and the loss functions L m , L p of the domain-specific feature extractor 324 shown in the above formulas (5) and (6), the loss function L UFD of the entire unsupervised feature decomposition module 320 can be expressed as:
[0087] L UFD= αL s + βL m + γL p (7)
[0088] where α, β, and γ are the weights of L s , L m , L p respectively, and their values can be set by those skilled in the art, and the present invention does not limit this. In one embodiment, for example, α, β, and γ can be set to 1, 0.2, and 1 respectively. The training objective is to minimize the above loss function L UFD , that is, to minimize the weighted sum of the negative of the mutual information between the input and output of the domain-invariant feature extractor and the mutual information between the outputs of each pair of the first processing unit and the second processing unit.
[0089] The classification module 330 is adapted to determine the category to which the text to be classified belongs according to the domain-invariant features and domain-specific features output by the unsupervised feature decomposition module 320.
[0090] According to one embodiment, the classification module 330 includes a linear layer and a Softmax feed-forward layer. In the classification module 330, first, the linear layer is used to map the concatenated vector of the domain-invariant features and domain-specific features into an intermediate vector. Then, a simple feed-forward layer with a Softmax activation function is used to map the intermediate vector into a class label.
[0091] According to one embodiment, the training objective of the classification module 330 is to minimize the cross-entropy loss L t .
[0092] As described above, the text classification model 300 of the present invention includes three modules: a multilingual embedding module 310, an unsupervised feature decomposition module 320, and a classification module 330. Based on these three modules, the text classification model 300 is trained in the following phased manner:
[0093] First, obtain the pre-trained multilingual embedding module 310. The multilingual embedding module 310 is pre-trained and remains unchanged throughout the training process of the text classification model 300.
[0094] Subsequently, use the text data of the source language and target domain without labeled categories as training samples to train the unsupervised feature decomposition module 320. Denote the text data set of the source language and target domain without labeled categories as D s,t (s represents source, t represents target), and aim to minimize the loss function L UFD , and continuously optimize the values of the parameters to be trained in the unsupervised feature decomposition module 320, that is, adjust the parameters {ω s , ψ s, ω m , ψ p , ω p} value, denote the finally trained parameters as
[0095] After the unsupervised feature decomposition module 320 finishes training, keep the value of the finally trained parameters unchanged, use the text data of the source language and source domain with labeled categories as training samples to train the classification module 330. Denote the text data set of the source language and source domain with labeled categories as D s,s , aiming to minimize the loss function L t as the goal, continuously optimize the values of the parameters to be trained in the classification module 330, that is, adjust the node weights and biases in the linear layer and Softmax layer of the classification module 330.
[0096] When the classification module 330 finishes training, the entire text classification model 300 finishes training.
[0097] According to one embodiment, during the above training process, the number of training samples of the unsupervised feature decomposition module 320 is less than the number of training samples of the classification module 330.
[0098] The text classification model of the present invention is trained by using the text data of the source language and source domain with labeled categories and the text data of the source language and target domain with unlabeled categories as training samples. The trained text classification model can be migrated and applied to the target language and target domain, that is, cross - language and cross - domain model migration is realized, greatly reducing the acquisition and annotation work of text data.
[0099] In the embodiment of the present invention, the trained text classification model 300 can be used for text classification. Figure 4 The flowchart of a text classification method 400 according to an embodiment of the present invention is shown. Method 400 uses the trained text classification model to classify the text to be classified. Method 400 can be executed, for example, in a user computing device (such as the aforementioned user computing device 9110) or a server computing system (such as the aforementioned server computing system 9130). As Figure 4 shown, method 400 starts from step S410.
[0100] In step S410, obtain the text to be classified, where the language of the text to be classified is any one in the source language and target language set, and the field to which the text to be classified belongs is any one in the source field and target field set.
[0101] It should be noted that the source language, target language, source field, and target field in step S410 are relative to the training process of the text classification model.
[0102] Specifically, the source language refers to the language to which the training samples of the text classification model belong; the target language is the language to which the text classification model can be applied for migration, that is, the languages other than the above-mentioned source language in the languages of the training samples of the multi-language embedding module 310. The multi-language embedding module 310 is usually trained with text data in multiple (sometimes up to more than 100) languages. Correspondingly, there are multiple target languages, and the target language set is relatively large. In some cases, it can be approximately considered that the text classification model can be applied to text classification problems in any language, that is, the target language can be any language other than the source language. Correspondingly, the text to be classified in step S410 can be in any language.
[0103] The source domain refers to the domain to which the labeled training samples of the text classification model belong; the target domain refers to the domain to which the unlabeled training samples of the text classification model belong, and there may be multiple target domains.
[0104] Subsequently, in step S420, the text to be classified is input into a preset text classification model so that the text classification model outputs the category to which the text to be classified belongs. Among them, the text classification model is trained with the text data in the source language and source domain with labeled categories and the text data in the source language and target domain with unlabeled categories as training samples.
[0105] In step S420, the text classification model takes the text to be classified as input, performs a series of processing and calculations on the text to be classified, and outputs the category label of the text to be classified, thereby determining the category to which the text to be classified belongs. For the specific structure and training method of the text classification model, reference can be made to Figure 3 and the corresponding written description, which will not be elaborated here.
[0106] The text classification model of the present invention can be applied to various natural language processing scenarios, such as comment sentiment analysis, spam filtering, fraud SMS identification, news classification, intelligent customer service, etc. Those skilled in the art can understand that the text classification models adopted in different application scenarios have the same structure, but only the data used for training the model is different.
[0107] According to an embodiment, the text classification model of the present invention can be applied to the scenario of comment sentiment analysis, that is, the present invention also provides a sentiment analysis model and a comment sentiment analysis method.
[0108] The sentiment analysis model has the same structure as the aforementioned text classification model, that is, the sentiment analysis model includes: a multilingual embedding module, which is suitable for processing the review text to generate a cross-language representation vector of the review text; an unsupervised feature decomposition module, which is suitable for extracting domain-invariant features and domain-specific features from the cross-language representation vector; and a classification module, which is suitable for determining the sentiment polarity of the review text according to the domain-invariant features and domain-specific features.
[0109] The sentiment analysis model is trained according to the following steps: obtaining a pre-trained multilingual embedding module; using the review data of the source language and the target domain without labeled sentiment polarity as training samples to train the unsupervised feature decomposition module; after the unsupervised feature decomposition module is trained, using the review data of the source language and the source domain with labeled sentiment polarity as training samples to train the classification module.
[0110] The sentiment analysis model is trained using the data in the review sentiment analysis scenario (i.e., review data). Specifically, it is trained using the review data of the source language and the source domain with labeled sentiment polarity and the review data of the source language and the target domain without labeled sentiment polarity. The trained sentiment analysis model can also be used to classify the review text of the target language and the target domain and determine the sentiment polarity (positive or negative) of the review text.
[0111] The sentiment analysis model is trained and tested on a review dataset. The review dataset can be, for example, a labeled multilingual and multi-domain Amazon review dataset (see Peter Prettenhofer and Benno Stein. Cross-language text classification using structural correspondence learning. In ACL, pages 1118–1127, 2010). This dataset includes four languages: English, German, French, and Japanese. Each language includes three domains, namely books, DVDs, and music. Each domain of each language has a training set and a test set, and they both include the same number of positive and negative reviews.
[0112] According to one embodiment, during the training process of the sentiment analysis model, English can be used as the only source language and attempts can be made to adapt to the other three languages respectively, i.e., the other three languages are target languages. Since each language includes three domains, 3*2 source-target pairs can be constructed between English and a certain target language. For example, taking German as the target language, six source-target pairs can be constructed between English and German, namely English books-German DVDs, English books-German music, English DVDs-German books, English DVDs-German music, English music-German books, and English music-German DVDs. Considering there are three target languages, there are a total of 18 source-target pairs.
[0113] The unlabeled review data in the fields of books, DVDs, and music can be extracted from the unlabeled dataset released in 2016 (see Ruining He and Julian McAuley. Ups and downs: Modeling the visual evolution of fashion trends with one-class collaborative filtering. In WWW, pages 507–517. International World Wide Web Conferences Steering Committee, 2016.).
[0114] In the training stage of the sentiment analysis model, first, the unlabeled data D from the source language (English) and the target domain is used s,t to train the unsupervised feature decomposition module. For example, if one wants to transfer English DVDs to German books, the unlabeled data in English books is used to train the unsupervised feature decomposition module. Subsequently, the labeled data D in the source language and the source domain is used s,s to train the classification module. In the testing stage, the trained model is evaluated on the test set D in the target language and the target domain t,t .
[0115] Based on the trained sentiment analysis model, the present invention also provides a review sentiment analysis method 500. The method 500 uses the trained sentiment analysis model to determine the sentiment polarity of the review text to be analyzed. The method 500 can be executed, for example, in a user computing device (such as the aforementioned user computing device 9110) or a server computing system (such as the aforementioned server computing system 9130). As Figure 5 shown, the method 500 starts from step S510.
[0116] In step S510, the review text to be analyzed is obtained, where the language of the review text is any one of the source language and the target language set, and the field to which the review text belongs is any one of the source field and the target field set.
[0117] Subsequently, in step S520, the review text is input into a preset sentiment analysis model so that the sentiment analysis model outputs the sentiment polarity of the review text. The sentiment analysis model is trained with the review data of the source language and source field with labeled sentiment polarity and the review data of the source language and target field without labeled sentiment polarity as training samples.
[0118] In step S520, the sentiment analysis model takes the review text to be analyzed as input, performs a series of processing and calculations on the review text, and outputs the sentiment polarity label of the review text, thereby determining the category to which the text to be classified belongs. The specific structure of the sentiment analysis model can be referred to above and will not be elaborated here.
[0119] Figure 6 Shows a comparison chart of the classification effects of the text classification model of the present invention and other text classification models (both applied to the review sentiment analysis scenario).
[0120] Figure 6 The left column in shows the models participating in the comparison. Among them, XLM-UFD is the text classification model of the present invention, and other models such as CL-RL and Bi-PV are migration models proposed by predecessors. Figure 6 The numbers in are the classification accuracies (i.e., the adaptabilities of the models), where Avg represents the average value. The classification accuracy of each target language and target field is the average value of the English classification accuracies of the two source fields. For example, the classification accuracy under German-Books is the average value of the accuracies of the models trained on English DVDs and English music. As Figure 6 can be seen, the migration classification accuracy of the text classification model of the present invention is higher and is superior to all other models.
[0121] In addition, in order to understand the functions of the various modules in the text classification model of the present invention, the present invention conducted a thorough model ablation experiment. As Figure 7 shown, we use the XLM model (multilingual embedding module) as the benchmark model.
[0122] First, check the impact of maximizing the mutual information (MI) of the input and output of the domain-invariant feature extractor on the model, that is, Max MI w / o Res (w / o Res means without using the residual connection layer). The classification accuracy shows that Max-MI-w / o-Res with only domain-invariant features reduces the performance of XLM.
[0123] With the enhancement of the residual connection between the input and output of each processing unit in the domain-invariant feature extractor, Max-MI (with a residual connection layer) has shown a significant performance improvement compared to Max-MI-w / o-Res, but it is still slightly lower than XLM.
[0124] By supplementing the domain-specific feature extractor and the final minimized MI objective (i.e., L p ), the performance of MaxMin-MI has been significantly improved compared to MaxMI and is better than XLM, demonstrating that unsupervised feature decomposition can support dynamic domain-specific and domain-invariant feature combinations to improve classification performance.
[0125] By adding an intermediate minimized MI objective (i.e., L m ), MAX-2Min-MI performs best in German and French and can be used as the main model for other comparisons and ablations. We also experimented with the impact of different sizes of unlabeled data in the source language on the model performance. As Figure 7 can be seen, 2K (i.e., 2000) unlabeled original texts have already brought about a meaningful performance improvement of the text classification model of the present invention over XLM. Further increasing the unlabeled original texts will continuously improve the model performance, and when the size of the original data is greater than 10K, the performance improvement becomes insignificant.
[0126] The text classification method of the present invention is executed in a computing device. The computing device can be implemented, for example, as the aforementioned user computing device 9110 and server computing system 9130.
[0127] Figure 8 A schematic diagram of a computing device 600 according to an embodiment of the present invention is shown. As Figure 8 shown, the computing device 600 includes at least one processor 610 and a memory 620 storing program instructions. Among them, the memory 620 stores a text classification device 700 and / or a comment sentiment analysis device 800.
[0128] The device 700 includes program instructions for executing the text classification method 400 of the present invention. When the program instructions are read and executed by the processor 610, the computing device 600 executes the text classification method 400 of the present invention.
[0129] The device 800 includes program instructions for executing the comment sentiment analysis method 500 of the present invention. When the program instructions are read and executed by the processor 610, the computing device 600 executes the comment sentiment analysis method 500 of the present invention.
[0130] Figure 9 A schematic diagram of a text classification device 700 according to an embodiment of the present invention is shown. As Figure 7As shown, the apparatus 700 includes a text acquisition module 710 and a category determination module 720.
[0131] The text acquisition module 710 is adapted to acquire the text to be classified. Among them, the language of the text to be classified is any one of the source language and the target language set, and the field to which the text to be classified belongs is any one of the source field and the target field set. The specific functions and processing logics of the text acquisition module 710 can refer to the relevant descriptions in step S410 above and will not be elaborated here.
[0132] The category determination module 720 is adapted to input the text to be classified into a preset text classification model so that the text classification model outputs the category to which the text to be classified belongs. Among them, the text classification model is trained with the text data of the source language and source field with labeled categories and the text data of the source language and target field without labeled categories as training samples. The specific functions and processing logics of the category determination module 720 can refer to the relevant descriptions in step S420 above and will not be elaborated here.
[0133] Figure 10 The schematic diagram of a comment sentiment analysis apparatus 800 according to an embodiment of the present invention is shown. As Figure 10 shown, the comment sentiment analysis apparatus 800 includes a comment acquisition module 810 and a sentiment polarity determination module 820.
[0134] The comment acquisition module 810 is adapted to acquire the comment text to be analyzed. Among them, the language of the comment text is any one of the source language and the target language set, and the field to which the comment text belongs is any one of the source field and the target field set. The specific functions and processing logics of the comment acquisition module 810 can refer to the relevant descriptions in step S510 above and will not be elaborated here.
[0135] The sentiment polarity determination module 820 is adapted to input the comment text into a preset sentiment analysis model so that the sentiment analysis model outputs the sentiment polarity of the comment text. Among them, the sentiment analysis model is trained with the comment data of the source language and source field with labeled sentiment polarities and the comment data of the source language and target field without labeled sentiment polarities as training samples. The specific functions and processing logics of the sentiment polarity determination module 820 can refer to the relevant descriptions in step S520 above and will not be elaborated here.
[0136] The present invention also provides a readable storage medium storing program instructions. When the program instructions are read and executed by a computing device, the computing device is caused to execute the text classification method 400 and / or the comment sentiment analysis method 500 of the present invention.
[0137] The various technologies described herein can be implemented in combination with hardware or software, or a combination thereof. Thus, the methods and devices of the present invention, or certain aspects or portions of the methods and devices of the present invention, may take the form of program code (i.e., instructions) embedded in a tangible medium, such as a removable hard disk, a USB flash drive, a floppy disk, a CD-ROM, or any other machine-readable storage medium, wherein when the program is loaded into and executed by a machine such as a computer, the machine becomes a device for practicing the present invention.
[0138] In the case where the program code is executed on a programmable computer, the computing device generally includes a processor, a processor-readable storage medium (including volatile and non-volatile memories and / or storage elements), at least one input device, and at least one output device. Among them, the memory is configured to store the program code; the processor is configured to execute the text classification method or the comment sentiment analysis method of the present invention according to the instructions in the program code stored in the memory.
[0139] By way of example and not limitation, the readable medium includes a readable storage medium and a communication medium. The readable storage medium stores information such as computer-readable instructions, data structures, program modules, or other data. The communication medium generally embodies computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and includes any information delivery medium. A combination of any of the above is also included within the scope of the readable medium.
[0140] In the specification provided herein, the algorithms and displays are not inherently related to any particular computer, virtual system, or other device. Various general-purpose systems can also be used in conjunction with the examples of the present invention. Based on the above description, the structure required to construct such systems is obvious. In addition, the present invention is not directed to any particular programming language. It should be understood that the content of the present invention described herein can be implemented using various programming languages, and the description of the specific language above is for disclosing the preferred embodiments of the present invention.
[0141] In the specification provided herein, a large number of specific details are set forth. However, it can be understood that the embodiments of the present invention can be practiced without these specific details. In some instances, well-known methods, structures, and technologies have not been shown in detail so as not to obscure the understanding of this specification.
[0142] Similarly, it should be understood that, in order to streamline the present disclosure and assist in understanding one or more of the various inventive aspects, in the above description of the exemplary embodiments of the present invention, the various features of the present invention are sometimes grouped together into a single embodiment, figure, or description thereof. However, the disclosed method should not be construed as reflecting an intention that the claimed invention requires more features than are expressly recited in each claim. Rather, as reflected in the following claims, the inventive aspects lie in less than all the features of the single embodiments disclosed previously. Thus, the claims following the detailed description are hereby expressly incorporated into the detailed description, where each claim stands on its own as a separate embodiment of the present invention.
[0143] Those skilled in the art should understand that the modules or units or components of the devices in the examples disclosed herein can be arranged in the devices as described in this embodiment, or alternatively can be located in one or more devices different from the devices in this example. The modules in the foregoing examples can be combined into one module or further divided into multiple sub-modules.
[0144] Those skilled in the art can understand that the modules in the devices of the embodiments can be adaptively changed and arranged in one or more devices different from this embodiment. The modules or units or components in the embodiments can be combined into one module or unit or component, and further can be divided into multiple sub-modules or sub-units or sub-components. Except that at least some of such features and / or processes or units are mutually exclusive, any combination can be used to combine all the features disclosed in this specification (including the accompanying claims, abstract, and drawings) and all the processes or units of any method or device so disclosed. Unless otherwise expressly stated, each feature disclosed in this specification (including the accompanying claims, abstract, and drawings) can be replaced by an alternative feature that provides the same, equivalent, or similar purpose.
[0145] In addition, those skilled in the art can understand that although some of the embodiments described herein include certain features included in other embodiments rather than other features, the combination of the features of different embodiments means that it is within the scope of the present invention and forms different embodiments. For example, in the following claims, any one of the claimed embodiments can be used in any combination.
[0146] In addition, some of the embodiments described herein are described as a method or a combination of method elements that can be implemented by a processor of a computer system or by other devices performing the functions. Therefore, a processor having the necessary instructions for implementing the method or method elements forms a device for implementing the method or method elements. In addition, the elements described herein in the device embodiments are examples of devices for implementing the functions performed by the elements for the purpose of implementing the invention.
[0147] As used herein, unless otherwise specified, the use of ordinal numbers "first", "second", "third", etc. to describe ordinary objects only indicates different instances of similar objects and is not intended to imply that the objects so described must have a given order in time, space, ranking, or in any other way.
[0148] Although the present invention has been described in terms of a limited number of embodiments, those skilled in the art in this technical field will appreciate that other embodiments can be contemplated within the scope of the invention as thus described. In addition, it should be noted that the language used in this specification has been principally selected for readability and teaching purposes rather than for the purpose of explaining or limiting the subject matter of the invention. Therefore, many modifications and variations will be apparent to those of ordinary skill in the art in this technical field without departing from the scope and spirit of the appended claims. For the scope of the present invention, the disclosure of the present invention is illustrative rather than restrictive, and the scope of the present invention is defined by the appended claims.
Claims
1. A text classification method, comprising: Obtaining a text to be classified, wherein the language of the text to be classified is any one of a source language and a target language set, and the field to which the text to be classified belongs is any one of a source field and a target field set; Inputting the text to be classified into a preset text classification model so that the text classification model outputs the category to which the text to be classified belongs, wherein the text classification model is trained with text data in a source language and a source field with labeled categories and text data in a source language and a target field without labeled categories as training samples; the text classification model includes: a multilingual embedding module adapted to process the text to be classified to generate a cross-language representation vector of the text to be classified; an unsupervised feature decomposition module including a domain-invariant feature extractor and a domain-specific feature extractor adapted to extract domain-invariant features and domain-specific features from the cross-language representation vector, wherein the domain-invariant feature extractor includes at least two first processing units, each first processing unit includes at least one feedforward processing layer and a residual connection layer, the residual connection layer is adapted to add the input of the first feedforward processing layer of the corresponding first processing unit to the output of the last feedforward processing layer, the domain-specific feature extractor includes at least two second processing units, each second processing unit includes at least one feedforward processing layer, the number of first processing units included in the domain-invariant feature extractor is the same as the number of second processing units included in the domain-specific feature extractor, and the training objective of the unsupervised feature decomposition module is: to minimize the negative value of the mutual information between the input and output of the domain-invariant feature extractor and the weighted sum of the mutual information between the outputs of each pair of first processing units and second processing units; a classification module adapted to determine the category to which the text to be classified belongs according to the domain-invariant features and domain-specific features.
2. The method according to claim 1, wherein, The text classification model is trained according to the following steps: Obtaining a pre-trained multilingual embedding module; Using the text data in a source language and a target field without labeled categories as training samples to train the unsupervised feature decomposition module; After the unsupervised feature decomposition module is trained, using the text data in a source language and a source field with labeled categories as training samples to train the classification module.
3. The method according to claim 2, wherein The number of training samples of the unsupervised feature decomposition module is less than the number of training samples of the classification module.
4. The method according to any one of claims 1-3, wherein The classification module includes a linear layer and a Softmax feedforward layer.
5. The method according to claim 4, wherein, The training objective of the classification module is: to minimize the cross-entropy loss.
6. A review sentiment analysis method, comprising: Obtaining a review text to be analyzed, wherein the language of the review text is any one of a source language and a target language set, and the field to which the review text belongs is any one of a source field and a target field set; Input the review text into a preset sentiment analysis model so that the sentiment analysis model outputs the sentiment polarity of the review text. Among them, the sentiment analysis model is trained with review data in the source language and source domain with labeled sentiment polarity and review data in the source language and target domain without labeled sentiment polarity as training samples. The sentiment analysis model includes: a multilingual embedding module suitable for processing the review text to generate a cross-language representation vector of the review text; an unsupervised feature decomposition module including a domain-invariant feature extractor and a domain-specific feature extractor, suitable for extracting domain-invariant features and domain-specific features from the cross-language representation vector. Among them, the domain-invariant feature extractor includes at least two first processing units, and each first processing unit includes at least one feed-forward processing layer and a residual connection layer. The residual connection layer is suitable for adding the input of the first feed-forward processing layer of the corresponding first processing unit to the output of the last feed-forward processing layer. The domain-specific feature extractor includes at least two second processing units, and each second processing unit includes at least one feed-forward processing layer. The number of first processing units included in the domain-invariant feature extractor is the same as the number of second processing units included in the domain-specific feature extractor. The training objective of the unsupervised feature decomposition module is to minimize the negative value of the mutual information between the input and output of the domain-invariant feature extractor and the weighted sum of the mutual information between the outputs of each pair of first processing units and second processing units; a classification module suitable for determining the sentiment polarity of the review text according to the domain-invariant features and domain-specific features.
7. The method according to claim 6, wherein, The sentiment analysis model is trained according to the following steps: Obtain a pre-trained multilingual embedding module; Use the review data in the source language and target domain without labeled sentiment polarity as training samples to train the unsupervised feature decomposition module; After the unsupervised feature decomposition module is trained, use the review data in the source language and source domain with labeled sentiment polarity as training samples to train the classification module.
8. A text classification device, including: A text acquisition module suitable for acquiring the text to be classified. Among them, the language of the text to be classified is any one of the source language and the target language set, and the field to which the text to be classified belongs is any one of the source field and the target field set; A category determination module, adapted to input the text to be classified into a preset text classification model, so that the text classification model outputs the category to which the text to be classified belongs. Among them, the text classification model is trained with text data in the source language and source domain with labeled categories and text data in the source language and target domain without labeled categories as training samples; the text classification model includes: a multilingual embedding module, adapted to process the text to be classified to generate a cross-language representation vector of the text to be classified; an unsupervised feature decomposition module, including a domain-invariant feature extractor and a domain-specific feature extractor, adapted to extract domain-invariant features and domain-specific features from the cross-language representation vector. Among them, the domain-invariant feature extractor includes at least two first processing units, and each first processing unit includes at least one feed-forward processing layer and a residual connection layer. The residual connection layer is adapted to add the input of the first feed-forward processing layer of the corresponding first processing unit to the output of the last feed-forward processing layer. The domain-specific feature extractor includes at least two second processing units, and each second processing unit includes at least one feed-forward processing layer. The number of first processing units included in the domain-invariant feature extractor is the same as the number of second processing units included in the domain-specific feature extractor. The training objective of the unsupervised feature decomposition module is: to minimize the negative value of the mutual information between the input and output of the domain-invariant feature extractor and the weighted sum of the mutual information between the outputs of each pair of first processing units and second processing units; a classification module, adapted to determine the category to which the text to be classified belongs according to the domain-invariant features and domain-specific features.
9. A comment sentiment analysis device, comprising: A comment acquisition module, adapted to acquire a comment text to be analyzed. Among them, the language of the comment text is any one in the set of source language and target language, and the domain to which the comment text belongs is any one in the set of source domain and target domain; An emotional polarity determination module, adapted to input the review text into a preset sentiment analysis model, so that the sentiment analysis model outputs the emotional polarity of the review text, wherein the sentiment analysis model is trained with the review data of the source language and source domain with the labeled emotional polarity and the review data of the source language and target domain without the labeled emotional polarity as training samples; the sentiment analysis model includes: a multilingual embedding module, adapted to process the review text to generate a cross-language representation vector of the review text; an unsupervised feature decomposition module, including a domain-invariant feature extractor and a domain-specific feature extractor, adapted to extract domain-invariant features and domain-specific features from the cross-language representation vector, wherein the domain-invariant feature extractor includes at least two first processing units, each first processing unit includes at least one feed-forward processing layer and a residual connection layer, the residual connection layer is adapted to add the input of the first feed-forward processing layer of the corresponding first processing unit to the output of the last feed-forward processing layer, the domain-specific feature extractor includes at least two second processing units, each second processing unit includes at least one feed-forward processing layer, the number of the first processing units included in the domain-invariant feature extractor is the same as the number of the second processing units included in the domain-specific feature extractor, and the training objective of the unsupervised feature decomposition module is to minimize the negative value of the mutual information between the input and output of the domain-invariant feature extractor and the weighted sum of the mutual information between the outputs of each pair of the first processing unit and the second processing unit; a classification module, adapted to determine the emotional polarity of the review text according to the domain-invariant features and the domain-specific features.
10. A computing device, comprising: at least one processor and a memory storing program instructions; When the program instructions are read and executed by the processor, the computing device is caused to execute the text classification method according to any one of claims 1-5 and / or the review sentiment analysis method according to any one of claims 6-7.
11. A readable storage medium storing program instructions, when the program instructions are read and executed by a computing device, causing the computing device to execute the text classification method according to any one of claims 1-5 and / or the review sentiment analysis method according to any one of claims 6-7.
Citation Information
Patent Citations
Cross-domain emotion classification method and a related device
CN109492229A
Text classification method and device and storage medium
CN110796160A