Text Sentiment Classification Method, Device, Storage Medium and Electronic Device

The binary classifier processes text emotions and judges emotion categories based on preset probability thresholds, which solves the uncertainty problem of text emotion classification model in the prior art, and improves classification accuracy and efficiency.

CN114168735BActive Publication Date: 2025-07-22BEIJING YOUZHUJU NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111491713.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-08
Publication Date
2025-07-22
Estimated Expiration
2041-12-08

AI Technical Summary

Technical Problem

The text emotion classification model under the prior art has reduced classification ability and increased learning uncertainty because a large number of texts are divided into neutral emotion categories.

Method used

The binary classification method is adopted, and the positive and negative emotions categories are processed respectively through two binary classifiers, and the target emotions category of the text is judged based on the preset probability threshold, thereby reducing the uncertainty of the three-classification model.

Benefits of technology

The binary classification processing reduces the uncertainty of the text emotion classification model, improves classification accuracy and efficiency, and can more accurately judge the emotional categories of the text.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114168735B_ABST
    Figure CN114168735B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a method, apparatus, storage medium, and electronic device for text sentiment classification. The method includes: obtaining the text to be classified; performing first classification processing and second classification processing on the text to be classified respectively to obtain a first classification result and a second classification result, where the first classification result includes a first probability for characterizing that the text to be classified is a positive sentiment category, and the second classification result includes a second probability for characterizing that the text to be classified is a negative sentiment category; determining the target sentiment category of the text to be classified according to the first probability, the second probability, a preset positive probability threshold, and a preset negative probability threshold, where the target sentiment category is a positive sentiment category, a negative sentiment category, or a neutral sentiment category. Through this solution, the classification uncertainty of the three-class machine model is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of natural language processing, and specifically, to a method, apparatus, storage medium, and electronic device for text sentiment classification. Background Art

[0002] Language is a tool for human communication and can be the main medium for humans to express emotions. In the Internet, text sentiment classification is an important topic. For example, in a live broadcast room, if the emotional state of the commenting users during the interaction can be judged, for some comments with negative emotions, certain measures can be taken to handle them in order to create a good emotional communication atmosphere for the live broadcast room.

[0003] Under the existing technology, text sentiment classification uses a machine model with a softmax function to perform three-class classification of positive sentiment, negative sentiment, and neutral sentiment. However, in data annotation, a considerable part is classified into the category of "neutral sentiment". In this way, the classification ability of the model will be potentially damaged, and the uncertainty of model learning will be increased. Summary of the Invention

[0004] This Summary of the Invention section is provided to introduce concepts in a brief form, which will be described in detail in the following Detailed Description section. This Summary of the Invention section is not intended to identify key features or essential features of the claimed technical solution, nor is it intended to be used to limit the scope of the claimed technical solution.

[0005] In a first aspect, the present disclosure provides a method for text sentiment classification, including:

[0006] Obtain a text to be classified;

[0007] Perform first classification processing and second classification processing on the text to be classified respectively to obtain a first classification result and a second classification result. The first classification result includes a first probability for characterizing that the text to be classified is in the positive sentiment category, and the second classification result includes a second probability for characterizing that the text to be classified is in the negative sentiment category;

[0008] Determine the target sentiment category of the text to be classified according to the first probability, the second probability, a preset positive probability threshold, and a preset negative probability threshold. The target sentiment category is a positive sentiment category, a negative sentiment category, or a neutral sentiment category.

[0009] In a second aspect, the present disclosure provides a text sentiment classification apparatus, including:

[0010] A first acquisition module, configured to obtain a text to be classified;

[0011] A first processing module, configured to perform a first classification process and a second classification process on the text to be classified respectively, to obtain a first classification result and a second classification result, where the first classification result includes a first probability for characterizing that the text to be classified is of a positive sentiment category, and the second classification result includes a second probability for characterizing that the text to be classified is of a negative sentiment category;

[0012] A first determination module, configured to determine a target sentiment category of the text to be classified according to the first probability, the second probability, a preset positive probability threshold, and a preset negative probability threshold, where the target sentiment category is a positive sentiment category, a negative sentiment category, or a neutral sentiment category.

[0013] In a third aspect, the present disclosure provides a computer-readable medium, on which a computer program is stored, and when the program is executed by a processing device, the steps of the method in the first aspect are implemented.

[0014] In a fourth aspect, the present disclosure provides an electronic device, including:

[0015] A storage device, on which a computer program is stored;

[0016] A processing device, configured to execute the computer program in the storage device to implement the steps of the method in the first aspect.

[0017] Through the above technical solutions, the text to be classified is obtained; the text to be classified is respectively subjected to a first classification process and a second classification process to obtain a first classification result and a second classification result. The first classification result includes a first probability for characterizing that the text to be classified is of a positive sentiment category, and the second classification result includes a second probability for characterizing that the text to be classified is of a negative sentiment category; according to the first probability, the second probability, a preset positive probability threshold, and a preset negative probability threshold, the target sentiment category of the text to be classified is determined, and the target sentiment category is a positive sentiment category, a negative sentiment category, or a neutral sentiment category. Among them, the first classification process is used to classify the text to be classified into a positive sentiment category and a non-positive sentiment category, and the second classification process is used to classify the text to be classified into a negative sentiment category and a non-negative sentiment category. In this way, by means of binary classification and combining the classification results of two binary classifications to determine the target sentiment category of the text to be classified, the classification uncertainty of the three-class machine model is reduced.

[0018] Other features and advantages of the present disclosure will be described in detail in the subsequent specific implementation part. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In conjunction with the accompanying drawings and with reference to the following detailed description, the above and other features, advantages, and aspects of the embodiments of the present disclosure will become more apparent. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic and that the original elements and elements are not necessarily drawn to scale. In the drawings:

[0020] Figure 1 is a flowchart of a text sentiment classification method shown according to an exemplary embodiment of the present disclosure.

[0021] Figure 2 is a schematic diagram of the network structure of a first sentiment classifier and a second sentiment classifier shown according to an exemplary embodiment of the present disclosure.

[0022] Figure 3 is another flowchart of a text sentiment classification method shown according to an exemplary embodiment of the present disclosure.

[0023] Figure 4 is a block diagram of a text sentiment classification device shown according to an exemplary embodiment of the present disclosure.

[0024] Figure 5 is a schematic diagram of the structure of an electronic device shown according to an exemplary embodiment of the present disclosure. Detailed Description

[0025] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Instead, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.

[0026] It should be understood that the various steps recited in the method embodiments of the present disclosure can be executed in a different order and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard.

[0027] As used herein, the term "including" and its variations are open-ended, i.e., "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.

[0028] It should be noted that the concepts such as "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order of functions executed by these devices, modules or units or their interdependent relationships.

[0029] It should be noted that the modification of "one" and "multiple" mentioned in this disclosure is illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly specified in the context, it should be understood as "one or more".

[0030] The names of the messages or information exchanged between multiple devices in the embodiments of this disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.

[0031] As described in the background art, directly performing three-class classification using a machine model with the softmax function will potentially damage the classification ability of the model and increase the uncertainty of model learning.

[0032] In view of this, this disclosure provides a text sentiment classification method, device, storage medium and electronic device. By means of binary classification and combining the classification results of two binary classifications, the target sentiment category of the text to be classified is determined, reducing the classification uncertainty of the three-class machine model.

[0033] The following introduces and illustrates the detailed technical solutions of this disclosure with specific examples.

[0034] Figure 1 It is a flowchart of a text sentiment classification method shown according to an exemplary embodiment of this disclosure. This text sentiment classification method can be applied to an electronic device. Referring to Figure 1 this text sentiment classification method may include the following steps:

[0035] S101, obtain the text to be classified.

[0036] In some embodiments, the text to be classified may be the text content obtained from an image. Among them, obtaining the text content based on optical character recognition of the image can refer to related technologies, which will not be elaborated in this embodiment.

[0037] In some embodiments, the text to be classified may be the text content obtained from speech. Among them, obtaining the text content based on speech recognition technology can refer to related technologies, which will not be elaborated in this embodiment.

[0038] In some embodiments, the text to be classified may be the comment text obtained from a live broadcast room. Correspondingly, obtaining the comment text from the live broadcast room can be based on an image acquisition method or a voice acquisition method, and this embodiment does not limit the text content acquisition method.

[0039] S102. Perform first classification processing and second classification processing on the text to be classified respectively to obtain a first classification result and a second classification result. The first classification result includes a first probability for characterizing that the text to be classified is a positive sentiment category, and the second classification result includes a second probability for characterizing that the text to be classified is a negative sentiment category.

[0040] It should be noted that the first classification processing is used to classify the text to be classified into a positive sentiment category and a non-positive sentiment category. Correspondingly, the first classification result includes a first probability for characterizing that the text to be classified is a positive sentiment category and a third probability for characterizing that the text to be classified is a non-positive sentiment category. It can be understood that the non-positive sentiment category includes a negative sentiment category and a neutral sentiment category. The second classification processing is used to classify the text to be classified into a negative sentiment category and a non-negative sentiment category. Correspondingly, the second classification result includes a second probability for characterizing that the text to be classified is a negative sentiment category and a fourth probability for characterizing that the text to be classified is a non-negative sentiment category. It can be understood that the non-negative sentiment category includes a positive sentiment category and a neutral sentiment category.

[0041] In some embodiments, performing the first classification processing and the second classification processing on the text to be classified in the above step S102 to obtain the first classification result and the second classification result may include: calculating the text feature vector by using the feature extraction layer shared by the first sentiment classifier and the second sentiment classifier for the text to be classified; inputting the text feature vector into the network layer other than the feature extraction layer in the first sentiment classifier for calculation to obtain the first classification result; inputting the text feature vector into the network layer other than the feature extraction layer in the second sentiment classifier for calculation to obtain the second classification result.

[0042] Refer to Figure 2 , Figure 2 which is a schematic diagram of the network structure of a first sentiment classifier and a second sentiment classifier. Perform the first classification processing and the second classification processing on the text to be classified through the Figure 2 shown network structure. In Figure 2 , the feature extraction layer and the first network layer form the first sentiment classifier. The first network layer is the network layer other than the feature extraction layer in the first sentiment classifier. The first sentiment classifier is used to perform the first classification processing to obtain the first classification result; the feature extraction layer and the second network layer form the second sentiment classifier. The second network layer is the network layer other than the feature extraction layer in the second sentiment classifier. The second sentiment classifier is used to perform the second classification processing to obtain the second classification result. Among them, sharing the feature extraction network by the first sentiment classifier and the second sentiment classifier can reduce the classification time cost of the text to be classified and improve the classification efficiency.

[0043] In some embodiments, the feature extraction layer may input a text feature vector of 768 dimensions. In some embodiments, the feature extraction layer may be a BERT (Bidirectional Encoder Representation from Transformers) model for capturing the semantic information representation of the text to be classified. Referring to the related art for extracting the text feature vector of the text by using the BERT model, this embodiment will not elaborate herein.

[0044] In some embodiments, the first network layer and the second network layer may be fully connected layers using the sigmoid function. In some embodiments, the sigmoid function may be as shown in the following formula:

[0045] P(y = 1|x) = exp(w * x) / (1 + exp(w * x));

[0046] where P represents the probability of y = 1, 1 corresponds to a category, x is the text feature vector of the text to be classified, and w is the weight coefficient. Taking the first classification process as an example, the sigmoid function is used to predict the probabilities that the text to be classified belongs to the positive sentiment category and the non-positive sentiment category respectively. Then y = 1 can represent the probability that the text to be classified is in the positive sentiment category. Correspondingly, the probability that the text to be classified is in the non-positive sentiment category is 1 - P. In the training stage, y = 1 indicates that the label is valid.

[0047] In some embodiments, the first sentiment classifier may be trained in the following manner: obtaining a first training sample set, where the first training sample set includes positive sample texts representing the positive sentiment category and negative sample texts representing the non-positive sentiment category; training the first sentiment classifier according to the first training sample set.

[0048] It should be noted that the negative sample texts representing the non-positive sentiment category include samples of the negative sentiment category and samples of the neutral sentiment category. Digital labels may be used to represent the positive sample texts and the negative sample texts. Exemplarily, the positive sample text belonging to the positive sentiment category may be represented by the number 1, and the negative sample text belonging to the non-positive sentiment category may be represented by the number 0.

[0049] For example, training the first sentiment classifier according to the first training sample set may include: iteratively updating the weight parameters of the first initial sentiment classifier by using the first training sample set, and outputting the first sentiment classifier when the preset iteration stop condition is satisfied. In some embodiments, the preset iteration stop condition may be that the number of iterations reaches a preset number.

[0050] Continuing with the example of the sigmoid function above, the training process of the first sentiment classifier in the embodiments of the present disclosure will be further explained. In the case where the input is a positive sample text with a label of 1, P represents the probability that the positive sample text is classified as belonging to the positive sentiment category by the first initial sentiment classifier. Correspondingly, the probability that the positive sample text belongs to a non-positive sentiment category is 1 - P. And when P = 1, it means the prediction is correct. When P is not equal to 1, the weight parameters can be updated, and the sample can be used for the next classification prediction. This process is repeated until the iteration stop condition is met.

[0051] In the embodiments of the present disclosure, the training sample data of the neutral sentiment category participates in the first sentiment classifier, so that the first sentiment classifier will not be confused by fuzzy data, enhancing the classification ability of the first sentiment classifier and reducing the uncertainty of model learning.

[0052] In some embodiments, the second sentiment classifier is trained as follows: obtain a second training sample set, where the second training sample set includes positive sample texts representing the negative sentiment category and negative sample texts representing the non-negative sentiment category; train the second sentiment classifier according to the second training sample set.

[0053] It should be noted that the negative sample texts representing the non-negative sentiment category include samples of the positive sentiment category and samples of the neutral sentiment category. Similar to the example of training the first sentiment classifier, digital labels can be used to represent positive sample texts and negative sample texts. For example, the number 1 can be used to represent the positive sample text belonging to the negative sentiment category, and the number 0 can be used to represent the negative sample text belonging to the non-negative sentiment category.

[0054] For example, training the second sentiment classifier according to the second training sample set may include: iteratively updating the weight parameters of the second initial sentiment classifier using the second training sample set, and outputting the second sentiment classifier when the preset iteration stop condition is met. Similar to the example of training the first sentiment classifier, the preset iteration stop condition may be that the number of iterations reaches a preset number.

[0055] Continuing with the example of the sigmoid function above, the training process of the second sentiment classifier in the embodiments of the present disclosure will be further explained. In the case where the input is a positive sample text with a label of 1, P represents the probability that the positive sample text is classified as belonging to the negative sentiment category by the first initial sentiment classifier. Correspondingly, the probability that the positive sample text belongs to a non-negative sentiment category is 1 - P. And when P = 1, it means the prediction is correct. When P is not equal to 1, the weight parameters can be updated, and the sample can be used for the next classification prediction. This process is repeated until the iteration stop condition is met.

[0056] In the embodiments of the present disclosure, the training sample data of the neutral emotion category participates in the second emotion classifier and is used as negative samples for training, so that the second emotion classifier will not be confused by fuzzy data, enhancing the classification ability of the first emotion classifier and reducing the uncertainty of model learning.

[0057] S103. Determine the target emotion category of the text to be classified according to the first probability, the second probability, a preset positive probability threshold, and a preset negative probability threshold, where the target emotion category is a positive emotion category, a negative emotion category, or a neutral emotion category.

[0058] It should be noted that the preset positive probability threshold and the preset negative probability threshold are parameters used to assist in judging the target emotion category of the text to be classified. In some embodiments, the preset negative probability threshold and the preset positive probability threshold can be set according to actual situations.

[0059] In some embodiments, the above step S103 may include: when a first preset condition is met, determining that the target emotion category of the text to be classified is a positive emotion category, where the first preset condition is that the first probability is greater than the second probability and the first probability is greater than the preset positive probability threshold; when a second preset condition is met, determining that the target emotion category of the text to be classified is a negative emotion category, where the second preset condition is that the second probability is greater than the first probability and the second probability is greater than the preset negative probability threshold; when the first preset condition or the second preset condition is not met, determining that the target emotion category of the text to be classified is a neutral emotion category.

[0060] In the above manner, a first classification process and a second classification process are performed on the text to be classified. Among them, the first classification process is used to classify the text to be classified into a positive emotion category and a non-positive emotion category, and the second classification process is used to classify the text to be classified into a negative emotion category and a non-negative emotion category. In this way, by means of binary classification and combining the classification results of two binary classifications, the target emotion category of the text to be classified is judged, reducing the classification uncertainty of the three-class machine model.

[0061] Considering that when the text to be classified is a review text obtained from a live broadcast scenario, the possible public opinion in the live broadcast room can be judged according to the target sentiment category of the text to be classified, and the review text can be processed in some cases. For example, for some review texts with negative sentiment categories (such as vulgar review texts), in order to create a good interaction atmosphere in the live broadcast room, the display attributes of the review texts with negative sentiment categories in the live broadcast room (such as display time, whether to display, etc.) can be controlled; for some review texts with positive sentiment categories, in order to avoid the situation that a large number of fake review texts are posted by live broadcast users and cause misguidance to other viewing users, the display attributes of the comments of the live broadcast user in the live broadcast room can be controlled; for some review texts with neutral sentiment categories, since they do not affect the interaction in the live broadcast room and mislead users, such comments may not be further processed. On this basis, in order to avoid misprocessing caused by misjudging the review text, the review texts with negative sentiment categories and the texts with positive sentiment categories can be handed over to humans for re-review. Further, in order to save the cost of manual review, the review volume handed over to humans for re-review can be controlled. Based on this consideration, by setting a preset positive probability threshold and a preset negative probability threshold, the number of texts judged as neutral sentiment categories by combining the first sentiment classifier and the second sentiment classifier can be controlled, thereby reducing the number of texts to be manually reviewed. For example, Figure 3 is another flowchart of a text sentiment classification method shown according to an exemplary embodiment of the present disclosure. Refer to Figure 3 , the method includes the following steps:

[0062] Step S301, obtain the text to be classified within a preset time period.

[0063] In some embodiments, the preset time period can be set according to the actual situation. The text to be classified in this embodiment is similar to the text to be classified shown above Figure 1 and will not be elaborated here in this embodiment.

[0064] Step S302, determine a preset positive probability threshold and a preset negative probability threshold according to the number of texts to be classified within the preset time period.

[0065] In this embodiment, the preset positive probability threshold and the preset negative probability threshold are determined by the number distribution of the texts to be classified within the preset time period. It can be understood that the positive probability threshold and the preset negative probability threshold are parameters for assisting in determining whether the text to be classified is of neutral sentiment category. Therefore, by setting this threshold, the number of texts to be classified as neutral sentiment category can be controlled.

[0066] In some embodiments, a correspondence table of the number of text to be classified within a preset time period with a preset positive probability threshold and a preset negative probability threshold can be set. According to this correspondence table, a uniquely corresponding preset positive probability threshold and preset negative probability threshold are determined based on the number of text to be classified within the preset time period. Among them, the correspondence table can be specifically set according to the actual situation, and this embodiment does not limit it here.

[0067] Step S303: Perform first classification processing and second classification processing on the text to be classified respectively to obtain a first classification result and a second classification result. The first classification result includes a first probability indicating that the text to be classified is a positive sentiment category, and the second classification result includes a second probability indicating that the text to be classified is a negative sentiment category.

[0068] Step S304: Determine the target sentiment category of the text to be classified according to the first probability, the second probability, the preset positive probability threshold, and the preset negative probability threshold.

[0069] Steps S303 and S304 can refer to Figure 1 Steps S102 and S103 shown, and this embodiment will not elaborate here.

[0070] In the above manner, the preset positive probability threshold and the preset negative probability threshold are determined by the number of text to be classified within the preset time period, so as to control the number of central sentiment categories, thereby saving the time cost of manual review.

[0071] Figure 4 is a block diagram of a text sentiment classification device shown according to an exemplary embodiment. Referring to Figure 4 , the device 400 includes:

[0072] A first acquisition module 401, configured to acquire the text to be classified;

[0073] A first processing module 402, configured to perform first classification processing and second classification processing on the text to be classified respectively to obtain a first classification result and a second classification result. The first classification result includes a first probability indicating that the text to be classified is a positive sentiment category, and the second classification result includes a second probability indicating that the text to be classified is a negative sentiment category;

[0074] A first determination module 403, configured to determine the target sentiment category of the text to be classified according to the first probability, the second probability, the preset positive probability threshold, and the preset negative probability threshold. The target sentiment category is a positive sentiment category, a negative sentiment category, or a neutral sentiment category.

[0075] Optionally, the first processing module 402 includes:

[0076] A feature extraction sub-module, configured to calculate the text to be classified by using a feature extraction layer shared by a first sentiment classifier and a second sentiment classifier, so as to obtain a text feature vector;

[0077] A first calculation sub-module, configured to input the text feature vector into a network layer other than the feature extraction layer in the first sentiment classifier for calculation, so as to obtain the first classification result;

[0078] A second calculation sub-module, configured to input the text feature vector into a network layer other than the feature extraction layer in the second sentiment classifier for calculation, so as to obtain the second classification result.

[0079] Optionally, the apparatus 400 includes:

[0080] A second acquisition module, configured to acquire a first training sample set, where the first training sample set includes positive sample texts representing positive sentiment categories and negative sample texts representing non-positive sentiment categories;

[0081] A first training module, configured to train the first sentiment classifier according to the first training sample set.

[0082] Optionally, the apparatus 400 includes:

[0083] A third acquisition module, configured to acquire a second training sample set, where the second training sample set includes positive sample texts representing negative sentiment categories and negative sample texts representing non-negative sentiment categories;

[0084] A second training module, configured to train the second sentiment classifier according to the second training sample set.

[0085] Optionally, the first determination module 403 includes:

[0086] A first determination sub-module, configured to determine that the target sentiment category of the text to be classified is a positive sentiment category when a first preset condition is met, where the first preset condition is that the first probability is greater than the second probability, and the first probability is greater than the preset positive probability threshold;

[0087] A second determination sub-module, configured to determine that the target sentiment category of the text to be classified is a negative sentiment category when a second preset condition is met, where the second preset condition is that the second probability is greater than the first probability, and the second probability is greater than the preset negative probability threshold;

[0088] A third determination sub-module, configured to determine that the target sentiment category of the text to be classified is a neutral sentiment category when the first preset condition or the second preset condition is not met.

[0089] Optionally, the first obtaining module 401 is specifically configured to obtain the text to be classified within a preset time period;

[0090] The apparatus 400 further includes:

[0091] A second determining module, configured to determine a preset positive probability threshold and a preset negative probability threshold according to the number of texts of the text to be classified within the preset time period;

[0092] The first determining module 403 is specifically configured to, after determining the preset positive probability threshold and the preset negative probability threshold, perform a first classification process and a second classification process on the text to be classified respectively, to obtain a first classification result and a second classification result.

[0093] An embodiment of the present disclosure further provides a computer-readable medium, on which a computer program is stored, and when the program is executed by a processing device, the steps of the above-mentioned method are implemented.

[0094] An embodiment of the present disclosure further provides an electronic device, including:

[0095] A storage device, on which a computer program is stored;

[0096] A processing device, configured to execute the computer program in the storage device to implement the steps of the above-mentioned method.

[0097] Next, refer to Figure 5 , which shows a schematic structural diagram of an electronic device 500 suitable for implementing an embodiment of the present disclosure. The terminal device in the embodiment of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 5 The electronic device shown is only an example, and should not impose any limitation on the functions and usage scope of the embodiment of the present disclosure.

[0098] As Figure 5 shown, the electronic device 500 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 501, which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage device 508 into a random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the electronic device 500 are also stored. The processing device 501, the ROM 502, and the RAM 503 are connected to each other through a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0099] Typically, the following devices can be connected to the I / O interface 505: an input device 506 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 507 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 508 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 509. The communication device 509 can allow the electronic device 500 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 5 the electronic device 500 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices can be implemented or had.

[0100] Specifically, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product that includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program codes for performing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through the communication device 509, or installed from the storage device 508, or installed from the ROM 502. When the computer program is executed by the processing device 501, the above functions defined in the method of the embodiment of the present disclosure are performed.

[0101] It should be noted that the above-mentioned computer-readable medium in the present disclosure can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present disclosure, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which computer-readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable signal medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0102] In some embodiments, the electronic device can communicate using any currently known or future-developed network protocol such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed network.

[0103] The above-mentioned computer-readable medium can be included in the above-mentioned electronic device; it can also exist separately without being assembled into the electronic device.

[0104] The above computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to: obtain the text to be classified; perform a first classification process and a second classification process on the text to be classified respectively to obtain a first classification result and a second classification result, where the first classification result includes a first probability for characterizing that the text to be classified is a positive sentiment category, and the second classification result includes a second probability for characterizing that the text to be classified is a negative sentiment category; determine the target sentiment category of the text to be classified according to the first probability, the second probability, a preset positive probability threshold, and a preset negative probability threshold, where the target sentiment category is a positive sentiment category, a negative sentiment category, or a neutral sentiment category.

[0105] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages or combinations thereof. The programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0106] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that, in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0107] The modules involved in the embodiments of the present disclosure can be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation on the module itself in some cases. For example, the first acquisition module can also be described as "the module for acquiring the text to be classified".

[0108] The functions described above herein can be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Arrays (FPGA), Application Specific Integrated Circuits (ASIC), Application Specific Standard Products (ASSP), System on a Chip (SOC), Complex Programmable Logic Devices (CPLD), and so on.

[0109] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media would include electrical connections based on one or more wires, portable computer disks, hard disks, Random Access Memory (RAM), Read Only Memory (ROM), Erasable Programmable Read Only Memory (EPROM or Flash Memory), optical fibers, portable compact disc read only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0110] According to one or more embodiments of the present disclosure, Example 1 provides a method for text sentiment classification, including:

[0111] Acquire the text to be classified;

[0112] Perform a first classification process and a second classification process on the text to be classified respectively to obtain a first classification result and a second classification result. The first classification result includes a first probability for characterizing that the text to be classified is in the positive sentiment category, and the second classification result includes a second probability for characterizing that the text to be classified is in the negative sentiment category;

[0113] Determine the target sentiment category of the text to be classified according to the first probability, the second probability, a preset positive probability threshold, and a preset negative probability threshold. The target sentiment category is the positive sentiment category, the negative sentiment category, or the neutral sentiment category.

[0114] According to one or more embodiments of the present disclosure, Example 2 provides the method of Example 1. The method of separately performing a first classification process and a second classification process on the text to be classified to obtain a first classification result and a second classification result includes:

[0115] Calculating the text to be classified by using a feature extraction layer shared by a first sentiment classifier and a second sentiment classifier to obtain a text feature vector;

[0116] Inputting the text feature vector into a network layer other than the feature extraction layer in the first sentiment classifier for calculation to obtain the first classification result;

[0117] Inputting the text feature vector into a network layer other than the feature extraction layer in the second sentiment classifier for calculation to obtain the second classification result.

[0118] According to one or more embodiments of the present disclosure, Example 3 provides the method of Example 2. The first sentiment classifier is trained in the following manner:

[0119] Obtaining a first training sample set, where the first training sample set includes positive sample texts representing positive sentiment categories and negative sample texts representing non-positive sentiment categories;

[0120] Training the first sentiment classifier according to the first training sample set.

[0121] According to one or more embodiments of the present disclosure, Example 4 provides the method of Example 2. The second sentiment classifier is trained in the following manner:

[0122] Obtaining a second training sample set, where the second training sample set includes positive sample texts representing negative sentiment categories and negative sample texts representing non-negative sentiment categories;

[0123] Training the second sentiment classifier according to the second training sample set.

[0124] According to one or more embodiments of the present disclosure, Example 5 provides the method of Example 1. Determining the target sentiment category of the text to be classified according to the first probability, the second probability, a preset positive probability threshold, and a preset negative probability threshold includes:

[0125] When a first preset condition is satisfied, determining the target sentiment category of the text to be classified as a positive sentiment category. The first preset condition is that the first probability is greater than the second probability, and the first probability is greater than the preset positive probability threshold;

[0126] When the second preset condition is satisfied, it is determined that the target sentiment category of the text to be classified is the negative sentiment category, where the second preset condition is that the second probability is greater than the first probability and the second probability is greater than the preset negative probability threshold;

[0127] When the first preset condition or the second preset condition is not satisfied, it is determined that the target sentiment category of the text to be classified is the neutral sentiment category.

[0128] According to one or more embodiments of the present disclosure, Example 6 provides the method of Example 1, where obtaining the text to be classified includes: obtaining the text to be classified within a preset time period;

[0129] The method further includes:

[0130] According to the number of texts to be classified within the preset time period, determining a preset positive probability threshold and a preset negative probability threshold;

[0131] After determining the preset positive probability threshold and the preset negative probability threshold, perform the steps of respectively performing a first classification process and a second classification process on the text to be classified to obtain a first classification result and a second classification result.

[0132] According to one or more embodiments of the present disclosure, Example 7 provides a text sentiment classification device, including:

[0133] A first acquisition module, configured to acquire the text to be classified;

[0134] A first processing module, configured to respectively perform a first classification process and a second classification process on the text to be classified to obtain a first classification result and a second classification result, where the first classification result includes a first probability for characterizing that the text to be classified is a positive sentiment category, and the second classification result includes a second probability for characterizing that the text to be classified is a negative sentiment category;

[0135] A first determination module, configured to determine the target sentiment category of the text to be classified according to the first probability, the second probability, the preset positive probability threshold, and the preset negative probability threshold, where the target sentiment category is a positive sentiment category, a negative sentiment category, or a neutral sentiment category.

[0136] According to one or more embodiments of the present disclosure, Example 8 provides the device of Example 7, where the first processing module includes:

[0137] A feature extraction sub-module, configured to calculate the text to be classified by using a feature extraction layer shared by a first sentiment classifier and a second sentiment classifier to obtain a text feature vector;

[0138] The first calculation sub-module is configured to input the text feature vector into the network layers other than the feature extraction layer in the first sentiment classifier for calculation to obtain the first classification result;

[0139] The second calculation sub-module is configured to input the text feature vector into the network layers other than the feature extraction layer in the second sentiment classifier for calculation to obtain the second classification result.

[0140] According to one or more embodiments of the present disclosure, Example 9 provides a computer-readable medium having a computer program stored thereon, and when the program is executed by a processing device, the steps of the method according to any one of Examples 1-6 are implemented.

[0141] According to one or more embodiments of the present disclosure, Example 10 provides an electronic device, including:

[0142] A storage device having a computer program stored thereon;

[0143] A processing device configured to execute the computer program in the storage device to implement the steps of the method according to any one of Examples 1-6.

[0144] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in the present disclosure.

[0145] In addition, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order. In certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments.

[0146] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. On the contrary, the specific features and acts described above are merely example forms for implementing the claims. Regarding the devices in the above embodiments, the specific manner in which each module performs operations has been described in detail in the embodiments related to the method, and will not be elaborated herein.

Claims

1. A text sentiment classification method, characterized in that, Including: Obtain the text to be classified; Perform first classification processing and second classification processing on the text to be classified respectively to obtain a first classification result and a second classification result. The first classification result includes a first probability for characterizing that the text to be classified is a positive sentiment category, and the second classification result includes a second probability for characterizing that the text to be classified is a negative sentiment category; Determine the target sentiment category of the text to be classified according to the first probability, the second probability, a preset positive probability threshold, and a preset negative probability threshold. The target sentiment category is a positive sentiment category, a negative sentiment category, or a neutral sentiment category; The determining the target sentiment category of the text to be classified according to the first probability, the second probability, a preset positive probability threshold, and a preset negative probability threshold includes: when the first preset condition is satisfied, determining that the target sentiment category of the text to be classified is a positive sentiment category. The first preset condition is that the first probability is greater than the second probability and the first probability is greater than the preset positive probability threshold; when the second preset condition is satisfied, determining that the target sentiment category of the text to be classified is a negative sentiment category. The second preset condition is that the second probability is greater than the first probability and the second probability is greater than the preset negative probability threshold; when the first preset condition or the second preset condition is not satisfied, determining that the target sentiment category of the text to be classified is a neutral sentiment category.

2. The method according to claim 1, wherein The performing first classification processing and second classification processing on the text to be classified respectively to obtain a first classification result and a second classification result includes: Calculate the text to be classified by using a feature extraction layer shared by a first sentiment classifier and a second sentiment classifier to obtain a text feature vector; Input the text feature vector into a network layer other than the feature extraction layer in the first sentiment classifier for calculation to obtain the first classification result; Input the text feature vector into a network layer other than the feature extraction layer in the second sentiment classifier for calculation to obtain the second classification result.

3. The method according to claim 2, wherein The first sentiment classifier is trained in the following manner: Obtain a first training sample set, where the first training sample set includes positive sample texts representing positive sentiment categories and negative sample texts representing non-positive sentiment categories; Train the first sentiment classifier according to the first training sample set.

4. The method according to claim 2, wherein The second sentiment classifier is trained in the following manner: Obtain a second training sample set, where the second training sample set includes positive sample texts representing negative sentiment categories and negative sample texts representing non-negative sentiment categories; Train the second sentiment classifier according to the second training sample set.

5. The method according to claim 1, wherein The obtaining the text to be classified includes: obtaining the text to be classified within a preset time period; The method further includes: Determine a preset positive probability threshold and a preset negative probability threshold according to the number of texts to be classified within the preset time period; After determining a preset positive probability threshold and a preset negative probability threshold, perform the steps of respectively performing a first classification process and a second classification process on the text to be classified to obtain a first classification result and a second classification result.

6. A text sentiment classification device, characterized in that, Including: A first acquisition module, configured to acquire the text to be classified; A first processing module, configured to respectively perform a first classification process and a second classification process on the text to be classified to obtain a first classification result and a second classification result, where the first classification result includes a first probability for characterizing that the text to be classified is of a positive sentiment category, and the second classification result includes a second probability for characterizing that the text to be classified is of a negative sentiment category; A first determination module, configured to determine the target sentiment category of the text to be classified according to the first probability, the second probability, the preset positive probability threshold, and the preset negative probability threshold, where the target sentiment category is a positive sentiment category, a negative sentiment category, or a neutral sentiment category; The first determination module includes: A first determination sub-module, configured to determine that the target sentiment category of the text to be classified is a positive sentiment category when a first preset condition is satisfied, where the first preset condition is that the first probability is greater than the second probability, and the first probability is greater than the preset positive probability threshold; A second determination sub-module, configured to determine that the target sentiment category of the text to be classified is a negative sentiment category when a second preset condition is satisfied, where the second preset condition is that the second probability is greater than the first probability, and the second probability is greater than the preset negative probability threshold; A third determination sub-module, configured to determine that the target sentiment category of the text to be classified is a neutral sentiment category when the first preset condition or the second preset condition is not satisfied.

7. The device according to claim 6, characterized in that, The first processing module includes: A feature extraction sub-module, configured to calculate the text to be classified by using a feature extraction layer shared by a first sentiment classifier and a second sentiment classifier to obtain a text feature vector; A first calculation sub-module, configured to input the text feature vector into a network layer other than the feature extraction layer in the first sentiment classifier for calculation to obtain the first classification result; A second calculation sub-module, configured to input the text feature vector into a network layer other than the feature extraction layer in the second sentiment classifier for calculation to obtain the second classification result.

8. A computer-readable medium having a computer program stored thereon, characterized in that, When the program is executed by a processing device, the steps of the method according to any one of claims 1-5 are implemented.

9. An electronic device, characterized in that, Including: A storage device, on which a computer program is stored; A processing device, configured to execute the computer program in the storage device to implement the steps of the method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Text classification method and device, electronic equipment and storage medium

    CN111339305A

  • Text sentiment classification method and device

    CN113761186A