Text multi-classification method and device

By constructing a standard comparison table and calculating conditional probability values ​​and prior probability values, the limitation problems of the prior art in text classification tasks under various restrictions are solved, and higher classification accuracy is achieved.

CN120067332APending Publication Date: 2025-05-30CHINA EVERBRIGHT BANK
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510146427.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Existing deep learning and traditional machine learning methods have limitations in dealing with text classification tasks with multiple constraints, making it difficult to adapt to scenarios with multiple constraints.

Method used

The target category is determined by constructing a standard comparison table based on multiple sets of historical texts, and calculating the conditional probability values ​​and prior probability values ​​between the keyword set and the elements of the historical category.

Benefits of technology

It has achieved the improvement of the accuracy of text classification and can better adapt to scenarios with multiple restrictions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067332A_ABST
    Figure CN120067332A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a text multi-classification method and device, and the method comprises the steps: obtaining a standard comparison table based on a plurality of groups of historical texts; determining a word set based on the to-be-tested text; taking an intersection of the word set and multiple groups of historical semantic description word subsets in the standard comparison table as a keyword set; calculating a first class conditional probability value of the keyword set relative to the first historical class element, a second class conditional probability value of the keyword set relative to the non-first historical class element, a first prior probability value of the first historical class element and a second prior probability value of the non-first historical class element; obtaining at least one first probability value based on the first-class conditional probability value, the second-class conditional probability value, the first prior probability value and the second prior probability value; determining a maximum value in the at least one first probability value; determining the first historical category element corresponding to the maximum value as a target category; and determining that the to-be-tested text corresponding to the keyword set belongs to the target category.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and in particular, to a method and device for text multi-classification. Background Art

[0002] With the rapid development of artificial intelligence technology, text classification technology has been widely applied in the field of natural language processing. Text classification technology mainly adopts deep learning and traditional machine learning methods. The deep learning method constructs a multi-layer neural network model and uses a large amount of labeled data for training to achieve text classification. The traditional machine learning method extracts text features and constructs a classification model to achieve text classification. However, these methods have certain limitations when dealing with text classification tasks with multiple restrictive conditions.

[0003] The principles of deep learning and traditional machine learning methods in text classification are as follows: The deep learning method constructs a multi-layer neural network model and uses a large amount of labeled data for training to achieve text classification. Specifically, the deep learning method first preprocesses the input text data, such as word segmentation, stop word removal, etc., and then converts the text into a vector representation through word embedding technology. Then, the vector representation is input into the neural network model, and through the non-linear transformation of the multi-layer neural network, the extraction and classification of text features are achieved. The traditional machine learning method extracts text features and constructs a classification model to achieve text classification. Specifically, the traditional machine learning method first preprocesses the input text data, such as word segmentation, stop word removal, etc., and then extracts text features, such as word frequency, TF-IDF, etc. Then, the extracted features are input into the classification model, such as support vector machine, random forest, etc., to achieve text classification.

[0004] However, these methods have certain limitations when dealing with text classification tasks with multiple restrictive conditions. The restrictive conditions of deep learning and traditional machine learning methods are relatively fixed and it is difficult to adapt to scenarios with multiple restrictive conditions. Summary of the Invention

[0005] Embodiments of the present invention provide a method and device for text multi-classification, which at least solve the problem that the classification method using related technologies is difficult to adapt to scenarios with multiple restrictive conditions.

[0006] According to an embodiment of the present invention, a text multi-classification method is provided, including: obtaining a standard comparison table based on multiple groups of historical texts, wherein the standard comparison table includes multiple historical category elements and multiple groups of historical semantic descriptor subsets, and a mapping relationship between each historical category element and each group of historical semantic descriptor subsets is set, and each group of historical semantic descriptor subsets includes at least one historical semantic descriptor element; determining a word set based on the text to be tested; taking the intersection of the word set and the multiple groups of historical semantic descriptor subsets in the standard comparison table as a keyword set; calculating a first type of conditional probability value of the keyword set relative to a first historical category element, a second type of conditional probability value of the keyword set relative to non-first historical category elements, a first prior probability value of the first historical category element, and a second prior probability value of the non-first historical category elements, wherein the first historical category element is any one of the multiple historical category elements; obtaining at least one first probability value based on the first type of conditional probability value, the second type of conditional probability value, the first prior probability value, and the second prior probability value; determining the maximum value among the at least one first probability value; determining the first historical category element corresponding to the maximum value as the target category; and determining that the text to be tested corresponding to the keyword set belongs to the target category.

[0007] According to another embodiment of the present application, a text multi-classification system is provided, including: a knowledge base construction module for obtaining a standard comparison table based on multiple groups of historical texts, wherein the standard comparison table includes multiple historical category elements and multiple groups of historical semantic descriptor subsets, and a mapping relationship between each historical category element and each group of historical semantic descriptor subsets is set, and each group of historical semantic descriptor subsets includes at least one historical semantic descriptor element; a classification module for determining a word set based on the text to be tested; and taking the intersection of the word set and the multiple groups of historical semantic descriptor subsets in the standard comparison table as a keyword set; and calculating a first type of conditional probability value of the keyword set relative to a first historical category element, a second type of conditional probability value of the keyword set relative to non-first historical category elements, a first prior probability value of the first historical category element, and a second prior probability value of the non-first historical category elements, wherein the first historical category element is any one of the multiple historical category elements; and obtaining at least one first probability value based on the first type of conditional probability value, the second type of conditional probability value, the first prior probability value, and the second prior probability value; and determining the maximum value among the at least one first probability value; and determining the first historical category element corresponding to the maximum value as the target category; and determining that the text to be tested corresponding to the keyword set belongs to the target category.

[0008] According to another embodiment of the present application, there is also provided a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any one of the above method embodiments when running.

[0009] According to another embodiment of the present application, there is also provided an electronic device including a memory and a processor, where the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.

[0010] According to another embodiment of the present application, there is also provided a computer program product including computer instructions, and the computer instructions implement the steps in any one of the above method embodiments when executed by a processor.

[0011] Through one of the embodiments of the present invention, since the embodiment of the present invention constructs a standard comparison table based on multiple groups of historical texts and determines the target category by calculating the conditional probability value and the prior probability value between the keyword set and the historical category elements, rather than relying on fixed-mode deep learning or traditional machine learning methods, the problem that the classification method using related technologies is difficult to adapt to scenarios with multiple limiting conditions is solved, and thus the effect of improving the accuracy of classification is achieved. Description of the Drawings

[0012] The drawings described herein are used to provide a further understanding of the present invention, form a part of the present application, and the illustrative embodiments and descriptions of the present invention are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:

[0013] Figure 1 is a flowchart of a text multi-classification method according to an embodiment of the present invention;

[0014] Figure 2 is a flowchart of a method for calculating a first type of conditional probability value of a keyword set relative to a first historical category element and a second type of conditional probability value of the keyword set relative to non-first historical category elements according to an embodiment of the present invention;

[0015] Figure 3 is a flowchart of a method for calling a third type of conditional probability value of each keyword element relative to a first historical category element and a fourth type of conditional probability value of each keyword element relative to non-first historical category elements according to an embodiment of the present invention;

[0016] Figure 4 is a schematic structural diagram of a text multi-classification system according to an embodiment of the present application;

[0017] Figure 5It is a schematic structural diagram of a computer terminal for implementing the text multi-classification method of the embodiments of the present invention. Detailed implementation manners

[0018] In the following, the present invention will be described in detail with reference to the accompanying drawings and in combination with embodiments. It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other.

[0019] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. The embodiments of the present invention can be applied to scenarios for identifying and classifying text content. The following takes the scenario of a bank screening and classifying customer complaint content as an example for explanation. Of course, the scenarios of the embodiments of the present invention are not limited to the following scenarios.

[0020] In this embodiment, a text multi-classification method is provided. Figure 1 It is a flowchart of the text multi-classification method according to the embodiments of the present invention. As Figure 1 shown, the process includes:

[0021] Step S101, obtaining a standard comparison table based on multiple groups of historical texts, where the standard comparison table includes a combination of multiple historical category elements and multiple groups of historical semantic description word subsets, and a mapping relationship between each historical category element and each group of historical semantic description word subsets is set, and each group of historical semantic description word subsets includes at least one historical semantic description word element;

[0022] In an exemplary implementation manner, for example, historical text 1 is "The customer called to feedback problems with the mobile banking sunshine value exchange activity. By the end of the year, the available gifts have also decreased, resulting in the inability to exchange for the desired gifts, and the sunshine value has expired and cannot be used." Historical text 2 is "The customer complained that there are quota restrictions on the gifts that can be exchanged for the sunshine value of his mobile banking and cannot be used. Please assist in handling." Historical text 3 is "The customer complained that the current mobile banking application often experiences forced termination, network anomalies, system busyness, network latency, etc., and often experiences forced termination, inability to open, and automatic exit of the mobile banking application. Please assist in handling." Historical text 4 is "The customer complained that he found some overcharged amounts in the receipts of online payment transactions or electronic payment transactions of the mobile banking application, and the logistics of the gifts exchanged at the end of the year was abnormally returned. The individual doubts the authenticity of these records and there is currently a dispute. Please assist in handling." Historical text 5 is "The customer complained that he has doubts about some credits and receipts in the receipts of online payment transactions or electronic payment transactions of the mobile banking application. Please assist in handling."

[0023] Then, a first historical semantic descriptor subset {participate, mobile banking, sunshine value, exchange} can be extracted from historical text 1, and it is set to belong to the first historical text category element "rights and activities category"; a first historical semantic descriptor subset {mobile banking, sunshine value, exchange} can be extracted from historical text 2, and it is set to belong to the first historical text category element "rights and activities category"; a first historical semantic descriptor subset {mobile banking, forced termination, unable to open} can be extracted from historical text 3, and it is set to belong to the first historical text category element "mobile banking & system problems"; a first historical semantic descriptor subset {online payment, electronic payment} can be extracted from historical text 4, and it is set to belong to the first historical text category element "electronic payment & transaction disputes"; a first historical semantic descriptor subset {online payment, electronic payment, credit, arrival} can be extracted from historical text 5, and it is set to belong to the first historical text category element "electronic payment & transaction credit".

[0024] For example, based on the above content, the following initial comparison table (Table 1) is obtained:

[0025] Table 1

[0026]

[0027]

[0028] Since there are the same first historical text category elements (such as "rights and activities category") in the initial comparison table. Therefore, the same first historical text category elements (the two "rights and activities category" in Table 1) can be merged into the same historical category element (one "rights and activities category" in Table 2), and the union of the first historical semantic descriptor subsets corresponding to the same first historical text category elements is taken (for example, from {participate, mobile banking, sunshine value, exchange}, {mobile banking, sunshine value, exchange} to get {participate, mobile banking, sunshine value, exchange}) to obtain a historical semantic descriptor subset. Therefore, a standard comparison table can be obtained, as shown in Table 2.

[0029] Table 2

[0030] Historical category element Subset of historical semantic descriptor words Rights and activities category Participation, mobile banking, sunshine value, redemption Mobile banking & system problems Mobile banking, forced termination, unable to open Electronic payment & transaction disputes Online payment, electronic payment Electronic payment & transaction posting Online payment, electronic payment, posting, arrival of funds

[0031] Step S102, determining a word set based on the text to be measured;

[0032] In an exemplary implementation manner, for example, during a classification process, a text to be measured received is "The customer called to feedback that through mobile banking, they participated in the sunshine value activity, but many of the sunshine values expired, and there was a forced termination situation. Please assist in handling." Then, the text to be measured can be transformed into a word set W' by means of techniques such as word segmentation and stop word deletion:

[0033] {Participation, Mobile Banking, Sunshine Value, Exchange, Processing, Forced End}

[0034] Step S103: Take the intersection of the word set and multiple subsets of historical semantic description words in the standard comparison table as the keyword set;

[0035] In an exemplary implementation, take the intersection of the word set W': {Participation, Mobile Banking, Sunshine Value, Exchange, Processing, Forced End} and all subsets of historical semantic description words in Table 2 to obtain the keyword set W: {Participation, Mobile Banking, Sunshine Value, Exchange, Forced End}.

[0036] Step S104: Calculate the first conditional probability value of the keyword set with respect to the first historical category element, the second conditional probability value of the keyword set with respect to non-first historical category elements, the first prior probability value of the first historical category element, and the second prior probability value of non-first historical category elements, where the first historical category element is any one of multiple historical category elements;

[0037] Step S105: Obtain at least one first probability value based on the first conditional probability value, the second conditional probability value, the first prior probability value, and the second prior probability value;

[0038] In an exemplary implementation, the algorithms corresponding to Step S104 and Step S105 can be converted into Formula 1:

[0039]

[0040] Where W is the keyword set {Participation, Mobile Banking, Sunshine Value, Exchange, Forced End}. C j is the first historical category element and is any one of the historical category elements in Table 2. For example, it can be "Rights and Activities", or "Mobile Banking & System Issues", or "Electronic Payment & Transaction Dispute", or "Electronic Payment & Transaction Posting". P(W|C j ) is the probability of the keyword set W appearing under the condition of the first historical category element C j . P(W|-C j ) is the probability of the keyword set W appearing under the condition of non-first historical category element -C j . P(C j |W) is the probability that the keyword set W is classified as the first historical category element C j .

[0041] Step S106: Determine the maximum value among at least one first probability value;

[0042] Step S107, determine that the first historical category element corresponding to the maximum value is the target category;

[0043] In an exemplary embodiment, the algorithms corresponding to steps S106 and S107 can be converted into Formula 2:

[0044] max{P(C 1 |W), P(C 2 |W), P(C 3 |W),..., P(C m |W)} = P(C x |W) (Formula 2).

[0045] Where m is the total number of historical category elements in Table 2, and x ∈ (1, m).

[0046] Step S108, determine that the text to be tested corresponding to the keyword set belongs to the target category.

[0047] In an exemplary embodiment, for the maximum value P(C x |W) obtained in step S107. For example, if the first historical category element C x is "Rights and Activities", then it is determined that the keyword set W belongs to "Rights and Activities". That is, it is determined that the text to be tested corresponding to the keyword set W ("The customer called to feedback that they participated in the Sunshine Value activity through mobile banking, but many of the Sunshine Values have expired and there is a situation of forced termination. Please assist in handling.") belongs to "Rights and Activities".

[0048] Through the above steps S101 to S108, since the embodiments of the present invention construct a standard comparison table based on multiple groups of historical texts and determine the target category by calculating the conditional probability value and the prior probability value between the keyword set and the historical category element, rather than relying on fixed-mode deep learning or traditional machine learning methods, therefore, the problem that the classification methods of related technologies are difficult to adapt to scenarios with multiple limiting conditions is solved, and the effect of improving the classification accuracy is achieved.

[0049] In one embodiment, Figure 2 is a flowchart of the method for calculating the first type of conditional probability value of the keyword set relative to the first historical category element and the second type of conditional probability value of the keyword set relative to non-first historical category elements according to the embodiments of the present invention. As Figure 2 shown, calculating the first type of conditional probability value of the keyword set relative to the first historical category element and the second type of conditional probability value of the keyword set relative to non-first historical category elements includes:

[0050] Step S201: Invoke the third - type conditional probability values of each keyword element relative to the first - historical - category elements and the fourth - type conditional probability values of each keyword element relative to non - first - historical - category elements.

[0051] In an exemplary implementation, for example, the third - type conditional probability values can be P(w 1 |C j ), P(w 2 |C j ),..., P(w n |C j ), and the fourth - type conditional probability values can be P(w 1 |-C j ), P(w 2 |-C j ),..., P(w n |-C j ). Here, n is the total number of keyword elements in the keyword set W.

[0052] Step S202: Take the product value of multiple third - type conditional probability values as the first - type conditional probability value, and take the product value of multiple fourth - type conditional probability values as the second - type conditional probability value.

[0053] In an exemplary implementation, for example, the first - type conditional probability value can be P(W|C j ), where the calculation formula for P(W|C j ) can be:

[0054] P(W|C j ) = P(w 1 , w 2 ,......, w n |C j ) = P(w 1 |C j ) × P(w 2 |C j ) ×...... × P(w n |C j )

[0055] (Formula 3).

[0056] Similarly, the second - type conditional probability value can be P(W|-C j ), where the calculation formula for P(W|-C j ) can be:

[0057] P(W|-C j ) = P(w 1 , w 2 ,..., w n |-Cj ) = P(w 1 |-C j ) × P(w 2 |-C j ) ×... × P(w n |-C j )

[0058] (Formula 4).

[0059] In one embodiment, Figure 3 is a flowchart of a method for calling the third type of conditional probability value of each keyword element relative to the first historical category element and the fourth type of conditional probability value of each keyword element relative to non-first historical category elements according to an embodiment of the present invention. As Figure 3 shown, calling the third type of conditional probability value of each keyword element relative to the first historical category element and the fourth type of conditional probability value of each keyword element relative to non-first historical category elements includes:

[0060] Step S301, calculating the first likelihood probability value of the first historical semantic descriptor element relative to the second historical category element based on the standard comparison table;

[0061] In an exemplary embodiment, for example, when obtaining the above standard comparison table, P(participation|rights and activities category), P(mobile banking|rights and activities category), P(sunshine value|rights and activities category), P(exchange|rights and activities category), P(mobile banking|mobile banking & system problems), P(forced termination|mobile banking & system problems), P(cannot open|mobile banking & system problems), P(online payment|electronic payment & transaction disputes), P(electronic payment|electronic payment & transaction disputes), P(online payment|electronic payment & transaction posting), P(electronic payment|electronic payment & transaction posting), P(posting|electronic payment & transaction posting), and P(arrival|electronic payment & transaction posting) can be calculated respectively, and the above probability values are used as the first likelihood probability values and stored before classification calculation and directly called during the classification calculation.

[0062] Step S302, calculating the second likelihood probability value of the first historical semantic descriptor element relative to the third historical category element based on the standard comparison table;

[0063] In an exemplary embodiment, for example, when obtaining the above standard comparison table, P(participation|non-rights and activities category), P(mobile banking|non-rights and activities category), P(sunshine value|non-rights and activities category), P(exchange|non-rights and activities category), P(mobile banking|non-mobile banking & system problems), P(forced termination|non-mobile banking & system problems), P(cannot open|non-mobile banking & system problems), P(online payment|non-electronic payment & transaction disputes), P(electronic payment|non-electronic payment & transaction disputes), P(online payment|non-electronic payment & transaction posting), P(electronic payment|non-electronic payment & transaction posting), P(posting|non-electronic payment & transaction posting), and P(arrival|non-electronic payment & transaction posting) can be calculated respectively, and the above probability values are used as the second likelihood probability values and stored before classification, and can be directly called during the classification calculation process.

[0064] Step S303, when the keyword element is the same as the first historical semantic description word element and the first historical category element is the same as the second historical category element, set the first likelihood probability value as the third type of conditional probability;

[0065] In an exemplary embodiment, for example, during the classification process, when the keyword element is "participation", the first historical semantic description word element is "participation", the first historical category element is "rights and activities category", and the second historical category element is "rights and activities category", then assign P(participation|rights and activities category) to the third type of conditional probability.

[0066] Step S305, when the keyword element is the same as the first historical semantic description word element and the non-first historical category element is the same as the third historical category element, set the second likelihood probability value as the fourth type of conditional probability;

[0067] Among them, the multiple historical category elements include the first historical category element, the second historical category element, and the third historical category element, and the second historical category element is different from the third historical category element.

[0068] In an exemplary embodiment, for example, during the classification process, when the keyword element is "participation", the first historical semantic description word element is "participation", the first historical category element is "non-rights and activities category", and the second historical category element is "non-rights and activities category", then assign P(participation|non-rights and activities category) to the fourth type of conditional probability.

[0069] In one embodiment, calling the third type of conditional probability value of each keyword element relative to the first historical category element and the fourth type of conditional probability value of each keyword element relative to the non-first historical category element further includes:

[0070] When the keyword element is the same as the first historical semantic descriptor element and the first historical category element is different from the second historical category element, the preset threshold is set to the third type of conditional probability; wherein, the preset threshold is less than the first likelihood probability value.

[0071] In an exemplary embodiment, for example, for the keyword set W: {participate, mobile banking, sunshine value, exchange, forced termination}, there is a keyword element "forced termination", and there is a first historical semantic descriptor element "forced termination" in the historical semantic descriptor subset in Table 2, that is, the keyword element is the same as the first historical semantic descriptor element. However, for example, in one step of the classification calculation process of the text to be tested ("The customer called to feedback that they participated in the sunshine value activity through mobile banking, but a lot of sunshine values expired, and there was a situation of forced termination. Please assist in handling."), the first historical category element determined is "rights and activities category", and the second historical category element corresponding to the first historical semantic descriptor element "forced termination" in Table 2 is "mobile banking & system problems", that is, the first historical category element is different from the second historical category element. Since the system does not store the first likelihood probability value of the first historical semantic descriptor element "forced termination" relative to the first historical category element "rights and activities category", the third type of conditional probability corresponding to the keyword element "forced termination" can be automatically assigned to the preset threshold, where the preset threshold is less than the first likelihood probability value (that is, the preset threshold is a small value). Thus, the keyword element "forced termination" also participates in the calculation process. Thereby further ensuring the accuracy of the classification result.

[0072] In one embodiment, calculating the first likelihood probability value of the first historical semantic descriptor element relative to the second historical category element based on the standard comparison table includes:

[0073] Taking the ratio of the first information gain and the second information gain as the first likelihood probability value;

[0074] Wherein, the first information gain is the information gain of a first historical semantic descriptor element relative to the first historical category element in the first dataset, the second information gain is the sum value of the information gains of multiple first historical semantic descriptor elements relative to the first historical category element in the first dataset, a first historical semantic descriptor element is any one of the multiple first historical semantic descriptor elements, and the first dataset is obtained based on the standard comparison table.

[0075] In an exemplary embodiment, for example, the calculation formula of the first likelihood probability value can be:

[0076]

[0077] Where k is C jThe total number of historical semantic descriptor elements, i ∈ (1, k). w i is a first historical semantic descriptor element in Table 2. Gain(D, w i , C j ) is the first information gain.

[0078] is the second information gain.

[0079] In an exemplary embodiment, the first data set may be as shown in Table 3:

[0080] Table 3

[0081]

[0082]

[0083] In one embodiment, the first information gain is obtained based on the difference between the first information entropy and the second information entropy; wherein, the calculation formula of the first information entropy is:

[0084]

[0085] wherein, D is the first data set, and m is the total number of multiple historical category elements in the first data set;

[0086] The second information entropy is obtained based on the sum of the third information entropy and the fourth information entropy; wherein, the third information entropy is based on the information entropy of the second data set and the first weight, and the fourth information entropy is based on the information entropy of the third data set and the second weight. The second data set is a data set composed of at least one historical semantic descriptor subset that does not include the first historical semantic descriptor element (such as w i ) in the first data set. The third data set is a data set composed of at least one historical semantic descriptor subset that includes the first historical semantic descriptor element (such as w i ) in the first data set. The first weight is the ratio of the total number of category elements in the second data set to the total number of category elements in the first data set, and the second weight is the ratio of the total number of elements in the third data set to the total number of elements in the first data set. In an exemplary embodiment, for example, the calculation formula of the first information gain is:

[0087] The calculation formula of the second information entropy is:

[0088] wherein, D 0 is the second data set, D 1 is the third data set, Dv It includes a second data set and a third data set.

[0089] Based on Table 3, the following will be explained with examples:

[0090] Example 1, the first data set is: In one calculation, “participate” is used as the first historical semantic descriptor element w i , “rights and activities category” is used as the first historical category element C j , “mobile banking & system problems, electronic payment & transaction disputes, electronic payment & transaction posting” is used as the non-first historical category element -C j , to calculate P(participate|rights and activities category).

[0091] Then, the third data set is: D 1 =(1111000000), (for example, corresponding to the historical semantic descriptor subset: {participate, mobile banking, sunshine value, redemption}).

[0092] The second data set is: (for example, corresponding to the historical semantic descriptor subsets: {mobile banking, forced termination, unable to open}, {online payment, electronic payment}, {online payment, electronic payment, posting, arrival of funds}).

[0093] Then,

[0094]

[0095] Therefore,

[0096]

[0097] For example,

[0098] Gain(D, participate, rights and activities category)=0.8,

[0099]

[0100] Gain(D, mobile banking, rights and activities category)=0.6,

[0101]

[0102] Gain(D, sunshine value, rights and activities category)=0.6,

[0103]

[0104] Gain(D, redemption, rights and activities category)=0.2,

[0105]

[0106] Example 2, the first data set is: In one calculation, "online payment" is used as the first historical semantic descriptor element w i . "Electronic payment & transaction objection" is used as the first historical category element C j , and "rights and activities, mobile banking & system problems, electronic payment & transaction posting" are used as non-first historical category elements -C j , to calculate P(online payment|electronic payment & transaction objection). Similarly, in another calculation, "online payment" is used as the first historical semantic descriptor element w i . "Electronic payment & transaction posting" is used as the first historical category element C j , and "rights and activities, mobile banking & system problems, electronic payment & transaction objection" are used as non-first historical category elements -C j , to calculate P(online payment|electronic payment & transaction posting).

[0107] Then,

[0108] The third data set is: (For example, corresponding historical semantic descriptor subsets: {online payment, electronic payment}, {online payment, electronic payment, posting, arrival}), the second data set is (For example, corresponding historical semantic descriptor subsets: {participate, mobile banking, sunshine value, exchange}, {mobile banking, forced termination, unable to open}).

[0109] Then,

[0110]

[0111] Therefore,

[0112]

[0113] Similarly, the calculation formula for the second likelihood probability value can be:

[0114]

[0115] where k is the total number of historical semantic descriptor elements in -C j , i ∈ (1, k). w i is a first historical semantic descriptor element in Table 2. Gain(D, w i , -C j ) corresponds to the first information gain. Corresponds to the second information gain.

[0116]

[0117] Among them, it should be noted that when calculating P(w i |-C j ), the difference from calculating P(w i |C j ) is that: when calculating P(w i |-C j ), the D 0 involved is the category including w i in the data set D, and the D 1 involved is the category not including w i in the data set D. The calculation process is similar and will not be elaborated in the embodiments of the present invention.

[0118] In an exemplary implementation manner, for example, the text to be measured: "The customer called to feedback that by participating in the sunshine value activity through mobile banking, but a lot of sunshine values have expired and there is a situation of forced termination. Please assist in handling."

[0119] For example, the probability of calculating that the text to be measured is of the rights and interests and activity category is as follows:

[0120]

[0121] Through the description of the above implementation manners, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of adding the necessary general hardware platform through software. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation manner. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc), and includes several instructions to enable a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in the various embodiments of the present application.

[0122] In this embodiment, a text multi-classification system is further provided. This system is used to implement the above embodiments and preferred implementation manners, and those that have been described will not be elaborated again. As the term "module" used below can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented by software, the implementation of hardware, or a combination of software and hardware, is also possible and contemplated.

[0123] Figure 4 is a schematic structural diagram of the text multi-classification system according to the embodiments of the present application. As Figure 4 shown, the system includes:

[0124] A knowledge base construction module 41, configured to obtain a standard comparison table based on multiple groups of historical texts, where the standard comparison table includes a combination of multiple historical category elements and multiple groups of historical semantic descriptor subsets, and a mapping relationship between each historical category element and each group of historical semantic descriptor subsets is set, and each group of historical semantic descriptor subsets includes at least one historical semantic descriptor element;

[0125] A classification module 42, configured to determine a word set based on the text to be tested;

[0126] And, taking the intersection of the word set and multiple groups of historical semantic descriptor subsets in the standard comparison table as the keyword set;

[0127] And, calculating a first type of conditional probability value of the keyword set with respect to the first historical category element, a second type of conditional probability value of the keyword set with respect to non-first historical category elements, a first prior probability value of the first historical category element, and a second prior probability value of non-first historical category elements, where the first historical category element is any one of the multiple historical category elements;

[0128] And, obtaining at least one first probability value based on the first type of conditional probability value, the second type of conditional probability value, the first prior probability value, and the second prior probability value;

[0129] And, determining the maximum value among the at least one first probability value;

[0130] And, determining the first historical category element corresponding to the maximum value as the target category;

[0131] And, determining that the text to be tested corresponding to the keyword set belongs to the target category.

[0132] In one implementation, the system is further configured to: call a third type of conditional probability value of each keyword element with respect to the first historical category element and a fourth type of conditional probability value of each keyword element with respect to non-first historical category elements;

[0133] Taking the product value of multiple third type of conditional probability values as the first type of conditional probability value, and taking the product value of multiple fourth type of conditional probability values as the second type of conditional probability value.

[0134] In one implementation, the system is further configured to: calculate a first likelihood probability value of the first historical semantic descriptor element with respect to the second historical category element based on the standard comparison table;

[0135] Calculate a second likelihood probability value of the first historical semantic descriptor element with respect to the third historical category element based on the standard comparison table;

[0136] When the keyword element is the same as the first historical semantic descriptor element and the first historical category element is the same as the second historical category element, set the first likelihood probability value to the third type of conditional probability;

[0137] When the keyword element is the same as the first historical semantic descriptor element and the non-first historical category element is the same as the third historical category element, set the second likelihood probability value to the fourth type of conditional probability;

[0138] Among them, the multiple historical category elements include the first historical category element, the second historical category element, and the third historical category element, and the second historical category element is different from the third historical category element.

[0139] In one implementation, the system is also used for: when the keyword element is the same as the first historical semantic descriptor element and the first historical category element is different from the second historical category element, set the preset threshold to the third type of conditional probability; where the preset threshold is less than the first likelihood probability value.

[0140] In one implementation, the system is also used for: taking the ratio of the first information gain and the second information gain as the first likelihood probability value;

[0141] Among them, the first information gain is the information gain of a first historical semantic descriptor element relative to the first historical category element in the first data set, the second information gain is the sum value of the information gains of multiple first historical semantic descriptor elements relative to the first historical category element in the first data set, a first historical semantic descriptor element is any one of the multiple first historical semantic descriptor elements, and the first data set is obtained based on the standard comparison table.

[0142] In one implementation, in the system, the first information gain is obtained based on the difference between the first information entropy and the second information entropy; where the calculation formula of the first information entropy is:

[0143]

[0144] Among them, D is the first data set, and m is the total number of multiple historical category elements in the first data set;

[0145] The second information entropy is obtained based on the sum value of the third information entropy and the fourth information entropy; where the third information entropy is based on the information entropy of the second data set and the first weight, the fourth information entropy is based on the information entropy of the third data set and the second weight, the second data set is a data set composed of at least one historical semantic descriptor subset that does not include the first historical semantic descriptor element (such as w i ) in the first data set, and the third data set is a data set that includes the first historical semantic descriptor element (such as w iA data set composed of at least one historical semantic descriptor subset of (), where the first weight is the ratio of the total number of category elements in the second data set to the total number of category elements in the first data set, and the second weight is the ratio of the total number of category elements in the third data set to the total number of category elements in the first data set.

[0146] It should be noted that the above-mentioned various modules can be implemented by software or hardware. For the latter, it can be implemented in the following ways, but not limited to: all the above-mentioned modules are located in the same processor; or, the above-mentioned various modules are located in different processors in any combination form.

[0147] The method embodiments provided in the embodiments of the present invention can be executed on a mobile terminal, a computer terminal, or a similar computing device. Taking running on a computer terminal as an example, Figure 5 is a schematic structural diagram of a computer terminal for executing the text multi-classification method of the embodiments of the present invention, as Figure 5 shown, the computer terminal may include one or more ( Figure 5 only one is shown in the figure) processors 102 (the processor 102 may include, but is not limited to, processing devices such as a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data. Optionally, the above computer terminal may further include a transmission device 106 for communication functions and an input / output device 108. Those of ordinary skill in the art can understand that Figure 5 the structure shown is only schematic and does not limit the structure of the above computer terminal. For example, the computer terminal may further include more or fewer components than Figure 5 shown in the figure, or have a different configuration from Figure 5 shown in the figure.

[0148] The memory 104 can be used to store computer programs. For example, software programs and modules of application software, such as the computer program corresponding to the text multi-classification method in the embodiments of the present invention. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implements the above method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely set relative to the processor 102, and these remote memories can be connected to the computer terminal through a network. Examples of the above network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and their combinations.

[0149] The transmission device 106 is used to receive or send data via a network. Specific examples of the above-mentioned network may include a wireless network provided by a communication provider of a computer terminal. In one example, the transmission device 106 includes a network adapter (Network Interface Controller, abbreviated as NIC), which can be connected to other network devices through a gateway so as to communicate with the Internet. In one example, the transmission device 106 can be a Radio Frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0150] An embodiment of the present application also provides a computer-readable storage medium, in which a computer program is stored. Wherein, the computer program is configured to execute the steps in any one of the above method embodiments when running.

[0151] In an exemplary embodiment, the above computer-readable storage medium may include but is not limited to: USB flash drive, Read-Only Memory (ROM), Random Access Memory (RAM), mobile hard disk, magnetic disk or optical disc and other various media that can store computer programs.

[0152] An embodiment of the present application also provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.

[0153] In an exemplary embodiment, the above electronic device may further include a transmission device and input / output devices. Wherein, the transmission device is connected to the above processor, and the input / output device is connected to the above processor.

[0154] An embodiment of the present application also provides a computer program product, including a computer program, which implements the steps in any one of the above method embodiments when executed by a processor.

[0155] Specific examples in this embodiment may refer to the examples described in the above embodiments and exemplary embodiments, and will not be repeated here.

[0156] Obviously, those skilled in the art should understand that the various modules or steps of the present application described above can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. They can be implemented by program codes executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a sequence different from that here, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module for implementation. In this way, the present application is not limited to any specific combination of hardware and software.

[0157] The foregoing is only a preferred embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the principle of the present application shall be included within the protection scope of the present application.

Claims

1. A text multi-classification method, characterized in that: include: A standard comparison table is obtained based on multiple groups of historical texts, wherein the standard comparison table includes multiple historical category elements and multiple groups of historical semantic description word subsets and is provided with a mapping relationship between each of the historical category elements and each group of the historical semantic description word subsets, and each group of the historical semantic description word subsets includes at least one historical semantic description word element; Determining a word set based on the text to be tested; The intersection of the word set and the plurality of historical semantic description word subsets in the standard comparison table is used as a keyword set; Calculate a first conditional probability value of the keyword set relative to a first historical category element, a second conditional probability value of the keyword set relative to a non-first historical category element, a first priori probability value of the first historical category element, and a second priori probability value of the non-first historical category element, wherein the first historical category element is any one of the multiple historical category elements; Obtain at least one first probability value based on the first type of conditional probability value, the second type of conditional probability value, the first priori probability value, and the second priori probability value; determining a maximum value among the at least one first probability value; Determine the first historical category element corresponding to the maximum value as the target category; It is determined that the text to be tested corresponding to the keyword set belongs to the target category.

2. The method according to claim 1, characterized in that: Calculating the first conditional probability value of the keyword set relative to the first historical category element and the second conditional probability value of the keyword set relative to the non-first historical category element, including: Calling the third conditional probability value of each keyword element relative to the first historical category element and the fourth conditional probability value of each keyword element relative to the non-first historical category element; The product value of the plurality of third-category conditional probability values ​​is used as the first-category conditional probability value, and the product value of the plurality of fourth-category conditional probability values ​​is used as the second-category conditional probability value.

3. The method according to claim 2, characterized in that Calling the third conditional probability value of each keyword element relative to the first historical category element and the fourth conditional probability value of each keyword element relative to the non-first historical category element, comprises: Calculate a first likelihood probability value of the first historical semantic description word element relative to the second historical category element based on the standard comparison table; Calculating a second likelihood probability value of the first historical semantic description word element relative to the third historical category element based on the standard comparison table; When the keyword element is the same as the first historical semantic description word element, and the first historical category element is the same as the second historical category element, setting the first likelihood probability value as the third type conditional probability; When the keyword element is the same as the first historical semantic description word element, and the non-first historical category element is the same as the third historical category element, setting the second likelihood probability value as the fourth type conditional probability; The multiple history category elements include the first history category element, the second history category element, and the third history category element, and the second history category element is different from the third history category element.

4. The method according to claim 2, characterized in that: Calling the third conditional probability value of each keyword element relative to the first historical category element and the fourth conditional probability value of each keyword element relative to the non-first historical category element also includes: When the keyword element is the same as the first historical semantic description word element, and the first historical category element is different from the second historical category element, the preset threshold is set to the third type conditional probability; wherein the preset threshold is less than the first likelihood probability value.

5. The method according to claim 3, characterized in that: Calculating a first likelihood probability value of the first historical semantic description word element relative to the second historical category element based on the standard comparison table includes: Taking the ratio of the first information gain to the second information gain as the first likelihood probability value; Among them, the first information gain is the information gain of a first historical semantic description word element relative to the first historical category element in the first data set, and the second information gain is the sum of the information gains of multiple first historical semantic description word elements relative to the first historical category element in the first data set, the one first historical semantic description word element is any one of the multiple first historical semantic description word elements, and the first data set is obtained based on the standard comparison table.

6. The method according to claim 5, characterized in that The first information gain is obtained based on the difference between the first information entropy and the second information entropy; wherein the calculation formula of the first information entropy is: Wherein, D is the first data set, and m is the total number of the plurality of historical category elements in the first data set; The second information entropy is obtained based on the sum of the third information entropy and the fourth information entropy; wherein the third information entropy is obtained based on the information entropy of the second data set and the first weight, and the fourth information entropy is obtained based on the information entropy and the second weight of the third data set, the second data set is a data set composed of at least one subset of the historical semantic description words that does not include the first historical semantic description word element in the first data set, and the third data set is a data set composed of at least one subset of the historical semantic description words that includes the first historical semantic description word element in the first data set, the first weight is the ratio of the total number of category elements in the second data set to the total number of category elements in the first data set, and the second weight is the ratio of the total number of category elements in the third data set to the total number of category elements in the first data set.

7. A text multi-classification system, characterized in that: include: A knowledge base construction module, used to obtain a standard comparison table based on multiple groups of historical texts, wherein the standard comparison table includes multiple historical category elements and multiple groups of historical semantic description word subsets and is provided with a mapping relationship between each of the historical category elements and each group of the historical semantic description word subsets, and each group of the historical semantic description word subsets includes at least one historical semantic description word element; A classification module, used to determine a word set based on the text to be tested; and, taking the intersection of the word set and the plurality of historical semantic description word subsets in the standard comparison table as a keyword set; and calculating a first conditional probability value of the keyword set relative to a first historical category element, a second conditional probability value of the keyword set relative to a non-first historical category element, a first priori probability value of the first historical category element, and a second priori probability value of the non-first historical category element, wherein the first historical category element is any one of the multiple historical category elements; And, obtaining at least one first probability value based on the first type of conditional probability value, the second type of conditional probability value, the first priori probability value, and the second priori probability value; and, determining a maximum value among the at least one first probability value; and determining the first historical category element corresponding to the maximum value as a target category; And, determining whether the text to be tested corresponding to the keyword set belongs to the target category.

8. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, wherein the computer program executes the method described in any one of claims 1 to 6 when executed by a processor.

9. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to run the computer program to perform the method according to any one of claims 1 to 6.

10. A computer program product comprising computer instructions, characterized in that When the computer instructions are executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.