Text classification method and device, electronic equipment and storage medium
By extracting and matching the subject and address data in the complaint work ticket text, combined with the preset complaint work ticket text library, the problem of low classification efficiency in the existing technology is solved, and more efficient complaint handling is achieved.
Patent Information
- Application Number
- CN202411983752.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-27
- Publication Date
- 2025-05-09
AI Technical Summary
The existing technology has low efficiency in classification of complaint work order texts, which has led to huge governance pressure and service bottlenecks in handling citizen complaints.
By obtaining the complaint work ticket text set, the subject data and address data of each complaint work ticket text are extracted, and the preset complaint work ticket text library is used for matching processing, and the types of complaints are determined, including multiple complaints about the same incident and multiple complaints about the same incident by the same person.
The accurate classification of the complaint work order text has been achieved, the efficiency of urban management departments in handling complaints has been improved, and the dependence on the personal experience and attitudes of staff is reduced.
Smart Images

Figure CN119961451A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of smart city management, and in particular to a text classification method, device, electronic equipment and storage medium. Background Art
[0002] In the process of focusing on the construction of "civilized cities", many cities are facing the challenge of citizens' increasing requirements for living environment. Citizens not only hope to enjoy a safe and convenient living environment, but also have higher expectations for the efficiency of urban management and service quality. This dual pressure has caused departments such as urban management and problem guidance to bear huge governance pressure and service bottlenecks in the operation of the existing business system.
[0003] By establishing a public opinion appeal platform, city managers can directly reach public opinion, but the number of complaints received by the public opinion appeal platform has increased sharply. These complaints are not only huge in number, but also diverse in type, covering all aspects of citizens' daily lives. The current working model is highly dependent on the personal experience and work attitude of the staff, which makes the work content cumbersome and monotonous. The division of labor between departments is relatively complex, which leads to low classification efficiency of complaint work order texts. Therefore, how to provide a text classification method that can improve the classification efficiency of complaint work order texts has become an urgent problem to be solved. Summary of the invention
[0004] The embodiment of the present invention provides a text classification method, which aims to solve the problem of low classification efficiency of complaint work order texts by existing text classification methods. When a complaint work order text set is obtained, the main body data and address data and other core information of each complaint work order text can be accurately determined based on the complaint work order text set, and then the complaint type of each complaint work order text can be accurately determined based on the main body data and core data. There is no need to rely on the personal experience and work attitude of the staff, thereby improving the classification efficiency of the complaint work order text.
[0005] In a first aspect, an embodiment of the present invention provides a text classification method, the method comprising the following steps:
[0006] Obtaining a complaint work order text set, wherein the complaint work order text set includes at least one complaint work order text;
[0007] Based on the complaint work order text set, determining the main body data and address data of each complaint work order text;
[0008] Based on the subject data and the address data, the complaint type of each complaint ticket text in the complaint ticket text set is determined, and the complaint type includes a first type and a second type. The first type is that multiple people complain about the same event, and the second type is that the same person complains about the same event multiple times.
[0009] Optionally, determining the complaint type of each complaint work order text in the complaint work order text set based on the subject data and the address data includes:
[0010] Obtain a preset complaint work order text library, wherein the preset complaint work order text library includes a first type of clustering clusters, each of the first type of clustering clusters includes multiple historical complaint work order texts complaining about the same event, and each of the first type of clustering clusters corresponds to at least one representative complaint work order text, and the representative complaint work order text is one of the multiple historical complaint work order texts;
[0011] Performing a first matching process based on each of the complaint work order texts in the complaint work order text set and each of the representative complaint work order texts in the preset complaint work order text library to obtain a first matching result;
[0012] Based on the first matching result, the main data of the complaint work order text, and the address data, determine whether the complaint type of the complaint work order text includes the first type.
[0013] Optionally, the determining whether the complaint type of the complaint work order text includes the first type based on the first matching result, the subject data of the complaint work order text, and the address data comprises:
[0014] If the first matching result is that the complaint work order text matches the representative complaint work order text successfully, then the complaint type of the complaint work order text includes the first type;
[0015] If the first matching result is that the complaint work order text fails to match all the representative complaint work order texts, the complaint work order text is used as the complaint work order text to be clustered;
[0016] Clustering the complaint work order texts to be clustered based on the subject data and the address data corresponding to all the complaint work order texts to be clustered to obtain a target complaint work order text clustering cluster;
[0017] Based on the target complaint ticket text clustering cluster, determine whether the complaint type of the complaint ticket text to be clustered includes the first type.
[0018] Optionally, clustering the complaint work order texts to be clustered based on the subject data and the address data corresponding to all the complaint work order texts to be clustered to obtain a target complaint work order text clustering cluster includes:
[0019] Based on the subject data, performing a first grouping process on the complaint work order text to be clustered to obtain at least one subject group;
[0020] Based on the address data, performing a second grouping process on the subject group to obtain at least one target group;
[0021] Clustering is performed based on each of the target groups to obtain a target complaint work order text cluster corresponding to each of the target groups.
[0022] Optionally, determining the complaint type of each complaint work order text in the complaint work order text set based on the subject data and the address data includes:
[0023] Obtain a preset complaint work order text library, wherein the preset complaint work order text library includes multiple historical complaint work order texts, and the historical complaint work order texts correspond to subject data and address data;
[0024] Performing a second matching process based on the complaint work order text and the historical complaint work order text in the preset complaint work order text library to obtain a second matching result;
[0025] Based on the second matching result, the main data of the complaint work order text, and the address data, determine whether the complaint type of each complaint work order text includes the second type.
[0026] Optionally, determining whether the complaint type of each complaint work order text includes the second type based on the second matching result, the subject data of the complaint work order text, and the address data comprises:
[0027] If the second matching result is a match failure, the complaint type in the complaint ticket text does not include the second type;
[0028] If the second matching result is a successful match, based on the main data and address data of the complaint ticket text and the main data and address data of the target historical complaint ticket text, it is determined whether the complaint type of the complaint ticket text includes the second type, and the target historical complaint ticket text is the historical complaint ticket text corresponding to the second matching result.
[0029] Optionally, the determining whether the complaint type of the complaint work order text includes the second type based on the subject data and the address data of the complaint work order text and the subject data and the address data of a target historical complaint work order text, wherein the target historical complaint work order text is a historical complaint work order text corresponding to the second matching result, includes:
[0030] Performing subject matching processing based on the subject data of the complaint work order text and the subject data of the target historical complaint work order text to obtain a subject matching result of the complaint work order text;
[0031] Performing address matching processing based on the address data of the complaint work order text and the address data of the target historical complaint work order text to obtain an address matching result of the complaint work order text;
[0032] Based on the similarity matching between the complaint work order text and the target historical complaint work order text, a similarity matching result of the complaint work order text is obtained;
[0033] Based on the subject matching result, the address matching result and the similarity matching result, determine whether the complaint type of the complaint work order text includes the second type.
[0034] In a second aspect, an embodiment of the present invention further provides a text classification device, the text classification device comprising:
[0035] A first acquisition module is used to acquire a complaint work order text set, wherein the complaint work order text set includes at least one complaint work order text;
[0036] A first determination module, used to determine the main body data and address data of each complaint work order text based on the complaint work order text set;
[0037] The second determination module is used to determine the complaint type of each complaint work order text in the complaint work order text set based on the subject data and the address data, and the complaint type includes a first type and a second type. The first type is that multiple people complain about the same event, and the second type is that the same person complains about the same event multiple times.
[0038] In a third aspect, an embodiment of the present invention provides an electronic device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps in the text classification method provided in the embodiment of the present invention when executing the computer program.
[0039] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps in the text classification method provided in the embodiment of the invention are implemented.
[0040] In an embodiment of the present invention, a complaint work order text set is obtained, wherein the complaint work order text set includes at least one complaint work order text; based on the complaint work order text set, the main body data and address data of each of the complaint work order texts are determined; based on the main body data and the address data, the complaint type of each of the complaint work order texts in the complaint work order text set is determined, wherein the complaint type includes a first type and a second type, wherein the first type is that multiple people complain about the same event, and the second type is that the same person complains about the same event multiple times. When the complaint work order text set is obtained, the core information such as the main body data and address data of each complaint work order text can be accurately determined based on the complaint work order text set, and then the complaint type of each complaint work order text can be accurately determined based on the main body data and the core data. There is no need to rely on the personal experience and work attitude of the staff, thereby improving the classification efficiency of the complaint work order texts. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0042] Figure 1 is a flowchart of a text classification method provided by an embodiment of the present invention;
[0043] Figure 2 is a flow chart of a data preprocessing method provided by an embodiment of the present invention;
[0044] Figure 3 is a flowchart of a method for determining a first type of text provided by an embodiment of the present invention;
[0045] Figure 4 is a flowchart of a method for determining a second type of text provided by an embodiment of the present invention;
[0046] Figure 5 is a structural schematic diagram of a text classification device provided in an embodiment of the present invention;
[0047] Figure 6 It is a structural schematic diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0048] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0049] like Figure 1 As shown, Figure 1 : is a flowchart of a text classification method provided by an embodiment of the present invention, including:
[0050] 101. Get the complaint ticket text set.
[0051] In an embodiment of the present invention, the above-mentioned text classification method can be applied to a city management platform, and the above-mentioned city management platform can be constructed by a server or a server cluster. The above-mentioned server or server cluster can be any electronic device with functions such as text processing, text analysis, text classification, data storage, and data transmission.
[0052] The above-mentioned complaint work order text set includes at least one complaint work order text, which can be a complaint text provided by citizens regarding any management or problem in the city, and a corresponding complaint work order text can be constructed based on the above-mentioned complaint text, and the above-mentioned complaint text can be stored as a complaint work order text in the above-mentioned complaint work order text set.
[0053] The above-mentioned city management platform can directly obtain the complaint work order text uploaded by the citizens through the above-mentioned data transmission function, and collect the obtained complaint work order text into the above-mentioned complaint work order text set. Alternatively, the above-mentioned city management platform may not directly connect with the citizens, and the users may upload it to the transfer platform. The transfer platform can regularly package the complaint work order text within a certain period of time into a complaint work order text set and transmit it to the above-mentioned city management platform.
[0054] It should be noted that the above-mentioned citizens are only for illustrative purposes. Correspondingly, they may also refer to villagers, citizens, and other people within any administrative area.
[0055] 102. Based on the complaint ticket text set, determine the main body data and address data of each complaint ticket text.
[0056] In the embodiment of the present invention, the subject data may include individuals, organizations, etc., and the address data may include place names, regions, etc. The subject data and address data may be extracted from the complaint ticket text by NER entity recognition or large language model.
[0057] Specifically, when conducting text analysis, it is possible to clearly identify the entity categories that need to be identified, including subjects (such as individuals, organizations), addresses (such as place names, regions), times (such as dates, time periods), etc., so as to use NER technology for annotation, automatically identify and classify key information in the complaint ticket text, and then extract the subject and address data.
[0058] In a possible embodiment, the subject data and address data can also be obtained by Figure 2 The flowchart of a data preprocessing method shown further illustrates that it includes the following steps:
[0059] The first step is to obtain sample data, for example, data on people's livelihood demands can be collected through various social media, customer feedback, questionnaires and other sources;
[0060] The second step is to label the key information in the sample data (such as entities, emotions, topics, etc.). After labeling, the sorted sample data and the corresponding labels are input into the big model for pre-training. After the training is completed, the big model is used to recognize and process the complaint ticket text to output features, generate summaries, etc.
[0061] The third step is to build a NER model using sample data, and then provide the complaint ticket text set to the NER model to determine the main data and address data of each complaint ticket text.
[0062] It should be noted that the summary and features obtained through the above-mentioned large model (i.e., large language model) can be used as the cover after the complaint type is determined through the subject data and address data, and the cover and the corresponding complaint ticket text can be stored in the corresponding database.
[0063] 103. Based on the subject data and address data, determine the complaint type of each complaint ticket text in the complaint ticket text set.
[0064] In the embodiment of the present invention, the above complaint types include a first type and a second type. The first type is that multiple people complain about the same event, and the above second type is that the same person complains about the same event multiple times.
[0065] The first type mentioned above can be specifically understood as a group litigation incident, that is, a complaint incident jointly initiated by multiple individuals or groups due to similar demands or problems. This situation usually occurs in the fields of public services, environmental protection, consumer rights, etc. The emergence of group litigation incidents reflects the collective concern of the group on a certain issue, and often requires coordination and handling by relevant departments to ensure timely response to public demands.
[0066] The second type mentioned above can be specifically understood as repeated work orders, that is, multiple work orders submitted by the same user or multiple users for the same or similar issues in livelihood complaints. This may involve areas such as substandard services, facility failures, and environmental issues.
[0067] Specifically, a preset complaint ticket text library can be obtained, and historical complaint ticket texts can be stored in the complaint ticket text library. When the historical complaint ticket texts are stored in the complaint ticket text library, corresponding summaries and features can also be generated according to the large model, and the summaries and features can be used as covers for storage.
[0068] The main body data and address data of each complaint work order text in the above complaint work order text set are compared with the main body data and address data of each historical complaint work order text in the above preset complaint work order text library to obtain a first comparison result. At the same time, when there are multiple complaint work order texts in the complaint work order text set, the main body data and address data between the multiple complaint work order texts can also be compared to obtain a second comparison result. Thus, the complaint type of the complaint work order text is determined based on the above first comparison result and the second comparison result.
[0069] It can be understood that after comprehensively considering the second comparison results between complaint ticket texts within the same time period and the first comparison results between complaint ticket texts within different time periods, the accuracy of judging the complaint type of the complaint ticket text can be further improved.
[0070] In an embodiment of the present invention, a complaint work order text set is obtained, wherein the complaint work order text set includes at least one complaint work order text; based on the complaint work order text set, the main body data and address data of each of the complaint work order texts are determined; based on the main body data and the address data, the complaint type of each of the complaint work order texts in the complaint work order text set is determined, wherein the complaint type includes a first type and a second type, wherein the first type is that multiple people complain about the same event, and the second type is that the same person complains about the same event multiple times. When the complaint work order text set is obtained, the core information such as the main body data and address data of each complaint work order text can be accurately determined based on the complaint work order text set, and then the complaint type of each complaint work order text can be accurately determined based on the main body data and the core data. There is no need to rely on the personal experience and work attitude of the staff, thereby improving the classification efficiency of the complaint work order texts.
[0071] It can be understood that in the specific implementation of the present application, it involves complaint ticket text sets, subject data, address data and other related data. When the embodiments in the present application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of relevant data and the construction and use of the complaint ticket text library need to comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0072] It should be noted that the text classification method provided in the embodiment of the present invention can be applied to devices such as computers and servers that can perform text classification.
[0073] Optionally, in the step of determining the complaint type of each complaint ticket text in the complaint ticket text set based on the subject data and the address data, a preset complaint ticket text library can also be obtained; each complaint ticket text in the complaint ticket text set is matched with each representative complaint ticket text in the preset complaint ticket text library to obtain a first matching result; based on the first matching result, the subject data and the address data of the complaint ticket text, it is determined whether the complaint type of the complaint ticket text includes the first type.
[0074] In an embodiment of the present invention, the preset complaint work order text library includes a first type of clustering clusters, each of the first type of clustering clusters includes multiple historical complaint work order texts for complaints about the same event, and each of the first type of clustering clusters corresponds to at least one representative complaint work order text, and the representative complaint work order text is one of the multiple historical complaint work order texts. The representative complaint work order text can specifically be the central node in the corresponding first type of clustering cluster, and each complaint work order text corresponds to a complaint made by a citizen. Therefore, multiple historical complaint work order texts for complaints about the same event are the first type of complaints made by multiple people about the same event.
[0075] Specifically, in the above-mentioned preset complaint ticket text library, each historical complaint ticket text corresponds to a cover (i.e., the summary and features obtained through large model recognition), and the above-mentioned first matching process can be to match each complaint ticket text with each cover representing the complaint ticket text (If the first matching result is a successful match, it means that the complaint ticket text is also a complaint about the same event corresponding to the first type clustering cluster, so the complaint ticket text can be added to the corresponding first type clustering cluster. It can also be determined that the complaint type of the complaint ticket text includes the first type. Conversely, if the matching fails, the complaint ticket text can be used as the complaint ticket text to be clustered, and clustering processing can be performed based on the main data and address data between the complaint ticket texts to be clustered to determine whether there are multiple complaint ticket texts complaining about the same event between the complaint ticket texts in the complaint ticket text set, and then determine whether the complaint ticket text to be clustered has a complaint type including the first type of complaint ticket text.
[0076] Optionally, in the step of determining whether the complaint type of the complaint work order text includes the first type based on the first matching result, the main data of the complaint work order text and the address data, if the first matching result is that the complaint work order text successfully matches the representative complaint work order text, then the complaint type of the complaint work order text includes the first type; if the first matching result is that the complaint work order text fails to match all representative complaint work order texts, then the complaint work order text is used as the complaint work order text to be clustered; the complaint work order texts to be clustered are clustered based on the main data and address data corresponding to all complaint work order texts to be clustered to obtain a target complaint work order text cluster cluster; based on the target complaint work order text cluster cluster, it is determined whether the complaint type of the complaint work order text to be clustered includes the first type.
[0077] In an embodiment of the present invention, if the first matching result is that the complaint work order text successfully matches the representative complaint work order text, then the complaint type of the complaint work order text includes the first type, and the above-mentioned complaint work order text can be archived into the first type cluster corresponding to the above-mentioned representative complaint work order text to achieve the update of the above-mentioned first type cluster. That is, the above-mentioned first type cluster is re-clustered, and the TPO N complaint work order texts at the center of the cluster are determined as new representative complaint work order texts, thereby improving the classification accuracy of subsequent text classification processing.
[0078] If the first matching result is that the complaint work order text fails to match all representative complaint work order texts, the complaint work order text is used as the complaint work order text to be clustered; the complaint work order text to be clustered is clustered based on the subject data and address data corresponding to all complaint work order texts to be clustered to obtain the target complaint work order text clustering cluster; based on the target complaint work order text clustering cluster, determine whether the complaint type of the complaint work order text to be clustered includes the first type. In order to determine whether there are multiple complaint work order texts complaining about the same event among the complaint work order texts in the complaint work order text set, and then determine whether there are complaint work order texts to be clustered whose complaint types include the first type of complaint work order text.
[0079] The clustering process may be to use the subject data and the address data as clustering targets, thereby determining a target complaint work order text clustering cluster with the same subject and the same address, and with a high degree of similarity between the texts. It may then be determined that the complaint types of the complaint work order texts to be clustered in the target complaint work order text clustering cluster include the first type, while the complaint types of the complaint work order texts that have not been successfully clustered, i.e., the separate complaint work order texts to be clustered, do not include the first type.
[0080] In a possible embodiment, whether the complaint type of the complaint work order text includes the first text type can be determined by Figure 3 The flowchart of a method for determining a first type of text is further described, comprising the following steps:
[0081] The first step is to determine whether the group litigation event classification will be performed for the first time when the group litigation event interface request is obtained (i.e., whether the preset complaint ticket text library is empty);
[0082] In the second step, if it is the first time to classify a group litigation event, the complaint ticket texts in the complaint ticket text set corresponding to the request are grouped and clustered to obtain the same subject grouping, as well as address similarity calculation and graph construction address grouping under different subjects;
[0083] The third step is to perform N:N feature comparison in each group, and perform infomap clustering on the complaint ticket text in each group to obtain multiple target complaint ticket text clusters;
[0084] Step 4: The top K complaint ticket texts in the center of each target complaint ticket text cluster are used as the cover of the cluster, and each cluster and cover are output to the user and stored in the preset complaint ticket text library;
[0085] In the fifth step, if it is not the first time to classify a group complaint event, each complaint ticket text is compared with the cover of the cluster (i.e., the cover representing the complaint ticket text, i.e., the summary and features generated by the large model);
[0086] Step 6: If the text of the complaint ticket is successfully compared with the cover (i.e., higher than the archiving threshold), the complaint type of the complaint ticket text includes the first type, so the complaint type of the complaint ticket text can be output, and the complaint ticket text can also be added to the corresponding cluster;
[0087] Step 7. If the comparison between the complaint ticket text and the cover fails, the complaint ticket text is returned to the process of step 2 to determine whether the complaint type of the complaint ticket text includes the first type.
[0088] Optionally, in the step of clustering the complaint work order texts to be clustered based on the subject data and address data corresponding to all complaint work order texts to be clustered to obtain a target complaint work order text clustering cluster, the complaint work order texts to be clustered can also be subjected to a first grouping process based on the subject data to obtain at least one subject group; based on the address data, the subject group can be subjected to a second grouping process to obtain at least one target group; and clustering can be performed based on each target group to obtain a target complaint work order text clustering cluster corresponding to each target group.
[0089] In an embodiment of the present invention, the above-mentioned first grouping processing can be implemented based on the above-mentioned subject data, and the complaint work order texts of the same subject data are classified into the same subject group. The above-mentioned second grouping processing can construct a spatial data model based on the address information. By calculating the address similarity between different complaint contents, a similarity matrix is generated. Using the similarity matrix, a graph is constructed in which nodes represent addresses and edges represent similarity relationships. Subsequently, the connectivity of the graph is calculated, and similar addresses are grouped through graph algorithms (such as DFS or BFS) to reveal the relationship and structure between addresses, and then determine the address grouping. Both address grouping and subject grouping are used as target groups. Then, an N:N feature comparison is performed based on the text data within the group. After completing the feature comparison, Infomap clustering analysis is performed on each group.
[0090] After clustering through Infomap, you can also analyze the central node of each cluster to determine its representative characteristics. Select the top 3 data centers after clustering as the cover data to be used for classification of the base database of the next group litigation event.
[0091] Optionally, in the step of determining the complaint type of each complaint work order text in the complaint work order text set based on the subject data and the address data, a preset complaint work order text library can also be obtained, the preset complaint work order text library includes multiple historical complaint work order texts, and the historical complaint work order texts correspond to the subject data and the address data; a second matching process is performed based on the complaint work order text and the historical complaint work order text in the preset complaint work order text library to obtain a second matching result; based on the second matching result, the subject data and the address data of the complaint work order text, it is determined whether the complaint type of each complaint work order text includes the second type.
[0092] In an embodiment of the present invention, the preset complaint work order text library includes multiple historical complaint work order texts, and the historical complaint work order texts correspond to subject data and address data. In the preset complaint work order text library, each historical complaint work order text corresponds to a cover (i.e., the summary and features obtained through large model recognition).
[0093] The second matching process can also be to match the complaint ticket text with the cover corresponding to the historical complaint ticket text (i.e., the summary and features obtained through large model recognition) to obtain a second matching result. If the second matching result is that the complaint ticket text fails to match all historical complaint ticket texts, there is no need to perform subsequent address data matching, indicating that the complaint type of the complaint ticket text does not include the second type.
[0094] On the contrary, the historical complaint work order text corresponding to the second matching result can be determined as the target historical complaint work order text, and the main data and address data of the target historical complaint work order text are matched with the main data and address data of the complaint work order text to determine whether the complaint type of the complaint work order text includes the second type.
[0095] Optionally, in the step of determining whether the complaint type of each complaint ticket text includes the second type based on the second matching result, the main data of the complaint ticket text, and the address data, if the second matching result is a matching failure, the complaint type of the complaint ticket text does not include the second type; if the second matching result is a matching success, then based on the main data and address data of the complaint ticket text and the main data and address data of the target historical complaint ticket text, it is determined whether the complaint type of the complaint ticket text includes the second type, and the target historical complaint ticket text is the historical complaint ticket text corresponding to the second matching result.
[0096] In an embodiment of the present invention, if the second matching result is a matching failure (ie, the complaint ticket text fails to match the cover pages of all historical complaint ticket texts), the complaint type of the complaint ticket text does not include the second type.
[0097] On the contrary, if the second matching result is a successful match (i.e. the complaint ticket text successfully matches the cover of any historical complaint ticket text), the historical complaint ticket text corresponding to the second matching result can be determined as the target historical complaint ticket text. And based on the main data and address data of the complaint ticket text and the main data and address data of the target historical complaint ticket text, matching processing is performed to determine whether the complaint type of the complaint ticket text includes the second type
[0098] For example, if both the subject data and the address data match successfully, the complaint type of the complaint ticket text includes the second type; if either the subject data or the address data fails to match, the complaint type of the complaint ticket text does not include the second type.
[0099] Optionally, in the step of determining whether the complaint type of the complaint work order text includes the second type based on the subject data, address data of the complaint work order text and the subject data, address data of the target historical complaint work order text, and the target historical complaint work order text is the historical complaint work order text corresponding to the second matching result, a subject matching process can also be performed based on the subject data of the complaint work order text and the subject data of the target historical complaint work order text to obtain a subject matching result of the complaint work order text; an address matching process can be performed based on the address data of the complaint work order text and the address data of the target historical complaint work order text to obtain an address matching result of the complaint work order text; a similarity matching process can be performed based on the complaint work order text and the target historical complaint work order text to obtain a similarity matching result of the complaint work order text; and based on the subject matching result, the address matching result and the similarity matching result, it can be determined whether the complaint type of the complaint work order text includes the second type.
[0100] In an embodiment of the present invention, the main body data of the complaint ticket text and the target historical complaint ticket text may be compared for similarity. If the similarity between the two is not less than 1.0, it means that the main body matching result is a successful match, otherwise it is a failed match.
[0101] Similarly, the address data of the complaint ticket text can be compared with the address data of the target historical complaint ticket text for similarity. If the similarity between the two is greater than 0.9, it means that the address matching result is a successful match, otherwise it means that the match fails.
[0102] Similarly, the complaint ticket text and the target historical complaint ticket text can be compared for similarity (eg, cosine similarity comparison). If the similarity between the two is greater than 0.97, then the similarity matching result is a successful match, otherwise it is a failed match.
[0103] After obtaining the above-mentioned subject matching results, address matching results and similarity matching results, the number of target historical complaint ticket texts in the second matching results can be counted. If the above number is only 1 and any of the three fails to match, it means that the complaint type of the complaint ticket text does not include the second type. Conversely, if the above number is only 1 and all three are successfully matched, it means that the complaint type of the complaint ticket text includes the second type.
[0104] Furthermore, if there are multiple target historical complaint work order texts, the number of successful matches between each target historical complaint work order text and the complaint work order text in terms of subject matching results, address matching results, and similarity matching results can be counted, and the ratio of the number of matches to the number of all target historical complaint work order texts can be compared with a preset ratio threshold. If it is greater than the ratio threshold, it means that the complaint type of the complaint work order text includes the second type. If it is not greater than the ratio threshold, it means that the complaint type of the complaint work order text does not include the second type.
[0105] In a possible embodiment, whether the above complaint work order text includes the second type can also be determined by Figure 4 The flowchart of a method for determining the second type of text shown in the figure further illustrates that the method comprises the following steps:
[0106] The first step is to extract the work order information after obtaining the work order content and store it in the database, that is, the preset complaint work order text database;
[0107] The second step is to Figure 2 A data preprocessing method is shown, which extracts the subject involved (i.e., subject data), the address involved (i.e., address data), the work order content (i.e., the complaint work order text), and the features (i.e., the cover such as the summary and features) of the complaint work order text;
[0108] The third step is to check whether the input parameter contains repeated event set information;
[0109] In the fourth step, if the set of repeated events is not included, it can be determined that the complaint type of the complaint ticket text does not include the second type, and the summary and features of the complaint ticket text are identified by the large model as cover information and stored in the preset complaint ticket text library;
[0110] Step 5: If the set of repeated events contains information, the preset complaint ticket text library can be traversed (i.e., the second matching process, whether the complaint ticket text successfully matches the cover information of the historical complaint ticket text);
[0111] In the sixth step, if the complaint ticket text fails to match the cover information of the historical complaint ticket text, the process jumps to the fourth step to determine that the complaint type of the complaint ticket text does not include the second type, and uses the large model to identify the summary and features of the complaint ticket text as the cover information, and stores it in the preset complaint ticket text library;
[0112] Step 7: If the complaint ticket text successfully matches the cover information of the historical complaint ticket text, the number of target historical complaint ticket texts corresponding to the successful match (i.e., the number of covers in the match) is determined;
[0113] In the eighth step, if it is a single cover (i.e., the number of target historical complaint work order texts is 1), the main body data of the complaint work order text can be directly compared with the main body data of the target historical complaint work order text, the address data of the complaint work order text can be compared with the address data of the target historical complaint work order text, and the complaint work order text can be compared with the target historical complaint work order text for cosine similarity to obtain the similarity between the three; if the address similarity is greater than 0.9, the main body similarity is completely consistent (i.e., greater than 1.0), and the feature similarity is greater than 0.97 (i.e., cosine similarity), then the complaint type of the complaint work order text includes the second type, otherwise jump to the fourth step to determine that the complaint type of the complaint work order text does not include the second type, and use the large model to identify the summary and features of the complaint work order text as cover information, and store it in the preset complaint work order text library;
[0114] The ninth step, if there are multiple covers (that is, the number of target historical complaint work order texts is multiple), then traverse multiple target historical complaint work order texts, and count the number of successful matches between each target historical complaint work order text and the subject matching result, address matching result, and similarity matching result of the complaint work order text. The ratio of the number of matches to the number of all target historical complaint work order texts is compared with the preset ratio threshold. If it is greater than the ratio threshold, it means that the complaint type of the complaint work order text includes the second type. If it is not greater than the ratio threshold, jump to the fourth step to determine that the complaint type of the complaint work order text does not include the second type, and use the large model to identify the summary and features of the complaint work order text as cover information, and store it in the preset complaint work order text library.
[0115] like Figure 5 As shown, an embodiment of the present invention further provides a text classification device, comprising:
[0116] A first acquisition module 501 is used to acquire a complaint work order text set, wherein the complaint work order text set includes at least one complaint work order text;
[0117] A first determination module 502, configured to determine the body data and address data of each complaint work order text based on the complaint work order text set;
[0118] The second determination module 503 is used to determine the complaint type of each complaint ticket text in the complaint ticket text set based on the subject data and the address data, and the complaint type includes a first type and a second type. The first type is that multiple people complain about the same event, and the second type is that the same person complains about the same event multiple times.
[0119] Optionally, the second determining module 503 includes:
[0120] A first acquisition submodule is used to acquire a preset complaint work order text library, wherein the preset complaint work order text library includes a first type of clustering clusters, each of the first type of clustering clusters includes a plurality of historical complaint work order texts complaining about the same event, and each of the first type of clustering clusters corresponds to at least one representative complaint work order text, and the representative complaint work order text is one of the plurality of historical complaint work order texts;
[0121] A first matching submodule, configured to perform a first matching process based on each of the complaint work order texts in the complaint work order text set and each of the representative complaint work order texts in a preset complaint work order text library to obtain a first matching result;
[0122] The first determination submodule is used to determine whether the complaint type of the complaint work order text includes the first type based on the first matching result, the main data of the complaint work order text, and the address data.
[0123] Optionally, the first determining submodule includes:
[0124] A first processing unit, configured to, if the first matching result is that the complaint work order text and the representative complaint work order text are successfully matched, then the complaint type of the complaint work order text includes the first type;
[0125] A second processing unit is configured to use the complaint work order text as the complaint work order text to be clustered if the first matching result is that the complaint work order text fails to match all the representative complaint work order texts;
[0126] A first clustering unit, configured to perform clustering processing on the complaint work order texts to be clustered based on the subject data and the address data corresponding to all the complaint work order texts to be clustered, to obtain a target complaint work order text clustering cluster;
[0127] The first determination unit is used to determine whether the complaint type of the complaint ticket text to be clustered includes the first type based on the target complaint ticket text clustering cluster.
[0128] Optionally, the first clustering unit includes:
[0129] A first grouping subunit is used to perform a first grouping process on the complaint work order text to be clustered based on the subject data to obtain at least one subject group;
[0130] A second grouping subunit is used to perform a second grouping process on the subject group based on the address data to obtain at least one target group;
[0131] The first clustering subunit is used to perform clustering processing based on each of the target groups to obtain a target complaint work order text cluster corresponding to each of the target groups.
[0132] Optionally, the second determining module 503 includes:
[0133] The second acquisition submodule is used to acquire a preset complaint work order text library, wherein the preset complaint work order text library includes a plurality of historical complaint work order texts, and the historical complaint work order texts correspond to subject data and address data;
[0134] A second matching submodule is used to perform a second matching process based on the complaint work order text and the historical complaint work order text in the preset complaint work order text library to obtain a second matching result;
[0135] The second determination submodule is used to determine whether the complaint type of each complaint work order text includes the second type based on the second matching result, the main data of the complaint work order text, and the address data.
[0136] Optionally, the second determining submodule includes:
[0137] A third processing unit, configured to, if the second matching result is a matching failure, cause the complaint type of the complaint work order text to not include the second type;
[0138] The second determination unit is used to determine whether the complaint type of the complaint work order text includes the second type based on the main data and the address data of the complaint work order text and the main data and the address data of the target historical complaint work order text if the second matching result is a successful match, and the target historical complaint work order text is the historical complaint work order text corresponding to the second matching result.
[0139] Optionally, the second determining unit includes:
[0140] A first matching subunit is used to perform subject matching processing based on the subject data of the complaint work order text and the subject data of the target historical complaint work order text to obtain a subject matching result of the complaint work order text;
[0141] A second matching subunit is used to perform address matching processing based on the address data of the complaint work order text and the address data of the target historical complaint work order text to obtain an address matching result of the complaint work order text;
[0142] A third matching subunit is used to perform similarity matching based on the complaint work order text and the target historical complaint work order text to obtain a similarity matching result of the complaint work order text;
[0143] The first determination subunit is used to determine whether the complaint type of the complaint work order text includes the second type based on the subject matching result, the address matching result and the similarity matching result.
[0144] like Figure 6 As shown, an embodiment of the present invention further provides an electronic device, characterized in that it includes a processor, and the processor can execute any one of the above-mentioned text classification methods.
[0145] Specifically, it includes a processor 601 and a memory 602, and a computer program for executing the text classification method stored in the memory 602 and capable of running on the processor 601, wherein:
[0146] The processor 601 runs the computer program of the text classification method stored in the memory 602 to perform the following steps:
[0147] Obtaining a complaint work order text set, wherein the complaint work order text set includes at least one complaint work order text;
[0148] Based on the complaint work order text set, determining the main body data and address data of each complaint work order text;
[0149] Based on the subject data and the address data, the complaint type of each complaint ticket text in the complaint ticket text set is determined, and the complaint type includes a first type and a second type. The first type is that multiple people complain about the same event, and the second type is that the same person complains about the same event multiple times.
[0150] Optionally, the determining of the complaint type of each complaint work order text in the complaint work order text set based on the subject data and the address data performed by the processor 601 includes:
[0151] Obtain a preset complaint work order text library, wherein the preset complaint work order text library includes a first type of clustering clusters, each of the first type of clustering clusters includes multiple historical complaint work order texts complaining about the same event, and each of the first type of clustering clusters corresponds to at least one representative complaint work order text, and the representative complaint work order text is one of the multiple historical complaint work order texts;
[0152] Performing a first matching process based on each of the complaint work order texts in the complaint work order text set and each of the representative complaint work order texts in the preset complaint work order text library to obtain a first matching result;
[0153] Based on the first matching result, the main data of the complaint work order text, and the address data, determine whether the complaint type of the complaint work order text includes the first type.
[0154] Optionally, the determining, based on the first matching result, the subject data of the complaint work order text, and the address data, whether the complaint type of the complaint work order text includes the first type, performed by the processor 601, includes:
[0155] If the first matching result is that the complaint work order text matches the representative complaint work order text successfully, then the complaint type of the complaint work order text includes the first type;
[0156] If the first matching result is that the complaint work order text fails to match all the representative complaint work order texts, the complaint work order text is used as the complaint work order text to be clustered;
[0157] Clustering the complaint work order texts to be clustered based on the subject data and the address data corresponding to all the complaint work order texts to be clustered to obtain a target complaint work order text clustering cluster;
[0158] Based on the target complaint ticket text clustering cluster, determine whether the complaint type of the complaint ticket text to be clustered includes the first type.
[0159] Optionally, the processor 601 performs clustering processing on the complaint work order texts to be clustered based on the subject data and the address data corresponding to all the complaint work order texts to be clustered to obtain a target complaint work order text clustering cluster, including:
[0160] Based on the subject data, performing a first grouping process on the complaint work order text to be clustered to obtain at least one subject group;
[0161] Based on the address data, performing a second grouping process on the subject group to obtain at least one target group;
[0162] Clustering is performed based on each of the target groups to obtain a target complaint work order text cluster corresponding to each of the target groups.
[0163] Optionally, the determining of the complaint type of each complaint work order text in the complaint work order text set based on the subject data and the address data performed by the processor 601 includes:
[0164] Obtain a preset complaint work order text library, wherein the preset complaint work order text library includes multiple historical complaint work order texts, and the historical complaint work order texts correspond to subject data and address data;
[0165] Performing a second matching process based on the complaint work order text and the historical complaint work order text in the preset complaint work order text library to obtain a second matching result;
[0166] Based on the second matching result, the main data of the complaint work order text, and the address data, determine whether the complaint type of each complaint work order text includes the second type.
[0167] Optionally, the determining, based on the second matching result, the subject data of the complaint work order text, and the address data, whether the complaint type of each complaint work order text includes the second type, performed by the processor 601, includes:
[0168] If the second matching result is a match failure, the complaint type in the complaint ticket text does not include the second type;
[0169] If the second matching result is a successful match, based on the main data and address data of the complaint ticket text and the main data and address data of the target historical complaint ticket text, it is determined whether the complaint type of the complaint ticket text includes the second type, and the target historical complaint ticket text is the historical complaint ticket text corresponding to the second matching result.
[0170] Optionally, the processor 601 determines whether the complaint type of the complaint work order text includes the second type based on the subject data, the address data of the complaint work order text and the subject data and address data of the target historical complaint work order text, and the target historical complaint work order text is the historical complaint work order text corresponding to the second matching result, including:
[0171] Performing subject matching processing based on the subject data of the complaint work order text and the subject data of the target historical complaint work order text to obtain a subject matching result of the complaint work order text;
[0172] Performing address matching processing based on the address data of the complaint work order text and the address data of the target historical complaint work order text to obtain an address matching result of the complaint work order text;
[0173] Based on the similarity matching between the complaint work order text and the target historical complaint work order text, a similarity matching result of the complaint work order text is obtained;
[0174] Based on the subject matching result, the address matching result and the similarity matching result, determine whether the complaint type of the complaint work order text includes the second type.
[0175] An embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the computer program implements the various processes of the text classification method or the application-side text classification method provided in the embodiment of the present invention, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.
[0176] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing related hardware through a computer program, and the above-mentioned computer program can be stored in a computer-readable storage medium, and when the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, the above-mentioned computer-readable storage medium can be a disk, an optical disk, a read-only memory (ROM) or a random access memory (RAM), etc.
[0177] The above disclosure is only the preferred embodiment of the present invention, which certainly cannot be used to limit the scope of the present invention. Therefore, equivalent changes made according to the claims of the present invention are still within the scope of the present invention.
Claims
1. A text classification method, characterized in that: The method comprises the following steps: Obtaining a complaint work order text set, wherein the complaint work order text set includes at least one complaint work order text; Based on the complaint work order text set, determining the main body data and address data of each complaint work order text; Based on the subject data and the address data, the complaint type of each complaint ticket text in the complaint ticket text set is determined, and the complaint type includes a first type and a second type. The first type is that multiple people complain about the same event, and the second type is that the same person complains about the same event multiple times.
2. The text classification method according to claim 1, characterized in that: The determining, based on the subject data and the address data, the complaint type of each complaint work order text in the complaint work order text set comprises: Obtain a preset complaint work order text library, wherein the preset complaint work order text library includes a first type of clustering clusters, each of the first type of clustering clusters includes multiple historical complaint work order texts complaining about the same event, and each of the first type of clustering clusters corresponds to at least one representative complaint work order text, and the representative complaint work order text is one of the multiple historical complaint work order texts; Performing a first matching process based on each of the complaint work order texts in the complaint work order text set and each of the representative complaint work order texts in the preset complaint work order text library to obtain a first matching result; Based on the first matching result, the main data of the complaint work order text, and the address data, determine whether the complaint type of the complaint work order text includes the first type.
3. The text classification method according to claim 2, characterized in that: The determining, based on the first matching result, the subject data of the complaint work order text, and the address data, whether the complaint type of the complaint work order text includes the first type comprises: If the first matching result is that the complaint work order text matches the representative complaint work order text successfully, then the complaint type of the complaint work order text includes the first type; If the first matching result is that the complaint work order text fails to match all the representative complaint work order texts, the complaint work order text is used as the complaint work order text to be clustered; Clustering the complaint work order texts to be clustered based on the subject data and the address data corresponding to all the complaint work order texts to be clustered to obtain a target complaint work order text clustering cluster; Based on the target complaint ticket text clustering cluster, determine whether the complaint type of the complaint ticket text to be clustered includes the first type.
4. The text classification method according to claim 3, characterized in that: The step of clustering the complaint work order texts to be clustered based on the subject data and the address data corresponding to all the complaint work order texts to be clustered to obtain a target complaint work order text clustering cluster includes: Based on the subject data, performing a first grouping process on the complaint work order text to be clustered to obtain at least one subject group; Based on the address data, performing a second grouping process on the subject group to obtain at least one target group; Clustering is performed based on each of the target groups to obtain a target complaint work order text cluster corresponding to each of the target groups.
5. The text classification method according to claim 1, characterized in that: The determining, based on the subject data and the address data, the complaint type of each complaint work order text in the complaint work order text set comprises: Obtain a preset complaint work order text library, wherein the preset complaint work order text library includes multiple historical complaint work order texts, and the historical complaint work order texts correspond to subject data and address data; Performing a second matching process based on the complaint work order text and the historical complaint work order text in the preset complaint work order text library to obtain a second matching result; Based on the second matching result, the main data of the complaint work order text, and the address data, determine whether the complaint type of each complaint work order text includes the second type.
6. The text classification method according to claim 5, characterized in that: The determining, based on the second matching result, the subject data of the complaint work order text, and the address data, whether the complaint type of each complaint work order text includes the second type comprises: If the second matching result is a match failure, the complaint type in the complaint ticket text does not include the second type; If the second matching result is a successful match, based on the main data and address data of the complaint ticket text and the main data and address data of the target historical complaint ticket text, it is determined whether the complaint type of the complaint ticket text includes the second type, and the target historical complaint ticket text is the historical complaint ticket text corresponding to the second matching result.
7. The text classification method according to claim 6, characterized in that: The determining whether the complaint type of the complaint work order text includes the second type based on the subject data and the address data of the complaint work order text and the subject data and the address data of the target historical complaint work order text, wherein the target historical complaint work order text is the historical complaint work order text corresponding to the second matching result, includes: Performing subject matching processing based on the subject data of the complaint work order text and the subject data of the target historical complaint work order text to obtain a subject matching result of the complaint work order text; Performing address matching processing based on the address data of the complaint work order text and the address data of the target historical complaint work order text to obtain an address matching result of the complaint work order text; Based on the similarity matching between the complaint work order text and the target historical complaint work order text, a similarity matching result of the complaint work order text is obtained; Based on the subject matching result, the address matching result and the similarity matching result, determine whether the complaint type of the complaint work order text includes the second type.
8. A text classification device, characterized in that: The text classification device comprises: A first acquisition module is used to acquire a complaint work order text set, wherein the complaint work order text set includes at least one complaint work order text; A first determination module, used to determine the main body data and address data of each complaint work order text based on the complaint work order text set; The second determination module is used to determine the complaint type of each complaint work order text in the complaint work order text set based on the subject data and the address data, and the complaint type includes a first type and a second type. The first type is that multiple people complain about the same event, and the second type is that the same person complains about the same event multiple times.
9. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps in the text classification method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the text classification method according to any one of claims 1 to 7 are implemented.