A method, device, electronic device and storage medium for classifying feedback information

By performing word segmentation processing and word frequency-inverse document frequency analysis on the feedback information, the multi-level classification results of the feedback information are determined, and the problem of inaccurate classification of feedback information in the prior art is solved, and the accuracy and effectiveness of classification are improved.

CN115129860BActive Publication Date: 2025-06-17TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210357691.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-06
Publication Date
2025-06-17
Estimated Expiration
2042-04-06

AI Technical Summary

Technical Problem

When classifying feedback information, the prior art has a short text length of the feedback information, resulting in limited information extracted from the text by the deep learning model, inaccurate classification results and poor effectiveness.

Method used

By performing word segmentation processing on the feedback information, the word frequency-inverse document frequency of each word segmentation word is calculated, the confidence level of each first category is determined, and the target first category and sub-level classification category are then determined, and a multi-level classification result is generated.

Benefits of technology

The classification accuracy and effectiveness of feedback information are improved, and an effective basis for judgment is provided for the subsequent processing of feedback information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115129860B_ABST
    Figure CN115129860B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, apparatus, electronic device and storage medium for classifying feedback information. The method includes: performing word segmentation processing on the feedback text corresponding to the newly added feedback information; determining the confidence levels of the newly added feedback information belonging to each first category according to the term frequency-inverse document frequency of each segmented word in a plurality of segmented words, and determining the target first category to which the newly added feedback information belongs according to the first category with the highest confidence level; performing vector embedding on the plurality of segmented words according to the target first category, the sentence corresponding to each segmented word in the feedback text, and the position in the sentence; encoding the result of vector embedding, and classifying the encoded result to obtain the target second category to which the newly added feedback information belongs; generating a multi-level classification result corresponding to the newly added feedback information according to the hierarchical relationship between the target second category and the target first category. The present invention improves the classification accuracy and effectiveness of feedback information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and particularly to a method, apparatus, electronic device and storage medium for classifying feedback information. Background Art

[0002] With the development of Internet technology, when users use application programs in a terminal, they often give feedback based on actual experience. For example, if it is very slow to generate a video using a certain application program, the feedback content based on this experience can be "Why is it so slow for me to generate a video? It has been a whole day" and other feedback information.

[0003] In related technologies, when analyzing feedback information, usually a deep learning model is first used to classify the feedback information. However, since the text length of feedback information is usually only 10 - 15 characters, the information extracted from the text based solely on the deep learning model is limited, resulting in inaccurate classification results and poor effectiveness for feedback information in related technologies, and unable to provide an effective judgment basis for subsequent processing of feedback information. Summary of the Invention

[0004] To solve the problems of the prior art, embodiments of the present invention provide a method, apparatus, electronic device and storage medium for classifying feedback information. The technical solutions are as follows:

[0005] On the one hand, a method for classifying feedback information is provided, and the method includes:

[0006] Determine the feedback text corresponding to the newly added feedback information, and perform word segmentation processing on the feedback text to obtain a plurality of segmented words;

[0007] According to the term frequency - inverse document frequency of each segmented word in the plurality of segmented words corresponding to a plurality of first categories, determine the confidence of the newly added feedback information belonging to each of the first categories, and determine the target first category to which the newly added feedback information belongs according to the first category with the highest confidence;

[0008] According to the target first category, each segmented word corresponding sentence in the feedback text, and the position in the sentence, perform vector embedding on the plurality of segmented words;

[0009] Encode the result of the vector embedding, and perform classification according to the result of the encoding to obtain the target second category to which the newly added feedback information belongs; the target second category is a sub - classification category of the target first category;

[0010] According to the target second category, the target first category, and the hierarchical relationship between the target second category and the target first category, generate a multi - level classification result corresponding to the newly added feedback information.

[0011] On the other hand, a feedback information classification device is provided, and the device includes:

[0012] A word segmentation module, configured to determine a feedback text corresponding to the newly added feedback information, and perform word segmentation processing on the feedback text to obtain a plurality of segmented words;

[0013] A first classification module, configured to determine the confidence of the newly added feedback information belonging to each of the first categories according to the term frequency-inverse document frequency of each segmented word corresponding to a plurality of first categories among the plurality of segmented words, and determine the target first category to which the newly added feedback information belongs according to the first category with the highest confidence;

[0014] A first vector embedding module, configured to perform vector embedding on the plurality of segmented words according to the target first category, each sentence corresponding to each segmented word in the feedback text, and the position in the sentence;

[0015] A second classification module, configured to encode the result of the vector embedding, and perform classification according to the result of the encoding to obtain the target second category to which the newly added feedback information belongs; the target second category is a sub-classification category of the target first category;

[0016] A hierarchical classification result generation module, configured to generate a multi-level classification result corresponding to the newly added feedback information according to the target second category, the target first category, and the hierarchical relationship between the target second category and the target first category.

[0017] In an exemplary embodiment, the second classification module includes:

[0018] A first convolution module, configured to perform convolution processing on the result of the vector embedding to obtain a feature vector;

[0019] A first encoding module, configured to perform encoding processing on the feature vector based on a self-attention mechanism to obtain an encoded feature vector;

[0020] A first classification sub-module, configured to perform classification according to the encoded feature vector to obtain the target second category to which the newly added feedback information belongs.

[0021] In an exemplary embodiment, the first classification module includes:

[0022] A term frequency determination module, configured to determine the term frequency of each segmented word in the feedback text among the plurality of segmented words;

[0023] An inverse document frequency determination module, configured to determine the inverse document frequency of each segmented word corresponding to each of the first categories according to the classification word information corresponding to each of the first categories;

[0024] The term frequency - inverse document frequency determination module is used to determine the term frequency - inverse document frequency of each segmented word corresponding to each of the first categories according to the term frequency of each segmented word and the inverse document frequency corresponding to each of the first categories.

[0025] The category confidence determination module is used to, for each of the first categories, determine the sum value of the term frequency - inverse document frequencies of multiple segmented words corresponding to the first category, and obtain the confidence that the new feedback information belongs to the first category.

[0026] The second classification sub - module is used to determine the target first category to which the new feedback information belongs according to the first category with the highest confidence among multiple first categories.

[0027] In an exemplary embodiment, the device further includes:

[0028] The sample segmentation module is used to perform word segmentation processing on the historical feedback texts in the historical feedback text set to obtain the segmented words corresponding to each historical feedback text.

[0029] The hierarchical clustering module is used to perform hierarchical clustering on the historical feedback text set according to the segmented words corresponding to each historical feedback text to obtain a hierarchical clustering result; the hierarchical clustering result includes the multiple first categories and the clustering clusters of each first category.

[0030] The first determination module is used to determine the classification word information of each first category according to the segmented words corresponding to the historical feedback texts in the clustering cluster of each first category.

[0031] In an exemplary embodiment, the device further includes:

[0032] The category prediction module is used to input the segmented words of the historical feedback texts in the historical feedback text set and the corresponding first categories of the historical feedback texts into a classification model for category prediction to obtain a category prediction result; wherein, the classification model is used to perform vector embedding on the historical feedback text according to the first category of the historical feedback text, each segmented word of the historical feedback text corresponding sentence in the historical feedback text and the position in the sentence, and predict the category of the historical feedback text according to the result of vector embedding.

[0033] The model training module is used to adjust the model parameters of the classification model for iterative training until a preset training end condition is met according to the difference between the category prediction result corresponding to each historical feedback text and the reference category corresponding to the historical feedback text; wherein, the reference category corresponding to the historical feedback text is a sub - classification category of the corresponding first category.

[0034] In an exemplary embodiment, the category prediction module includes:

[0035] A second vector embedding module, configured to input the segmented words of the historical feedback texts in the historical feedback text set and the first category corresponding to the historical feedback texts into the embedding layer of a classification model for vector embedding;

[0036] A convolutional encoding module, configured to perform convolution on the result of the vector embedding through the convolutional layer of the classification model, and input the result of the convolution into the self-attention encoding layer of the classification model for encoding;

[0037] A category prediction sub-module, configured to perform category prediction on the result of the encoding through the classification layer of the classification model to obtain a category prediction result.

[0038] In an exemplary embodiment, the device further includes:

[0039] A display module, configured to display the classification word information of each of the first categories;

[0040] An update module, configured to update the classification word information of any one of the first categories in response to a configuration instruction for the classification word information of any one of the first categories.

[0041] On the other hand, an electronic device is provided, including a processor and a memory, where at least one instruction or at least one segment of program is stored in the memory, and the at least one instruction or the at least one segment of program is loaded and executed by the processor to implement the above feedback information classification method.

[0042] On the other hand, a computer-readable storage medium is provided, where at least one instruction or at least one segment of program is stored in the computer-readable storage medium, and the at least one instruction or the at least one segment of program is loaded and executed by a processor to implement the feedback information classification method as described above.

[0043] On the other hand, a computer program product or a computer program is provided, the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the electronic device executes the above feedback information classification method.

[0044] In the embodiments of the present invention, aiming at the short text characteristics of feedback information, word segmentation is performed on the feedback text corresponding to the newly added feedback information to obtain a plurality of segmented words. Then, according to the term frequency-inverse document frequency of each segmented word corresponding to a plurality of first categories among the plurality of segmented words, the confidence of the newly added feedback information belonging to each first category is determined. And according to the first category with the highest confidence, the target first category to which the newly added feedback information belongs is determined. Furthermore, according to the target first category, the sentence corresponding to each segmented word in the above-mentioned feedback text, and the position in the sentence, vector embedding is performed on the plurality of segmented words, and the result of vector embedding is encoded and the encoded result is classified to obtain the sub-classification category of the newly added feedback information. According to the hierarchical relationship between the sub-classification category and the target first category, a multi-level classification result of the newly added feedback information is generated, thereby improving the classification accuracy and effectiveness of feedback information and providing an effective judgment basis for the subsequent processing of feedback information. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0046] Figure 1 It is a schematic diagram of an implementation environment provided by an embodiment of the present invention;

[0047] Figure 2 It is a schematic flowchart of a feedback information classification method provided by an embodiment of the present invention;

[0048] Figure 3 It is an example of displaying the multi-level classification results of each newly added feedback information within a preset time period provided by an embodiment of the present invention;

[0049] Figure 4 It is a schematic flowchart of a training process provided by an embodiment of the present invention;

[0050] Figure 5 It is a schematic diagram of a configuration page provided by an embodiment of the present invention;

[0051] Figure 6 It is a structural block diagram of a feedback information classification device provided by an embodiment of the present invention;

[0052] Figure 7 It is a hardware structural block diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0053] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0054] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order different from those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or server comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0055] It can be understood that in the specific implementation manners of the present application, when it comes to relevant data such as user information, when the above embodiments of the present application are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.

[0056] Please refer to Figure 1 , which shows a schematic diagram of an implementation environment provided by an embodiment of the present invention. The implementation environment may include a terminal 110 and a server 120, where the terminal 110 and the server 120 can communicate through a wired connection or a wireless connection.

[0057] The terminal 110 includes, but is not limited to, a mobile phone, a computer, a smart voice interaction device, a smart home appliance, a vehicle-mounted terminal, an aircraft, etc. The terminal 110 runs client software with a human-computer interaction function, such as an application program (abbreviated as App). The application program can be an independent application program or a subroutine in an application program. Exemplarily, the application program can be a game application program, a shopping application program, etc. The user of the terminal 110 can log in to the application program with pre-registered user information, which can include an account and a password. During the operation of the application program, the user of the terminal 110 can provide information feedback.

[0058] Server 120 can be a server that provides background services for applications in terminal 110. Specifically, it can provide classification services for feedback information. Server 120 can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.

[0059] In an exemplary embodiment, both terminal 110 and server 120 can be node devices in a blockchain system, capable of sharing the information obtained and generated with other node devices in the blockchain system, so as to achieve information sharing among multiple node devices. Multiple node devices in the blockchain system can be configured with the same blockchain, which is composed of multiple blocks, and adjacent blocks before and after have an association relationship, so that when the data in any block is tampered with, it can be detected by the next block, thereby avoiding the data in the blockchain from being tampered with and ensuring the security and reliability of the data in the blockchain.

[0060] The feedback information classification method according to the embodiments of the present invention can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, intelligent transportation, assisted driving, etc.

[0061] Artificial Intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning, and decision-making. Artificial intelligence technology is an interdisciplinary subject, involving a wide range of fields, including both hardware-level technologies and software-level technologies. Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technologies, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0062] Among them, Natural Language Processing (NLP) is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between humans and computers in natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, the research in this field will involve natural language, that is, the language people use in daily life, so it has a close connection with the research of linguistics. Natural language processing technologies usually include text processing, semantic understanding, machine translation, robot question answering, knowledge graph and other technologies.

[0063] Machine Learning (ML) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.

[0064] Intelligent Transportation System (ITS), also known as Intelligent Transportation System, effectively integrates advanced scientific and technological means (information technology, computer technology, data communication technology, sensor technology, electronic control technology, automatic control theory, operations research, artificial intelligence, etc.) into transportation, service control, and vehicle manufacturing, strengthening the connection among vehicles, roads, and users, thereby forming a comprehensive transportation system that ensures safety, improves efficiency, improves the environment, and saves energy.

[0065] Please refer to Figure 2 , which shows a schematic flowchart of a feedback information classification method provided by an embodiment of the present invention. This method can be applied to Figure 1 the server in. It should be noted that this specification provides method operation steps as described in the embodiment or flowchart, but based on routine or non-creative labor, there may be more or fewer operation steps. The step order listed in the embodiment is only one of the many step execution orders and does not represent the only execution order. In actual system or product execution, it can be executed in the order shown in the embodiment or the drawing or in parallel (for example, in an environment of parallel processors or multi-threaded processing). Specifically, as Figure 2 shown, the method may include:

[0066] S201. Determine the feedback text corresponding to the newly added feedback information, and perform word segmentation on the feedback text to obtain a plurality of segmented words.

[0067] Among them, the newly added feedback information refers to the feedback information that has not been classified yet. The newly added feedback information may include text and pictures.

[0068] In a specific implementation, after obtaining the newly added feedback information, preprocessing can be performed on the newly added feedback information to obtain the corresponding feedback text. Among them, the preprocessing may include English case conversion, Chinese simplified and traditional conversion, text recognition of pictures, etc. For text recognition of pictures, optical character recognition (OCR) can be used but is not limited to it.

[0069] For word segmentation of the feedback text, a word segmentation tool (such as the Chinese word segmentation tool Jieba) can be used. For the word segmentation process of the word segmentation tool, relevant descriptions can be referred to and will not be elaborated here.

[0070] S203. According to the term frequency–inverse document frequency (TF-IDF) of each segmented word in the plurality of segmented words corresponding to a plurality of first categories, determine the confidence level of the newly added feedback information belonging to each of the first categories, and determine the target first category to which the newly added feedback information belongs according to the first category with the highest confidence level.

[0071] Among them, the term frequency–inverse document frequency (TF-IDF) is used to evaluate the importance of a word for a document set or a single document in a corpus. The term frequency–inverse document frequency has two meanings: one is the term frequency, and the importance of a word increases in direct proportion to the number of times it appears in the document; the other is the inverse document frequency, and the importance of a word decreases in inverse proportion to the frequency of its appearance in the corpus. Based on this, the above step S203 can include when implemented:

[0072] Determine the term frequency of each segmented word in the plurality of segmented words in the feedback text;

[0073] According to the classification word information corresponding to each first category, determine the inverse document frequency of each segmented word corresponding to each first category;

[0074] According to the term frequency of each segmented word and the inverse document frequency corresponding to each first category, determine the term frequency–inverse document frequency of each segmented word corresponding to each first category;

[0075] For each of the first categories, determine the sum of the term frequency-inverse document frequencies of the multiple segmented words corresponding to the first category, and obtain the confidence level of the new feedback information belonging to the first category;

[0076] According to the first category with the highest confidence level among the multiple first categories, determine the target first category to which the new feedback information belongs.

[0077] Specifically, the term frequency represents the frequency of a given word appearing in the text, which can be calculated by the following formula (1):

[0078]

[0079] where n i,j represents the number of occurrences of the segmented word i in the feedback text j; ∑ k n k,j represents the total number of all segmented words corresponding to the feedback text j; tf i,j represents the term frequency of the segmented word i in the feedback text j.

[0080] The inverse document frequency represents the importance of a given word in the current classification. The inverse document frequency of a specific word can be obtained by dividing the total number of documents by the number of documents containing the word, and then taking the logarithm of the obtained quotient. Specifically, it can be calculated by the following formula (2):

[0081]

[0082] where |D| represents the total number of classification words corresponding to the classification; |{j:t i ∈d j}| represents the total number of historical feedback texts of the first category containing the given word.

[0083] Then the term frequency-inverse document frequency TF-IDF = tf i,j *idf i .

[0084] In the embodiments of the present invention, the classification word information of the first category indicates multiple classification words corresponding to the first category, as well as the total number and distribution of each classification word. The distribution of each classification word refers to the number of historical feedback texts of the first category containing the classification word. In a specific implementation, the classification word information of the multiple first categories can be obtained through unsupervised machine learning of historical feedback information. The determination of the classification word information of the multiple first categories will be described in detail in the subsequent embodiments of the present invention.

[0085] As can be seen from the above formula (2), based on the total number and distribution of each classification term in the above-mentioned classification term information of the first category, the inverse document frequency of the classification term can be calculated, so that it is not necessary to count the classification terms in real time, which is beneficial to improving the determination speed of the target first category described in the newly added feedback information, and further improving the overall classification efficiency.

[0086] For example, let tf i represent the word frequency of the segmented word i in the feedback text. For any first category C p among multiple first categories, search for the classification term matching the segmented word i in the classification term information of this first category C p , and calculate the inverse document frequency idf i based on the quantity and distribution corresponding to the matching classification term, and then calculate tf i *idf i to obtain the value corresponding to the segmented word i for C p of In this way, the value corresponding to each first category for the segmented word i can be calculated. For any first category C p , calculate the sum of the TF-IDF values of all segmented words corresponding to this C p to obtain the confidence level that the newly added feedback information belongs to this first category C p , that is Furthermore, the target first category can be determined according to the first category with the highest confidence level among multiple first categories.

[0087] In a specific implementation, when determining the target first category to which the newly added feedback information belongs according to the first category with the highest confidence level among multiple first categories, the first category with the highest confidence level can be determined as the target first category to which the newly added feedback information belongs.

[0088] It should be noted that in the embodiments of the present invention, multiple first categories can be multiple categories at the same level, or multiple subcategories corresponding to the same upper category. When multiple first categories are categories at the first level, the target first category includes the target first-level category. When multiple first categories are multiple second-level categories corresponding to a certain first-level category, the target first category can include the target second-level category. That is to say, when implementing step S203, the target first-level category to which the newly added feedback information belongs can be determined first from multiple first-level categories, and then the target second-level category to which the newly added feedback information belongs can be determined from the multiple second-level categories corresponding to the target first-level category. The determination processes of the target first-level category and the target second-level category are the same, and both can refer to the foregoing determination method for the target first category, so that the current classification of the newly added feedback information can be the target first-level category / target second-level category.

[0089] S205. Perform vector embedding on the multiple segmented words according to the target first category, the sentence corresponding to each segmented word in the feedback text, and the position in the sentence.

[0090] Specifically, through vector embedding, a word vector, a position vector, a sentence vector, and a parent classification vector of each segmented word can be obtained. Among them, the word vector represents the text information of the corresponding segmented word and is obtained by performing vector embedding on the corresponding segmented word; the position vector represents the context information and is obtained by performing vector embedding on the position information of the corresponding segmented word in the sentence.

[0091] The sentence vector represents the inter-sentence information. The segmented words in the same sentence correspond to the same sentence vector. For example, if the feedback text has two sentences: sentence 0 and sentence 1, assuming that the segmented words a and b are both words in sentence 0, and the segmented word c is a word in sentence 1, then the sentence vectors corresponding to the segmented words a and b are the same, which can be the embedding vector obtained by performing vector embedding on sentence number 0. The sentence vector corresponding to the segmented word c is different from those of the segmented words a and b, and it can be the embedding vector obtained by performing vector embedding on sentence number 1. The parent classification vector is the embedding vector obtained by performing vector embedding on the target first category.

[0092] In a specific implementation, for each segmented word, the corresponding word vector, position vector, sentence vector, and parent vector obtained by vector embedding are superimposed as the result of the vector embedding corresponding to the segmented word, thereby integrating various information multi-dimensionally.

[0093] S207. Encode the result of the vector embedding, and classify according to the result of the encoding to obtain the target second category to which the new feedback information belongs.

[0094] Among them, the target second category is a sub-classification category of the target first category. For example, if the target first category is a first-level category, the target second category can be the next-level category, i.e., the second-level category, of this first-level category, or it can be the next-next-level category, i.e., the third-level category, of this first-level category. That is to say, the target second category is not limited to the next-level category of the target first category, as long as it is a sub-classification category of the target first category.

[0095] In a specific implementation, encoding the result of the vector embedding can obtain the corresponding encoding vector, and then classify based on this encoding vector to obtain the target second category corresponding to the new feedback information.

[0096] To improve the classification efficiency and accuracy, in an exemplary implementation manner, the above step S207 may include when implemented:

[0097] Perform convolution processing on the result of the vector embedding to obtain a feature vector;

[0098] Perform encoding processing on the feature vector based on the self-attention mechanism to obtain an encoded feature vector;

[0099] Classify according to the encoded feature vector to obtain the target second category to which the new feedback information belongs.

[0100] Specifically, feature extraction on the result of vector embedding through convolution processing can further strengthen the context connection. The dimension of the vector after convolution drops significantly, which can effectively accelerate the running speed of the subsequent module, and thus is beneficial to improving the classification efficiency. The above encoding processing can adopt the Encoder unit of Transformer, which can solve the problems of parallel acceleration and enhancing word relevance, and improve the classification accuracy. Classification can use the Softmax activation function.

[0101] It should be noted that the above steps S205 and S207 can be implemented based on a pre-trained classification model, that is, the target first category, multiple segmented words, and the sentence information corresponding to each segmented word in the feedback text and the position information in the corresponding sentence can be input into the classification model, and the target second category is output through the classification model. This classification model can be a deep learning model including an embedding layer, a convolution layer, a self-attention encoding layer, and a classification layer. The training of this classification model will be described in detail in the subsequent content of this embodiment of the present invention.

[0102] S209. Generate a multi-level classification result corresponding to the new feedback information according to the target second category, the target first category, and the hierarchical relationship between the target second category and the target first category.

[0103] Specifically, when the target first category and the target second category are in an adjacent upper and lower hierarchical relationship, the multi-level classification result is the target first category and the target second category represented based on this upper and lower hierarchical relationship, such as target first category / target second category; when the target first category and the target second category are not in an adjacent upper and lower hierarchical relationship, it indicates that there are other hierarchical categories between the target first category and the target second category. Then, when generating the multi-level classification result, it is necessary to obtain the hierarchical categories located between the two according to the hierarchical relationship between the target first category and the target second category. Furthermore, the multi-level classification result is the target first category, the target second category, and the obtained hierarchical categories represented based on the hierarchical relationship, such as target first category / intermediate hierarchical category / target second category.

[0104] In an exemplary embodiment, after obtaining the multi-level classification results of the newly added feedback information, it can be displayed. In practical applications, in order to improve the efficiency of data processing, usually the newly added feedback information within a preset time period can be obtained, and then based on Figure 2 the method shown, determine the multi-level classification results of each piece of newly added feedback information within the preset time period. Then, when displaying, the multi-level classification results of each piece of newly added feedback information within the preset time period can be displayed, as Figure 3 shown is an example provided by an embodiment of the present invention for displaying the multi-level classification results of each piece of newly added feedback information within a preset time period. Among them, the preset time period can be set according to actual needs. For example, it can be 1 day, 3 days, etc.

[0105] In practical applications, when displaying the multi-level classification results of each piece of newly added feedback information within a preset time period, the number of newly added feedback information belonging to each category can also be counted for each category, and the above-mentioned quantities of each category can be displayed based on the hierarchical relationship between the categories, as Figure 3 shown, so as to facilitate the business to understand the situation of the newly added feedback information corresponding to each category.

[0106] It can be seen from the above technical solutions of the embodiments of the present invention that the embodiments of the present invention perform multi-level classification for the short text characteristics of feedback information, use the pre-classification to capture high-frequency information and act on the post-classification, thereby improving the classification accuracy and effectiveness of feedback information, and providing an effective judgment basis for the subsequent processing of feedback information.

[0107] The classification word information of the foregoing multiple first categories in the embodiments of the present invention can be obtained by performing unsupervised machine learning on historical feedback information. Specifically, the unsupervised learning can adopt a hierarchical clustering algorithm. The following introduces the unsupervised learning of historical feedback information and the training process of the classification model.

[0108] Please refer to Figure 4 , which shows a schematic flowchart of the training process provided by an embodiment of the present invention. The training process includes training to obtain the classification word information of multiple first categories based on the hierarchical clustering algorithm, and the process of training the classification model.

[0109] Among them, training to obtain the classification word information of multiple first categories based on the hierarchical clustering algorithm includes the following steps:

[0110] Perform word segmentation processing on the historical feedback texts in the historical feedback text set to obtain the segmented words corresponding to each historical feedback text;

[0111] Perform hierarchical clustering on the historical feedback text set according to the segmented words corresponding to each historical feedback text; the hierarchical clustering result includes the multiple first categories and the clustering clusters of each first category.

[0112] Determine the classification word information of each first category according to the segmented words corresponding to the historical feedback texts in the clustering cluster of each first category.

[0113] Hierarchical clustering is to perform clustering layer by layer, which can include divisive clustering and agglomerative clustering. Among them, divisive clustering means that initially all samples are grouped into one cluster, and then gradually split according to a certain criterion until a certain condition is reached or the set number of classifications is reached; agglomerative clustering means that initially each sample point is regarded as a cluster, so the size of the original cluster is equal to the number of sample points, and then these initial clusters are merged according to a certain criterion until a certain condition is reached or the set number of classifications is reached.

[0114] In a specific implementation, for the segmented words corresponding to each historical feedback text, the term frequency-inverse document frequency TF-IDF of the segmented words is calculated in combination with the historical feedback text set, and the TF-IDF of the segmented words corresponding to the historical feedback text constitutes the eigenvalue of the historical feedback text. Taking agglomerative clustering as an example: (1) Each historical feedback text is used as a cluster, and then the similarity between different clusters is calculated based on the eigenvalues of the historical feedback texts in the historical feedback text set; (2) The two clusters with the closest similarity are merged into one cluster; (3) Determine whether the clustering end condition is reached. If not, repeat the above (1) to (3). If so, stop clustering to obtain the hierarchical clustering result. Among them, the clustering end condition can be that the number of clustering clusters reaches the preset number, or the number of clustering times reaches the preset number, which can be specifically set according to actual experience and needs.

[0115] It can be understood that through the above agglomerative clustering method, a hierarchical nested clustering tree can be created from bottom to top. After the hierarchical tree is constructed, hierarchical division can be performed according to the distance between tree nodes. The smaller the distance, the higher the similarity between tree nodes. Select the position with the largest distance between tree nodes for division to obtain a hierarchical clustering result with small intra-class distance and large inter-class distance. Among them, the distance between tree nodes can be, but is not limited to, the Euclidean distance.

[0116] As Figure 4 shown is the hierarchical clustering result obtained by performing hierarchical clustering based on agglomerative clustering, where the multiple first categories include the first-level categories (i.e., classification A and classification B) and the second-level categories (i.e., classification A / a, classification A / b, classification B / a, classification B / b).

[0117] After the classification division, for each clustering cluster of the first category, the word frequency of the segmented words corresponding to the historical feedback texts can be calculated, and the segmented words with a word frequency greater than the preset word frequency are used as the classification words of the first category, and the total number and distribution of each classification word are stored, where the distribution of the classification word is the number of historical feedback texts containing the classification word in the clustering cluster of the first category, so that the classification word information of each first category can be obtained.

[0118] In an exemplary embodiment, to improve the accuracy of classification, after determining the classification word information of each first category, the method may further include:

[0119] Display the classification word information of each of the first categories;

[0120] In response to a configuration instruction for the classification word information of any one of the first categories, update the classification word information of any one of the first categories.

[0121] Among them, the configuration instruction may include adding, modifying, deleting classifications and classification words. As Figure 5 shown in the schematic diagram of the configuration page provided by the embodiment of the present invention, business personnel can configure the classification word information of the first category displayed in the configuration page according to actual needs, which is beneficial to improving the accuracy of the segmented information of the first category, and further beneficial to improving the accuracy and flexibility of online classification.

[0122] Please continue to refer to Figure 4 , where the training of the classification model may include the following steps:

[0123] Input the segmented words of the historical feedback texts in the historical feedback text set and the first category corresponding to the historical feedback texts into the classification model for category prediction to obtain a category prediction result;

[0124] According to the difference between the category prediction result corresponding to each historical feedback text and the reference category corresponding to the historical feedback text, adjust the model parameters of the classification model for iterative training until the preset training end condition is met;

[0125] Among them, the reference category corresponding to the historical feedback text is the sub-classification category of the corresponding first category. In addition, when the classification model performs category prediction, it is used to perform vector embedding on the historical feedback text according to the first category of the historical feedback text, each segmented word of the historical feedback text corresponding sentence in the historical feedback text and the position in the sentence, and predict the category of the historical feedback text according to the result of the vector embedding.

[0126] In a specific implementation, the loss value can be calculated based on the difference between the category prediction result corresponding to each historical feedback text and the reference category corresponding to the historical feedback text. Then, based on the loss value, the model parameters of the classification model are adjusted in the reverse direction, and the iterative training is continued according to the adjusted model parameters. Among them, the preset training end condition can be that the loss value reaches the minimum, or the number of iterations reaches a preset iteration number threshold. For example, when the number of iterations reaches 100 times, the iteration stops.

[0127] In the above embodiment, the machine learning algorithm is used to effectively capture the high-frequency information in the parent classification. Then, when training the classification model based on the information of the parent classification, the influence of the long-tail data distribution form on the training effect of the classification model can be weakened, and the accuracy of the classification model is greatly improved by combining the fusion of multi-dimensional information.

[0128] In an exemplary implementation manner, in order to accelerate the training speed of the classification model and improve the accuracy of the classification after training, as Figure 4 shown, when inputting the segmented words of the historical feedback text in the historical feedback text set and the first category corresponding to the historical feedback text into the classification model for category prediction to obtain the category prediction result, it may include:

[0129] Input the segmented words of the historical feedback text in the historical feedback text set and the first category corresponding to the historical feedback text into the embedding layer of the classification model for vector embedding;

[0130] Perform convolution on the result of the vector embedding through the convolution layer of the classification model, and input the result of the convolution into the self-attention encoding layer of the classification model for encoding;

[0131] Perform category prediction on the result of the encoding through the classification layer of the classification model to obtain the category prediction result.

[0132] Among them, the embedding layer can use the Embedding mapping relationship of the BERT model for vector embedding, the convolution layer can use the convolution layer of the TextCNN model, the self-attention encoding layer can use the Transformer encoding module, and the classification layer can use the Softmax activation function. Among them, the TextCNN model is the application of the convolutional neural network CNN in text tasks, which converts the discrete segmented results into continuous vectors in the way of word embedding, and retains the structure of the convolutional layer and pooling layer of the traditional neural network; the BERT model is based on the bidirectional encoding representation of Transformer, uses mask pre-training to enhance the attention of words, and uses sentence prediction to enhance the inter-sentence relationship.

[0133] Corresponding to the feedback information classification methods provided in the above several embodiments, an embodiment of the present invention further provides a feedback information classification device. Since the feedback information classification device provided in the embodiment of the present invention corresponds to the feedback information classification methods provided in the above several embodiments, the implementation manners of the foregoing feedback information classification methods are also applicable to the feedback information classification device provided in this embodiment, and will not be described in detail in this embodiment.

[0134] Please refer to Figure 6 , which shows a schematic structural diagram of a feedback information classification device provided in an embodiment of the present invention. The feedback information classification device 600 has the function of implementing the feedback information classification method in the above method embodiment. The function can be implemented by hardware or by hardware executing corresponding software. As Figure 6 shown, the feedback information classification device 600 may include:

[0135] A word segmentation module 610, configured to determine a feedback text corresponding to the newly added feedback information, and perform word segmentation processing on the feedback text to obtain a plurality of segmented words;

[0136] A first classification module 620, configured to determine the confidence of the newly added feedback information belonging to each of the first categories according to the term frequency-inverse document frequency of each segmented word corresponding to a plurality of first categories among the plurality of segmented words, and determine the target first category to which the newly added feedback information belongs according to the first category with the highest confidence;

[0137] A vector embedding module 630, configured to perform vector embedding on the plurality of segmented words according to the target first category, each sentence corresponding to each segmented word in the feedback text, and the position in the sentence;

[0138] A second classification module 640, configured to encode the result of the vector embedding, and perform classification according to the result of the encoding to obtain the target second category to which the newly added feedback information belongs; the target second category is a sub-classification category of the target first category;

[0139] A hierarchical classification result generation module 650, configured to generate a multi-level classification result corresponding to the newly added feedback information according to the target second category, the target first category, and the hierarchical relationship between the target second category and the target first category.

[0140] In an exemplary implementation manner, the second classification module 640 includes:

[0141] A first convolution module, configured to perform convolution processing on the result of the vector embedding to obtain a feature vector;

[0142] The first encoding module is used to encode the feature vector based on the self-attention mechanism to obtain an encoded feature vector;

[0143] The first classification sub-module is used to classify according to the encoded feature vector to obtain the target second category to which the new feedback information belongs.

[0144] In an exemplary embodiment, the first classification module 620 includes:

[0145] The word frequency determination module is used to determine the word frequency of each segmented word in the feedback text among the multiple segmented words;

[0146] The inverse document frequency determination module is used to determine the inverse document frequency of each segmented word corresponding to each first category according to the classification word information corresponding to each first category;

[0147] The word frequency-inverse document frequency determination module is used to determine the word frequency-inverse document frequency of each segmented word corresponding to each first category according to the word frequency of each segmented word and the inverse document frequency corresponding to each first category;

[0148] The category confidence determination module is used to, for each first category, determine the sum value of the word frequency-inverse document frequencies of the multiple segmented words corresponding to the first category to obtain the confidence that the new feedback information belongs to the first category;

[0149] The second classification sub-module is used to determine the target first category to which the new feedback information belongs according to the first category with the highest confidence among the multiple first categories.

[0150] In an exemplary embodiment, the device further includes:

[0151] The sample segmentation module is used to perform segmentation processing on the historical feedback texts in the historical feedback text set to obtain the segmented words corresponding to each historical feedback text;

[0152] The hierarchical clustering module is used to perform hierarchical clustering on the historical feedback text set according to the segmented words corresponding to each historical feedback text to obtain a hierarchical clustering result; the hierarchical clustering result includes the multiple first categories and the clustering clusters of each first category;

[0153] The first determination module is used to determine the classification word information of each first category according to the segmented words corresponding to the historical feedback texts in the clustering cluster of each first category.

[0154] In an exemplary embodiment, the device further includes:

[0155] A category prediction module, which is used to input the segmented words of the historical feedback texts in the historical feedback text set and the first category corresponding to the historical feedback text into a classification model for category prediction to obtain a category prediction result; wherein, the classification model is used to perform vector embedding on the historical feedback text according to the first category of the historical feedback text, each segmented word of the historical feedback text corresponding sentence in the historical feedback text, and its position in the sentence, and predict the category of the historical feedback text according to the result of the vector embedding;

[0156] A model training module, which is used to adjust the model parameters of the classification model according to the difference between the category prediction result corresponding to each historical feedback text and the reference category corresponding to the historical feedback text, and perform iterative training until a preset training end condition is met; wherein, the reference category corresponding to the historical feedback text is a sub-classification category of the corresponding first category.

[0157] In an exemplary embodiment, the category prediction module includes:

[0158] A second vector embedding module, which is used to input the segmented words of the historical feedback texts in the historical feedback text set and the first category corresponding to the historical feedback text into the embedding layer of the classification model for the vector embedding;

[0159] A convolutional encoding module, which is used to perform convolution on the result of the vector embedding through the convolutional layer of the classification model, and input the result of the convolution into the self-attention encoding layer of the classification model for encoding;

[0160] A category prediction sub-module, which is used to perform category prediction on the result of the encoding through the classification layer of the classification model to obtain a category prediction result.

[0161] In an exemplary embodiment, the device further includes:

[0162] A display module, which is used to display the classification word information of each first category;

[0163] An update module, which is used to update the classification word information of any first category in response to a configuration instruction for the classification word information of any first category.

[0164] It should be noted that for the device provided in the above embodiments, when implementing its functions, only the above-mentioned division of each functional module is used for illustration. In practical applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the device provided in the above embodiments and the method embodiments belong to the same concept, and the specific implementation process is detailed in the method embodiments, which will not be repeated here.

[0165] An embodiment of the present invention provides an electronic device, which includes a processor and a memory. At least one instruction or at least one program segment is stored in the memory, and the at least one instruction or the at least one program segment is loaded and executed by the processor to implement the feedback information classification method provided in the above method embodiment.

[0166] The memory can be used to store software programs and modules. By running the software programs and modules stored in the memory, the processor can execute various functional applications and the classification of feedback information. The memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for functions, etc.; the data storage area can store data created according to the use of the device. In addition, the memory can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices. Correspondingly, the memory can also include a memory controller to provide the processor with access to the memory.

[0167] The method embodiment provided by the embodiment of the present invention can be executed on a computer terminal, a server, or a similar computing device. Taking the operation on a server as an example, Figure 7 is a hardware structure block diagram of a server for running a feedback information classification method provided by an embodiment of the present invention. As Figure 7 shown, the server 700 may vary greatly due to configuration or performance differences. It may include one or more central processing units (CPUs) 710 (the processor 710 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 730 for storing data, and one or more storage media 720 for storing application programs 723 or data 722 (such as one or more mass storage devices). Among them, the memory 730 and the storage media 720 can be short-term storage or persistent storage. The program stored in the storage media 720 may include one or more modules, and each module may include a series of instruction operations on the server. Further, the central processor 710 can be set to communicate with the storage media 720 and execute a series of instruction operations in the storage media 720 on the server 700. The server 700 may also include one or more power supplies 760, one or more wired or wireless network interfaces 750, one or more input / output interfaces 740, and / or one or more operating systems 721, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.

[0168] The input / output interface 740 can be used to receive or send data via a network. Specific examples of the above-mentioned network may include a wireless network provided by a communication provider of the server 700. In one example, the input / output interface 740 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the input / output interface 740 can be a Radio Frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0169] Those of ordinary skill in the art can understand that Figure 7 The structure shown is only schematic and does not limit the structure of the above-mentioned electronic device. For example, the server 700 may also include more or fewer components than those shown Figure 7 shown, or have a different configuration from that Figure 7 shown.

[0170] Embodiments of the present invention also provide a computer-readable storage medium, which can be disposed in an electronic device to store at least one instruction or at least one segment of a program related to implementing an image detection method. The at least one instruction or the at least one segment of the program is loaded and executed by the processor to implement the feedback information classification method provided by the above method embodiments.

[0171] Embodiments of the present invention also provide a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. The processor of the electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the electronic device executes the above-mentioned feedback information classification method.

[0172] Optionally, in this embodiment, the above storage medium may include, but is not limited to: various media such as USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks or optical discs that can store program codes.

[0173] It should be noted that: The above order of the embodiments of the present invention is only for description and does not represent the superiority or inferiority of the embodiments. And the above specific embodiments of this specification have been described. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the drawings do not necessarily require the specific order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0174] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the apparatus embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and for the relevant parts, reference can be made to the partial description of the method embodiments.

[0175] Those of ordinary skill in the art can understand that all or part of the steps to implement the above embodiments can be completed by hardware, or can be completed by a program instructing the relevant hardware. The program can be stored in a computer-readable storage medium. The above-mentioned storage medium can be a read-only memory, a magnetic disk, an optical disk, etc.

[0176] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for classifying feedback information, characterized in that, The method includes: Determine the feedback text corresponding to the newly added feedback information, and perform word segmentation on the feedback text to obtain multiple segmented words; Determine the word frequency of each segmented word in the multiple segmented words in the feedback text; according to the classification word information corresponding to each first category, determine the inverse document frequency of each segmented word corresponding to each first category; according to the word frequency of each segmented word and the inverse document frequency corresponding to each first category, determine the word frequency-inverse document frequency of each segmented word corresponding to each first category; For each first category, determine the sum value of the word frequency-inverse document frequencies of the multiple segmented words corresponding to the first category to obtain the confidence level that the newly added feedback information belongs to the first category; according to the first category with the highest confidence level among the multiple first categories, determine the target first category to which the newly added feedback information belongs; According to the target first category, each sentence corresponding to each segmented word in the feedback text, and the position in the sentence, perform vector embedding on the multiple segmented words; Encode the result of the vector embedding, and perform classification according to the result of the encoding to obtain the target second category to which the newly added feedback information belongs; the target second category is a sub-classification category of the target first category; Generate a multi-level classification result corresponding to the newly added feedback information according to the target second category, the target first category, and the hierarchical relationship between the target second category and the target first category.

2. The method for classifying feedback information according to claim 1, characterized in that, The encoding the result of the vector embedding and performing classification according to the result of the encoding to obtain the target second category to which the newly added feedback information belongs includes: Perform convolution processing on the result of the vector embedding to obtain a feature vector; Encode the feature vector based on the self-attention mechanism to obtain an encoded feature vector; Perform classification according to the encoded feature vector to obtain the target second category to which the newly added feedback information belongs.

3. The method for classifying feedback information according to claim 1, characterized in that, The method further includes: Perform word segmentation on the historical feedback texts in the historical feedback text set to obtain the segmented words corresponding to each historical feedback text; Perform hierarchical clustering on the historical feedback text set according to the segmented words corresponding to each historical feedback text to obtain a hierarchical clustering result; the hierarchical clustering result includes the multiple first categories and the clustering clusters of each first category; Determine the classification word information of each first category according to the segmented words corresponding to the historical feedback texts in the clustering cluster of each first category.

4. The method for classifying feedback information according to claim 3, characterized in that, The method further includes: Input the segmented words of the historical feedback texts in the historical feedback text set and the first category corresponding to the historical feedback texts into a classification model for category prediction to obtain a category prediction result; wherein, the classification model is used to perform vector embedding on the historical feedback text according to the first category of the historical feedback text, each sentence corresponding to each segmented word in the historical feedback text, and the position in the sentence, and predict the category of the historical feedback text according to the result of the vector embedding; Adjust the model parameters of the classification model according to the difference between the predicted category result corresponding to each historical feedback text and the reference category corresponding to the historical feedback text, and perform iterative training until the preset training end condition is met; wherein, the reference category corresponding to the historical feedback text is a sub-classification category of the corresponding first category.

5. The method for classifying feedback information according to claim 4, characterized in that, The step of inputting the segmented words of the historical feedback texts in the historical feedback text set and the corresponding first category of the historical feedback texts into a classification model for category prediction to obtain a category prediction result includes: Input the segmented words of the historical feedback texts in the historical feedback text set and the corresponding first category of the historical feedback texts into the embedding layer of the classification model for vector embedding; Perform convolution on the result of the vector embedding through the convolution layer of the classification model, and input the result of the convolution into the self-attention encoding layer of the classification model for encoding; Perform category prediction on the result of the encoding through the classification layer of the classification model to obtain a category prediction result.

6. The method for classifying feedback information according to claim 4 or 5, characterized in that, Perform vector embedding on the plurality of segmented words according to the target first category, the sentence corresponding to each segmented word in the feedback text, and the position in the sentence; And determine the target second category to which the new feedback information belongs according to the result of the vector embedding, including: Input the target first category and the plurality of segmented words into the classification model to obtain a classification result; The classification result indicates the target second category to which the new feedback information belongs.

7. The method for classifying feedback information according to claim 3, characterized in that, After determining the classification word information of each first category, the method further includes: Display the classification word information of each first category; In response to a configuration instruction for the classification word information of any one of the first categories, update the classification word information of any one of the first categories.

8. A device for classifying feedback information, characterized in that, The device includes: A word segmentation module, configured to determine a feedback text corresponding to new feedback information, and perform word segmentation processing on the feedback text to obtain a plurality of segmented words; A first classification module, configured to determine the word frequency of each segmented word in the plurality of segmented words in the feedback text; according to the classification word information corresponding to each first category, determine the inverse document frequency of each segmented word corresponding to each first category; according to the word frequency of each segmented word and the inverse document frequency corresponding to each first category, determine the term frequency-inverse document frequency of each segmented word corresponding to each first category; for each first category, determine the sum value of the term frequency-inverse document frequencies of the plurality of segmented words corresponding to the first category to obtain the confidence level that the new feedback information belongs to the first category; and determine the target first category to which the new feedback information belongs according to the first category with the highest confidence level among the plurality of first categories; A vector embedding module, configured to perform vector embedding on the plurality of segmented words according to the target first category, the sentence corresponding to each segmented word in the feedback text, and the position in the sentence; A second classification module, configured to encode the result of the vector embedding, and classify according to the result of the encoding to obtain a target second category to which the new feedback information belongs; the target second category is a sub-classification category of the target first category; A hierarchical classification result generation module, configured to generate a multi-level classification result corresponding to the new feedback information according to the target second category, the target first category, and the hierarchical relationship between the target second category and the target first category.

9. An electronic device, characterized in that, It includes a processor and a memory, and at least one instruction or at least one program segment is stored in the memory, and the at least one instruction or the at least one program segment is loaded and executed by the processor to implement the feedback information classification method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, At least one instruction or at least one program segment is stored in the computer-readable storage medium, and the at least one instruction or the at least one program segment is loaded and executed by a processor to implement the feedback information classification method according to any one of claims 1 to 7.

11. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the feedback information classification method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Text classification method and text classification device

    CN105005589A

  • Text classification method and device

    CN110399488A

  • Text data multi-level classification method and device, electronic equipment and storage medium

    CN110781292A