Text Classification Method, Apparatus, Device, and Storage Medium

By segmenting and vectorizing the call data between customers and agents, and using the BERT language model and the fully connected neural network model for fusion processing, the problem of failing to fully utilize interactive information and semantic information in the prior art is solved, and the accuracy of text classification is improved.

CN115248862BActive Publication Date: 2025-05-27PING AN TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210968439.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-12
Publication Date
2025-05-27
Estimated Expiration
2042-08-12

AI Technical Summary

Technical Problem

The prior art fails to fully utilize the interactive information and front-and-back semantic information of customers and agents in the text classification of dialogue content, resulting in insufficient classification accuracy.

Method used

By obtaining the call data between customers and agents, segmenting and adding time tags, using pre-trained text classification models, including BERT language model and fully connected neural network model, vectorized encoding and fusion processing are performed, and classification results are corrected using interactive information and semantic information.

Benefits of technology

The accuracy of text classification results is improved, and the interactive information between customers and agents is analyzed by correlation, and the classification results are corrected using the semantic information before and after.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115248862B_ABST
    Figure CN115248862B_ABST
Patent Text Reader

Abstract

This application relates to the field of natural language processing, and specifically discloses a text classification method, apparatus, device, and storage medium. The text classification method includes: obtaining call data, splitting the call data into first call data of a customer and second call data of an agent, where the first call data and the second call data include time tags; performing vector quantization encoding on the first call data and the second call data in a BERT language model to obtain a first representation matrix and a second representation matrix including token vectors; performing fusion processing on the first representation matrix and the second representation matrix according to the token vectors in a fully connected neural network model to obtain a first target vector corresponding to the first representation matrix and a second target vector corresponding to the second representation matrix; determining a classification label for the first call data according to the first target vector, and determining a classification label for the second call data according to the second target vector. In this way, the accuracy of text classification is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of natural language processing, and in particular, to a text classification method, apparatus, device, and storage medium. Background Art

[0002] Currently, with the explosive growth of information, the text classification method by manual annotation takes a long time and is easily affected by the subjective consciousness of the annotator. Therefore, it becomes practically significant to automatically implement text classification by computer. With the development of computer technology, people can conduct quantitative research on language text information with the support of a computer, and hand over the repetitive and boring language text annotation task to the computer, thereby effectively overcoming the above problems.

[0003] In the text classification of conversation content, not only the customer's words need to be text classified, but also the agent's words need to be text classified, and there is a correlation between the customer's words and the agent's words. Existing technical solutions separate the agent's words and the customer's words and perform text classification separately. This approach does not make full use of the interaction information between the customer and the agent, and at the same time does not utilize the semantic information before and after the words spoken by the customer and the agent, resulting in insufficient accuracy. Summary of the Invention

[0004] This application provides a text classification method, apparatus, device, and storage medium, which are used to perform correlation analysis on the interaction information between a customer and an agent in a computer, and can correct the classification result by using the semantic information before and after the words spoken by the customer and the agent, thereby improving the accuracy of the text classification result.

[0005] In a first aspect, this application provides a text classification method, the method includes: obtaining call data between a customer and an agent, splitting the call data and adding time tags to obtain first call data of the customer and second call data of the agent, where the first call data and the second call data include time tags; inputting the first call data and the second call data into a pre-trained text classification model, where the text classification model at least includes a BERT language model and a fully connected neural network model; wherein, the BERT language model performs vector quantization encoding on the first call data and the second call data to obtain a first representation matrix and a second representation matrix, where the first representation matrix and the second representation matrix include the marker vectors corresponding to the time tags; the fully connected neural network model performs fusion processing on the first representation matrix and the second representation matrix according to the marker vectors to obtain a first target vector corresponding to the first representation matrix and a second target vector corresponding to the second representation matrix; determining a classification label for the first call data according to the first target vector, and determining a classification label for the second call data according to the second target vector.

[0006] In a second aspect, the present application provides a text classification device, which includes: a data classification module, a data processing module, and a label determination module;

[0007] The data classification module is configured to obtain the call data between the customer and the agent, segment the call data and add time tags to obtain the first call data of the customer and the second call data of the agent, where the first call data and the second call data include time tags;

[0008] The data processing module is configured to input the first call data and the second call data into a pre-trained text classification model, where the text classification model at least includes a BERT language model and a fully connected neural network model; wherein, the BERT language model performs vector quantization encoding on the first call data and the second call data to obtain a first representation matrix and a second representation matrix, and the first representation matrix and the second representation matrix include the marker vectors corresponding to the time tags; the fully connected neural network model performs fusion processing on the first representation matrix and the second representation matrix according to the marker vectors to obtain a first target vector corresponding to the first representation matrix and a second target vector corresponding to the second representation matrix;

[0009] The label determination module is configured to determine the classification label of the first call data according to the first target vector, and determine the classification label of the second call data according to the second target vector.

[0010] In a third aspect, the present application provides a computer device, which includes a memory and a processor;

[0011] The memory is used to store a computer program;

[0012] The processor is configured to execute the computer program and implement any one of the text classification methods provided in the embodiments of the present application when executing the computer program.

[0013] In a fourth aspect, the present application provides a computer-readable storage medium, characterized in that the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor is enabled to implement any one of the text classification methods provided in the embodiments of the present application.

[0014] The present application discloses a text classification method, apparatus, device, and storage medium. The method includes: obtaining call data between a customer and an agent, segmenting the call data and adding time tags to obtain the first call data of the customer and the second call data of the agent, where the first call data and the second call data include time tags; inputting the first call data and the second call data with time tags into a pre-trained text classification model, where the text classification model at least includes a BERT language model and a fully connected neural network model; wherein, the BERT language model performs vectorized encoding on the first call data and the second call data to obtain a first representation matrix and a second representation matrix, and the first representation matrix and the second representation matrix include marker vectors corresponding to the time tags; the fully connected neural network model performs fusion processing on the first representation matrix and the second representation matrix according to the marker vectors to obtain a first target vector corresponding to the first representation matrix and a second target vector corresponding to the second representation matrix; determining a classification label for the first call data according to the first target vector, and determining a classification label for the second call data according to the second target vector. Through the text classification method provided by the embodiments of the present application, after digitizing and encoding the call data of the customer and the agent respectively, the front and back semantic information in the interaction information between the customer and the agent is fused and processed in a pre-trained text classification model, and the classification result is corrected according to the interaction information between the customer and the agent, that is, the automatic text classification of the call data is realized, and the accuracy of text classification is also improved. Description of the Drawings

[0015] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0016] Figure 1 It is an application scenario diagram of a text classification method provided by an embodiment of the present application;

[0017] Figure 2 It is a schematic flowchart of a text classification method provided by an embodiment of the present application;

[0018] Figure 3 It is a schematic block diagram of a text classification model provided by an embodiment of the present application;

[0019] Figure 4 It is a schematic flowchart of a text classification method provided by an embodiment of the present application;

[0020] Figure 5 It is a schematic block diagram of a text classification model provided by an embodiment of the present application;

[0021] Figure 6 It is a schematic block diagram of a text classification device provided by an embodiment of the present application;

[0022] Figure 7 It is a schematic structural block diagram of a computer device provided by an embodiment of the present application. Detailed implementation manners

[0023] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0024] The flowcharts shown in the accompanying drawings are only illustrative, not necessarily including all contents and operations / steps, nor necessarily executed in the described order. For example, some operations / steps can also be decomposed, combined, or partially merged. Therefore, the actual execution order may be changed according to the actual situation.

[0025] It should be understood that the terms used in the specification of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the specification of the present application and the appended claims, unless otherwise clearly specified in the context, the singular forms "a", "an", and "the" are intended to include the plural forms.

[0026] It should also be understood that the term " / and / " used in the specification of the present application and the appended claims refers to any combination and all possible combinations of one or more of the related listed items, and includes these combinations.

[0027] Next, some embodiments of the present application will be described in detail in conjunction with the accompanying drawings. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0028] Please refer to Figure 1 , Figure 1 which shows an application scenario diagram of a text classification method provided by an embodiment of the present application. As Figure 1 shown, the method of the embodiment of the present application can be applied to a server, specifically to the server side of an application program for text classification. The server side runs in the server and is used to obtain call data uploaded by a client installed with the corresponding application program, and is also used to send the target information generated by text classification according to the call data to the terminal. In addition, the client runs in the terminal. Among them, the terminal and the server can be communicatively connected through a wireless network.

[0029] The project provider obtains the text classification result through an application program, such as generating formatted text based on call data or adding classification labels to call data. The application program includes a client and a server. The client is installed in the terminal device for the project provider to use, and the server is installed in the server. The terminal device and the server are connected through network communication.

[0030] When installing the text classification application program in the terminal device, the terminal device needs to authorize corresponding permissions. For example, permissions such as obtaining basic attribute information, device number, call data, and network information can be obtained.

[0031] Among them, the server can be an independent server, a server cluster, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. The terminal can be an electronic device such as a mobile phone, a tablet computer, a notebook computer, a desktop computer, a personal digital assistant, and a wearable device.

[0032] It should also be noted that the embodiments of the present application can obtain and process relevant data based on artificial intelligence technology. For example, the first representation matrix and the second representation matrix are fused according to time tags through artificial intelligence. Among them, Artificial Intelligence (AI) is a theory, method, technology, and application system that uses a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.

[0033] Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technologies, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0034] Please refer to Figure 2 , Figure 2 is a schematic flowchart of a text classification method provided by an embodiment of the present application. As Figure 2 shown, the specific steps of the text classification method provided by the embodiment of the present application include: S101 - S103.

[0035] S101. Obtain the call data between the customer and the agent, segment the call data and add time tags to obtain the customer's first call data and the agent's second call data. The first call data and the second call data include time tags.

[0036] Exemplarily, in project cooperation, the project provider usually needs to collect call data of the project conversation process between the customer and the agent. Based on this call data, the project provider can evaluate the feasibility of the project, standardize the process of project cooperation, and can also adjust the management system of the agent to improve the business ability of the agent. Therefore, after obtaining the recording permission of the customer, the project provider records and stores the call data generated in the communication of project cooperation. To improve the readability of the call data, the call data is segmented into the first call data of the customer and the second call data of the agent through voice segmentation technology. Voice segmentation can be performed when recording the call data. For example, the call data is segmented according to the vibration signals of the microphones of the customer and the agent. Voice segmentation can also be performed after the call data is collected. For example, voice segmentation is performed according to the timbres of the customer and the agent. After voice segmentation, each sentence of the customer is a piece of the first call data, and each sentence of the agent is a piece of the second call data. During the voice segmentation process, time tags also need to be added to the first call data and the second call data in chronological order to record the sequence of the first call data and the second call data.

[0037] In some embodiments, some common and clearly described questions and their corresponding answers are configured for the agent. In this way, it is beneficial for voice segmentation and also beneficial for the computer to understand the content to be expressed in the agent's words.

[0038] In some embodiments, the official time corresponding to the call process is recorded in the call data. During the voice segmentation process, time tags are generated according to the official time. In this way, it is beneficial to save time information so as to save more comprehensive transaction information for the project provider.

[0039] In the embodiments of the present application, the call data is segmented into the first call data and the second call data, and the segmented data is more conducive to being imported into the processing model, improving the overall processing speed of the system.

[0040] S102. Input the first call data and the second call data into a pre-trained text classification model, where the text classification model at least includes a BERT language model and a fully connected neural network model; wherein, the BERT language model performs vector quantization encoding on the first call data and the second call data to obtain a first representation matrix and a second representation matrix, and the first representation matrix and the second representation matrix include the marked vectors corresponding to the time tags; the fully connected neural network model performs fusion processing on the first representation matrix and the second representation matrix according to the time tags to obtain a first target vector corresponding to the first representation matrix and a second target vector corresponding to the second representation matrix.

[0041] Please refer to Figure 3 , Figure 3Shows a schematic block diagram of a text classification model. As Figure 3 shown, the text classification model at least includes a BERT (Bidirectional Encoder Representation from Transformers) language model and a fully connected neural network model. The embedding layer of the BERT language model is connected to the fully connected neural network model.

[0042] In an embodiment of the present application, it is also necessary to train the text classification model. First, it is necessary to construct a training sample set for model training. The training sample set is historical call data including manually labeled classification labels. The classification labels include first-level labels, second-level labels, and third-level labels. Among them, the first-level labels include multiple second-level labels, and the second-level labels include multiple third-level labels.

[0043] Input the training sample set into the text classification model. The BERT language model converts each word in the historical call data into a vector through the embedding layer. Specifically, each word is converted into a 768-dimensional vector, and a corresponding first representation matrix or second representation matrix is generated according to each sentence in the historical call data. The fully connected neural network model calculates the first mapping relationship among the first representation matrix, the second representation matrix, and the manually labeled classification labels. The first mapping relationship is used to reflect the label weight distribution corresponding to the first representation matrix or the second representation matrix.

[0044] Import the first call data and the second call data into the pre-trained BERT language model. The first call data and the second call data consist of a complete sentence. The embedding layer of the BERT language model performs mask replacement on the words in the sentences of the first call data and the second call data. The steps of mask replacement are as follows: Extract 15% of the words from the first call data corresponding to each sentence for replacement. There are three replacement options for the selected words: 1. Replace with MASK, with a probability of 80%; 2. Replace with other words, with a probability of 10%; 3. Remain unchanged, with a probability of 10%. After completing the mask replacement, output the first mask vector matrix corresponding to the first call data and the second mask vector matrix corresponding to the second call data. Also use the pre-trained BERT language model to generate a marker vector according to the time labels of the first call data and the second call data, and implant the marker vector at the beginning of the corresponding first vector matrix or second vector matrix. It is also possible to implant the marker vector at the end or other positions of the first vector matrix or the second vector matrix to generate the first representation matrix corresponding to the first call data with dimensions of n×768 and the second representation matrix corresponding to the second call data, where n is the number of words in each sentence.

[0045] Input the first representation matrix and the second representation matrix output by the embedding layer of the BERT model into a pre-trained fully connected neural network model. The fully connected neural network model performs fusion processing on the first representation matrix and the second representation matrix according to the time tag to correct the weight values of each vector in the first representation matrix and the second representation matrix, and determines the first target vector corresponding to the first representation matrix and the second target vector corresponding to the second representation matrix according to the weight values of each vector in the first representation matrix and the second representation matrix.

[0046] Specifically, obtain the first correlation range of the time tag. For example, the first correlation range includes the three adjacent time tags before and after the current time tag in chronological order. The fully neural network judges the correlation relationship between each matrix according to the marked vector, and there is a corresponding relationship between the marked vector and the time tag. Load the pre-trained fully connected neural network model, and perform fusion processing on the first representation matrix and the second representation matrix corresponding to the target time tag within the first correlation range. Generally, determine the current time tag according to chronological order, and determine the target time tag according to the current time tag and the first correlation range.

[0047] Obtain the target first representation matrix and the target second representation matrix corresponding to the target time tag; perform fusion processing on the target first representation matrix and the target second representation matrix in the fully connected neural network model. Specifically, correct the weight values of each vector in the target first representation matrix or the target second representation matrix corresponding to the current time tag according to the target first representation matrix and the target second representation matrix corresponding to the non-current time tag. After the fusion processing, the vector with the largest weighted value in the first representation matrix is the first target vector, and the vector with the largest weighted value in the second representation matrix is the second target vector.

[0048] In this application, the call data is processed by a pre-trained text classification model to digitize the words and sentences in the call data, so that the computer can interpret natural language, and it is also corrected according to chronological order to improve the efficiency and accuracy of text classification.

[0049] S103. Determine the classification label of the first call data according to the first target vector, and determine the classification label of the second call data according to the second target vector.

[0050] Specifically, obtain the pre-edited vector-label relationship table; find the corresponding classification label from the vector relationship table according to the first target vector or the second target vector.

[0051] In some embodiments, after obtaining the classification label, mark the classification label at the corresponding position of the first call data or the second call data.

[0052] In some other embodiments, after obtaining the classification labels, according to the classification labels, the words in the first call data or the second call data are input into a specification document. Exemplarily, the first-level label is location address verification, the second-level labels include video location and nearby landmarks, and the third-level labels include home, car, company, and landmark names. A specification text is generated based on the above first-level, second-level, and third-level labels, and the corresponding content is filled into the specification text.

[0053] The text classification method provided in this application is used to perform correlation analysis on the interaction information between the customer and the agent in a computer, and can correct the classification result by using the semantic information before and after the words spoken by the customer and the agent, thereby improving the accuracy of the text classification result.

[0054] In order to further improve the accuracy of text classification, in this application, another text classification method is also provided. Please refer to Figure 4 , Figure 4 which shows a schematic flowchart of a text classification method. As Figure 4 shown, the specific steps of this text classification method include: S201 - S203.

[0055] S201. Obtain the call data between the customer and the agent, segment the call data and add time tags to obtain the first call data of the customer and the second call data of the agent. The first call data and the second call data include time tags.

[0056] The process of S201 is the same as that of S101. Please refer to the above text and will not be elaborated here.

[0057] S202. Input the first call data and the second call data with time tags into a pre-trained text classification model. The text classification model at least includes a BERT language model, an LSTM (Long Short Term Memory) model, and a fully connected neural network model. Among them, the BERT language model performs vector quantization encoding on the first call data and the second call data to obtain a first representation matrix corresponding to the first call data and a second representation matrix corresponding to the second call data, and performs linear quantization encoding on the second representation matrix to obtain a third representation matrix. The first representation matrix, the second representation matrix, and the third representation matrix include token vectors. The LSTM model obtains the correlation relationship between each third representation matrix according to the token vectors and corrects the third representation matrix according to the correlation relationship to obtain a fourth representation matrix. The fully connected neural network model performs fusion processing on the first representation matrix and the second representation matrix according to the token vectors to obtain a first target vector corresponding to the first representation matrix and a second target vector corresponding to the second representation matrix.

[0058] Please refer to Figure 5 , Figure 5Shows a schematic block diagram of a text classification model. As Figure 5 shown, the text classification model includes a BERT language model, an LSTM model, and a fully connected neural network model. The BERT language model includes an embedding layer and an attention layer. The embedding layer of the BERT language model is connected to the fully connected neural network model, and the attention layer of the BERT language model is connected to the hidden layer of the LSTM model; the hidden layer of the LSTM model is also connected to the fully connected neural network model.

[0059] In the embodiments of the present application, it is also necessary to train the text classification model. First, it is necessary to construct a training sample set for model training. The training sample set is historical call data including manually annotated classification labels. The BERT language model, the LSTM model, and the fully connected neural network model are trained using the training sample set.

[0060] The training sample set is input into the text classification model. The BERT language model converts each word in the historical call data into a vector through the embedding layer. Specifically, each word is converted into a 768-dimensional vector, and a corresponding first representation matrix or second representation matrix is generated according to each statement in the historical call data. Since the statements of the agent in the call data are often statements describing clear problems, the vector dimension of the second representation matrix of the second call data can be further reduced. With the help of the attention layer of the BERT language model, the vector dimension of the second representation matrix is reduced. For example, the vector dimension of the agent statement is reduced to 256 dimensions. The LSTM model obtains the correlation relationship between each third representation matrix according to the marked vector and corrects the third representation matrix according to the correlation relationship to obtain a fourth representation matrix. The fully connected neural network model calculates the second mapping relationship among the first representation matrix, the second representation matrix, the third representation matrix, and the manually annotated classification label. The second mapping relationship is used to reflect the label weight distribution corresponding to the representation matrix.

[0061] The LSTM model in the embodiments of the present application is a bidirectional LSTM neural network model. Therefore, the pre-trained LSTM model can correct the third representation matrix according to the correlation relationship between the third representation matrices.

[0062] Specifically, the third representation matrix is input into a pre-trained LSTM model. In the hidden layer of the LSTM model, the association relationship between the second representation matrices is determined according to the labeled vectors in the third matrix. Since the association relationship between each sentence in the conversation and the several sentences before and after it is the strongest, a second association range of time tags can be preset. For example, the first association range includes the time tags adjacent to the current time tag in the chronological order before and after it. The third representation matrices corresponding to the time tags within the second association range are the target third representation matrices. The hidden layer of the pre-trained LSTM model is used to perform contrastive correction on multiple target third representation matrices, and the corrected target third representation matrix is the fourth representation matrix.

[0063] In the fully connected neural network model, the target first representation matrix, the target second representation matrix, and the target fourth matrix are obtained according to the labeled vectors, and the above three are subjected to fusion processing. Specifically, the weight values of the vectors of the target first representation matrix or the target second representation matrix corresponding to the non-current time tag are corrected according to the target first representation matrix, the target second representation matrix, and the target fourth matrix corresponding to the non-current time tag. After the fusion processing, the vector with the largest weighted value in the first representation matrix is the first target vector, and the vector with the largest weighted value in the second representation matrix is the second target vector.

[0064] S203. Determine the classification label of the first call data according to the first target vector, and determine the classification label of the second call data according to the second target vector.

[0065] The process of S203 is the same as that of S103. Please refer to the above text and will not be elaborated here.

[0066] In the embodiments of the present application, a pre-trained LSTM model is added, the dimension of the second representation matrix is shrunk in the BERT model, the third representation matrix is corrected in the LSTM model, and the input vectors of the fully connected neural network model are increased, thereby improving the accuracy of text classification.

[0067] Please refer to Figure 6 , Figure 6 FIG. is a schematic block diagram of a text classification device provided by an embodiment of the present application. The text classification device 300 is used to execute the foregoing text classification method. Among them, the text classification device can be configured in a server or a terminal.

[0068] Among them, the server can be an independent server, a server cluster, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. The terminal can be an electronic device such as a mobile phone, a tablet computer, a laptop computer, a desktop computer, a customer digital assistant, and a wearable device.

[0069] As Figure 6 shown, the text classification device 300 includes: a data classification module 301, a data processing module 302, and a label determination module 303.

[0070] The data classification module 301 is configured to obtain the call data between the customer and the agent, segment the call data and add time tags to obtain the first call data of the customer and the second call data of the agent, and the first call data and the second call data include time tags.

[0071] The data processing module 302 is configured to input the first call data and the second call data into a pre-trained text classification model, and the text classification model at least includes a BERT language model and a fully connected neural network model; wherein, the BERT language model performs vector quantization encoding on the first call data and the second call data to obtain a first representation matrix corresponding to the first call data and a second representation matrix corresponding to the second call data, and the first representation matrix and the second representation matrix include marker vectors corresponding to time tags; the fully connected neural network model performs fusion processing on the first representation matrix and the second representation matrix according to the marker vectors to obtain a first target vector corresponding to the first representation matrix and a second target vector corresponding to the second representation matrix.

[0072] In some embodiments, the data processing module 302 is further configured to input the first call data and the second call data with time tags added thereto into a pre-trained text classification model, where the text classification model at least includes a BERT language model, an LSTM model, and a fully connected neural network model; wherein, the BERT language model performs vector quantization encoding on the first call data and the second call data to obtain a first representation matrix corresponding to the first call data and a second representation matrix corresponding to the second call data, and performs linear quantization encoding on the second representation matrix to obtain a third representation matrix; the LSTM model obtains the correlation relationship between the third representation matrices according to the token vectors and corrects the third representation matrices according to the correlation relationship, and the first representation matrix, the second representation matrix, and the third representation matrix include the token vectors corresponding to the time tags to obtain a fourth representation matrix; the fully connected neural network model performs fusion processing on the first representation matrix and the second representation matrix according to the token vectors to obtain a first target vector corresponding to the first representation matrix and a second target vector corresponding to the second representation matrix.

[0073] In some embodiments, the data processing module 302 is further configured to use the embedding layer of the pre-trained BERT language model to perform mask replacement on the words of the first call data and the second call data, output a first mask vector matrix corresponding to the first call data and a second mask vector matrix corresponding to the second call data, generate token vectors according to the time tags of the first call data and the second call data, generate a first representation matrix according to the first mask vector matrix and the token vectors, and generate a second representation matrix according to the second mask vector matrix and the token vectors.

[0074] The label determination module 303 is configured to determine the classification label of the first call data according to the first target vector, and determine the classification label of the second call data according to the second target vector.

[0075] In some embodiments, the label determination module 303 is further configured to obtain a pre-edited vector label relationship table; and look up the corresponding classification label from the vector relationship table according to the first target vector or the second target vector.

[0076] It should be noted that those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the above-described text classification device and each module can refer to the corresponding processes in the foregoing text classification method embodiments, and will not be elaborated herein.

[0077] It should be noted that those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the above-described model training device and each module can refer to the corresponding processes in the foregoing text classification method embodiments, and will not be elaborated herein.

[0078] The above-mentioned text classification device can be implemented in the form of a computer program, which can run on a computer device as shown in Figure 7 the following.

[0079] Please refer to Figure 7 , Figure 7 which is a schematic block diagram of the structure of a computer device provided by an embodiment of the present application. The computer device can be a server or a terminal.

[0080] Referring to Figure 7 , the computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the memory can include a storage medium and an internal memory.

[0081] The storage medium can store an operating system and a computer program. The computer program includes program instructions, which when executed, can cause the processor to execute any text classification method provided by an embodiment of the present application.

[0082] The processor is used to provide computing and control capabilities to support the operation of the entire computer device.

[0083] The internal memory provides an environment for the operation of the computer program in the storage medium. When the computer program is executed by the processor, the processor can execute any text classification method. The storage medium can be non-volatile or volatile.

[0084] The network interface is used for network communication, such as sending assigned tasks, etc. Those skilled in the art can understand that Figure 7 the structure shown in

[0085] is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different component layout.

[0086] Exemplarily, in one embodiment, the processor is used to run a computer program stored in the memory to implement the following steps: obtaining call data between a customer and an agent, segmenting the call data and adding time tags to obtain first call data of the customer and second call data of the agent, where the first call data and the second call data include time tags; inputting the first call data and the second call data into a pre-trained text classification model, where the text classification model at least includes a BERT language model and a fully connected neural network model; wherein, the BERT language model performs vector quantization encoding on the first call data and the second call data to obtain a first representation matrix corresponding to the first call data and a second representation matrix corresponding to the second call data, and the first representation matrix and the second representation matrix include marker vectors corresponding to the time tags; the fully connected neural network model performs fusion processing on the first representation matrix and the second representation matrix according to the marker vectors to obtain a first target vector corresponding to the first representation matrix and a second target vector corresponding to the second representation matrix; determining a classification label of the first call data according to the first target vector, and determining a classification label of the second call data according to the second target vector.

[0087] In some embodiments, when the processor is used to implement the BERT language model to perform vector quantization encoding on the first call data and the second call data to obtain a first representation matrix and a second representation matrix, it is further specifically used to implement using the embedding layer of the pre-trained BERT language model to perform mask replacement on the words of the first call data and the second call data, outputting a first mask vector matrix corresponding to the first call data and a second mask vector matrix corresponding to the second call data, generating marker vectors according to the time tags of the first call data and the second call data, generating the first representation matrix according to the first mask vector matrix and the marker vectors, and generating the second representation matrix according to the second mask vector matrix and the marker vectors.

[0088] In some embodiments, the processor is further specifically used to implement using the attention layer of the pre-trained BERT language model to perform linearization encoding on the multi-dimensional vectors of each second representation matrix to generate multiple third representation matrices.

[0089] In some embodiments, when the processor implements determining the classification label of the first call data according to the first target vector and determining the classification label of the second call data according to the second target vector, it is further specifically used to implement obtaining a pre-edited vector label relationship table; looking up the corresponding classification label from the vector relationship table according to the first target vector or the second target vector.

[0090] In some embodiments, the processor is further specifically used to implement obtaining a training sample set, where the training sample set includes historical call data and manually annotated classification labels; using the training sample set to train the BERT language model, the LSTM model and the fully connected neural network model.

[0091] As described above, it is only the specific implementation manner of the present application. However, the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.

Claims

1. A text classification method, characterized in that, the method includes: obtaining the call data between the customer and the agent, segmenting the call data and adding time tags to obtain the first call data of the customer and the second call data of the agent, where the first call data and the second call data include time tags; inputting the first call data and the second call data into a pre-trained text classification model, where the text classification model at least includes a BERT language model and a fully connected neural network model; wherein, the BERT language model performs vector quantization encoding on the first call data and the second call data to obtain a first representation matrix and a second representation matrix, where the first representation matrix and the second representation matrix include the token vectors corresponding to the time tags; the fully connected neural network model performs fusion processing on the first representation matrix and the second representation matrix according to the token vectors to obtain a first target vector corresponding to the first representation matrix and a second target vector corresponding to the second representation matrix; determining the classification label of the first call data according to the first target vector, and determining the classification label of the second call data according to the second target vector; the text classification model further includes an LSTM model; the pre-trained BERT language model also performs linearization encoding on the multi-dimensional vectors of each second representation matrix to obtain a plurality of third representation matrices; the LSTM model obtains the correlation relationship between the third representation matrices according to the token vectors and corrects the third representation matrices according to the correlation relationship to obtain a fourth representation matrix; the fully connected neural network model performs fusion processing on the first representation matrix, the second representation matrix and the fourth representation matrix according to the token vectors, and outputs a first target vector corresponding to the first representation matrix and a second target vector corresponding to the second representation matrix; the BERT language model includes an embedding layer and an attention layer, the embedding layer of the BERT language model is connected to the fully connected neural network model, and the attention layer of the BERT language model is connected to the LSTM model; the LSTM model is also connected to the fully connected neural network model.

2. The text classification method according to claim 1, characterized in that, the BERT language model performs vector quantization encoding on the first call data and the second call data to obtain a first representation matrix and a second representation matrix, where the first representation matrix and the second representation matrix include the token vectors corresponding to the time tags, including: Using the embedding layer of the pre-trained BERT language model, perform masked replacement on the words of the first call data and the second call data, output a first masked vector matrix corresponding to the first call data and a second masked vector matrix corresponding to the second call data, generate the token vector according to the time tag, generate the first representation matrix according to the first masked vector matrix and the token vector, and generate the second representation matrix according to the second masked vector matrix and the token vector.

3. The text classification method according to claim 1, wherein, the method further includes: Using the attention layer of the pre-trained BERT language model, linearly encode the multi-dimensional vectors of each of the second representation matrices to generate a plurality of third representation matrices.

4. The text classification method according to claim 1, wherein, the determining the classification label of the first call data according to the first target vector and the classification label of the second call data according to the second target vector includes: Obtain a pre-edited vector label relationship table; Find the corresponding classification label from the vector relationship table according to the first target vector or the second target vector.

5. The text classification method according to claim 1, wherein, before inputting the first call data and the second call data into a pre-trained text classification model, the method further includes: Obtain a training sample set, where the training sample set includes historical call data and the classification label; Use the training sample set to train the BERT language model, the LSTM model, and the fully connected neural network model.

6. A text classification device, wherein, for performing the text classification method according to any one of claims 1 to 5, the text classification device includes: A data classification module, configured to obtain call data between a customer and an agent, segment the call data and add time tags to obtain first call data of the customer and second call data of the agent, where the first call data and the second call data include time tags; A data processing module, configured to input the first call data and the second call data into a pre-trained text classification model, where the text classification model at least includes a BERT language model and a fully connected neural network model; wherein, the BERT language model performs vectorization encoding on the first call data and the second call data to obtain a first representation matrix and a second representation matrix, and the first representation matrix and the second representation matrix include token vectors corresponding to the time tags; the fully connected neural network model performs fusion processing on the first representation matrix and the second representation matrix according to the token vectors to obtain a first target vector corresponding to the first representation matrix and a second target vector corresponding to the second representation matrix; A label determination module, configured to determine the classification label of the first call data according to the first target vector and the classification label of the second call data according to the second target vector.

7. A computer device, It is characterized in that the computer device includes a memory and a processor; the memory is used for storing a computer program; the processor is used for executing the computer program and implementing the text classification method according to any one of claims 1 to 5 when executing the computer program.

8. A computer-readable storage medium It is characterized in that the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor is caused to implement the text classification method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Content classification method and device, computer equipment and storage medium

    CN110737801A

  • Telephone traffic quality inspection method and device

    CN112580367A