A data processing method, apparatus, electronic device and computer-readable medium
The method uses deep learning models to analyze customer service interactions, predicting positive and negative factors, and classifying them for effective training of new staff, addressing limitations in existing rule-based filtering.
Patent Information
- Application Number
- CN202211261184.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-14
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-10-14
AI Technical Summary
Excellent response cases discovered through rules in the prior art have high limitations, and truly good service cases cannot be effectively screened out, resulting in poor learning results for newly hired customer service personnel.
By receiving data processing requests, obtaining session data and role identification, extracting target role session data from session data based on role identification, and using the percentage model, positive factor model and negative factor model to predict their attribution probability and the probability of containing factors, screening out excellent session data for newcomers to learn.
It has achieved rapid and accurate mining of excellent customer service conversation data, improved the learning effect of newly hired customer service personnel, and improved their work ability.
Smart Images

Figure CN115705359B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular, to a data processing method, apparatus, electronic device, and computer-readable medium. Background Art
[0002] In the Internet era, every company needs customer service to answer various questions for customers. For customer service staff, the improvement of work ability is not only related to personal growth and performance, but also related to the company's image. If the customer service is too unprofessional and often fails to solve problems, new customer service staff are required to view the response session details of other senior customer service staff to learn response skills.
[0003] In the process of implementing this application, the inventors found that there are at least the following technical problems in the prior art:
[0004] The excellent response cases mined by rules have high limitations. For example, by filtering based on service duration or the total number of words during the service, only cases with fewer conversation rounds or some meaningless service cases can be filtered out, and truly good service cases may not be extracted, resulting in poor learning effects for new customer service staff. Summary of the Invention
[0005] In view of this, embodiments of this application provide a data processing method, apparatus, electronic device, and computer-readable medium, which can solve the problem that the excellent response cases mined by existing rules have high limitations. For example, by filtering based on service duration or the total number of words during the service, only cases with fewer conversation rounds or some meaningless service cases can be filtered out, and truly good service cases may not be extracted, resulting in poor learning effects for new customer service staff.
[0006] To achieve the above objective, according to one aspect of the embodiments of this application, a data processing method is provided, including:
[0007] Receiving a data processing request, and obtaining corresponding session data and a role identifier;
[0008] Extracting target role session data from the session data according to the role identifier;
[0009] Predicting the probabilities that the target role session data belongs to a preset category, contains a positive factor, and contains a negative factor;
[0010] Determining and outputting target session data from the target role session data according to the probabilities that it belongs to a preset category, contains a positive factor, and contains a negative factor.
[0011] Optionally, determine and output target session data from the target role session data according to the probability of belonging to a preset category, the probability of including a positive factor, and the probability of including a negative factor, including:
[0012] For each piece of target role session data, determine the corresponding score according to the probability of belonging to a preset category and the preset category score;
[0013] In response to the probability of including a positive factor being greater than a preset threshold and the probability of including a negative factor being less than a preset threshold, determine and output the target session data in the target role session data according to the scores.
[0014] Optionally, determine and output the target session data in the target role session data according to the scores, including:
[0015] In response to there being a score greater than a preset score threshold among the scores, determine and output the target role session data corresponding to the scores greater than the preset score threshold as the target session data.
[0016] Optionally, the data processing method further includes:
[0017] In response to the probability of including a positive factor being less than a preset threshold for each or the probability of including a negative factor being greater than a preset threshold for each, determine that there is no target session data in the target role session data, and end the data processing process.
[0018] Optionally, predict the probability of the target role session data belonging to a preset category, the probability of including a positive factor, and the probability of including a negative factor, including:
[0019] Call a percentile model to input the target role session data into the percentile model, and then output the probability of the target role session data belonging to a preset category;
[0020] Call a positive factor model to input the target role session data into the positive factor model, and then output the probability of including a positive factor in the target role session data;
[0021] Call a negative factor model to input the target role session data into the negative factor model, and then output the probability of including a negative factor in the target role session data.
[0022] Optionally, extract target role session data from the session data according to the role identifier, including:
[0023] Determine the target role according to the role identifier, and then extract the target role session data corresponding to the target role from the session data.
[0024] Optionally, before predicting the probabilities that the session data of the target role belongs to a preset category, contains a positive factor, and contains a negative factor, the method further includes:
[0025] Determine the time corresponding to the session data of the target role, and then splice the target session data according to the time to generate spliced session data;
[0026] Update the session data of the target role by using the spliced session data.
[0027] In addition, the present application further provides a data processing apparatus, including:
[0028] A receiving unit, configured to receive a data processing request and obtain corresponding session data and a role identifier;
[0029] A data extraction unit, configured to extract the session data of the target role from the session data according to the role identifier;
[0030] A probability prediction unit, configured to predict the probabilities that the session data of the target role belongs to a preset category, contains a positive factor, and contains a negative factor;
[0031] An output unit, configured to determine and output the target session data from the session data of the target role according to the probability of belonging to a preset category, the probability of containing a positive factor, and the probability of containing a negative factor.
[0032] Optionally, the output unit is further configured to:
[0033] For each piece of session data of the target role, determine the corresponding score according to the probability of belonging to a preset category and the preset category score;
[0034] In response to the probability of containing a positive factor being greater than a preset threshold and the probability of containing a negative factor being less than a preset threshold, determine and output the target session data in the session data of the target role according to the scores.
[0035] Optionally, the output unit is further configured to:
[0036] In response to there being a score greater than a preset score threshold among the scores, determine and output the session data of the target role corresponding to the score greater than the preset score threshold as the target session data.
[0037] Optionally, the output unit is further configured to:
[0038] In response to the probability of each containing a positive factor being less than a preset threshold or the probability of each containing a negative factor being greater than a preset threshold, determine that there is no target session data in the session data of the target role, and end the data processing process.
[0039] Optionally, the probability prediction unit is further configured to:
[0040] Invoke a percentage model to input the target role session data into the percentage model, and then output the probability that the target role session data belongs to a preset category;
[0041] Invoke a positive factor model to input the target role session data into the positive factor model, and then output the probability that the target role session data contains positive factors;
[0042] Invoke a negative factor model to input the target role session data into the negative factor model, and then output the probability that the target role session data contains negative factors.
[0043] Optionally, the data extraction unit is further configured to:
[0044] Determine the target role according to the role identifier, and then extract the target role session data corresponding to the target role from the session data.
[0045] Optionally, the data processing device further includes a data splicing unit, configured to:
[0046] Determine the time corresponding to the target role session data, and then splice the target session data according to the time to generate spliced session data;
[0047] Update the target role session data with the spliced session data.
[0048] In addition, the present application also provides a data processing electronic device, including: one or more processors; a storage device for storing one or more programs, when the one or more programs are executed by the one or more processors, enabling the one or more processors to implement the data processing method as described above.
[0049] In addition, the present application also provides a computer-readable medium, on which a computer program is stored, and when the program is executed by a processor, it implements the data processing method as described above.
[0050] One embodiment of the above invention has the following advantages or beneficial effects: By receiving a data processing request, the present application obtains the corresponding session data and role identifier; extracts the target role session data from the session data according to the role identifier; predicts the probability that the target role session data belongs to a preset category, the probability of containing positive factors, and the probability of containing negative factors; determines and outputs the target session data from the target role session data according to the probability of belonging to the preset category, the probability of containing positive factors, and the probability of containing negative factors. The target session data is based on the combined screening results of multiple deep learning methods, and can quickly and accurately mine the target session data for new users to learn.
[0051] The further effects of the above non-conventional alternative manners will be described below in conjunction with specific embodiments. Brief Description of the Drawings
[0052] The drawings are used to better understand the present application and do not constitute an improper limitation to the present application. Among them:
[0053] Figure 1 is a schematic diagram of the main process of the data processing method according to the first embodiment of the present application;
[0054] Figure 2 is a schematic diagram of the main process of the data processing method according to the second embodiment of the present application;
[0055] Figure 3 is a schematic diagram of the application scenario of the data processing method according to the third embodiment of the present application;
[0056] Figure 4 is a schematic diagram of the execution logic of the percentage model of the data processing method according to the embodiment of the present application;
[0057] Figure 5 is a schematic diagram of the execution logic of the positive / negative factor model of the data processing method according to the embodiment of the present application;
[0058] Figure 6 is a schematic diagram of the main units of the data processing device according to the embodiment of the present application;
[0059] Figure 7 is an exemplary system architecture diagram to which the embodiment of the present application can be applied;
[0060] Figure 8 is a schematic diagram of the structure of a computer system of a terminal device or a server suitable for implementing the embodiment of the present application. Detailed Embodiments
[0061] The following describes exemplary embodiments of the present application in conjunction with the drawings, including various details of the embodiments of the present application to facilitate understanding. It should be considered that they are only exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present application. Similarly, for clarity and conciseness, the description of well-known functions and structures is omitted below. The acquisition, storage, use, processing, etc. of data in the technical solution of the present application all comply with the relevant regulations of national laws and regulations.
[0062] Figure 1 is a schematic diagram of the main process of the data processing method according to the first embodiment of the present application. As Figure 1 shown, the data processing method includes:
[0063] Step S101: Receive a data processing request and obtain the corresponding session data and role identifier.
[0064] In this embodiment, the execution subject of the data processing method (for example, it can be a server) can receive a data processing request through a wired connection or a wireless connection. Specifically, the data processing request can be a request to filter out excellent session data cases from the session data between the customer service and the customer. The session data to be processed is carried in the data processing request, as shown below:
[0065] Customer / CSR conversation content
[0066] Customer I want to transfer to a human agent
[0067] CSR Hello, yes, customer. Hello, I want to return this product
[0068] CSR Is your pick-up address the same as your delivery address? Have the contact person and phone number changed?
[0069] Customer No
[0070] Customer Are you going to submit the return on your side?
[0071] CSR If submitted before 18:00, it will be reviewed before 22:00 on the same day. If submitted after 18:00, it will be reviewed before 12:00 the next day
[0072] CSR Yes
[0073] CSR I'll submit it first and wait for review
[0074] CSR Please wait patiently. You can follow the processing progress through My - Return / After-sales - Application Record in the XX APP
[0075] CSR The application is done
[0076] Customer You've applied on your side. I don't need to do anything on my side, just wait for XX to come and pick up the goods, right?
[0077] CSR Yes, just wait for a phone call first
[0078] Customer Approximately when will you call?
[0079] CSR It will be reviewed before 22:00 today. It is expected to contact you tomorrow morning. Please keep your phone line clear
[0080] Customer Okay, thank you
[0081] CSR Yes
[0082] The role identifier carried in the data processing request is used to represent which role's session data is to be extracted for processing. The role identifier, for example, can be KF or 01, etc., indicating that the session data of the customer service is to be processed.
[0083] Step S102: Extract the target role session data from the session data according to the role identifier.
[0084] Specifically, extracting the target role session data from the session data according to the role identifier includes: determining the target role according to the role identifier, and then extracting the target role session data corresponding to the target role from the session data.
[0085] Exemplarily, the target role can be, for example, the role of "customer service". When the role identifier is KF, the execution entity can extract the session data of "customer service" from the session data according to the role identifier. Examples of some customer session data (i.e., target role session data) are as follows:
[0086] Customer service: Hello, I'm here.
[0087] Customer service: Is your pick-up address the same as your delivery address? Have there been any changes to the contact person and number?
[0088] Customer service: Submit before 18:00, review before 22:00 on the same day. Submit after 18:00, review before 12:00 the next day.
[0089] Customer service: Yes.
[0090] Customer service: I'll submit it first and wait for review.
[0091] Customer service: Please be patient. You can follow the processing progress through XXAPP - My - Return / After-sales - Application Record.
[0092] Customer service: The application has been completed.
[0093] Customer service: Just wait for a phone call first.
[0094] Customer service: It will be reviewed before 22:00 today. It is expected to contact you tomorrow morning. Please keep your phone unblocked.
[0095] Customer service: Mhm.
[0096] Step S103: Predict the probabilities that the target role session data belongs to a preset category, contains a positive factor, and contains a negative factor.
[0097] Preset categories can include very satisfied, satisfied, average, dissatisfied, and very dissatisfied. The probability that the execution subject predicts that the target role conversation data belongs to the preset category can be achieved in the following way: Segment the target role conversation data to obtain segmented data. For example, segment the target role conversation data "Review before 22:00 today, it is expected to contact you tomorrow morning, please keep the phone unblocked" to obtain "Review before 22:00", "Tomorrow morning", "Contact you", "Please keep the phone unblocked". After obtaining the segmented data, the execution subject can call the pre-set database storing the word-segmentation - category key-value pairs to obtain the corresponding category based on the obtained segmented data "Review before 22:00", "Tomorrow morning", "Contact you", "Please keep the phone unblocked", and then determine the probability of belonging to the preset category based on the obtained corresponding category. Suppose the categories corresponding to the segmented data "Review before 22:00", "Tomorrow morning", "Contact you", "Please keep the phone unblocked" are respectively: "Review before 22:00" - "Very satisfied", "Tomorrow morning" - "Very satisfied", "Contact you" - "Very satisfied", "Please keep the phone unblocked" - "Satisfied", then the execution subject can accordingly obtain that the probability that the target role conversation data "Review before 22:00 today, it is expected to contact you tomorrow morning, please keep the phone unblocked" belongs to the preset category is respectively: the probability of belonging to "Very satisfied" is 75%, the probability of belonging to "Satisfied" is 25%, the probability of belonging to "Average" is 0, the probability of belonging to "Dissatisfied" is 0, and the probability of belonging to "Very dissatisfied" is 0.
[0098] The positive factor can be a word, phrase or expression with a positive meaning. The negative factor can be a word, phrase or expression with a negative meaning. For example:
[0099] Positive factors: Words, phrases or expressions such as sweet talk, empathy, upgrade feedback plan, soothing / apologizing to the user's emotions, having a clear plan time limit, and having a sense of responsibility.
[0100] Negative factors: Words, phrases or expressions with semantic meanings such as giving an irrelevant answer, improper communication, shifting responsibility, not soothing emotions, unfulfilled promise, and not responding to the customer's request for compensation.
[0101] The execution subject can segment the target role conversation data to obtain segmented data. For example: "Review before 22:00", "Tomorrow morning", "Contact you", "Please keep the phone unblocked", "You can ask XX customer service, I don't quite understand this", "This is not within my scope of responsibility", "If you think so, I can't help it", "It's not my fault", "This compensation can't be given to you now", "Sorry, it's off-duty time now".
[0102] In this embodiment, "review before 22:00", "tomorrow morning", "contact you", and "please keep the phone unblocked" are positive factors, and the probability in the obtained word segmentation data is 40%, that is, the probability that the target role conversation data contains positive factors is 40%; "You can ask XX customer service, I don't quite understand this", "This is not within the scope of my responsibility", "I can't do anything if you think so", "It's not my fault", "This compensation cannot be given to you now", and "Sorry, it's off-duty time now" are not positive factors, and the probability is 60%.
[0103] In this embodiment, "review before 22:00", "tomorrow morning", "contact you", and "please keep the phone unblocked" are not negative factors, and the probability in the obtained word segmentation data is 40%; "You can ask XX customer service, I don't quite understand this", "This is not within the scope of my responsibility", "I can't do anything if you think so", "It's not my fault", "This compensation cannot be given to you now", and "Sorry, it's off-duty time now" are negative factors, that is, the probability that the target role conversation data contains negative factors is 60%.
[0104] Step S104, determine and output the target conversation data from the target role conversation data according to the probability belonging to the preset category, the probability of containing positive factors, and the probability of containing negative factors.
[0105] Specifically, before predicting the probability that the target role conversation data belongs to the preset category, the probability of containing positive factors, and the probability of containing negative factors, the method further includes: determining the time corresponding to the target role conversation data, and then splicing the target conversation data according to the time to generate spliced conversation data; using the spliced conversation data to update the target role conversation data.
[0106] Example of the splicing method: string splicing, with a space in the middle.
[0107] For example, before splicing:
[0108] Customer service: Hello, I'm here
[0109] Customer service: May I help you?
[0110] After splicing: Hello, I'm here, May I help you?
[0111] By splicing each target conversation data, the target conversation data can be made more complete, the context semantics can be clearer, and it helps to improve the accuracy of various probability predictions.
[0112] Specifically, predicting the probabilities that the session data of the target role belongs to a preset category, contains positive factors, and contains negative factors includes: invoking a percentile model to input the session data of the target role into the percentile model, and then outputting the probability that the session data of the target role belongs to the preset category; invoking a positive factor model to input the session data of the target role into the positive factor model, and then outputting the probability that the session data of the target role contains positive factors; invoking a negative factor model to input the session data of the target role into the negative factor model, and then outputting the probability that the session data of the target role contains negative factors.
[0113] In a session, all parts belonging to the customer service are extracted and concatenated, and then used as the input to the percentile model. Model structure: The model uses the fasttxt model for classification, and other classification models can also be used here. The specific execution logic of the percentile model is as Figure 4 shown. First, the customer service words are concatenated, and then the category prediction is performed through the fasttext model to obtain the probabilities of 5 categories. Then, the probabilities are multiplied by the corresponding scores respectively to obtain the final score. For example:
[0114] The probability of being very satisfied is 0.9, the probability of being satisfied is 0.05, the probability of being average is 0.01, the probability of being dissatisfied is 0.02, and the probability of being very dissatisfied is 0.02. The final score is:
[0115] 100 * 0.9 + 80 * 0.05 + 60 * 0.01 + 40 * 0.02 + 20 * 0.02 = 95.8
[0116] Among them, the model training corpus of the percentile model: 10,000 artificial customer service sessions are selected for annotation, and each session is annotated with a category according to the satisfaction level. Annotation categories: very satisfied, satisfied, average, dissatisfied, very dissatisfied, one of these five. The annotation rule is to manually classify the customer service answers in these 10,000 sessions into five categories, namely very satisfied, satisfied, average, dissatisfied, and very dissatisfied.
[0117] Positive factor model and negative factor model: The processing logics used by the two models are the same and will be described together here. Taking the positive factor model as an example below:
[0118] Model input: The part of the customer service in the session details between the customer service and the customer. All parts belonging to the customer service are extracted and concatenated, and then used as the model input. The logic here is the same as that of the percentile model.
[0119] Model Output: Two categories are output. The first category is the probability of containing positive factors, and the second category is the probability of not containing positive factors. That is to say, if the probability of the first category is greater than the probability of the second category, the customer service response in this reply contains positive factors.
[0120] Model Training Corpus: 2000 artificial customer service conversations are selected for annotation, and each conversation is annotated according to whether it contains positive factors.
[0121] The model structure of the positive factor model is as Figure 5 shown. In Figure 5 , the execution logic of the model is as follows: First, splice the customer service words, then perform category prediction through the fasttext model, and finally determine whether it contains positive factors according to the size of the category probability.
[0122] The negative factor model also has the same steps and model. The difference is that the training corpora of the two are different. One is annotated with positive and non - positive, and the other is annotated with negative and non - negative.
[0123] Based on the probability of belonging to a preset category, the probability of containing positive factors, and the probability of containing negative factors, determine and output the target conversation data from the target role conversation data. The target conversation data is the selected excellent response cases. The excellent response cases are jointly mined according to the percentile model, the positive factor model, and the negative factor model. The specific rules are as follows: (1) The result of the percentile model is greater than 80; (2) The result of the positive factor model is to contain positive factors (the probability of containing positive factors is greater than 50%); (3) The result of the negative factor model is not to contain negative factors (the probability of containing negative factors is less than or equal to 50%).
[0124] Among them, the screening condition of the percentile model (the result of the percentile model is greater than 80) can be adjusted at any time according to the final mining result. Screening excellent customer service cases jointly based on multiple deep - learning models can make the selected excellent customer service cases (i.e., the target conversation data) more accurate, which helps to quickly improve the working ability of new employees.
[0125] In this embodiment, by receiving a data - processing request, the corresponding conversation data and role identifier are obtained; the target role conversation data is extracted from the conversation data according to the role identifier; the probability of the target role conversation data belonging to a preset category, the probability of containing positive factors, and the probability of containing negative factors are predicted; based on the probability of belonging to a preset category, the probability of containing positive factors, and the probability of containing negative factors, the target conversation data is determined and output from the target role conversation data. The target conversation data is the result of joint screening based on multiple deep - learning methods, which can quickly and accurately mine the target conversation data for new employees to learn.
[0126] Figure 2It is a schematic diagram of the main process of the data processing method according to the second embodiment of the present application. As Figure 2 shown, the data processing method includes:
[0127] Step S201: Receive a data processing request, and obtain the corresponding session data and role identifier.
[0128] The session data corresponding to the data processing request can be the session data of the same day, or the session data separated by several days, such as the session data on Monday, Wednesday, and Friday. The session data can also be the session data of discontinuous time periods, such as the session data from 10:01 to 10:15 in the morning and from 14:05 to 15:00 in the afternoon on Monday. The embodiments of the present application do not make specific limitations on the time corresponding to the session data.
[0129] Step S202: Extract the target role session data from the session data according to the role identifier.
[0130] Step S203: Predict the probabilities that the target role session data belongs to the preset categories, the probability of including positive factors, and the probability of including negative factors.
[0131] Step S204: For each piece of target role session data, determine the corresponding score according to the probability of belonging to the preset category and the preset category score.
[0132] The execution subject can multiply the probability of belonging to the preset category and the corresponding preset category score to obtain the final score of each piece of target role session data (which can be the session data of customer service 1, or the session data of customer service 2, or the session data of customer service N).
[0133] Example: The probability of belonging to the very satisfied category is 0.9, the probability of belonging to the satisfied category is 0.05, the probability of belonging to the general category is 0.01, the probability of belonging to the dissatisfied category is 0.02, and the probability of belonging to the very dissatisfied category is 0.02. The final score is:
[0134] 100 * 0.9 + 80 * 0.05 + 60 * 0.01 + 40 * 0.02 + 20 * 0.02 = 95.8
[0135] Step S205: In response to the probability of including positive factors being greater than the preset threshold and the probability of including negative factors being less than the preset threshold, determine the target session data in the target role session data according to each score and output it.
[0136] The executing entity can determine the target session data of the target role with the result of the percentile model being greater than 80 points, the result of the positive factor model being the target role session data that includes positive factors and the result of the negative factor model not including negative factors as the target session data, that is, the excellent customer service cases, and can push them to the corresponding terminals of the new employees via methods such as emails and text messages for the new employees to learn from.
[0137] For example, the screening result: for example, in a session, the result of the percentile model is 90 points, and at the same time, the positive factors are: being responsible, and at the same time, there are no negative factors. In this case, the screening passes. If the score of a session is 70 points, the screening fails because the requirement is that the score must be > 80, or in a session, there are negative factors such as "answering off-topic" and the occurrence probability is greater than 50%, then the screening also fails.
[0138] Specifically, determining and outputting the target session data in the target role session data according to each score includes:
[0139] In response to there being scores greater than the preset score threshold among each score, determining and outputting the target role session data corresponding to the scores greater than the preset score threshold among each score as the target session data.
[0140] On the premise that the probability of including positive factors is greater than the preset threshold and the probability of including negative factors is less than the preset threshold, determine which target role session data is the target session data by judging the score output by the percentile model. Specifically, determine and output the target role session data corresponding to the scores greater than the preset score threshold, such as 80 points, as the target session data.
[0141] Specifically, the method further includes:
[0142] In response to the probability of each including positive factors being less than the preset threshold or the probability of each including negative factors being greater than the preset threshold, determine that there is no target session data in the target role session data, and end the data processing process.
[0143] As long as the target role session data does not include positive factors (that is, the probability of each including positive factors is less than the preset threshold) or includes negative factors (that is, the probability of each including negative factors is greater than the preset threshold), then the target role session data does not meet the requirements of the excellent customer service cases and cannot be output as the target session data, and end the determination process of the target session data.
[0144] Figure 3 It is a schematic diagram of the application scenario of the data processing method according to the third embodiment of the present application. The data processing method of the embodiments of the present application is applied to the scenario of selecting excellent customer service cases from numerous customer service session data. Such as Figure 3As shown, after receiving the entire conversation, the execution entity can extract the customer's words (i.e., the customer service conversation data, which is also the target role conversation data in this application), and splice the extracted customer service words. Then, the spliced customer service words are respectively input into the percentile model, the positive factor model, and the negative factor model, and the scores, the probability of containing positive factors, and the probability of containing negative factors are respectively output. Then, the execution entity can perform result screening on the input orthodox conversation based on the output scores, the probability of containing positive factors, and the probability of containing negative factors. The specific screening rules are as follows: (1) The result of the percentile model is greater than 80; (2) The result of the positive factor model is to contain positive factors (the probability of containing positive factors is greater than 50%); (3) The result of the negative factor model is not to contain negative factors (the probability of containing negative factors is less than or equal to 50%). Finally, the excellent response answer cases (i.e., the target conversation data) that meet the screening rules are obtained and output.
[0145] The embodiment of the present application can map the performance of the customer service to a score from 0 to 100, with a finer granularity, which is easy to perform statistical analysis and quantification; positive and negative factors are customized and corresponding models are trained as the screening rules for excellent response answer cases. It realizes quickly and accurately mining excellent customer service cases for new employees to learn.
[0146] Figure 6 is a schematic diagram of the main units of the data processing device according to the embodiment of the present application. As Figure 6 shown, the data processing device 600 includes a receiving unit 601, a data extraction unit 602, a probability prediction unit 603, and an output unit 604.
[0147] The receiving unit 601 is configured to receive a data processing request and obtain the corresponding conversation data and role identifier;
[0148] The data extraction unit 602 is configured to extract the target role conversation data from the conversation data according to the role identifier;
[0149] The probability prediction unit 603 is configured to predict the probability that the target role conversation data belongs to a preset category, the probability of containing positive factors, and the probability of containing negative factors;
[0150] The output unit 604 is configured to determine and output the target conversation data from the target role conversation data according to the probability of belonging to a preset category, the probability of containing positive factors, and the probability of containing negative factors.
[0151] In some embodiments, the output unit 604 is further configured to: for each piece of target role session data, determine a corresponding score according to the probability of belonging to a preset category and the preset category score; in response to the probability of including a positive factor being greater than a preset threshold and the probability of including a negative factor being less than the preset threshold, determine the target session data in the target role session data according to the scores and output it.
[0152] In some embodiments, the output unit 604 is further configured to: in response to there being a score greater than a preset score threshold among the scores, determine the target role session data corresponding to the score greater than the preset score threshold among the scores as the target session data and output it.
[0153] In some embodiments, the output unit 604 is further configured to: in response to the probability of each including a positive factor being less than the preset threshold or the probability of each including a negative factor being greater than the preset threshold, determine that there is no target session data in the target role session data and end the data processing process.
[0154] In some embodiments, the probability prediction unit 603 is further configured to: call a percentage model to input the target role session data into the percentage model, and then output the probability of the target role session data belonging to a preset category; call a positive factor model to input the target role session data into the positive factor model, and then output the probability of the target role session data including a positive factor; call a negative factor model to input the target role session data into the negative factor model, and then output the probability of the target role session data including a negative factor.
[0155] In some embodiments, the data extraction unit 602 is further configured to: determine a target role according to the role identifier, and then extract the target role session data corresponding to the target role from the session data.
[0156] In some embodiments, the data processing device further includes Figure 6 a data splicing unit not shown in the figure, configured to: determine the time corresponding to the target role session data, and then splice the target session data according to the time to generate spliced session data; update the target role session data with the spliced session data.
[0157] It should be noted that there is a corresponding relationship between the data processing method and the data processing device in the present application in terms of specific implementation content, so the repeated content will not be described again.
[0158] Figure 7 Exemplary system architecture 700 to which the data processing method or data processing device of the embodiments of the present application can be applied is shown.
[0159] As Figure 7As shown, the system architecture 700 may include terminal devices 701, 702, 703, a network 704, and a server 705. The network 704 is used to provide a medium for communication links between the terminal devices 701, 702, 703 and the server 705. The network 704 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0160] Users can use the terminal devices 701, 702, 703 to interact with the server 705 through the network 704 to receive or send messages, etc. Various communication client applications may be installed on the terminal devices 701, 702, 703, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (for example only).
[0161] The terminal devices 701, 702, 703 may be various electronic devices with a data processing screen and supporting web browsing, including but not limited to smart phones, tablets, laptop computers, and desktop computers, etc.
[0162] The server 705 may be a server that provides various services, such as a background management server that supports data processing requests submitted by users using the terminal devices 701, 702, 703 (for example only). The background management server may receive a data processing request, obtain corresponding session data and a role identifier; extract target role session data from the session data according to the role identifier; predict the probabilities of the target role session data belonging to a preset category, including a positive factor, and including a negative factor; determine and output target session data from the target role session data according to the probabilities of belonging to the preset category, including the positive factor, and including the negative factor. The target session data is based on the combined screening results of multiple deep learning methods, and can quickly and accurately mine the target session data for new people to learn.
[0163] It should be noted that the data processing method provided by the embodiments of the present application is generally executed by the server 705. Correspondingly, the data processing device is generally set in the server 705.
[0164] It should be understood that Figure 7 the numbers of the terminal devices, the network, and the server in
[0165] are merely illustrative. According to the implementation requirements, there may be any number of terminal devices, networks, and servers.
[0165] Next, refer to Figure 8 , which shows a schematic structural diagram of a computer system 800 of a terminal device suitable for implementing the embodiments of the present application. Figure 8 The terminal device shown is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.
[0166] As Figure 8 shown, the computer system 800 includes a central processing unit (CPU) 801, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 802 or a program loaded from a storage section 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the computer system 800 are also stored. The CPU 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0167] The following components are connected to the I / O interface 805: an input section 806 including a keyboard, a mouse, etc.; an output section 807 including a cathode ray tube (CRT), a liquid crystal credit authorization query processor (LCD), etc. and a speaker, etc.; a storage section 808 including a hard disk, etc.; and a communication section 809 including a network interface card such as a LAN card, a modem, etc. The communication section 809 performs communication processing via a network such as the Internet. A drive 810 is also connected to the I / O interface 805 as needed. A removable medium 811, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 810 as needed so that a computer program read from it can be installed into the storage section 808 as needed.
[0168] Specifically, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 809, and / or installed from the removable medium 811. When the computer program is executed by the central processing unit (CPU) 801, the above-mentioned functions defined in the system of the present application are executed.
[0169] It should be noted that the computer-readable medium shown in this application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can, for example, include but is not limited to an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium can include but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this application, a computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in conjunction with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.
[0170] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks can occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0171] The units involved in the embodiments described in this application can be implemented in software or in hardware. The described units can also be provided in a processor. For example, it can be described as: A processor includes a receiving unit, a data extraction unit, a probability prediction unit, and an output unit. Among them, the names of these units do not constitute a limitation on the unit itself in some cases.
[0172] As another aspect, this application also provides a computer-readable medium. The computer-readable medium can be included in the device described in the above embodiments; or it can exist alone without being assembled into the device. The above computer-readable medium carries one or more programs. When the above one or more programs are executed by the device, the device receives a data processing request, obtains corresponding session data and a role identifier; extracts target role session data from the session data according to the role identifier; predicts the probabilities that the target role session data belongs to a preset category, contains a positive factor, and contains a negative factor; determines and outputs target session data from the target role session data according to the probabilities that it belongs to the preset category, contains the positive factor, and contains the negative factor.
[0173] According to the technical solution of the embodiments of this application, the target session data is based on the combined screening results of multiple deep learning methods, and can quickly and accurately mine the target session data for newbies to learn.
[0174] The above specific embodiments do not constitute a limitation on the protection scope of this application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can occur depending on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of this application shall be included within the protection scope of this application.
Claims
1. A data processing method, characterized in that, including: Receiving a data processing request, and obtaining corresponding session data and a role identifier; Extracting target role session data from the session data according to the role identifier; Predicting the probabilities that the target role session data belongs to a preset category, contains a positive factor, and contains a negative factor; Determining and outputting target session data from the target role session data according to the probabilities that the target role session data belongs to a preset category, contains a positive factor, and contains a negative factor.
2. The method according to claim 1, wherein The determining and outputting target session data from the target role session data according to the probabilities that the target role session data belongs to a preset category, contains a positive factor, and contains a negative factor includes: For each piece of target role session data, determining a corresponding score according to the probability that the target role session data belongs to a preset category and a preset category score; In response to the probability of containing a positive factor being greater than a preset threshold and the probability of containing a negative factor being less than the preset threshold, determining and outputting the target session data in the target role session data according to the scores.
3. The method according to claim 2, characterized in that, The determining and outputting the target session data in the target role session data according to the scores includes: In response to there being a score greater than a preset score threshold among the scores, determining and outputting the target role session data corresponding to the score greater than the preset score threshold among the scores as the target session data.
4. The method according to claim 2, characterized in that The method further includes: In response to the probability of containing a positive factor being less than the preset threshold for each of them or the probability of containing a negative factor being greater than the preset threshold for each of them, determining that there is no target session data in the target role session data, and ending the data processing process.
5. The method according to claim 1, wherein The predicting the probabilities that the target role session data belongs to a preset category, contains a positive factor, and contains a negative factor includes: Invoking a percentage model to input the target role session data into the percentage model, and then outputting the probability that the target role session data belongs to a preset category; Invoking a positive factor model to input the target role session data into the positive factor model, and then outputting the probability that the target role session data contains a positive factor; Invoking a negative factor model to input the target role session data into the negative factor model, and then outputting the probability that the target role session data contains a negative factor.
6. The method according to claim 1, wherein The extracting target role session data from the session data according to the role identifier includes: Determining a target role according to the role identifier, and then extracting the target role session data corresponding to the target role from the session data.
7. The method according to claim 1, wherein Before the predicting the probabilities that the target role session data belongs to a preset category, contains a positive factor, and contains a negative factor, the method further includes: Determining the time corresponding to the target role session data, and then splicing the target session data according to the time to generate spliced session data; Updating the target role session data by using the spliced session data.
8. A data processing device, characterized in that, including: A receiving unit configured to receive a data processing request and obtain corresponding session data and a role identifier; A data extraction unit, configured to extract target role session data from the session data according to the role identifier; A probability prediction unit, configured to predict the probabilities that the target role session data belongs to a preset category, includes a positive factor, and includes a negative factor; An output unit, configured to determine and output target session data from the target role session data according to the probability of belonging to the preset category, the probability of including a positive factor, and the probability of including a negative factor.
9. The device according to claim 8, characterized in that, The output unit is further configured to: For each piece of target role session data, determine a corresponding score according to the probability of belonging to the preset category and the preset category score; In response to the probability of including a positive factor being greater than a preset threshold and the probability of including a negative factor being less than the preset threshold, determine and output the target session data in the target role session data according to each score.
10. The device according to claim 9, characterized in that The output unit is further configured to: In response to there being a score greater than a preset score threshold among each score, determine and output the target role session data corresponding to the score greater than the preset score threshold among each score as the target session data.
11. A data processing electronic device, characterized in that, Comprising: One or more processors; A storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-7.
12. A computer-readable medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, the method according to any one of claims 1-7 is implemented.
Citation Information
Patent Citations
Session end prediction method and system, online customer service method and system, equipment and medium
CN111416728A
Session record searching method and device, electronic equipment and storage medium
CN111897943A