Text processing method, system, device and storage medium
By extracting key information from text data and determining the priority of its dimensions and business segmentation tags, the problem of high difficulty in identifying keywords in text was solved, and highly accurate tag identification and hierarchical confirmation were achieved.
Patent Information
- Application Number
- CN202211668949.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-24
- Publication Date
- 2026-01-02
- Estimated Expiration
- 2042-12-24
AI Technical Summary
In existing technologies, when textual information is scattered, keywords are difficult to discover, and multiple keywords cannot be linked to established business segmentation labels, resulting in low recognition accuracy.
By extracting key information from the preprocessed text data, the corresponding dimensions and business segmentation labels are determined based on the key information. The dimensions are combined according to preset feature elements to form a preset priority. The recognition model is used to identify the key information and output the corresponding prompts.
It improves the accuracy of identifying business segmentation tags, especially in scenarios where tags are added or changed. It can accurately identify tagged work orders and non-tagged work orders, providing support for the generation of new tags in the future.
Smart Images

Figure CN115952287B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of text processing, in particular to a text processing method, system, device and storage medium. BACKGROUND
[0002] Topic discovery is a commonly used technology in text processing, which is used to automatically discover topics from text. However, when the information in the text is relatively scattered, it is difficult to discover key topic words, and it is difficult to associate multiple topic words with established business sub-labels, and the recognition accuracy is low. SUMMARY
[0003] The technical problem to be solved by the present application is to overcome the defects of the prior art that the text recognition is difficult and the recognition accuracy is low, and to provide a text processing method, system, device and storage medium.
[0004] The present application solves the above technical problems by the following technical solutions:
[0005] The present application provides a text processing method, comprising: extracting key information in pre-processed text data; determining corresponding dimensions and corresponding business sub-labels according to the key information, wherein the dimensions are combined according to pre-set feature elements to form a pre-set priority; determining the accuracy of the business sub-labels according to the pre-set priority and outputting the corresponding prompt.
[0006] Optionally, the step of determining the corresponding dimensions and the corresponding business sub-labels according to the key information comprises: pre-setting each dimension and the corresponding key information of each dimension, and associating the key information with the corresponding business sub-label; using a recognition model to recognize the key information and determine the corresponding dimensions; calculating the similarity between the key information and the pre-set word information to determine the key information; outputting the key information with the pre-set label and the dimensions.
[0007] Optionally, the dimensions include an execution class, a phenomenon class and a reason class; the execution class includes execution class feature elements, the phenomenon class includes phenomenon class feature elements, and the reason class includes reason class feature elements; the dimensions are combined according to pre-set feature elements to form a pre-set priority, comprising: the execution class feature element priority is higher than the phenomenon class feature element priority, and the phenomenon class feature element priority is higher than the reason class feature element priority.
[0008] Optionally, the execution class feature element, the phenomenon class feature element and the reason class feature element correspond to different key information respectively; the determining the accuracy of the business sub-label according to the preset priority comprises: determining feature information corresponding to the key information; when the feature information simultaneously contains the execution class feature element, the phenomenon class feature element and the reason class feature element, and the execution class feature element, the phenomenon class feature element and the reason class feature element combine the same business sub-label, the business sub-label is determined as a high-accuracy label; when the feature information simultaneously contains the execution class feature element, the phenomenon class feature element and the reason class feature element, and the execution class feature element, the phenomenon class feature element and the reason class feature element cannot combine the same business sub-label, the business sub-label is labeled according to the preset priority, and the business sub-label is determined as a medium-accuracy label.
[0009] Optionally, the dimension further comprises a business class, and the business class further comprises a business class feature element.
[0010] When the feature information simultaneously lacks the execution class feature element, the phenomenon class feature element and the reason class feature element, the text processing method further comprises: re-extracting a business class feature element in the text data, and labeling a business sub-label according to a frequency and a position of the business class feature element, and determining the business sub-label as a low-accuracy label.
[0011] Optionally, the outputting the corresponding prompt comprises: obtaining a scene set of the business sub-label and determining the accuracy of the business sub-label; confirming a final business sub-label prompt and outputting a corresponding prompt.
[0012] Optionally, the text processing method further comprises: performing word segmentation and part-of-speech tagging on the input text; identifying that a current sentence of a user is a query sentence according to at least one of a preset symbol, a preset emotional word and a preset keyword, and identifying customer service reply information, and simultaneously labeling the customer service reply information; deleting useless content and saving customer service-user text to form preprocessed text data.
[0013] The application further provides a text processing system, comprising: a data preprocessing module configured to extract key information from preprocessed text data; a determination module configured to determine corresponding dimensions and corresponding business sub-labels according to the key information, wherein the dimensions are combined according to preset characteristic elements to form a preset priority; and an output module configured to determine the accuracy of the business sub-labels according to the preset priority and output corresponding prompt symbols.
[0014] The application further provides a computer readable storage medium having a computer program stored thereon, wherein the computer program is executed by a processor to implement the steps of any of the above methods.
[0015] The text processing method, system, device and storage medium of the application can link the key information to the established business sub-labels, and based on the combination of the business sub-labels and the dimensions, the business sub-labels can be accurately classified and confirmed, which facilitates manual modification and auditing, greatly improves the accuracy of business sub-label identification, and especially in the scene where the labels are constantly increasing or changing, the label work orders and non-label work orders can be accurately identified, and the work order information can be collected for subsequent new label generation. BRIEF DESCRIPTION OF DRAWINGS
[0016] Figure 1 Flowchart of the text processing method of one embodiment of the application;
[0017] Figure 2 Flowchart of the explainable label extraction method for dialogue text of one embodiment of the application;
[0018] Figures 3-5 Flowchart of the explainable label extraction method for dialogue text of one embodiment of the application; Figure 2
[0019] Figure 6 Module schematic diagram of the text processing system of one embodiment of the application;
[0020] Figure 7 Module schematic diagram of the explainable label extraction system for dialogue text of one embodiment of the application. DETAILED DESCRIPTION
[0021] The present application will be further illustrated by the following examples without limiting the present application to the examples.
[0022] As shown in the figure, the present application provides a flow chart of a text processing method, the text processing method comprising the following steps: Figure 1
[0023] Step S101, extracting key information in the pre-processed text data;
[0024] Step S102, determining corresponding dimensions and corresponding business sub-labels according to the key information, the dimensions being combined according to pre-set characteristic elements to form a pre-set priority;
[0025] Step S103, determining the accuracy of the business sub-labels according to the pre-set priority and outputting the corresponding prompt
[0026] The text processing method of the present embodiment extracts key information in the pre-processed text data; determines corresponding dimensions and corresponding business sub-labels according to the key information, the dimensions being combined according to pre-set characteristic elements to form a pre-set priority; determines the accuracy of the business sub-labels according to the pre-set priority and outputs the corresponding prompt, the key information can be linked to the established business sub-labels, and at the same time, based on the combination recognition method of the business sub-labels and the dimensions, the business sub-labels can be accurately classified and confirmed, which is convenient for manual modification and audit, greatly improves the accuracy of business sub-label recognition, especially in the scene where the labels are constantly increasing or changing, the label work order and non-label work order can be accurately identified, and the work order information can be collected for the generation of new labels.
[0027] In an optional embodiment, the text processing method further comprises: performing word segmentation and part-of-speech tagging on the input text; identifying that the current sentence of the user is a query sentence according to at least one of pre-set symbols, pre-set mood words and pre-set keywords, and identifying the customer service reply information, and at the same time, labeling the customer service reply information. The pre-set symbols include question marks "?", the pre-set mood words include "ma", "ba", "ne" and the like, and the pre-set keywords include "what", "want to ask", "have not" and the like.
[0028] Delete useless content and save the customer-service-user text to form pre-processed text data. Specifically, delete stop words and useless content irrelevant to the query dialogue; delete multiple mood auxiliaries such as "hmm" and "ah", and delete specific irrelevant dialogue content.
[0029] In the present embodiment, the above pre-processing process can improve the accuracy of key information extraction, and further improve the accuracy of the entire text processing result.
[0030] In an optional implementation, the determining the corresponding dimension and the corresponding business sub-label according to the key information in step S102 comprises: presetting the key information corresponding to each dimension and associating the key information with the corresponding business sub-label. In this embodiment, the key information comprises keywords / phrases, and the keywords / phrases of each dimension are preset and associated with specific business sub-labels.
[0031] The key information is identified by using the identification model, and the corresponding dimension is determined. In this embodiment, the identification model uses a Bilstm-CRF model, and in other optional implementations, other types of models such as neural network modules can be used. Specifically, the keywords / phrases of each dimension are labeled, and the Bilstm-CRF model is used for labeling training, and finally the keywords / phrases of each dimension are identified.
[0032] The key information is determined by similarity calculation between the key information and the preset word information. Specifically, the key information is calculated with the preset word / phrase, and a similarity threshold is set, the word / phrase is retained if the similarity is high, and the word / phrase is deleted if the similarity is low.
[0033] The key information with the preset label and dimension is output. In this embodiment, the preset label is a customer service / user label, and specifically, the keywords / phrases of each dimension with the customer service / user label are output.
[0034] In this embodiment, the accuracy of identification is improved by combining model recognition and threshold comparison, which is suitable for processing long text data without label dialogue / non-dialogue, and is suitable for processing scenarios where labels are constantly increasing or changing, and can accurately identify label work orders and non-label work orders, and can collect work order information for subsequent generation of new labels. In an optional implementation, the dimension comprises an execution class, a phenomenon class, and a reason class.
[0035] The execution class comprises execution class characteristic elements, the phenomenon class comprises phenomenon class characteristic elements, and the reason class comprises reason class characteristic elements; the dimensions are combined according to the preset characteristic elements to form a preset priority, comprising: the execution class characteristic element priority is higher than the phenomenon class characteristic element priority, and the phenomenon class characteristic element priority is higher than the reason class characteristic element priority. By combining the dimensions according to the preset characteristic elements to form a preset priority, the business sub-labels are classified, which can ensure the accuracy and explainability of each business sub-label.
[0036] In an optional implementation, the execution class feature element, the phenomenon class feature element, and the cause class feature element correspond to different key information respectively; and the determining the accuracy of the business sub-label according to the preset priority in step S103 comprises: determining feature information corresponding to the key information; when the feature information simultaneously contains the execution class feature element, the phenomenon class feature element, and the cause class feature element, and the execution class feature element, the phenomenon class feature element, and the cause class feature element combine the same business sub-label, the business sub-label is determined as a high-accuracy label; when the feature information simultaneously contains the execution class feature element, the phenomenon class feature element, and the cause class feature element, and the execution class feature element, the phenomenon class feature element, and the cause class feature element cannot combine the same business sub-label, the business sub-label is labeled according to the preset priority, and the business sub-label is determined as a medium-accuracy label.
[0037] The text processing method of the embodiment can accurately classify and confirm the business sub-labels, facilitate manual modification and auditing, and greatly improve the accuracy of business sub-label identification by determining the corresponding business sub-labels through key information and prioritizing the business sub-labels according to the accuracy.
[0038] In an optional implementation, the text processing method of the embodiment further comprises: sorting the accuracy of the business sub-labels according to the high-accuracy label and the medium-accuracy label, and outputting a prompt corresponding to the business sub-label of the high-accuracy label.
[0039] In an optional implementation, the dimension further comprises a business class, and the business class further comprises a business class feature element.
[0040] When the feature information simultaneously lacks the execution class feature element, the phenomenon class feature element, and the cause class feature element, the text processing method further comprises: re-extracting the business class feature element in the text data, labeling the business sub-label according to the frequency and position of the business class feature element, and determining the business sub-label as a low-accuracy label. Specifically, if the execution class feature element, the phenomenon class feature element, the cause class feature element, and the like are not extracted from the text, the business feature element of the text is re-extracted, and if multiple business feature elements are extracted from the text, the business sub-label is labeled depending on the frequency and position of the business feature element, and the label is a low-accuracy label, which can be checked and modified by a human. In an optional implementation, the text processing method further comprises: sorting the accuracy of the business sub-labels according to the high-accuracy label, the medium-accuracy label, and the low-accuracy label, and outputting a prompt corresponding to the business sub-label of the high-accuracy label.
[0041] In one optional implementation, step S103, outputting the corresponding prompt, includes: obtaining a set of scenarios for business segmentation labels and determining the accuracy of the business segmentation labels; confirming the final business segmentation label prompt and outputting the corresponding prompt. This method, based on element combination recognition of business segmentation scenarios, can accurately classify and confirm business segments, facilitating manual modification and review, and significantly improving the accuracy of business segmentation scenario recognition.
[0042] The present application will be further described in detail below with a specific embodiment:
[0043] like Figure 2 As shown, the specific steps of the interpretability tag extraction method for dialogue text in this embodiment are as follows:
[0044] Step S1: Input the dialogue text into the system and preprocess the data.
[0045] For example, the dialogue text is the voice-to-text information in the dialogue process, for example, "[agent] I [user] Oh, can I change the name of the invoice I developed? [agent] The name of the invoice. Is it the caller ID? [user] Yes. [user] Yes, the caller ID, and now the name is not the owner's name. Can I change it back to the owner's name? [agent] Hmm [agent] Hmm, you are now in, you first report the owner's name to me. [user] Zhang Jiejin. [user] Hmm, yes. [agent] Hmm, good, you mean you want to change the account name, right? I'll check it out, please wait a moment. [user] Hmm, good. [agent] Hmm [agent] Change the payment customer, [user] Hmm [user] I [agent] Hmm, I'm sorry, sir, I'll help you check it out. If you want to change it, it's possible. [user] Hmm [agent] Hmm [user] Hmm, how do I operate? [agent] Like this, sir, you can go to our business site to handle it, [agent] or, well, or you can also handle it through our online, [agent] online, you can also handle it. Please follow our WeChat public account, China Telecom Shanghai Customer Service, [agent] Hmm [user] Hmm, hmm [user] Hmm, public account, China Telecom Customer Service, and find the [agent] Oh, please wait a moment, I'll check it out. Oh [user] Which is the invoice header? [user] Still can't change it. [agent] No [user] Hmm [user] Hmm [user] Yes. [agent] You [agent] You now mean to change it to, well, change it back to, or change it back to Zhang Jiejin's name, right? [user] Yes, yes. [user] Yes [agent] I'll change it back to [user] Yes [user] Hmm [user] Get it. [agent] Sir, we can help you handle it here. Oh, I need to check your mobile phone number first, and the property customer, that is, the customer [user] Yes. [agent] The name and address of the paid customer, please report it. [agent] Hello [user] First report. [agent] Yes, please report your mobile phone number first. [user] Yes [user] Yes [user] Yes. [agent] Yes, good. Then, report the name of the property transfer."
[0046] As Figure 3 shown, step S1 specifically includes:
[0047] S11, performing word segmentation and part-of-speech tagging on the input dialogue text.
[0048] S12. Identify interrogative sentences and response sentences. Specifically, based on the symbol "?", or modal particles such as
吗
吧
呢
什么
想问
还没
[0049] For example:
[0050] [agent] I / rr
[0051] [user] Alas / e Hello / vl, / w
[0052] (AskTarget) Can / v the / rz name / n for / ude1 opening / ANCV invoice / n be / v changed / ANCV? / y / w
[0053] [agent] The / ude1 name / n for / ude1 opening / ANCV invoice / n. / w
[0054] (AskTarget) Is / v it / y the / rz call / n number / n? / y / w
[0055] [user](Confirm) Right / p. / w
[0056] [user](Confirm) Right / p, / w
[0057] Call / n number / n, / w
[0058] And / c currently / DATE it / udh is / c not / c the / ude1 name / n of / ude1 the / rz machine / ng owner / ag, / w
[0059] (AskTarget) Can / v it be / v changed / ANCV back / nz to / rz the / ude1 name / n of / ude1 the / rz machine / ng owner / ag? / y / w
[0060] #[Element - Negative] No, 8
[0061] [agent](Confirm) Mmm / e
[0062] [agent](Confirm) Mmm / e Well / rzv then / rz you / rr are / vshi currently / DATE here / p, / w please / d first / d report / n the / ude1 name / n of / ude1 the / rz machine / ng owner / ag to / p me / rr. / w
[0063] [user] Zhang Jie / NAME Jin / q. / w
[0064] [user](Confirm) Mmm / e Right / p. / w
[0065] [agent] Yes, / w
[0066] [agent] That's / r right, / w
[0067] [agent] Yes, / w
[0068] [agent] Let me / r check, / w
[0069] [agent] Please wait a moment, / w
[0070] [user] Hello, / w
[0071] [agent] Yes, / w
[0072] [agent] Change the / r payment / r client, / w
[0073] [user] Yes, / w
[0074] [user] I / r
[0075] [agent] Yes, / w
[0076] [agent] I'm / r very / r sorry, / w
[0077] [agent] Sir, / w
[0078] [agent] Let me / r help / r you / r check / r the / r conversation, / w
[0079] [agent] If / r you / r want / r to / r change / r the / r conversation, / w
[0080] [agent] Yes, / w
[0081] [user] Yes, / w
[0082] [agent] Yes, / w
[0083] [user] (AskTarget) Yes, / w
[0084] [agent] Yes, / w
[0085] Then you can go to one of our branches to apply for an ANCV.
[0086] #[Element-Judgment] Process, 16
[0087] [agent] or / c if using / udh, / w
[0088] Uh / y or / c you / rr want / v no / d also / d you can / v through / p us / rr a / mq line / n on / f, / w
[0089] #[Element-judgment] is acceptable, 9
[0090] [agent] Online / n / f / ude1 calls / n are also / d / vshi can / v process / ANCV / ude1, / w
[0091] You can follow this WeChat official account.
[0092] A / mq China / ns Telecom / n Shanghai / ns Customer Service / nz, / w
[0093] #[Element-Judgment] Process, 7
[0094] [agent](Confirm) Okay / e
[0095] [user](Confirm) Hmm / e Hmm / nz
[0096] [user](Confirm) Hmm / e Official Account / n, / w
[0097] China Telecom / N Customer Service / NZ, / w
[0098] Then find the / v / rz in / c / f.
[0099] [agent] Oh / e you / rr wait / d wait / udeng a moment / mq, / w
[0100] I'll take a look at the .mq.
[0101] Oh / e
[0102] [user](AskTarget) Which / ry is the invoice header / vi? / w
[0103] [user] is / c changing / ANCV not / d done / ule. / w
[0104] [agent] not / d
[0105] [user] (Confirm) hmm / e
[0106] [user] (Confirm) hmm / e
[0107] [user] s / ude1. / w
[0108] [agent] you / rr
[0109] [agent] you / rr now / DATE just / say / c want / v to / pba that / rz hmm / e change / ANCV uh / y just / v to / pba that / rz just / d change / nz back / nz to / v hmm / e is / c change / ANCV back / v to / p that / rz zhang jie jin / NAME that / rz name / JQ. / w
[0110] (AskConfirm) right / p yes / c? / w
[0111] [user] (Confirm) hmm / e right / p. / w
[0112] [user] (Confirm) hmm / e
[0113] [user] s / ude1. / w
[0114] [agent] i / rr go / v back / v
[0115] [user] s / ude1. / w
[0116] [user] (Confirm) hmm / e
[0117] [user] make / v. / w
[0118] [agent] mr / NAME, /
[0119] We / rr can / v also / d can / v help / v you / rr handle / ANCV it / v down / vf. / w
[0120] Oh / e that / rz i / rr want / v to / p check / v with you / nz first / d mobile phone / DEV. / w
[0121] Then / c property / n client / n hmm / e that / rz client / n
[0122] #[Element- judgment] check, 21
[0123] [user] hi / a. / w
[0124] [agent] that is / v pay / v customer / n of / ude1 a / mq name / n and / cc address / n, / w
[0125] You / rr report / n once / mq. / w
[0126] [agent] hello / v
[0127] [user] first / d report / n. / w
[0128] [agent] (Confirm) to / p, / w
[0129] Ah / y first / d put / p ba mobile phone / DEV please / v report / n once / mq. / w
[0130] [user] (Confirm) um / e
[0131] [user] hi / a. / w
[0132] [user] of / ude1. / w
[0133] [agent] (Confirm) um / um / nz good / y cc. / w
[0134] Then / c of / ude1 words / n is / v shichanquan / n transfer / JKW of / ude1 name / n report / n once / mq.
[0135] S13, delete stop words and useless sentence. Specifically, delete stop words and useless content irrelevant to the inquiry dialogue; delete um, ah and other multiple mood auxiliaries. Delete the dialogue content irrelevant to AskTarget and Confirm.
[0136] S14, keep the customer-service-user dialogue text, that is, save the final customer-service-user dialogue text.
[0137] For example: the dialogue content after data processing is as follows:
[0138] [user] Oh, hello, can I change the name of the invoice?
[0139] [agent] The name of the invoice. Is it the caller ID?
[0140] [user] Yes.
[0141] [user] Yes, the caller ID number, and now the name is not the owner, can change back to the owner's name?
[0142] [agent] Yes
[0143] [agent] Yes, and you are now in, you first report the owner's name to me.
[0144] [user] Zhang Jiejin.
[0145] [user] Yes.
[0146] [agent] Yes, good, you want to change the account name, right? I look at it, please wait a moment.
[0147] [user] Yes, good.
[0148] [user] How do I operate?
[0149] [agent] Yes, sir, you can go to our business network to handle,
[0150] [agent] or, or you can also through our online,
[0151] [agent] online is also available, you can follow the WeChat public number, China Telecom Shanghai customer service,
[0152] [user] Which is the billing name?
[0153] [user] Still can not change.
[0154] [agent] No
[0155] [agent] You now want to change the name to the Zhang Jiejin, right?
[0156] [user] Yes, right.
[0157] Step S2, set the business sub-label key elements, including execution class characteristic elements, phenomenon class characteristic elements, reason class characteristic elements and business class characteristic elements, and extract information from each type of characteristic elements. Step S2 contains a business sub-label key element rule, which defines common keywords or key phrases for each type of business sub-label element. Key elements include four types, namely execution class characteristic elements, phenomenon class characteristic elements, reason class characteristic elements and business class characteristic elements.
[0158] Specifically, in the composition sequence of the business sub-label key elements, the meanings represented by the four types of feature elements are as follows:
[0159] 1. Execution class feature element: execution class description features of the business sub-label, such as address change, bill change, penalty exemption, bill reissue, etc.
[0160] 2. Phenomenon class feature element: phenomenon class description features of the business sub-label, such as unpaid, incomplete electronic invoice, arrears, etc.
[0161] 3. Reason class feature element: reason class description features of the business sub-label, such as no bill received, out of town, exempt from reminders and stoppage, etc.
[0162] 4. Business class feature element: business class feature description of the business sub-label, such as bill, invoice, advertisement, mailbox, address, etc.
[0163] As shown in Figure 4 , step S2 specifically includes:
[0164] S21, set the key element dimension: execution class, phenomenon class, reason class, and business class.
[0165] S22, preset the key words / phrases for each dimension, and associate the preset key words / phrases to specific business sub-labels.
[0166] For example, taking the business sub-label
account information change
[0167] S23, label each dimension of the key words / phrases, and use Bilstm-CRF model for labeling training, and finally identify the key words / phrases of each dimension.
[0168] S24, calculate the similarity of the key words / phrases of each dimension identified in step S23 and the key words / phrases preset in step S22, set a similarity threshold, if the similarity is high, the identified key words / phrases are retained, if the similarity is low, the identified key words / phrases are deleted.
[0169] S25, output the key words / phrases of each dimension with customer service / user label. Dimension: execution class, phenomenon class, reason class, and business class.
[0170] For example,
[0171] [user] Well, can I change the name I invoiced to? [[Change account name]]
[0172] [agent] Yes, so you want to change the account name, right? Let me check. Please wait a moment. [[Change account name]]
[0173] Among them, the bold font and the italicized font in the bold font represent different dimensions, and different colors can also be used to represent different dimensions, such as marking the five words “change the name” in yellow, the four words “account name” in red font and yellow, and the four words “change it” in yellow.
[0174] Step S3, define the arrangement priority of the key elements of the business sub-label;
[0175] As shown in Figure 5 , step S3 specifically includes:
[0176] S31, set the priority of the combination of each dimension feature element, and divide the priority order into three layers.
[0177] S32, if the execution class, phenomenon class, and reason class are extracted from the dialogue text and combined into the same business sub-label, then the business sub-label is a high-accuracy label, and artificial can be completely relied on.
[0178] S33, if the execution class feature element, phenomenon class feature element, and reason class feature element are extracted from the dialogue text, but cannot form the same business sub-label, then according to the order of execution class feature priority > phenomenon class feature element > reason class feature element, the business sub-label is labeled, and this type of label is a medium-accuracy label, and artificial can be partially relied on.
[0179] For example: [RT-12] business sub-label-account information change, where [RT-12] belongs to the second priority
[0180] S34, if the execution class feature element, phenomenon class feature element, and reason class feature element are not extracted from the dialogue text, then the business feature elements of the text are re-extracted, and the business feature elements are extracted. If multiple business feature elements are extracted from the text, then the business sub-label is labeled according to the frequency and position of the business feature elements, and this type of label is a low-accuracy label, and artificial can be checked and modified.
[0181] Step S4, according to the business sub-label key element definition template, identify the feature elements in the dialogue text, and associate the business sub-label scene, get the business sub-label scene set. Associating business sub-label scene refers to the existing business sub-label type on the business side, such as penalty exemption, order supplement, account information change and other specific business sub-labels.
[0182] For example: The feature description of the relevant business sub-label is confirmed from two sentences, and the set is as follows:
[0183] [RT-12] Business sub-label - account information change
[0184] [RT-12] Business sub-label - account information change
[0185] Step S5, according to the accuracy of the priority determined in step S3, confirm the final business sub-label prompt, and display different prompts for customer service.
[0186] The application adopts the method of element combination recognition based on business sub-label scene, which can accurately classify and confirm the business sub-label, facilitate manual modification and audit, and greatly improve the accuracy of business sub-label scene recognition. Especially in the scene of increasing or changing labels, it can accurately identify label work orders and non-label work orders, and can collect work order information for subsequent new label generation.
[0187] As shown in Figure 6 , the application provides a text processing system, which comprises:
[0188] The data preprocessing module 1 is used to extract the key information in the preprocessed text data;
[0189] The determination module 2 is used to determine the corresponding dimension and the corresponding business sub-label according to the key information, and the dimension is composed of a preset priority according to a preset feature element combination;
[0190] The output module 3 is used to determine the accuracy of the business sub-label according to the preset priority and output the corresponding prompt.
[0191] The text processing system in this embodiment extracts key information from the preprocessed text data through the data preprocessing module 1; the determination module 2 determines the corresponding dimensions and corresponding business segmentation tags based on the key information, and the dimensions are combined according to preset feature elements to form a preset priority; the output module 3 determines the accuracy of the business segmentation tags based on the preset priority and outputs the corresponding prompts. The key information can be linked with the established business segmentation tags. At the same time, based on the combination recognition method of business segmentation tags and dimensions, the business segmentation tags can be accurately classified and confirmed, which is convenient for manual modification and review, and greatly improves the accuracy of business segmentation tag recognition. Especially in the scenario where the tags are constantly increasing or changing, it can accurately identify tagged work orders and non-tagged work orders, and can collect work order information for the generation of new tags.
[0192] The following is a specific example to illustrate this: Figure 7 As shown, an interpretability tag extraction system for dialogue text includes: a text data preprocessing unit 11, a key element identification unit 21, a business rule management unit 22, and a business segmentation tag identification unit 31. The data preprocessing module 1 in the text processing system includes the text data preprocessing unit 11, the determination module 2 includes the key element identification unit 21 and the business rule management unit 22, and the output module 3 includes the business segmentation tag identification unit 31.
[0193] Text data preprocessing unit 11: Provides data processing capabilities to clean up redundant text dialogue content. Specifically, it first performs word segmentation and part-of-speech tagging on the text data and identifies query statements, response statements, and declarative statements. Then, it removes stop words and declarative statements with a length of less than 10 characters, and finally obtains the cleaned dialogue text result.
[0194] Key Element Recognition Unit 22: Provides key element recognition capabilities. It is used to convert general word segmentation and tagging sequences into key element sequences. Specifically, this includes extraction using pre-defined keywords and extraction using named entity recognition. The results extracted using named entity recognition are matched with pre-defined words / phrases. Words with high similarity are replaced with pre-defined words / phrases, while words with low similarity are deleted.
[0195] Business Rule Management Unit 22: This unit manages business rules for use by other modules, including rules for identifying business elements and scenarios, and rules for generating work orders. These rules are manually created and maintained.
[0196] Business segmentation label recognition unit 31: Provides a set of business segmentation labels and, based on the business segmentation label rules, enables the final label confirmation.
[0197] In an optional implementation, the text data preprocessing unit 11 is further configured to perform word segmentation and part-of-speech tagging on the input text; identify that the current user sentence is a query sentence according to at least one of a preset symbol, a preset mood word, and a preset keyword, and identify the customer service reply information, while labeling the customer service reply information; delete useless content and save the customer-service-user text to form the preprocessed text data. Specifically, the preset symbol includes a question mark "?", the preset mood word includes "ma", "ba", "ne", and the like, and the preset keyword includes "what", "want to ask", "have not" and the like. Stop words and useless content irrelevant to the query dialogue are deleted; multiple mood auxiliaries such as "hmm" and "ah" are deleted, and specific irrelevant dialogue content is deleted.
[0198] In this embodiment, the preprocessing process of the text data preprocessing unit 11 can improve the accuracy of key information extraction, and further improve the accuracy of the entire text processing result.
[0199] In an optional implementation, the dimensions include an execution class, a phenomenon class, and a reason class; the execution class includes execution class feature elements, the phenomenon class includes phenomenon class feature elements, and the reason class includes reason class feature elements; the dimensions are combined according to preset feature elements to form preset priorities, including: the execution class feature element priority is higher than the phenomenon class feature element priority, and the phenomenon class feature element priority is higher than the reason class feature element priority. By combining the dimensions according to the preset feature elements to form the preset priorities to classify the business sub-labels, the accuracy and interpretability of each business sub-label can be ensured. The execution class feature elements, the phenomenon class feature elements, and the reason class feature elements correspond to different key information respectively.
[0200] The key element identification unit 21 is configured to preset each dimension and the key information corresponding to each dimension, and associate the key information to the corresponding business sub-label; use a recognition model to identify the key information and determine the corresponding dimension; determine the key information by performing similarity calculation on the key information and preset word information; and output the key information with preset labels and dimensions.
[0201] In this embodiment, the key information includes keywords / phrases, and the keywords / phrases of each dimension are preset and associated to specific business sub-labels. In this embodiment, the recognition model uses a Bilstm-CRF model, and in other optional implementations, other types of models such as neural network modules can be used. Specifically, the keywords / phrases of each dimension are labeled, and the Bilstm-CRF model is used for labeling training, and finally the keywords / phrases of each dimension are identified.
[0202] The key information is determined by similarity calculation between the key information and preset word information. Specifically, similarity calculation is performed between the key information and preset words / phrases, a similarity threshold is set, and the words / phrases are retained if the similarity is high, or the words / phrases are deleted if the similarity is low.
[0203] The key information with preset labels and dimensions is output. In this embodiment, the preset label is a customer / user label, and specifically, the key words / phrases of each dimension with the customer / user label are output.
[0204] In this embodiment, the accuracy of identification is improved by combining model recognition and threshold comparison, which is suitable for processing long text data without labeled dialogue / non-dialogue, and is suitable for processing scenarios with increasing or changing labels, and can accurately identify labeled work orders and non-labeled work orders, and can collect work order information for subsequent new label generation.
[0205] In an optional implementation, the business rule management unit 22 is configured to determine the accuracy of the business sub-label according to the preset priority, including: determining feature information corresponding to the key information; when the feature information simultaneously contains an execution class feature element, a phenomenon class feature element and a reason class feature element, and the execution class feature element, the phenomenon class feature element and the reason class feature element combine the same business sub-label, the business sub-label is determined as a high-accuracy label; when the feature information simultaneously contains an execution class feature element, a phenomenon class feature element and a reason class feature element, and the execution class feature element, the phenomenon class feature element and the reason class feature element cannot form the same business sub-label, the business sub-label is labeled according to the preset priority, and the business sub-label is determined as a medium-accuracy label.
[0206] The text processing system of this embodiment determines the business sub-label corresponding to the key information through the key element identification unit and the business rule management unit, and prioritizes the business sub-labels according to the accuracy, which can accurately classify and confirm the business sub-labels, facilitate manual modification and review, and greatly improve the accuracy of business sub-label identification.
[0207] In an optional implementation, the dimensions further include a business class, and the business class further includes a business class feature element.
[0208] When the feature information simultaneously lacks execution-type, phenomenon-type, and cause-type feature elements, the business rule management unit 22 further extracts business-type feature elements from the text data and labels them with business sub-categories based on their frequency and location. These sub-categories are then identified as low-accuracy labels. Specifically, if execution-type, phenomenon-type, and cause-type feature elements are not extracted from the text, the business feature elements are extracted again. If multiple business feature elements are extracted, they are labeled with business sub-categories based on their frequency and location. These labels are low-accuracy labels and can be manually verified and modified.
[0209] In one optional implementation, the business segmentation label recognition unit 31 is further configured to acquire a set of scenarios for business segmentation labels and determine the accuracy of the business segmentation labels; confirm the final business segmentation label prompt and output the corresponding prompt. This method, based on element combination recognition of business segmentation scenarios, can accurately classify and confirm business segments, facilitating manual modification and review, and significantly improving the accuracy of business segmentation scenario recognition.
[0210] To extract business segmentation tags from unlabeled long dialogue texts for automated annotation, this embodiment extracts features from various dimensions of business segmentation tags and combines them with business logic to identify these tags. A method based on element combination recognition within business segmentation tag scenarios is employed, accurately classifying and confirming the tags for easier manual review and modification, significantly improving the accuracy of business segmentation tag scenario recognition. Especially in scenarios where tags are constantly increasing or changing, it can accurately identify tagged and untagged work orders, collecting work order information for the generation of new tags. In one embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the steps of the above-described method embodiments.
[0211] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-described method embodiments.
[0212] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when executed, can include the processes of the above-mentioned embodiment methods. Any reference to memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. The non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical storage, high-density embedded non-volatile memory, resistive memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric memory (FRAM), phase change memory (PCM), graphene memory, etc. The volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration but not limitation, the RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The database involved in the embodiments provided in the present application can include at least one of a relational database and a non-relational database. The non-relational database can include a distributed database based on a block chain, etc., without being limited thereto. The processor involved in the embodiments provided in the present application can be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., without being limited thereto.
[0213] Any combination of the technical features of the above embodiments can be made, and in order to make the description concise, all possible combinations of the technical features in the above embodiments are not described, however, as long as the combination of the technical features does not exist, it should be considered as the scope of the present application.
[0214] Although the specific embodiments of the present application are described above, those skilled in the art should understand that these are only illustrative, and the protection scope of the present application is defined by the appended claims. Those skilled in the art can make various changes or modifications to these embodiments without departing from the principles and essence of the present application, and these changes and modifications all fall within the protection scope of the present application.
Claims
1. A text processing method characterized by, The method comprises the following steps: extracting key information from the preprocessed text data; determining corresponding dimensions and corresponding business sub-labels according to the key information, wherein the dimensions are combined according to preset characteristic elements to form a preset priority; determining the accuracy of the business sub-labels according to the preset priority and outputting corresponding prompts; the method of determining the corresponding dimensions and the corresponding business sub-labels according to the key information comprises: presetting each dimension and the corresponding key information of each dimension, and associating the key information with the corresponding business sub-labels; using a recognition model to recognize the key information and determine the corresponding dimensions; determining the key information by similarity calculation between the key information and preset word information; outputting the key information with preset labels and dimensions; the dimensions include execution class, phenomenon class and reason class; the execution class includes execution class characteristic elements, the phenomenon class includes phenomenon class characteristic elements, and the reason class includes reason class characteristic elements; the dimensions are combined according to preset characteristic elements to form a preset priority, which comprises: the priority of the execution class characteristic elements is higher than that of the phenomenon class characteristic elements, and the priority of the phenomenon class characteristic elements is higher than that of the reason class characteristic elements; the execution class characteristic elements, the phenomenon class characteristic elements and the reason class characteristic elements correspond to different key information respectively; determining the accuracy of the business sub-labels according to the preset priority, which comprises: determining the characteristic information corresponding to the key information; when the characteristic information simultaneously contains the execution class characteristic elements, the phenomenon class characteristic elements and the reason class characteristic elements, and the execution class characteristic elements, the phenomenon class characteristic elements and the reason class characteristic elements can form the same business sub-label, the business sub-label is determined as a high-accuracy label; when the characteristic information simultaneously contains the execution class characteristic elements, the phenomenon class characteristic elements and the reason class characteristic elements, and the execution class characteristic elements, the phenomenon class characteristic elements and the reason class characteristic elements cannot form the same business sub-label, the business sub-label is labeled according to the preset priority, and the business sub-label is determined as a medium-accuracy label; the dimensions further include a business class, and the business class further includes business class characteristic elements; when the characteristic information simultaneously lacks the execution class characteristic elements, the phenomenon class characteristic elements and the reason class characteristic elements, the text processing method further comprises: reextracting the business class characteristic elements in the text data, labeling the business sub-labels according to the frequency and position of the business class characteristic elements, and determining the business sub-labels as low-accuracy labels.
2. The text processing method of claim 1, wherein, the outputting of corresponding prompts comprises: obtaining a scene set of the business sub-labels and determining the accuracy of the business sub-labels; confirming the final business sub-label prompt and outputting the corresponding prompt.
3. The text processing method of claim 1, wherein, the text processing method further comprises: performing word segmentation and part-of-speech tagging on the input text; According to at least one of the preset symbols, the preset tone words and the preset keywords, it is identified that the current sentence of the user is an inquiry sentence, and the customer service reply information is identified, and the customer service reply information is labeled; Delete useless content and save the customer-service-user text to form preprocessed text data.
4. A text processing system, characterized by Comprise: The data preprocessing module is used for extracting the key information in the preprocessed text data; The determination module is used for determining the corresponding dimension and the corresponding business sub-label according to the key information, and the dimension is combined according to the preset characteristic element to form a preset priority; The output module is used for determining the accuracy of the business sub-label according to the preset priority and outputting the corresponding prompt; The determination of the corresponding dimension and the corresponding business sub-label according to the key information comprises: Predefine each dimension and the corresponding key information of each dimension, and associate the key information with the corresponding business sub-label; The key information is identified and the corresponding dimension is determined by using the identification model; The key information is determined by similarity calculation between the key information and the preset word information; Output the key information with the preset label and the dimension; The dimension includes execution class, phenomenon class and reason class; The execution class includes execution class characteristic elements, the phenomenon class includes phenomenon class characteristic elements, and the reason class includes reason class characteristic elements; The dimension is combined according to the preset characteristic element to form a preset priority, which comprises: The priority of the execution class characteristic element is higher than that of the phenomenon class characteristic element, and the priority of the phenomenon class characteristic element is higher than that of the reason class characteristic element; The execution class characteristic element, the phenomenon class characteristic element and the reason class characteristic element correspond to different key information respectively; The determination of the accuracy of the business sub-label according to the preset priority comprises: Determine the characteristic information corresponding to the key information; When the characteristic information simultaneously contains the execution class characteristic element, the phenomenon class characteristic element and the reason class characteristic element, and the execution class characteristic element, the phenomenon class characteristic element and the reason class characteristic element combine the same business sub-label, the business sub-label is determined as a high accuracy label; When the characteristic information simultaneously contains the execution class characteristic element, the phenomenon class characteristic element and the reason class characteristic element, and the execution class characteristic element, the phenomenon class characteristic element and the reason class characteristic element cannot form the same business sub-label, then the business sub-label is labeled according to the preset priority, and the business sub-label is determined as a medium accuracy label; The dimension also includes business class, and the business class also includes business class characteristic elements; When the characteristic information simultaneously lacks the execution class characteristic element, the phenomenon class characteristic element and the reason class characteristic element, the text processing system further comprises: Re-extract the business class characteristic elements in the text data, and label the business sub-label according to the frequency and position of the business class characteristic elements, and determine the business sub-label as a low accuracy label.
5. A computer device, comprising a memory and a processor, wherein the memory stores a computer program, and wherein the processor, when executing the computer program, implements the steps of the method according to any one of claims 1-3.
6. The computer program product, when executed by a processor, implements the steps of the method according to any one of claims 1-3.
6. A computer-readable storage medium having stored thereon a computer program, characterized in that, 7. The computer program product, when executed by a processor, implements the steps of the method according to any one of claims 1-3.
Citation Information
Patent Citations
Text label extraction method and device and storage medium
CN111563361A
Session text analysis method and device, computer equipment and storage medium
CN115374273A