Intention labeling method, model training method, intention recognition method and electronic equipment
By generating prompt words and clustering to remove duplicates based on the thought chain of reference words, the problem of time-consuming and labor-intensive training of traditional intent recognition models is solved, and efficient and accurate intent labeling and model training are achieved.
Patent Information
- Application Number
- CN202510295719.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-09-23
AI Technical Summary
Traditional intent recognition model training requires a lot of manpower and material resources, takes a long time to train, and large models have poor stability and accuracy.
A thought chain based on reference speech is used to generate prompt words, a large model is used to annotate the speech to be annotated with intent, intent annotation samples are constructed, and annotation efficiency and accuracy are improved through clustering and deduplication processing.
It improves the accuracy and stability of intent labeling, reduces manpower and material resources, shortens training time, and improves the performance of intent recognition models.
Smart Images

Figure CN120687539A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to an intent labeling method, a model training method, an intent recognition method, and an electronic device. Background Art
[0002] Intent recognition models are widely used in scenarios such as intelligent customer service, smart marketing, and smart elderly care. They solve customer questions by accurately identifying customer intent and providing corresponding answers. Traditional intent recognition models require manual labeling of large numbers of training samples, which are then used for training. This model training method consumes significant manpower and material resources and takes a long time. Summary of the Invention
[0003] The present disclosure provides an intent labeling method, a model training method, an intent recognition method, and an apparatus.
[0004] In a first aspect, the present disclosure provides an intent annotation method, comprising:
[0005] Obtaining prompt words, wherein the prompt words are generated based on a thought chain of a reference speech, and the thought chain is constructed based on speech segments and corresponding intentions after the reference speech is split;
[0006] The large model is used to annotate the intent of the speech to be annotated based on the prompt words to obtain intent annotation samples.
[0007] In a second aspect, the present disclosure provides an intent recognition model training method, comprising:
[0008] Acquire an intent-labeled sample, where the intent-labeled sample is labeled using any one of the intent-labeling methods provided in the embodiments of the present disclosure;
[0009] The intent-labeled samples are input into the model to be trained to obtain an intent recognition model.
[0010] In a third aspect, the present disclosure provides an intent recognition method, comprising:
[0011] Get the speech to be recognized;
[0012] An intention recognition model is used to perform intention recognition on the speech to be recognized to obtain an intention recognition result of the speech to be recognized, wherein the intention recognition model is trained using the intention recognition model training method provided by the embodiment of the present disclosure.
[0013] In a fourth aspect, the present disclosure provides an electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores one or more computer programs executable by the at least one processor, and one or more of the computer programs are executed by the at least one processor so that the at least one processor can execute the above-mentioned intent labeling method, and / or intent recognition model training method, and / or intent recognition method.
[0014] In the intention labeling method provided by the embodiment of the present disclosure, the prompt words input into the big model are generated based on the thought chain of the reference speech, and the thought chain is constructed based on the speech fragments and corresponding intentions after the reference speech is split. When the big model labels the intention, it also splits the speech to be labeled into speech fragments, and then uses the speech fragments and the corresponding intentions to construct the thought chain. This intention labeling process is the same as the human way of thinking. Therefore, the labeled intention is more accurate and comprehensive, which makes the intention recognition model trained with this intention labeling sample more accurate and stable. Moreover, the intention labeling method is fully automatic, which can improve the efficiency of intent labeling.
[0015] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The accompanying drawings are used to provide a further understanding of the present disclosure and constitute a part of the specification. Together with the embodiments of the present disclosure, they are used to explain the present disclosure and do not constitute a limitation of the present disclosure. The above and other features and advantages will become more apparent to those skilled in the art by describing detailed example embodiments with reference to the accompanying drawings. In the accompanying drawings:
[0017] Figure 1 A flowchart of a method for marking intent provided in an embodiment of the present disclosure;
[0018] Figure 2 A flowchart of obtaining prompt words provided in an embodiment of the present disclosure;
[0019] Figure 3 A flowchart of an intention mark provided in an embodiment of the present disclosure;
[0020] Figure 4 A flowchart of an intention mining annotation method provided by an embodiment of the present disclosure;
[0021] Figure 5 A flowchart of a deduplication process provided by an embodiment of the present disclosure;
[0022] Figure 6A flowchart of an intent recognition model training method provided in an embodiment of the present disclosure;
[0023] Figure 7 A flowchart of a method for identifying intent is provided for an embodiment of the present disclosure;
[0024] Figure 8 A schematic structural diagram of an intention marking device provided in an embodiment of the present disclosure;
[0025] Figure 9 A schematic diagram of the structure of an intent recognition model training device provided in an embodiment of the present disclosure;
[0026] Figure 10 A schematic diagram of the structure of an intention recognition device provided in an embodiment of the present disclosure;
[0027] Figure 11 A block diagram of an electronic device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0028] To enable those skilled in the art to better understand the technical solutions of the present disclosure, exemplary embodiments of the present disclosure are described below in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered merely exemplary. Therefore, those skilled in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0029] In the absence of conflict, the various embodiments of the present disclosure and the various features therein may be combined with each other.
[0030] As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.
[0031] The terms used herein are only used to describe specific embodiments and are not intended to limit the present disclosure. As used herein, the singular forms "a" and "the" are also intended to include the plural forms, unless the context clearly indicates otherwise. It will also be understood that when the terms "comprising" and / or "made of" are used in this specification, the presence of the features, wholes, steps, operations, elements and / or components is specified, but the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or groups thereof is not excluded. Similar words such as "connected" or "connected" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect.
[0032] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art. It will also be understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and the present disclosure, and will not be interpreted as having an idealized or overly formal meaning unless expressly defined as such herein.
[0033] In the technical solutions disclosed herein, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals. The use of user data in this technical solution complies with relevant national laws and regulations (for example, the "Information Security Technology Personal Information Security Specification", etc.). For example: corresponding prescribed measures are taken to control access to personal information; the display of personal information is subject to prescribed restrictions; the purpose of using personal information does not exceed the scope of direct or reasonable connection; when using personal information, clear identity reference is eliminated to avoid precise positioning of specific individuals.
[0034] In some related technologies, training intent recognition models requires a large number of training samples. The training samples are obtained by clustering the conversation text, manually defining the intent annotation of each category center and determining its meaning, and then manually annotating them based on pre-established annotation rules. This type of training sample requires a lot of manpower and material resources.
[0035] Pre-train the large model to obtain a pre-trained model, and then fine-tune the pre-trained model. Since the large model is a generative model, its stability and accuracy are poor, and fine-tuning the pre-trained model still requires a large number of labeled samples. Therefore, accurate labeled samples become an important factor in improving the accuracy and stability of the intent recognition model.
[0036] According to the embodiment of the present disclosure, the intent labeling method, model training method, intent recognition method and apparatus can be executed by an electronic device such as a terminal device or a server. The terminal device can be a user equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, a vehicle-mounted device, a wearable device, etc. The server can be an independent physical server, a server cluster composed of multiple physical servers, or a cloud server capable of cloud computing. The method can be implemented by a processor calling computer-readable program instructions stored in a memory.
[0037] In a first aspect, an embodiment of the present disclosure provides an intent labeling method.
[0038] Figure 1This is a flowchart of an intention annotation method provided by an embodiment of the present disclosure. Figure 1 , the intent annotation method includes:
[0039] Step S101, obtaining prompt words.
[0040] Among them, the prompt words are generated based on the thought chain of the reference speech, and the thought chain is constructed based on the speech fragments and corresponding intentions after the reference speech is split.
[0041] In some embodiments, as Figure 2 As shown, obtain prompt words, including:
[0042] Step S21, obtaining auxiliary elements corresponding to the reference speech.
[0043] In some embodiments, the auxiliary elements include one or more of roles, background information, and annotation requirements.
[0044] Step S22: split the reference speech to obtain multiple speech segments.
[0045] Step S23, generating an intention mining thought chain of a reference speech based on the intentions of multiple speech fragments.
[0046] Step S24, mining thought chains based on the reference words, auxiliary elements and the intentions corresponding to the reference words, and constructing prompt words.
[0047] In the disclosed embodiments, the role is to let the large model understand the role played, that is, the role of reference speech. For example, you are a senior telephone sales operations expert, and your current task is to analyze a given marketing text with reference to background information and provide corresponding intent annotation samples.
[0048] In the disclosed embodiments, background information allows the large model to understand the context of the referenced sales pitch, helping it understand the intended meaning of the referenced sales pitch. For example, the current marketing scenario is a financial loan scenario with a product called B Loan. The agent's goal is to call customers to inform them of current events or product information, guide them to participate in the event or use the product, answer questions they may have during the withdrawal process, and address any questions they may have about the product.
[0049] **The activities mainly include: increasing credit limits, reducing interest rates, issuing coupons, issuing interest-free coupons, participating in app lucky draws, and account upgrades.
[0050] **The main uses of the product include: reminding customers to use the B loan amount they applied for when it is approved or the contract is about to expire.
[0051] In the embodiment of the present disclosure, the marking requirements include but are not limited to the number of words and format of the marking.
[0052] For example, annotation requirement 1:
[0053] -When sorting out intent annotations, the given intent annotations should be as complete as possible, and the output intent annotations can represent the overall meaning of the marketing text.
[0054] -If the marketing text is short, just summarize it with an intent tag.
[0055] -When sorting out intent annotations, it is necessary to appropriately split the speech according to the semantic relationship and the logical relationship between the previous and next parts, and then summarize the intent of the split speech segments in turn. You can refer to the reference speech appropriately to give intent annotations.
[0056] -Intent annotations should be summarized in a concise and clear manner as much as possible, but key information should not be omitted. Variable information such as specific time, date, amount, etc. should be abstracted, and specific values should not be directly output.
[0057] -The intent annotations should include conversation details, such as the main content of the activity and what specific customer questions are being addressed.
[0058] Marking requirement 2:
[0059] -[Voice Input Timeout] means the client did not speak in the current turn.
[0060] - Each intention label should be as short as possible, no more than 6 words.
[0061] -You can refer to the background information to give the appropriate intent labels to the words. The intent labels need to be more detailed than the process, and try to avoid using the same labels as the process.
[0062] - Please only mark the current marketing language. Do not add any additional annotations that are not related to the reference language. Each intention annotation must be supported by a short sentence of marketing language and reflected in the thinking process.
[0063] The disclosed embodiment can split the reference script into multiple script segments according to natural sentence patterns. For example, the reference script is "Yes, our company has seen that your loan and repayment record is very good and you have never been overdue. We have opened up a credit limit of up to 200,000 yuan for your future use. If you need a higher credit limit, you only need to temporarily withdraw your remaining 500 yuan to your S bank card according to your personal needs, indicating that you really have no credit limit. After maintaining a good usage record, our company will give you an additional fixed credit limit of up to 50,000 yuan and a reduced interest rate activity opportunity, which will be given to you for your future use. Is that okay?" The reference script is split into:
[0064] Talk segment 1: Yes, our company has noticed that your loan and repayment record is very good and there has never been any overdue payment.
[0065] Talk segment 2: The credit limit here is up to 200,000 yuan for your convenience in the future.
[0066] Talk segment 3: If you need a higher credit limit.
[0067] Talk segment 4: You only need to temporarily withdraw your remaining 500 yuan to your bank card A according to your personal needs, indicating that you really have no available credit limit.
[0068] Talk segment 5: After you maintain a good usage record, our company will give you an additional fixed credit limit of up to 50,000 yuan and an opportunity to reduce interest rates, which will be given to you for your convenience in the future. Is that okay?
[0069] In the embodiment of the present disclosure, a thought chain of intention mining is generated based on the intention of the speech fragment. For example, the above speech fragments 1-5 are summarized respectively to obtain the intention of each speech fragment:
[0070] Talk segment 1: Yes, our company has noticed that you maintain an excellent loan and repayment record, with no overdue payments. Purpose: To build rapport and establish affinity by praising the customer's good usage record. Therefore, the intent of talk segment 1 can be summarized as: Praise the usage record.
[0071] Talk segment 2: The credit limit here is opened up to 200,000 yuan for your convenience in the future. Purpose: Introducing the highest credit limit. Therefore, the intention of talk segment 2 can be summarized as: Introducing the highest credit limit.
[0072] Talk segment 3: If you need a higher credit limit. Purpose: To suggest the possibility of a further increase in credit limit. Therefore, the purpose of talk segment 3 can be summarized as: exploring credit limit needs.
[0073] Talking Clip 4: You simply need to temporarily withdraw your remaining 500 yuan to your bank account A, based on your personal needs, to indicate that you truly have no available credit limit. Purpose: To guide the customer on how to meet the requirements for increasing their credit limit. Therefore, the intent of Talking Clip 4 can be summarized as: providing guidance on withdrawal procedures.
[0074] Talking Clip 5: "If you maintain a good usage record, our company will give you an additional fixed credit limit of up to 50,000 yuan and a reduced interest rate. This will be given to you for your convenience in the future. Is that okay?" Purpose: This promotion of increased credit limit and lowered interest rates also promotes customer conversion. Therefore, the purpose of Talking Clip 5 can be summarized as: increased credit limit and lowered interest rates.
[0075] In the embodiment of the present disclosure, a thought chain is mined based on reference words, roles, background information, annotation requirements, and the intention corresponding to the reference words to construct prompt words.
[0076] Based on the above reference words, roles, background information, and annotation requirements, the constructed prompt words are as follows:
[0077] ##Role
[0078] You are a senior telephone sales operations expert. Your current task is to analyze the given marketing text with reference to the background information and provide corresponding intent annotation samples.
[0079] Background
[0080] The current marketing scenario is a financial loan scenario. The product name is B Loan. The agent's purpose is to call customers to inform them of current activities or product information, guide them to participate in activities or use products, answer questions that arise during the withdrawal process, and answer questions about the product.
[0081] ##Marking requirements 1
[0082] -When sorting out intent annotations, the given intent annotations should be as complete as possible, and the output intent annotations can represent the overall meaning of the marketing text.
[0083] -If the marketing text is short, just summarize it with an intent tag.
[0084] -When sorting out intent annotations, it is necessary to appropriately split the speech according to the semantic relationship and the logical relationship between the previous and next parts, and then summarize the intent of the split speech segments in turn. You can refer to the reference speech appropriately to give intent annotations.
[0085] -Intent annotations should be summarized in a concise and clear manner as much as possible, but key information should not be omitted. Variable information such as specific time, date, amount, etc. should be abstracted, and specific values should not be directly output.
[0086] -The intent annotations should include conversation details, such as the main content of the activity and what specific customer questions are being addressed.
[0087] ##Marking requirement 2
[0088] -[Voice Input Timeout] means the client did not speak in the current turn.
[0089] - Each intention label should be as short as possible, no more than 6 words.
[0090] -You can refer to the background information to give the appropriate intent labels to the words. The intent labels need to be more detailed than the process, and try to avoid using the same labels as the process.
[0091] - Please only mark the current marketing language. Do not add additional annotations that are not related to the reference language. Each intention annotation must be supported by a short sentence of marketing language and reflected in the thinking process.
[0092] Let's think about this step by step based on an example:
[0093] ##Reference words
[0094] **Yes, our company has seen that you have maintained a very good record of borrowing and repayment, and you have never been overdue. We have opened up a credit limit of up to 200,000 yuan for your future use. If you need a higher credit limit, you only need to temporarily withdraw the remaining 500 yuan to your S bank card according to your personal needs, indicating that you really have no credit limit left. After maintaining a good usage record, our company will give you an additional fixed credit limit of up to 50,000 yuan and an opportunity to reduce interest rates, which will be given to you for your future use. Is that okay?
[0095] **Thought Process
[0096] First, we will divide the reference speech into the following segments according to semantics:
[0097] -Yes, our company has noticed that your loan and repayment record is very good and there has never been any overdue payment.
[0098] -The credit limit here is up to 200,000 yuan for your convenience in the future.
[0099] -If you need a higher limit.
[0100] -You only need to temporarily withdraw your remaining 500 yuan to your bank card A according to your personal needs, indicating that you really have no available credit limit.
[0101] -After you maintain a good usage record, our company will give you an additional fixed credit limit of up to 50,000 yuan and an opportunity to reduce interest rates, which will be given to you for your convenience in the future, okay?
[0102] So we finally get five intent annotations!
[0103] Then, we summarize each segment in turn:
[0104] -Yes, our company has noticed that you maintain an excellent loan and repayment record, with no overdue payments. Purpose: To build rapport and affinity by praising the customer's good usage record. Therefore, the first intention can be summarized as: Praise usage record.
[0105] -The credit limit here is opened up to 200,000 yuan for your convenience in the future. Purpose: To introduce the highest credit limit. Therefore, the second intention can be summarized as: to introduce the highest credit limit.
[0106] If you need a higher credit limit. Purpose: To suggest the possibility of a further credit limit increase. Therefore, the third intention can be summarized as: exploring credit limit needs.
[0107] - Based on your personal needs, you only need to temporarily withdraw your remaining 500 yuan to your bank account A, indicating that you truly have no available credit limit. Purpose: To guide customers on how to meet the requirements for increasing their credit limit. Therefore, the fourth intention can be summarized as: Guiding withdrawal operations.
[0108] -If you maintain a good usage record, our company will give you an additional fixed credit limit of up to 50,000 yuan and a reduced interest rate promotion opportunity. This will be given to you for your convenience in the future. Is that okay? Purpose: To introduce the promotion of increased credit limit and lower interest rates, while also promoting customer conversion. Therefore, the fifth intention can be summarized as: Increase credit limit and reduce interest rates promotion.
[0109] ##Output speech fragments and intention annotations
[0110] Oh, sir, the purpose of our call is to help you resolve this issue! Yes, our company has noticed that you maintain an excellent loan and repayment record, with no overdue payments ['Praise your record']. We've increased your credit limit to a maximum of 200,000 yuan for your future use ['Introducing the maximum credit limit']. If you require a higher credit limit ['Explore credit limit requirements']. Based on your personal needs, you can simply temporarily withdraw your remaining 500 yuan to your bank card A, indicating that you have no available credit ['Withdrawal instructions']. If you maintain a good credit record, our company will grant you an additional fixed credit limit of up to 50,000 yuan and a reduced interest rate for your future use. Is that okay ['Limit limit increase and interest rate reduction promotion']?
[0111] Make appropriate splits according to the 'semantic relationships' and 'logical relationships' of the current reference speech, and then summarize and annotate the intentions of the split speech segments in turn.
[0112] Okay, I got it. I will directly output the speech snippets and corresponding tags. Let’s get started!
[0113] **Current agent marketing techniques:
[0114] {}
[0115] ##Output the speech fragments and tags:
[0116] ""
[0117] In order to better understand this embodiment, the content of another prompt word is introduced below:
[0118] It should be noted that the roles, background information, and annotation requirements are the same as those in the above embodiment, and will not be repeated here to save space.
[0119] **Reference agent's words:
[0120] Hello, sir! I'm from Loan Company B. I'm calling to confirm that you're the one using Loan Company B, correct?
[0121] **Thought Process:
[0122] First, we will refer to the agent's speech and divide it into the following segments according to semantics:
[0123] -Hello, sir! I'm from B Loan Company.
[0124] -I'm calling you to confirm that you are the one using loan B, right?
[0125] So in the end there are two intent annotations!
[0126] Then, we summarize each segment in turn:
[0127] -'Hello, sir! I am an employee of B Loan Company.' This segment of speech is a self-introduction, so the first intention can be summarized as: self-introduction.
[0128] -'I'm calling you to confirm that you're the one using Loan B, right?' This segment of the conversation is confirming whether the customer is actually using the product, so the second intention can be summarized as: confirming the account user.
[0129] ##Output speech fragments and intention annotations
[0130] Hello, sir! I'm from Loan Company B ['Introduce Myself']. I'm calling to confirm that you're using Loan Company B, correct? ['Confirm Account Owner'].
[0131] Make appropriate splits according to the 'semantic relationships' and 'logical relationships' of the current reference speech, and then summarize and annotate the intentions of the split speech segments one by one.
[0132] Okay, I got it. I will directly output the speech snippets and corresponding tags. Let’s get started!
[0133] **Current agent marketing techniques:
[0134] {}
[0135] ##Output the speech fragments and tags:
[0136] ""
[0137] It should be noted that, although each of the above reference examples is used as a prompt word, this embodiment is not limited thereto. In fact, the above two reference examples may be used as one prompt word.
[0138] Step S102: Use the large model to annotate the intent of the speech to be annotated based on the prompt words to obtain intent annotation samples.
[0139] Since the large model labels the intent of the labeled speech based on prompt words, and the prompt words include the speech fragments of the reference speech and the thought chain constructed by its intent, that is, the prompt words include people's way of thinking, the large model uses people's way of thinking to label intent, and the output intent labeling samples are more accurate and consistent. Therefore, the intent recognition model trained using such intent labeling samples has better accuracy and consistency.
[0140] In some embodiments, before obtaining intent annotation samples, the method further includes: clustering the to-be-annotated speech in a speech library to obtain a first clustering result, wherein the first clustering result includes one or more cluster categories, each cluster category including multiple to-be-annotated speech. A reference speech corresponding to the to-be-annotated speech is determined based on the first clustering result.
[0141] Among them, the script library is a collection of scripts generated by agents in the process of handling actual business.
[0142] The embodiments of the present disclosure do not limit the clustering algorithm used to cluster the speech to be labeled. Each cluster category in the first clustering result is processed separately. Since each cluster category includes multiple speech to be labeled, when performing intent labeling on the speech to be labeled, the labeled speech in the cluster category can be used as a reference speech.
[0143] In some embodiments, the reference speech corresponding to the speech to be labeled is determined based on the first clustering result, including: calculating the distance between the speech to be labeled and the central speech within the clustering category, and sorting the speech to be labeled based on the distance to obtain a first sorting result; wherein the central speech is one of the speech to be labeled within the clustering category; and determining the reference speech of the speech to be labeled one by one based on the first sorting result.
[0144] In the embodiment of the present disclosure, a central speech can be selected for each cluster category. The central speech can be selected by the user, or a speech to be labeled can be arbitrarily selected from the center position of the cluster category. The embodiment of the present disclosure does not limit the method for determining the central speech.
[0145] The distance between other to-be-annotated phrases in the same cluster and the central phrase is calculated, and the to-be-annotated phrases are sorted based on the distance to obtain a first sorting result. In this first sorting result, to-be-annotated phrases with similar intent are ranked similarly. When annotating intents, reference phrases are determined based on the first sorting result. Phrases with similar intent can be used as reference phrases for the to-be-annotated phrase.
[0146] In some embodiments, reference words for the words to be marked are determined one by one based on the first sorting result, including: taking the words to be marked as the current words to be marked one by one according to the sorting order in the first sorting result; taking the marked words that are before the current words to be marked and within the marking window as reference words for the previous words to be marked, wherein the width of the marking window is preset.
[0147] In an embodiment of the present disclosure, a marking window is used to determine reference words for the current word to be marked one by one according to the first sorting result, wherein the size of the marking window can be pre-set by the user, and the embodiment of the present disclosure does not limit the size of the marking window.
[0148] The marked words within the marking window are used as reference words for the words to be marked. Assume that the width of the marking window is k, where k is a positive integer. When the nth word to be marked in the first sorting result is intentionally marked, the previous k marked words can be used as reference words, that is, the n-1th to be marked words to the nkth marked words are used as reference words. If there are less than k marked words before the word to be marked, no additional marked words will be added. If the number of marked words is large and exceeds the width of the marking window, then the marked words that are more than k away from the word to be marked will be discarded, and only the k closest marked words will be retained.
[0149] For example, when annotating the second speech to be annotated, the first annotated speech is used as a reference speech. When annotating the third speech to be annotated, the first annotated speech and the second annotated speech are used as reference speech. It should be noted that the first speech to be annotated can be annotated without a reference speech.
[0150] In the embodiment of the present disclosure, the to-be-annotated phrases within the same cluster category are placed into the to-be-annotated set, and the already-annotated phrases are placed into the already-annotated set. The to-be-annotated phrases in the to-be-annotated set are annotated in sequence according to the first sorting result. Each time one is annotated, it is placed into the already-annotated set. In the already-annotated set, the already-annotated phrases are sorted in the order of annotation. It can be understood that the sorting result according to the order of annotation is the same as the sorting result according to the distance.
[0151] Select a to-be-annotated phrase from the to-be-annotated set in order, select a certain number of annotated phrases from the annotated set as reference phrases according to the width of the annotation window, and finally select the annotated phrases that enter the annotated set and are within the annotation window as reference phrases for the to-be-annotated phrase.
[0152] In order to better understand the process of intent annotation, Figure 3 Further details are given.
[0153] Figure 3 This is a flowchart of an embodiment of the present disclosure. Figure 3 As shown in the figure, the process of intent annotation includes:
[0154] Step S301: cluster the to-be-annotated dialogues in the dialogue library to obtain a first clustering result.
[0155] Step S302: Divide the to-be-annotated speech into multiple batches according to the first clustering result, with each batch corresponding to a cluster category.
[0156] Step S303 : calculating the distance between the to-be-annotated speech and the central speech in the same cluster category, and sorting the to-be-annotated speech based on the distance to obtain a first sorting result.
[0157] Table 1 shows the identifier of the cluster category (cluster_id), the words to be labeled within the cluster category (text), the central words of the second cluster category (core_text), and the distance (distance) between each word to be labeled and the central word. Among them, the central word of the second cluster category is "Ah, Madam, we are a formal licensed financial institution."
[0158] Table 1
[0159]
[0160]
[0161]
[0162] Step S304: set a labeling window, use the big model to label the intended words one by one, and add the labeled words to the labeled set.
[0163] The large model performs intent annotation on the speech to be annotated based on the reference speech, and uses the reference speech to construct prompt words. The prompt words are constructed by splitting the reference speech into multiple speech segments, and then based on the speech segments and their corresponding intents.
[0164] When selecting reference phrases, if the number of tagged phrases in the tagged set exceeds the annotation window, the earliest tagged phrase is discarded, and only the most recently tagged phrase is used as the reference phrase. Tagged phrases serve as historical context in the prompt, helping the large model more accurately understand the intent of the phrase to be tagged.
[0165] Step S303 and step S304 are executed in a loop until all the to-be-annotated speech in all cluster categories are annotated.
[0166] In the embodiment of the present disclosure, the process of intent labeling is also a process of intent mining. The large model labels the labeled speech based on the prompt words. Since the large model is random and the labeling window cannot be set to infinitely large, the intent labeling output by the large model has omissions. In order to further refine the intent labeling and improve the accuracy of the intent labeling, the embodiment of the present disclosure can further mine the intent standards after the large model outputs the intent labeling samples.
[0167] In some embodiments, the cluster category includes a central speech segment and a second speech segment, and the second speech segment is any speech segment other than the central speech segment within the cluster category. After using the large model to annotate the intent of the speech to be annotated based on the prompt word and outputting the intent annotation sample, it also includes: obtaining multiple speech segments and their intent annotations in the intent annotation sample; clustering the multiple speech segments to obtain a second clustering result, and the second clustering result includes at least one cluster category; judging whether the semantics of the central speech segment of the cluster category are consistent with the second speech segment in the cluster category, and if the semantics of the central speech segment and the second speech segment are inconsistent, the intent annotation of the second speech segment is used as the intent annotation sample; if the semantics of the central speech segment and the second speech segment are consistent, the intent annotation of the central speech segment is used to update the intent annotation sample of the second speech segment.
[0168] Figure 4 This is a flowchart of a mining intention annotation provided by an embodiment of the present disclosure. Figure 4 As shown, the mining intention annotation provided by the embodiment of the present disclosure includes the following steps:
[0169] Step S401: Acquire multiple speech fragments and their intent annotations in the intent annotation sample.
[0170] Among them, the intent annotation samples are output by the large model, and the intent annotation samples output by the large model include the intent annotations of each speech fragment.
[0171] For example, the following snippets and their intended meanings are marked:
[0172] Hello, sir! I am an employee of B Loan Company ['Introduce myself'].
[0173] I'm calling you to confirm that you are the one using Loan B, correct? ['Confirm account owner'].
[0174] I see you haven't used your balance since you settled it. Are you dissatisfied with the interest rate plan? Or have you encountered some other issues? ['Explore reasons for non-use'].
[0175] Step S402: cluster the multiple speech segments to obtain a second clustering result, where the second clustering result includes at least one cluster category.
[0176] The embodiment of the present disclosure clusters multiple speech segments based on the intentions of the speech segments, and clusters speech segments with similar intentions into one cluster category.
[0177] The second clustering result includes at least one cluster category, and each cluster category includes a plurality of speech segments.
[0178] In step S403, the large model is used to determine whether the semantics of each speech segment in each cluster category are consistent with the central speech segment. If so, step S404 is executed; if not, step S405 is executed.
[0179] In step S403, the large model may be a qwen72b model. For each cluster category, the qwen72b model is used to determine whether the speech segment is semantically consistent with the central speech segment (core speech segment) within the cluster category. The central speech segment can be selected by the user from the cluster category or automatically from the cluster category.
[0180] Step S404: Mark the intention of the speech segment and replace it with the central speech segment.
[0181] It should be noted that during clustering, even if the intent labels of the speech fragments are not completely consistent, the intents are similar, so they are clustered in one cluster category. In other words, in a cluster category, the content of the intent labels may not be completely consistent, but the intents are similar.
[0182] Step S405: retain the speech fragment and intention annotation.
[0183] For example, the central speech segment in the cluster category is: "I am calling to recheck your credit limit and interest with you", with the intent annotation: ['check credit limit and interest'].
[0184] Other speech fragments within the cluster category: "I'm calling to help you recheck your credit limit and interest rate", intention label: ['State the purpose of the call'].
[0185] Through the above-mentioned intention annotation mining, the intention annotation samples can be expanded to make the intention annotation samples more comprehensive and accurate.
[0186] In the embodiment of the present disclosure, the speech library is a collection of speech that is generated during the actual use of the platform. Therefore, there are repetitions in the speech library, which affects the efficiency of generating intent annotation samples. Therefore, before generating intent annotation samples, it is necessary to deduplicate the speech to be annotated in the speech library.
[0187] In some embodiments, clustering is performed on the unlabeled speech in the speech library, and before obtaining the first clustering result, the process includes: sorting the candidate speech in the speech library according to the character encoding value of the candidate speech to obtain a second sorting result; adding the candidate speech ranked first in the second sorting result to the deduplication speech library, and the candidate speech in other positions are used as the current candidate speech in the sorting order, and performing the following steps: calculating the edit distance between the current candidate speech and the previous candidate speech added to the deduplication speech library; when the edit distance is less than a preset threshold, removing the current candidate speech; when the edit distance is greater than or equal to the preset threshold, adding the current candidate speech as the speech to be labeled to the deduplication speech library.
[0188] In an embodiment of the present disclosure, a Unicode encoder can be used to perform character encoding on candidate words to obtain character encoding values, and then the candidate words can be sorted according to the size of the character encoding values to obtain a second sorting result, and then deduplication processing can be performed based on the second sorting result.
[0189] For example, Figure 5 This is a flowchart of a deduplication process provided by an embodiment of the present disclosure. Figure 5 As shown, the deduplication process provided by the embodiment of the present disclosure includes the following steps:
[0190] Step S501: sort the candidate speech words according to the character code value.
[0191] A Unicode encoder is used to obtain the character code values of candidate phrases in the database of phrases to be deduplicated, and the candidate phrases are sorted by their character code values to obtain a second sorting result. The candidate phrases may be phrases obtained from a customer service system. The database of phrases to be deduplicated refers to a database of candidate phrases that have not yet been deduplicated.
[0192] In step S502, the first candidate speech is used as the deduplicated speech, and step S507 is executed.
[0193] The first candidate speech is the candidate speech ranked first in the second sorting result. For ease of description, in this embodiment, according to the second sorting result, the first candidate speech is referred to as the first candidate speech, the second candidate speech is referred to as the second candidate speech, and so on, the nth candidate speech is referred to as the nth candidate speech.
[0194] The first candidate phrase is placed in the deduplicated phrase library as the phrase to be labeled. In the disclosed embodiment, the phrase to be labeled is a phrase that has been deduplicated. At this point, the deduplicated phrase library contains only one phrase to be labeled, referred to as the first phrase to be labeled.
[0195] Step S503: Obtain the i-th candidate speech.
[0196] The i-th candidate speech is the candidate speech ranked at the i-th position in the second sorting result.
[0197] Step S504 , calculating the edit distance between the i-th candidate speech and the previous speech to be annotated that was added to the deduplication speech library.
[0198] The last unlabeled phrase added to the deduplication database is the most recent unlabeled phrase that entered the database. The deduplication database is a collection of candidate phrases that have been determined to be retained after deduplication. After each deduplication determination, similar candidate phrases are discarded and dissimilar candidate phrases are added to the deduplication database as unlabeled phrases. Therefore, the last unlabeled phrase to enter the deduplication database is the most recently determined and retained candidate phrase.
[0199] Step S505 , determining whether the edit distance is less than a preset threshold, if so, executing step S506 ; if not, executing step S507 .
[0200] Step S506: discard.
[0201] When the edit distance between the i-th candidate speech and the previous speech to be labeled added to the deduplication speech library is less than a preset threshold, it means that the i-th candidate speech is similar to the previous speech to be labeled added to the deduplication speech library, and the i-th candidate speech is discarded.
[0202] Step S507: put the i-th candidate speech into the deduplicated speech library.
[0203] When the edit distance between the i-th candidate speech and the previous candidate speech added to the deduplication speech library is greater than or equal to a preset threshold, it means that the i-th candidate speech is not similar to the previous speech to be labeled added to the deduplication speech library, and the i-th candidate speech is added to the deduplication speech library.
[0204] In the intention labeling method provided by the embodiment of the present disclosure, the prompt words input into the big model are generated based on the thought chain of the reference speech, and the thought chain is constructed based on the speech fragments and corresponding intentions after the reference speech is split. When the big model labels the intention, it also splits the speech to be labeled into speech fragments, and then uses the speech fragments and the corresponding intentions to construct the thought chain. This intention labeling process is the same as the human way of thinking. Therefore, the labeled intention is more accurate and comprehensive, which makes the intention recognition model trained with this intention labeling sample more accurate and stable. Moreover, the intention labeling method is fully automatic, which can improve the efficiency of intent labeling.
[0205] In a second aspect, an embodiment of the present disclosure provides a method for training an intent recognition model.
[0206] Figure 6 This is a flow chart of a method for training an intent recognition model provided by an embodiment of the present disclosure. Figure 6 As shown, the intent recognition model training method provided by the embodiment of the present disclosure includes:
[0207] Step S601: Obtain intent-labeled samples.
[0208] The intent-labeled samples are labeled using any of the intent-labeling methods provided in the embodiments of the present disclosure.
[0209] Step S602: Input the intent-labeled samples into the model to be trained to obtain an intent recognition model.
[0210] The model to be trained can be a BERT model. The intent annotation samples output by the large model are input into the BERT model to obtain an intent recognition model.
[0211] An embodiment of the present disclosure provides a method for training an intent recognition model, in which the intent labeling samples input to the model to be trained are samples labeled by a large model through prompt words, and the prompt words are produced based on the thinking chain of the reference speech. The thinking chain is constructed based on the speech fragments and corresponding intents after the reference speech is split, so that the large model labels the samples according to the human way of thinking. Therefore, these intent labeling samples are more accurate and comprehensive, and the accuracy and stability of the intent recognition model trained thereby are higher.
[0212] In a third aspect, an embodiment of the present disclosure provides a method for intent recognition.
[0213] Figure 7 A flowchart of an intention recognition method is provided for an embodiment of the present disclosure. Figure 7 As shown, the intention recognition method provided by the embodiment of the present disclosure includes:
[0214] Step S701: Obtain the speech to be recognized.
[0215] Step S702: Use the intention recognition model to perform intent recognition on the speech to be recognized, and obtain the intention recognition result of the speech to be recognized, wherein the intention recognition model is trained using the intention recognition model training method provided by the embodiment of the present disclosure.
[0216] The embodiment of the present disclosure provides an intent recognition method using an intent recognition model for recognition. The intent recognition model is trained through intent-labeled samples, and the intent-labeled samples are samples labeled by a large model through prompt words. The prompt words are generated based on the thinking chain of the reference speech, and the thinking chain is constructed based on the speech fragments and corresponding intentions after the reference speech is split. This allows the large model to label samples according to human thinking. These intent-labeled samples are more accurate and comprehensive, and the accuracy and stability of the intent recognition model trained thereby are higher. Therefore, the intent recognized by the intent recognition model is more accurate.
[0217] It is understood that the above-mentioned various method embodiments mentioned in this disclosure can be combined with each other to form combined embodiments without violating the principle logic. Due to space limitations, this disclosure will not go into details. It is understood by those skilled in the art that in the above-mentioned methods of specific implementation, the specific execution order of each step should be determined by its function and possible internal logic.
[0218] In a fourth aspect, an embodiment of the present disclosure provides an intention labeling device.
[0219] Figure 8 This is a schematic diagram of the structure of an intention marking device provided by an embodiment of the present disclosure. Figure 8 As shown, an embodiment of the present disclosure provides an intention labeling device, including:
[0220] The first acquisition module 801 is used to acquire prompt words, wherein the prompt words are generated based on the thought chain of the reference speech, and the thought chain is generated after splitting the reference speech into speech segments.
[0221] The annotation module 802 is used to use the large model to annotate the intention of the annotated speech based on the prompt words, and output the intention annotation sample.
[0222] In the embodiment of the present disclosure, the annotation module 802 can execute other steps in the intention annotation method provided in the embodiment of the present disclosure, which will not be repeated here to save space.
[0223] The intention labeling device provided by the embodiment of the present disclosure has a first acquisition module that acquires prompt words, which are generated based on the thought chain of the reference speech, and the thought chain is constructed based on the speech fragments and corresponding intentions after the reference speech is split. When the labeling module uses a large model to label the intention of the labeled speech based on the prompt words, the speech to be labeled is also split into speech fragments, and then the speech fragments and the corresponding intentions are used to construct the thought chain. This intention labeling process is the same as the way people think. Therefore, the intentions labeled by the intention labeling device are more accurate and comprehensive, so that the accuracy and stability of the intention recognition model trained using this intention labeling sample are higher. Moreover, the intention labeling method is fully automatic, which can improve the efficiency of intention labeling.
[0224] In a fifth aspect, an embodiment of the present disclosure provides an intent recognition model training device.
[0225] Figure 9 A schematic diagram of the structure of an intent recognition model training device provided in an embodiment of the present disclosure.
[0226] like Figure 9 As shown, an embodiment of the present disclosure provides an intent recognition model training device, comprising:
[0227] The second acquisition module 901 is used to acquire intent-labeled samples, where the intent-labeled samples are labeled using any one of the intent labeling methods provided in the embodiments of the present disclosure.
[0228] The training module 902 is used to input the intent-labeled samples into the model to be trained to obtain an intent recognition model.
[0229] An embodiment of the present disclosure provides an intent recognition model training device, in which a second acquisition module acquires intent labeled samples, and the intent labeled samples input to the training module into the model to be trained are samples labeled by a large model through prompt words, and the prompt words are produced based on the thinking chain of the reference speech, and the thinking chain is constructed based on the speech fragments and corresponding intentions after the reference speech is split, so that the large model labels samples according to the human way of thinking. Therefore, these intent labeled samples are more accurate and comprehensive, and the accuracy and stability of the intent recognition model trained thereby are higher.
[0230] In a sixth aspect, an embodiment of the present disclosure provides an intention recognition device.
[0231] Figure 10 This is a schematic diagram of the structure of an intention recognition device provided by an embodiment of the present disclosure. Figure 10 As shown, an embodiment of the present disclosure provides an intention recognition device, including:
[0232] The third acquisition module 1001 is used to acquire the speech to be recognized.
[0233] The recognition module 1002 is used to use the intention recognition model to perform intent recognition on the speech to be recognized and obtain the intention recognition result of the speech to be recognized, wherein the intention recognition model is trained using the intention recognition model training method provided by the embodiment of the present disclosure.
[0234] An embodiment of the present disclosure provides an intention recognition device, in which a third acquisition module obtains the speech to be recognized, and the recognition module uses an intention recognition model for recognition. The intention recognition model is trained through intention-labeled samples, and the intention-labeled samples are samples labeled by a large model through prompt words. The prompt words are generated based on the thinking chain of the reference speech, and the thinking chain is constructed based on the speech fragments and corresponding intentions after the reference speech is split, so that the large model labels samples according to the human way of thinking. These intention-labeled samples are more accurate and comprehensive, and the accuracy and stability of the intention recognition model trained thereby are higher. Therefore, the intention recognized by the intention recognition model is more accurate.
[0235] Each module in the above-mentioned apparatus may be implemented in whole or in part by software, hardware, or a combination thereof. Each module may be embedded in or independent of a processor in a computer device in the form of hardware, or may be stored in a memory in the computer device in the form of software, so that the processor can call and execute the operations corresponding to each module.
[0236] Figure 11 A block diagram of an electronic device provided in an embodiment of the present disclosure.
[0237] Reference Figure 11 , an embodiment of the present disclosure provides an electronic device, which includes: at least one processor 1101; a memory 1102 communicatively connected to the at least one processor 1101, and one or more I / O interfaces 1103 connected between the processor 1101 and the memory 1102; wherein the memory 1102 stores one or more computer programs that can be executed by the at least one processor 1101, and the one or more computer programs are executed by the at least one processor 1101 to enable the at least one processor 1101 to execute the above-mentioned intent labeling method, and / or execute the above-mentioned intent recognition model training method, and / or execute the above-mentioned intent recognition method.
[0238] Each module in the above-mentioned electronic device can be implemented in whole or in part through software, hardware, or a combination thereof. Each module can be embedded in or independent of the processor in the computer device in hardware form, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the corresponding operations of each module.
[0239] The present disclosure also provides a computer-readable storage medium having a computer program stored thereon. When executed by a processor / processing core, the computer program implements the aforementioned intent labeling method, and / or performs the aforementioned intent recognition model training method, and / or performs the aforementioned intent recognition method. The computer-readable storage medium may be volatile or non-volatile computer-readable storage medium.
[0240] An embodiment of the present disclosure also provides a computer program product, including a computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs in a processor of an electronic device, the processor in the electronic device executes the above-mentioned intent labeling method, and / or executes the above-mentioned intent recognition model training method, and / or executes the above-mentioned intent recognition method.
[0241] It will be understood by those skilled in the art that all or some of the steps, systems, and functional modules / units in the methods disclosed above may be implemented as software, firmware, hardware, and appropriate combinations thereof. In a hardware implementation, the division between the functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed by several physical components in cooperation. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or may be implemented as hardware, or may be implemented as an integrated circuit, such as an application-specific integrated circuit. Such software may be distributed on a computer-readable storage medium, which may include a computer storage medium (or non-transitory medium) and a communication medium (or temporary medium).
[0242] As is well known to those skilled in the art, the term computer storage media includes volatile and nonvolatile, removable and non-removable media implemented in any method or technology for storage of information (such as computer-readable program instructions, data structures, program modules or other data). Computer storage media includes, but is not limited to, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), static random access memory (SRAM), flash memory or other memory technology, portable compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical disc storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, as is well known to those skilled in the art, communication media typically contains computer-readable program instructions, data structures, program modules or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.
[0243] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions to be stored in the computer-readable storage medium in each computing / processing device.
[0244] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, and conventional procedural programming languages such as "C" language or similar programming languages. Computer-readable program instructions may be executed entirely on a user's computer, partially on a user's computer, as an independent software package, partially on a user's computer, partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., utilizing an Internet service provider to connect via the Internet). In some embodiments, an electronic circuit, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), may be personalized by utilizing the state information of the computer-readable program instructions. The electronic circuit may execute the computer-readable program instructions, thereby realizing various aspects of the present disclosure.
[0245] The computer program product described herein may be implemented in hardware, software, or a combination thereof. In one embodiment, the computer program product is implemented as a computer storage medium. In another embodiment, the computer program product is implemented as a software product, such as a software development kit (SDK).
[0246] Various aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0247] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the processor of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.
[0248] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more blocks in the flowchart and / or block diagram.
[0249] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple embodiments of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part of a module, program segment or instruction, and the part of the module, program segment or instruction contains one or more executable instructions for realizing the prescribed logical function. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the prescribed function or action, or can be implemented by a combination of dedicated hardware and computer instructions.
[0250] Example embodiments have been disclosed herein, and although specific terms are employed, they are used and should be interpreted only in a general illustrative sense and not for purposes of limitation. In some instances, it will be apparent to those skilled in the art that, unless otherwise expressly indicated, features, characteristics, and / or elements described in conjunction with a particular embodiment may be used alone or in combination with features, characteristics, and / or elements described in conjunction with other embodiments. Therefore, it will be understood by those skilled in the art that various changes in form and detail may be made without departing from the scope of the present disclosure as set forth in the appended claims.
Claims
1. A method for labeling intent, characterized in that: include: Obtaining prompt words, wherein the prompt words are generated based on a thought chain of a reference speech, and the thought chain is constructed based on speech segments and corresponding intentions after the reference speech is split; The large model is used to annotate the intent of the speech to be annotated based on the prompt words to obtain intent annotation samples.
2. The method according to claim 1, characterized in that The obtaining of prompt words includes: Obtaining auxiliary elements corresponding to the reference speech; Splitting the reference speech to obtain multiple speech segments; Generate an intention mining thought chain of the reference speech based on the intentions of the multiple speech fragments; The prompt word is constructed based on the reference speech, the auxiliary elements, and the intention mining thought chain corresponding to the reference speech.
3. The method according to claim 1, characterized in that The method further includes: labeling the intent of the to-be-labeled speech based on the prompt word using the large model, and obtaining the intent labeling sample; Clustering the to-be-annotated speech in the speech library to obtain a first clustering result, wherein the first clustering result includes one or more cluster categories; A reference speech corresponding to the speech to be labeled is determined based on the first clustering result.
4. The method according to claim 3, characterized in that The determining, based on the first clustering result, a reference speech corresponding to the speech to be labeled, includes: Calculating the distance between the to-be-annotated speech in the cluster category and the central speech, and sorting the to-be-annotated speech based on the distance to obtain a first sorting result, wherein the central speech is one of the to-be-annotated speech in the cluster category; Reference speech terms for the speech terms to be marked are determined one by one based on the first sorting result.
5. The method according to claim 4, characterized in that The step of determining reference speech words for the speech words to be marked one by one based on the first sorting result includes: According to the sorting order in the first sorting result, the to-be-marked speech phrases are taken as current to-be-marked speech phrases one by one; The marked speech that is before the current speech to be marked in the sorting order and is within the marking window is used as a reference speech for the current speech to be marked, wherein the width of the marking window is preset.
6. The method according to claim 3, characterized in that The clustering process of the to-be-annotated speech in the speech library, before obtaining the first clustering result, includes: Sorting the candidate speech phrases in the speech phrase library according to the character encoding values of the candidate speech phrases to obtain a second sorting result; The candidate speech ranked first in the second sorting result is added to the deduplication speech library as the speech to be marked, and the candidate speech in other positions is sequentially used as the current candidate speech in the order of the sorting in the second sorting result, and the following steps are performed: Calculating the edit distance between the current candidate speech and the speech to be marked that was previously added to the deduplication speech library; If the edit distance is less than a preset threshold, the current candidate speech is removed; When the edit distance is greater than or equal to the preset threshold, the current candidate speech is added to the deduplication speech library as the speech to be labeled.
7. The method according to claim 3, characterized in that The clustering category includes a central speech segment and a second speech segment. After the large model is used to label the intent of the speech to be labeled based on the prompt word and the intent labeling sample is output, the method further includes: Obtaining multiple speech fragments and their intent annotations in the intent annotation sample; performing clustering processing on the plurality of speech segments to obtain a second clustering result, wherein the second clustering result includes at least one cluster category; Determining whether the central speech segment of the cluster category is semantically consistent with a second speech segment within the cluster category; In the case where the semantics of the central speech segment and the second speech segment are inconsistent, taking the intention annotation of the second speech segment as an intention annotation sample; When the semantics of the central speech segment are consistent with those of the second speech segment, the intent annotation sample of the second speech segment is updated using the intent annotation of the central speech segment.
8. A method for training an intent recognition model, characterized in that: include: Obtaining an intent-annotated sample, wherein the intent-annotated sample is annotated using any one of the intent-annotating methods provided in claims 1 to 7; The intent-labeled samples are input into the model to be trained to obtain an intent recognition model.
9. A method for identifying intention, characterized in that: include: Get the speech to be recognized; An intention recognition model is used to perform intention recognition on the speech to be recognized to obtain an intention recognition result of the speech to be recognized, wherein the intention recognition model is trained using the intention recognition model training method described in claim 8.
10. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores one or more computer programs that can be executed by the at least one processor, and the one or more computer programs are executed by the at least one processor so that the at least one processor can execute the intent labeling method described in any one of claims 1 to 7, and / or execute the intent recognition model training method described in claim 8, and / or execute the intent recognition method described in claim 9.